What happens when GPUs get cheaper?

In our AI infrastructure market deck, you will find everything you need to understand the market
SUMMARY
Cheaper GPUs will make AI usage explode rather than shrink the industry, while shifting the real scarcity toward power, memory, networking, data-center capacity and software that keeps the hardware busy.
The price of a flagship system can rise at the same time as useful compute gets cheaper. Buyers are paying more for the rack, but much less for each experiment, token or completed AI task.
Lower unit costs are unlikely to reduce total AI spending soon. Companies are using every efficiency gain to add users, longer contexts, voice, video, agents and more ambitious research programs.
Frontier-model economics will split in two. Reproducing yesterday’s capability becomes cheaper, while leading laboratories reinvest the savings and keep pushing the cost of the next frontier model higher.
Inference pricing will divide into a commodity layer and a premium layer. Classification, extraction and routine writing become nearly invisible costs, while difficult reasoning, real-time media and high-reliability automation continue to command meaningful prices.
Agents are the biggest mechanism for absorbing cheaper compute. One visible task may trigger hundreds of hidden model calls, tool actions, retries and verification steps, so usage can grow much faster than the cost of an individual token falls.
Generic AI features will lose pricing power first. Products that own a valuable workflow, proprietary data, distribution, compliance or accountability can still charge far more than their underlying inference bill.
Open-weight models and on-device AI gain substantially because lower hardware requirements make private, customized and local deployment practical for more organizations. Most products will still use a hybrid approach, keeping routine work on the device and sending harder jobs to the cloud.
NVIDIA can continue growing even as compute gets cheaper, but stable workloads create room for custom chips from hyperscalers and large model providers. General-purpose GPUs remain strongest where flexibility, mature software and rapid architectural change matter most.
The broad result is both democratization and concentration. Many more developers can access capable AI, while ownership of the largest clusters remains concentrated among companies that can secure enormous amounts of capital, electricity, memory, networking and land.

This market map, featured in our AI infrastructure market deck, highlights top companies and startups in the AI infrastructure market
Are AI GPUs actually getting cheaper today?
Yes. Useful AI compute is getting much cheaper today, even though the newest GPU systems often carry higher price tags.
A better way to measure GPU prices is to ask how much work a buyer gets for each dollar. Epoch AI’s hardware database, updated recently with more than 170 accelerators, finds that AI chip performance per dollar has improved by about 37% a year since 2012. At that pace, a fixed calculation becomes roughly 80% cheaper over five years.
The hardware bill can still rise. A modern rack can cost far more than an older server because it contains more accelerators, high-bandwidth memory, networking and cooling. Epoch estimates that NVIDIA’s GB300 costs nearly nine times as much as a P100 did at launch, while delivering about 24 times more performance per dollar. Buyers spend more on the box and far less on each unit of useful computation.
Software pushes the cost down again. Better low-level code and serving methods let the same hardware produce more tokens. Cloud discounts can help too, although the hourly rental price usually falls more slowly because it also covers CPUs, storage, networking, buildings and the provider’s margin.
So when we say GPUs are getting cheaper, we mean that training one experiment, generating one token or serving one user requires fewer dollars than before. The sticker price alone answers the wrong question.
| What we measure | Current direction | What the buyer experiences |
|---|---|---|
| Price of a flagship AI system | Rising | A larger upfront bill |
| Compute delivered per dollar | Improving about 37% a year | Far cheaper fixed workloads |
| Software efficiency | Improving quickly | More output from the same chips |
| Cloud rental rates | Falling unevenly | Savings arrive more slowly |
| Cost per useful AI result | Falling fastest | Hardware and software gains compound |
Will cheaper GPU compute cut the AI industry’s total spending?
Cheaper GPU compute is expanding total AI spending right now because demand is growing much faster than the cost of each calculation is falling.
NVIDIA’s latest quarter makes that hard to dispute. Data-center revenue reached $75.2 billion, up 92% from a year earlier, even as Blackwell systems improved the economics of training and inference. Customers did not use the efficiency gain to hold their computing capacity flat. They bought far more capacity.
Alphabet is showing the same rebound from the customer side. Its latest quarterly capital spending reached $44.9 billion, mostly for AI infrastructure, and the company raised its full-year forecast to $195 billion to $205 billion. Google Cloud revenue climbed 82% to $24.8 billion, while its backlog reached $514 billion. The company also said it was using third-party capacity as a temporary bridge because its own supply could not keep up.
We can see why by comparing the rates. Epoch AI estimates that performance per dollar improves about 37% annually, while the total computing power in the global stock of AI chips has recently been growing about 3.4 times a year. Even after allowing for uncertainty in the chip estimates, installed capacity is rising several times faster than unit cost is falling.
A company rarely treats cheaper inference as a reason to run the same product for less money. Companies add more users, longer contexts, voice, images, video, agents and background automation. The budget may buy ten times more work, and the budget still grows.
If you want more recent data on this point, please see our latest AI infrastructure market report.

As this chart shows, and as featured in our AI infrastructure market deck, search interest in AI infrastructure has risen sharply
Why does lower GPU cost make companies use more AI?
Lower GPU cost makes frequent, uncertain and low-value AI tasks economical for the first time.
Consider customer support. An expensive model may be reserved for difficult cases, while a cheaper one can read every ticket, summarize every call, suggest every reply and check whether the issue was resolved. The feature moves from occasional assistance to a permanent layer inside the operation.
Software development follows the same pattern. A coding model used once per hour behaves like an autocomplete tool. Give it enough affordable compute and it can inspect a repository, run tests, search documentation, try several repairs and review its own work. One developer request turns into dozens or hundreds of model calls.
AI is already present in a large share of companies. Stanford’s 2026 AI Index reports that 88% of surveyed organizations used AI in at least one business function, while generative AI reached 70%. Agent deployment remained in the single digits across nearly every function, leaving a large gap between trying AI and running compute-heavy automated workflows.
Cheaper GPUs close part of that gap and make experimentation less painful. A team can test more prompts, compare more models, generate synthetic data and abandon weak ideas earlier. Most experiments fail, but lower failure costs encourage more attempts.
We have seen the same rebound in computing, storage and bandwidth. Falling unit prices widened the number of sensible uses so dramatically that total consumption kept climbing. AI is following the same path, only faster.
Do cheaper GPUs make frontier AI models cheaper to build?
Cheaper GPUs make yesterday’s model easier to reproduce and raise the spending ceiling for the next frontier model.
Epoch AI estimates that training compute for frontier language models has grown about fivefold per year since 2020. Over the same period, the estimated cost of the largest training runs rose roughly 3.5 times a year, from around $2 million for GPT-3 to nearly $390 million for the biggest runs measured in 2024. Hardware efficiency slowed the increase, but did not reverse it.
The published training run also understates the development bill. Laboratories spend heavily on experiments that never become products. Epoch’s reconstruction of OpenAI’s 2024 computing budget estimated roughly $5 billion for research compute, with only around $500 million going to the final training runs behind released models. Most of the money went into testing ideas, data recipes, architectures and post-training methods.
Copying gets cheaper before discovery for a simple reason. Once another laboratory knows which approach worked, it can skip many dead ends. The pioneer paid for the search; followers pay mainly for execution.
Today’s frontier teams also have more places to reinvest a hardware saving. They can train on more data, run longer reinforcement-learning programs, generate synthetic environments or use larger inference budgets during development. We should therefore expect the cost of matching an older capability to fall quickly while the cost of leading the field keeps rising.

This chart, included in our AI infrastructure market deck, shows annual VC investment in AI infrastructure startups
Will AI inference become almost free?
Basic AI inference is heading toward commodity pricing. Difficult reasoning, real-time media and high-reliability work will still carry meaningful costs.
Current API menus already show the split. OpenAI lists GPT-5.4 Nano at $0.10 per million input tokens and $0.625 per million output tokens, while GPT-5.5 Pro is listed at $15 and $90 respectively. The output price differs by a factor of 144. Anthropic also separates a lower-cost Sonnet tier from its more expensive Opus tier, and offers large discounts through caching and batch processing.
The price gaps reflect more than branding. A cheap model may classify a request, extract fields or draft a routine response. A premium model may use more parameters, more reasoning time, more tools and more verification. Customers will pay the premium when one wrong answer costs much more than the inference bill.
The cheapest tier should keep falling because several improvements stack together. Better GPUs lower the hardware cost. Quantization reduces memory and arithmetic. Prompt caching avoids processing the same context repeatedly. Batching spreads one hardware pass across many users. Smaller models keep improving, so fewer jobs need the largest system.
For everyday text tasks, the cost per interaction can become too small for the user to notice. Video generation, live voice, long-context analysis and autonomous agents consume far more resources, which gives providers room to keep charging for them.
| AI workload | Likely pricing direction | Why |
|---|---|---|
| Classification and extraction | Near-commodity | Small models handle the job well |
| Routine writing and summaries | Very cheap | Heavy competition and easy substitution |
| Live voice | Falling, but still visible | Continuous low-latency processing |
| Video generation | Material cost remains | Many pixels and frames must be produced |
| Frontier reasoning | Premium tier survives | More compute can improve the answer |
| High-stakes automation | Priced around the outcome | Verification and liability dominate |
If you want more recent data on this point, please see our latest AI infrastructure market report.
Will AI agents swallow the savings from cheaper GPUs?
AI agents will absorb a large share of cheaper compute because a useful agent may run hundreds of model steps before finishing one task.
A chatbot usually answers once. An agent can inspect files, build a plan, search several sources, operate software, run code, notice a failure, revise the plan and check the final result. Providers may also generate several candidate solutions and use another model to rank them. The user sees one outcome; the infrastructure handles a small tree of attempts.
Recent model launches make that direction explicit. OpenAI’s GPT-5.6 offers higher reasoning settings and parallel-agent modes for problems that reward more time and compute. Anthropic increased rate limits when it launched Sonnet 5 and described users selecting higher effort levels for demanding projects. These products now let customers spend more tokens when a task deserves it.
The economics depend on the whole task. Spending $5 on model calls is absurd for rewriting a sentence and trivial for repairing production software, researching an acquisition or recovering an unpaid invoice. As GPU costs fall, agents cross that economic threshold in more industries.
Agent use is still early, as Stanford’s latest AI Index shows. Today’s token volumes therefore understate how much compute widespread agents could consume. A move from a single response to 200 coordinated calls can erase years of hardware savings inside one workflow.

This chart, included in our AI infrastructure market deck, shows why CoreWeave is winning in AI infrastructure
Which AI products become viable when GPU compute gets cheaper?
Cheaper GPU compute makes always-on, highly personalized and media-heavy AI products practical for much larger audiences.
Real-time voice is one of the clearest examples. A voice agent must transcribe speech, understand intent, decide what to do and generate natural audio with little delay. Every second of the call keeps that pipeline running. Lower compute costs let businesses use it for routine bookings, sales qualification, technical support and language interpretation rather than a small set of premium calls.
Video has an even steeper compute bill. OpenAI’s current API pricing ranges from $0.10 per second for standard 720p generation to $0.70 per second for premium 1080p output. Those prices already make short clips accessible, yet a personalized ten-minute video remains expensive. Further GPU savings open advertising, training, entertainment and product demonstrations that can be generated for each viewer.
Robotics benefits before the robot even reaches a customer. Developers can simulate more environments, generate more training data and test more failures. During deployment, cheaper edge accelerators allow richer perception and planning without sending every sensor reading to the cloud.
Scientific search expands in the same way. Drug discovery, material design, weather forecasting and engineering optimization often involve screening huge numbers of possibilities. Cutting the cost of one simulation matters mainly because researchers can run far more of them.
Frequency is what changes the market. A feature used once a week may already be affordable. The bigger opportunity appears when AI can watch, listen, generate or decide continuously.
Will cheaper GPU compute crash AI software prices?
Cheaper GPU compute will crush prices for generic AI features, yet products tied to valuable workflows can keep charging well above their inference cost.
Writing, summarization, transcription and basic image generation are already being bundled into office suites, operating systems and business software. Once several providers can deliver a good-enough result for fractions of a cent, a standalone tool has little room to charge $20 a month for that feature alone.
The customer still pays for much more than tokens. Enterprise software must connect to internal systems, manage permissions, protect data, monitor quality, survive audits and support users. A legal review product may also carry specialized templates and responsibility for how the work is performed. Those costs decline more slowly than GPU prices.
Strong products can also charge according to the value they create. A system that recovers $100,000 in missed invoices can charge thousands of dollars even when its compute bill is $50. Customers compare the fee with the business result, not with the cost of the server.
The pressure will be brutal for thin wrappers. A company that owns no distribution, data or workflow can watch its gross margin improve briefly and then disappear as competitors pass the saving to customers. Better application companies can use cheaper compute to do more inside the same subscription, which makes the product harder to replace.
If you want more recent data on this point, please see our latest AI infrastructure market report.

This chart, included in our AI infrastructure market deck, shows annual funding in AI infrastructure startups
Are foundation models becoming commodities as GPU costs fall?
The middle of the foundation-model market is becoming a commodity, while the best models still earn a premium for capability, reliability and speed.
Stanford’s latest AI Index found that the top four closed models sat within fewer than 25 points on Arena, the public leaderboard built from head-to-head user votes. Such a tight grouping makes buyers compare price, latency, context limits, tool use and reliability instead of choosing one provider purely for raw intelligence.
The open-weight gap remains real. In the same report, the leading closed model was 3.4% ahead of the leading open model on Arena, compared with only 0.5% in 2024. New proprietary releases widened the gap again after open systems had nearly closed it.
A narrow benchmark difference can still be valuable on difficult work. An extra few percentage points in coding, finance or research may save many failed attempts. At the same time, most business tasks do not require the leader. A smaller model that is cheaper, faster and easier to control can win despite a lower benchmark score.
We are getting a market with short-lived premiums. A frontier provider can charge more while its model handles work others cannot. Within months, rivals and open-weight systems catch up on many of those tasks. The premium then moves to the next capability frontier.
Can AI startups compete more easily when GPUs get cheaper?
Cheaper GPUs give application startups a real opening, but they do little to make frontier-model development affordable.
A small team can now compare models through APIs, rent accelerators by the hour and deploy an open-weight model without buying a cluster. Falling inference costs also leave more room for customer support, sales and product development before each user becomes unprofitable.
Cheaper compute helps most during experimentation. Startups can test several architectures, fine-tune on customer data and run broader evaluations without committing millions of dollars. They can also keep sensitive workloads inside a private environment when API economics or data rules make that preferable.
The frontier remains concentrated because compute is only one part of the bill. Epoch estimates that research and inference compute together account for 54% to 62% of costs at three AI companies it examined. Staff, data, failed experiments and infrastructure add the rest. A project that falls from $500 million to $300 million is still inaccessible to almost every startup.
For startups, the sensible play becomes clearer as compute gets cheaper. Building a general model is usually a bad fight. Owning a narrow workflow, proprietary data or a hard-to-reach customer base gives the company something the model supplier cannot reproduce overnight.

This chart, included in our AI infrastructure market deck, compares the main business model options for AI cloud infrastructure providers
Does open-weight AI gain the most from cheaper GPUs?
Open-weight AI gains enormously from cheaper GPUs because organizations can run, adapt and own capable models at a steadily lower cost.
The improvement is already reaching ordinary hardware. Google says Gemma 4 12B can support local, multimodal agentic work on laptops with 16GB of memory. Its smaller Gemma 4 E2B model can run with a physical memory footprint of about 607MB on Apple mobile CPUs using Google’s optimized runtime. These are useful deployment targets, not giant research clusters.
Open models also attract optimization from outside the original developer. Quantized versions, faster runtimes, fine-tuning tools and task-specific variants spread across the ecosystem. Google reported more than 60 million Gemma 4 downloads within its first few weeks, giving developers a large base on which to improve deployment.
Closed providers still offer advantages because they operate the infrastructure, update the model and currently lead the open frontier on several broad evaluations. For teams with irregular usage, an API may remain cheaper than owning hardware and maintaining a serving stack.
Open-weight AI should gain the most where privacy, customization and predictable heavy usage matter. Banks, manufacturers, governments and software vendors can keep data under their control and avoid paying a margin on every token. Casual users will often prefer the convenience of a hosted service.
Will cheaper AI chips move inference from the cloud to phones and laptops?
Routine AI inference is already moving onto phones and laptops, and cheaper device accelerators will push far more of it to the edge.
Apple gives developers direct access to the on-device model behind Apple Intelligence through its Foundation Models framework. Google is taking a similar route with Gemma 4 and its AI Edge stack, including a 12-billion-parameter model designed for local laptop workflows.
Local processing avoids a cloud charge for every request, removes the network delay and keeps private data on the device. Those advantages suit summarization, message classification, personal search, translation and simple tool use.
Cloud models remain stronger for long contexts, difficult reasoning and large media jobs. Apple’s own developer guidance warns that on-device models are smaller and require more careful prompting than server-based frontier systems. Local hardware also has strict limits on memory, power and heat.
Most products will use a hybrid setup. A small model handles routine work and decides when a harder request needs the cloud. The user gets faster responses and more privacy, while the provider reserves expensive data-center compute for tasks that justify it.

This chart, featured in our AI infrastructure market deck, shows the share of revenue generated by each customer segment in the AI infrastructure market
Does NVIDIA lose when AI compute gets cheaper?
NVIDIA can keep winning as AI compute gets cheaper because customers are buying larger systems to obtain lower costs per result.
The company’s latest quarter delivered $81.6 billion in revenue, with data centers contributing about 92% of the total. That mix shows how completely NVIDIA’s business now depends on customers buying more AI capacity. Customers care about the cost of training a model or serving a billion tokens, rather than the price of one chip.
NVIDIA also sells the parts needed to reach the advertised performance. CUDA, networking, interconnects, libraries and optimized inference software reduce the time accelerators spend waiting. A cheaper rival can become expensive when porting work, weak tools or poor utilization erase its hardware discount.
The risk grows as workloads stabilize. A cloud company serving the same model billions of times can justify building a narrower chip and software stack around that workload. Large customers also gain bargaining power when they have credible alternatives from AMD, Google, Amazon, Microsoft or their own silicon programs.
For now, the market is growing fast enough for NVIDIA to expand even as alternatives gain share. The harder period arrives when capacity catches demand and buyers compare mature systems mainly on cost.
If you want more recent data on this point, please see our latest AI infrastructure market report.
Will custom AI chips replace NVIDIA GPUs?
Custom AI chips will capture predictable inference and training workloads first; NVIDIA GPUs will stay central to fast-changing research and mixed customer demand.
AWS says its Trainium2 instances deliver 30% to 40% better price-performance than comparable GPU-based P5e and P5en instances. Microsoft says Maia 200 gives 30% better performance per dollar than the latest hardware in its existing fleet. Google has moved beyond using TPUs internally and now sells TPU systems for customer data centers.
The newest example goes further. OpenAI and Broadcom recently unveiled Jalapeño, an inference processor designed around OpenAI’s models, kernels and serving patterns. The companies plan a multi-generation platform with gigawatt-scale deployment. The design aims to remove wasted movement between compute, memory and networking rather than merely reproduce a general GPU.
These programs make sense for companies with enormous, stable workloads. A 20% saving on several billion dollars of annual compute can pay for a custom chip program. Owning the silicon also reduces dependence on NVIDIA’s supply and pricing.
GPUs keep their edge when the workload changes every few months. Researchers value programmability, mature software and the ability to move from one architecture to another. Custom chips will take the repetitive work first, leaving NVIDIA strongest where flexibility still commands a premium.

This chart, included in our AI infrastructure market deck, shows how GPU cloud infrastructure technology has evolved over time
What becomes expensive after GPU compute gets cheap?
Power, high-bandwidth memory, networking, data-center space and software utilization become the real constraints as the calculations themselves get cheaper.
Memory is already tight. Micron said it had completed price and volume agreements for its entire 2026 HBM supply, and its latest presentation reported more than $1 billion in HBM4 revenue already shipped. The company expects data-center DRAM and NAND bit shipments this year to more than double their level from two years earlier. Faster processors simply wait when memory cannot feed them.
Networking becomes just as important in large clusters. Thousands of accelerators must exchange data constantly during training and inference. Delays leave expensive hardware idle, which explains why Broadcom’s AI opportunity spans custom accelerators and the networking chips connecting them.
Electricity and construction can be harder to secure than processors. Alphabet’s latest capital-spending breakdown put roughly 40% of technical infrastructure investment into data centers and networking rather than servers. The company also warned that depreciation, data-center operations and energy costs would pressure profits.
Utilization is the quiet bottleneck. A nominally cheaper chip saves little when weak scheduling, small batches or software incompatibility leave half its capacity unused. Providers are increasingly selling complete systems because customers need realized tokens per dollar, not a theoretical FLOP count.
| Scarce resource | Fresh evidence | What happens to its value |
|---|---|---|
| HBM memory | Micron contracted its full 2026 supply | Memory captures more of the system budget |
| Networking | Larger clusters need constant chip-to-chip communication | Fast interconnects become strategic |
| Electricity | Data-center demand is rising faster than grid expansion in many regions | Power access can determine project location |
| Buildings and cooling | A large share of hyperscaler spending now sits outside servers | Infrastructure owners gain leverage |
| Utilization software | Idle accelerators erase hardware savings | Compilers, schedulers and serving tools matter more |
| Proprietary data | Compute becomes easier to buy | Unique training and workflow data stand out |
Will more efficient GPUs reduce AI’s total electricity use?
More efficient GPUs will cut energy per AI calculation; total AI electricity use will still rise for years.
The International Energy Agency projects electricity supplied to data centers to climb from about 460 terawatt-hours in 2024 to more than 1,000 terawatt-hours in 2030. AI is one of the main drivers. The forecast already assumes ongoing improvements in chips, cooling and data-center efficiency.
Efficiency still prevents a much larger power problem. A model served on newer hardware may use far less electricity per token than the same model on an older cluster. Better utilization, lower precision and improved cooling reduce energy again.
Demand is outrunning those gains. Companies are deploying more chips, serving more users and giving each difficult request a larger compute budget. Video, voice and agents also use more resources than a short text answer.
Electricity use per unit of AI value should decline, while total consumption rises sharply. Energy becomes a bigger part of where data centers are built, how quickly projects open and which companies can scale.

In our AI infrastructure market deck, we identify pain points entrepreneurs should prioritize
Will cheaper GPUs democratize AI or concentrate it further?
Cheaper GPUs spread AI access widely and concentrate frontier infrastructure in fewer hands.
Developers almost anywhere can rent a model API, fine-tune an open-weight system or run smaller models locally. The Stanford AI Index estimates that generative AI reached 53% adoption in three years, faster than the personal computer or the internet. Falling costs help that spread continue.
Ownership of the largest infrastructure tells another story. Epoch AI estimates that the five major U.S. hyperscalers own more than 70% of global AI computing power, with NVIDIA-designed chips supplying over 60% of the total compute stock across chip designers. These estimates are imperfect, but the concentration is too large to dismiss as measurement noise.
The split is easy to understand. Cheap access lets many companies build products without owning hardware. Frontier training requires enormous power blocks, networking, memory, engineering teams and long-term capital. Only a small group can assemble all of them at once.
More people will therefore use and adapt advanced AI, while fewer companies control the biggest clusters and decide which frontier models get trained. Open weights and custom chips can soften that dependence, but they do not remove the underlying scale advantage.
If you want more recent data on this point, please see our latest AI infrastructure market report.
So what happens when GPUs get cheaper?
Cheaper GPUs will make AI usage explode, push frontier spending higher and move scarcity into power, memory, networking, data and customer distribution.
Routine intelligence becomes a low-cost feature inside almost every software product. Smaller models handle classification, summaries, search and simple automation for fractions of today’s cost. More of that work runs locally on phones and laptops.
Companies then spend the savings on heavier jobs. Agents take more steps, reasoning models think longer, video systems generate more frames and researchers test more possibilities. Total compute demand keeps rising because lower prices expand what is worth attempting.
Application builders get the widest opening. Startups, universities and companies worldwide can build with capable models without owning a giant cluster. At the frontier, costs continue climbing as the leading laboratories reinvest every efficiency gain into larger training and inference programs.
More chip suppliers can compete too. NVIDIA can remain the largest supplier while custom chips take predictable workloads and hyperscalers design more of their own stack. Infrastructure winners will deliver reliable results across compute, memory, networking and power at the lowest full-system cost.
The AI economy gets larger as GPUs get cheaper. More abundant computation increases the quantity of intelligence people consume and makes the surrounding bottlenecks more valuable.

This chart, included in our AI infrastructure market deck, shows the share of revenue by region across Europe, Asia, North America, Africa, and South America in the AI infrastructure market
OUR METHODOLOGY
This analysis examines what happens when GPUs get cheaper by separating the question into the areas where lower compute costs can produce very different outcomes: AI usage, total industry spending, frontier-model development, inference pricing, agents, application economics, startup competition, open-weight models, edge computing, chip competition, infrastructure constraints and electricity demand.
We define “cheaper GPUs” through useful work per dollar rather than the sticker price of one chip or rack. A new system can cost more upfront while making a fixed training run, token workload or completed AI task substantially cheaper. Hardware price-performance, software efficiency and realized utilization are therefore considered together.
For each part of the analysis, we reviewed recent operational evidence rather than relying on a single forecast. This included accelerator price-performance data, company financial results, capital-spending plans, API pricing, model launches, adoption surveys, custom-chip announcements, memory commitments and energy projections.
We prioritized first-hand disclosures and established research institutions. Company claims about price-performance were treated as vendor evidence rather than neutral benchmarks, while financial results, official pricing pages and infrastructure commitments were used to show what buyers and suppliers are doing in practice.
Several conclusions depend on outcomes that can occur at the same time. Compute can become cheaper per unit while total spending rises. Older model capabilities can become inexpensive to reproduce while the cost of reaching the next frontier increases. AI access can broaden even as ownership of the largest clusters becomes more concentrated.
The final conclusion comes from comparing those directions across the full system. We gave the greatest weight to repeated patterns visible in several independent datasets or company disclosures, especially the rebound in compute demand, the widening gap between commodity and premium inference, and the movement of scarcity toward memory, networking, electricity, buildings and utilization software.
Key sources include Epoch AI’s compute and hardware trends, Epoch AI’s machine-learning hardware database, NVIDIA’s first-quarter fiscal 2027 results, Alphabet’s investor disclosures, Stanford HAI’s 2026 AI Index, OpenAI’s API pricing, Anthropic’s pricing, AWS on Trainium2, Microsoft on Maia 200, Micron’s earnings and presentations, and the International Energy Agency’s Energy and AI report.

This chart, included in our AI infrastructure market deck, shows annual VC investment in AI infrastructure startups
Related blog posts
- GPU efficiency: which startup is ahead?
- Do AI data centers need fewer GPUs and more memory?
- Do AI agents need different GPUs?
- Are GPU shortages finally ending?
Who is the author of this content?
NEW MARKET PITCH TEAM
We track new markets so founders and investors can move fasterWe build living "market pitch" documents for emerging markets: AI, synthetic biology, new proteins, and more. Instead of outdated PDFs or hallucinated LLM answers, our clients get a clean, visual, always-updated view of what's really happening: key players, deals, regulations, and signals that matter. Learn more about us.