Are GPU shortages finally ending?

Last updated: 31 July 2026
market research pitch 2026 statistics AI chip market

In our AI chip market deck, you will find everything you need to understand the market

SUMMARY

GPU shortages are finally ending for ordinary buyers, but not for frontier AI companies that need the newest accelerators in very large, fully connected clusters.

The market has split into two. A developer can now rent a few H100s or B200s without special access, while a company seeking thousands of identical Blackwell GPUs still faces reservations, regional limits and infrastructure delays.

The strongest evidence of normalization is not that new GPUs exist in cloud catalogs. It is that older generations are becoming easier to rent and, in some cases, materially cheaper as frontier customers move toward Blackwell systems.

NVIDIA is shipping far more hardware without losing pricing power. Data-center revenue almost doubled year over year, margins remained high and the company guided to another record quarter, which suggests production has expanded but still has not overtaken demand.

The bottleneck has moved outward from the GPU die. HBM supply, advanced packaging, networking, electrical equipment, liquid cooling and grid connections now determine how quickly a usable AI cluster can come online.

That is why complete racks matter more than loose chips. A cloud provider can own accelerators and still be unable to sell useful capacity because one late switch, optical link, transformer or cooling system holds up the entire installation.

Gaming has largely returned to normal below the very top end. The RTX 5090 remains unusually expensive, but broad availability across lower RTX 50-series cards and AMD’s Radeon 9000 series means the pandemic-style consumer shortage has faded.

Alternative accelerators are helping, but they are not creating a glut. AMD, Amazon, Google and Microsoft are adding major new sources of compute, yet much of that capacity is reserved before deployment and is being used alongside NVIDIA rather than replacing it.

Inference is likely to keep the market tight for longer than training alone would have. Chatbots, coding tools, video generation and agents consume capacity continuously, and reasoning models can turn one user request into many separate model calls.

Efficiency gains will reduce the cost of each task, but hyperscalers are betting that cheaper tokens will expand usage even faster. Their combined 2026 capital-expenditure plans of roughly $710 billion to $740 billion are not the behavior of companies preparing for excess capacity.

The broad 2023-style scramble is therefore ending, but scarcity has narrowed rather than disappeared. It is now concentrated in the newest accelerators, the largest clusters and the power, memory, packaging and networking needed to operate them.

What does “GPU shortage” actually mean now?

The GPU shortage is over today only for buyers who are flexible about the chip, price and scale.

A gamer looking for a mid-range card, a startup renting eight H100s and a frontier lab seeking 20,000 identical Blackwell GPUs are dealing with three different markets. The first buyer may find stock immediately. The second can usually launch a cloud instance. The third still needs a long reservation, a suitable region and a data center with enough networking and power.

Price belongs in the definition too. A product can be technically available while remaining scarce in any practical sense. The same applies when a cloud lists a GPU but cannot provide a large cluster for the required weeks or months.

For this article, we call the shortage over only when buyers can get the preferred generation, in the required quantity, near a normal price and on a predictable schedule. GPU supply has improved on all four, but only ordinary purchases pass the full test today.

If you want more recent data on this point, please see our latest AI chip market report.

Why are GPUs easier to get today?

GPUs feel easier to obtain today because supply has expanded, cloud choice has widened and older accelerators have moved down to less demanding workloads.

Google Cloud now lists NVIDIA systems ranging from A100 and H100 machines to B200, GB200 and GB300 machines. AWS offers several Blackwell configurations, while specialist clouds let customers rent H100s and B200s without buying a server. A developer who needs a few GPUs has far more places to look than during the first H100 scramble.

Every new generation also pushes older machines toward cheaper, less demanding jobs. Frontier labs may prefer the newest systems, but H100 and A100 machines remain strong enough for fine-tuning, research and a large share of inference. Those older GPUs do not disappear when the biggest customers upgrade.

The improvement is real, although it reached small jobs, flexible buyers and previous-generation hardware first.

Market map chart showing top companies and startups in the AI chip market

This market map, featured in our AI chip market deck, highlights top companies and startups in the AI chip market

Is NVIDIA finally shipping enough AI GPUs?

NVIDIA is shipping far more AI hardware than a year ago, and buyers are still taking almost every additional system.

Its latest quarterly data-center revenue reached $75.2 billion, up 92% from $39.1 billion one year earlier. That adds $36.1 billion of quarterly revenue in twelve months. Blackwell systems cost more and include more networking, so the revenue jump does not tell us the exact number of GPUs. Still, adding $36.1 billion of quarterly sales means NVIDIA shipped many more complete systems.

Even after that ramp, NVIDIA kept roughly 75% non-GAAP gross margins and guided total company revenue to about $91 billion for the following quarter. NVIDIA has become an $80-billion-plus quarterly business, and prices still look firm.

We would expect weaker growth, discounting or rising unsold inventory if production had finally moved ahead of demand. Revenue and margins are still climbing.

Can startups get cloud GPUs without paying scarcity prices?

Cloud GPUs are now easy enough to access for many startups, although cheap and guaranteed capacity remains rare.

Lambda lets customers launch H100, B200 and older instances by the hour, while its self-service clusters run from 16 to more than 2,000 GPUs. Google Cloud lists six generations of NVIDIA data-center accelerators, and AWS lets customers reserve future capacity through Capacity Blocks. A small team can start useful work today without owning hardware or negotiating directly with NVIDIA.

Prices tell a less comfortable story. Lambda currently charges $6.69 per GPU-hour for an eight-B200 machine, $3.99 for an eight-H100 machine and $2.79 for an eight-A100 80GB machine. Keeping the B200 server busy for a month costs roughly $39,000 before taxes.

Older capacity is beginning to normalize. AWS cut on-demand prices by as much as 45% for P5 instances, 26% for P5en and 33% for P4d or P4de instances in 2025. Those cuts show that older capacity is finally getting cheaper. The newest systems still carry a large premium, and very large clusters usually require reservations rather than a casual click.

Eight-GPU Lambda instance Price per GPU-hour Approximate server cost per hour
NVIDIA B200 SXM6 $6.69 $53.52
NVIDIA H100 SXM $3.99 $31.92
NVIDIA A100 SXM 80GB $2.79 $22.32
Google Trends chart showing rising interest in AI chips

As this chart shows, and as featured in our AI chip market deck, search interest in AI chips has grown significantly

Are Blackwell GPUs still hard to get?

Yes, Blackwell capacity is still hard to secure when a customer needs thousands of GPUs together for months.

Cloud providers can now advertise B200, GB200 and GB300 machines because some capacity is live. That says little about how much is free in a chosen region. Large training runs need identical accelerators, fast interconnects and uninterrupted access, so a scattered pool of available machines does not solve the problem.

Microsoft gives us the clearest current answer. The company added another gigawatt of capacity in its latest reported quarter, shortened the time between receiving and activating new GPUs by nearly 20%, and still expects to remain capacity-constrained through 2026.

A market where one of the world’s largest buyers adds a gigawatt in a quarter and remains constrained has not reached comfortable supply at the frontier.

If you want more recent data on this point, please see our latest AI chip market report.

Is the gaming GPU shortage over?

The gaming GPU shortage is mostly over below the very top end, while the RTX 5090 still trades like a scarce luxury product.

Retailers currently show broad availability across RTX 5060, 5070 and 5080 cards, along with AMD’s Radeon 9000 series. Buyers can choose among several brands instead of chasing almost any card that appears, as they did during the pandemic and crypto-mining boom.

The RTX 5090 remains distorted. NVIDIA lists its Founders Edition at $1,999. Best Buy currently offers an ASUS ROG Astral model at $4,329.99, and several Newegg listings sit above $4,000. Premium cooling and factory overclocking explain some of the gap, but not a price more than twice NVIDIA’s reference level.

For gamers, the broad shortage has faded. Scarcity is concentrated in the top card, which also attracts creators and people running AI models locally.

Current RTX 5090 example Listed price Difference from NVIDIA’s $1,999 reference price
NVIDIA Founders Edition $1,999.00 Reference
ASUS ROG Astral at Best Buy $4,329.99 About 117% higher
Several current Newegg models Roughly $4,100 to $4,500 About 105% to 125% higher
Chart showing annual VC investment in AI chip startups

This chart, featured in our AI chip market deck, shows annual VC investment in AI chip startups

Has HBM become the real GPU bottleneck?

HBM has become one of the hardest limits on how many high-end AI accelerators the industry can finish.

High-bandwidth memory sits beside the processor and feeds it data far faster than ordinary server memory. Each new accelerator generation tends to need more HBM capacity and bandwidth, so memory output has to grow faster than GPU shipments just to keep pace.

Micron has already agreed the price and volume for its entire 2026 HBM supply, including HBM4. It expects the HBM market to rise from about $35 billion in 2025 to around $100 billion in 2028, which works out to roughly 42% annual growth. The company now expects that milestone two years earlier than it once did.

SK hynix is seeing the same squeeze. Its first-quarter revenue rose 198% year over year and operating profit rose 405%, while management said customer demand continued to exceed supply capacity. Memory companies do not post numbers like these when buyers have comfortable alternatives.

A GPU die without qualified HBM cannot be sold as a working accelerator. Memory allocation is now nearly as important as the processor itself.

Has advanced packaging caught up with GPU demand?

Advanced packaging is still behind demand, and TSMC says the gap is directly limiting its customers’ growth.

Packaging is where the GPU die, several HBM stacks and other components are joined into one working unit. The process is difficult, expensive and slow to expand. More wafer output cannot help if the finished dies are waiting for CoWoS or a similar packaging line.

In its latest earnings call, TSMC’s chief executive said packaging capacity was so tight that it was holding customers back. The company raised its 2026 capital budget to $60 billion to $64 billion and plans to spend 10% to 20% on advanced packaging, testing, mask-making and related work. That is roughly $6 billion to $12.8 billion for packaging, testing, masks and related work, although packaging is only part of that total.

TSMC is even welcoming competing packaging technologies because extra capacity would help its customers ship more products. A supplier rarely invites rivals to take work unless the backlog is genuinely painful.

Chart showing how Nvidia is leading in the AI chip market

This chart, featured in our AI chip market deck, shows how Nvidia is leading in AI chips

Are complete AI racks harder to get than GPUs?

Complete AI racks are now harder to deliver than loose GPU chips because every missing component can delay the whole system.

A modern AI rack needs GPUs, HBM, CPUs, NVLink switches, network cards, optical links, storage, power distribution and liquid cooling. The rack becomes useful only after all of it is installed and working together. A pile of accelerators in a warehouse has little value if the network or cooling loop is late.

NVIDIA’s latest results show how quickly the shortage has moved beyond the processor. Data-center compute revenue grew 77% year over year, while networking revenue grew 199%. Networking expanded about 2.6 times faster because customers are connecting larger numbers of GPUs into one machine.

These days, usable supply is best measured in powered and networked clusters rather than individual chips.

If you want more recent data on this point, please see our latest AI chip market report.

Are power and cooling now the bigger shortage?

Power and cooling now delay some AI projects more than the arrival of the GPUs themselves.

Eaton reported that its Electrical Americas data-center orders rose about 240% year over year, while related revenue increased around 50%. Orders are growing almost five times faster than delivered revenue, which points to a long queue of switchgear, power distribution and cooling work still waiting to be completed.

Vertiv is seeing the same rush from another angle. Its Americas organic sales grew 44% in the first quarter, total sales rose 30%, and it opened or expanded four manufacturing sites to build more power, rack and thermal equipment. Those are unusually high growth rates for the companies that furnish the less glamorous parts of a data center.

A cloud provider may own the GPUs and still wait for an electrical connection, transformers, backup equipment or a liquid-cooling system. For large clusters, the building can now take longer than the silicon.

Chart showing the projected CAGR of the AI chip market

This chart, featured in our AI chip market deck, shows annual funding in AI chip startups

Are hyperscalers buying every GPU they can find?

Yes, the largest technology companies are currently buying and building as much AI capacity as their supply chains can deliver.

Microsoft expects roughly $190 billion of capital expenditure in calendar 2026. Amazon expects about $200 billion. Alphabet has raised its range to $195 billion to $205 billion to speed up capacity delivery, while Meta lifted its plan to $125 billion to $145 billion partly because component prices were higher than expected.

Together, those four companies plan between $710 billion and $740 billion of capital expenditure in one year. Of course, only part of that money buys GPUs. Amazon also funds logistics, satellites and robotics, while every cloud company must pay for land, buildings, networking and power. Even with that caveat, the aggregate is enormous and heavily shaped by AI infrastructure.

Their budgets say more than any shortage claim. Companies do not commit three-quarters of a trillion dollars to capacity when they expect a near-term glut.

Company Current 2026 capital-expenditure plan What the company linked it to
Microsoft About $190 billion More capacity, GPUs, CPUs and higher component prices
Amazon About $200 billion AI, chips, robotics, satellites and other infrastructure
Alphabet $195 billion to $205 billion Faster capacity delivery to meet demand
Meta $125 billion to $145 billion AI infrastructure, data centers and higher component prices
Combined $710 billion to $740 billion Companywide spending, with AI infrastructure as a major driver

Is inference making the GPU shortage worse?

Inference is turning a bursty training shortage into a more permanent race for capacity.

Training a large model consumes an enormous cluster for a limited period. Inference runs every time someone asks a chatbot a question, generates a video, uses an AI coding tool or sends an agent to complete several steps. The machines have to stay available day and night once those products reach millions of users.

The workload is also becoming heavier. Reasoning models produce more internal tokens before answering, and agents may call a model repeatedly while searching, coding, checking and correcting. One user request can therefore create many separate inference jobs.

Microsoft’s AI business has already passed a $37 billion annual revenue run rate, up 123% year over year. Revenue at that scale gives providers a strong reason to keep adding inference hardware instead of slowing down after the next training run ends.

Chart comparing business model options for AI accelerator chip companies

This chart, featured in our AI chip market deck, compares the main business model options for AI accelerator chip companies

Will more efficient AI models reduce GPU demand?

More efficient models will cut the compute needed for each task, but total GPU demand is unlikely to fall soon.

Some of the efficiency gains are dramatic. NVIDIA says its Dynamo software can raise generative and agentic inference performance on Blackwell GPUs by as much as seven times for some workloads. Microsoft has also reported major throughput improvements from software and hardware optimization. Each improvement creates capacity without waiting for another data center.

Companies are using those savings to lower prices, support longer prompts, serve more users and run more complicated agents. A task that becomes ten times cheaper can spread to far more than ten times as many interactions, especially when it moves from a premium experiment into everyday software.

As we saw above, hyperscalers are pairing efficiency work with $710 billion to $740 billion of planned capital expenditure. Their budgets show that they expect cheaper tokens to expand usage faster than optimization reduces hardware needs.

If you want more recent data on this point, please see our latest AI chip market report.

Are AMD and custom AI chips ending NVIDIA’s shortage?

AMD and custom accelerators are widening supply, although most of the new capacity is being reserved before it arrives.

AMD has now announced agreements covering up to 14 gigawatts of future GPU deployments: six gigawatts for OpenAI, six for Meta and two for Anthropic. The first large MI450 deployments begin during the second half of 2026, with Anthropic’s first gigawatt planned for the first half of 2027. These deals are large enough to weaken NVIDIA’s grip over time, but much of the hardware is still on the roadmap.

Amazon’s Trainium shows what happens when an alternative succeeds. Almost one million Trainium2 chips are already training and serving Claude through Project Rainier, according to AWS and Anthropic. That already reduces dependence on NVIDIA, yet Anthropic expects to scale well beyond the current system with Trainium3.

Google’s TPUs and Microsoft’s Maia accelerators add still more capacity. They mainly sit beside NVIDIA systems rather than replacing them one for one, because AI companies keep expanding the total amount of compute they want.

Competition is easing the NVIDIA-specific shortage without creating excess AI compute.

Chart showing how revenue is split across customer segments in the AI chip market

This chart, featured in our AI chip market deck, shows how revenue is split across customer segments in the AI chip market

Can export controls create fake GPU surpluses?

Export controls can strand GPUs in one market while buyers elsewhere remain short of the products they actually need.

The clearest example was NVIDIA’s $4.5 billion H20 charge after new US licensing requirements disrupted sales to China. Those processors could not simply replace B200 or GB300 systems ordered for American and European data centers. Different products have different performance, memory, software and regulatory limits.

NVIDIA’s latest outlook assumes no data-center compute revenue from China, even as the company expects another record quarter elsewhere. So excess inventory tied to one restricted product can sit beside severe scarcity in permitted frontier systems.

Global supply totals blur these regional mismatches. What matters is whether the right chip can legally reach the customer and operate in a ready data center.

What would prove the GPU shortage is truly over?

We will know the GPU shortage is over when buyers regain choice instead of merely finding something available.

A large customer should be able to secure thousands of the newest accelerators in the preferred region without reserving years ahead. Previous-generation cloud prices should fall steadily. HBM producers should stop committing a full year of output in advance, and packaging companies should stop describing capacity as a limit on customer growth.

The surrounding infrastructure must loosen too. Power equipment orders should grow roughly in line with deliveries, data-center operators should have spare megawatts, and adding a rack should no longer depend on a long queue for cooling or grid connections.

We need several quarters in which hardware choice, cluster size, price and delivery time improve together.

Chart showing how AI accelerator chip technology has evolved over time

This chart, featured in our AI chip market deck, shows how AI accelerator chip technology has evolved over time

Are GPU shortages finally ending?

Partly. Everyday GPU access is improving fast, but frontier AI capacity is still short.

Gamers can find most cards, developers can rent several GPU generations and startups no longer need a special relationship to launch a modest cluster. Older cloud instances are getting cheaper, and the market now offers credible alternatives from AMD, Google, Amazon and Microsoft.

The frontier is still under pressure. The newest large clusters are reserved and power-equipment orders are outrunning deliveries. As seen above, HBM output is committed far ahead and advanced packaging remains tight. NVIDIA almost doubled data-center revenue in one year without forcing prices down.

Our judgment is that the broad 2023-style GPU scramble is ending. The shortage has narrowed toward the newest accelerators, the largest clusters and the infrastructure needed to run them. For most buyers, access is becoming normal. For frontier AI companies, the capacity race is still accelerating.

If you want more recent data on this point, please see our latest AI chip market report.

OUR METHODOLOGY

This analysis tests whether GPU shortages are genuinely ending or merely shifting toward the newest accelerators, the largest clusters and the infrastructure required to run them. We separate ordinary retail and cloud access from frontier-scale capacity because finding one card or a few instances is a very different test from securing thousands of identical GPUs in one region.

We assessed the market across product availability, price, accelerator generation, quantity, delivery predictability, cloud access, upstream component capacity and data-center readiness. The conclusion is based on the combined pattern across those dimensions rather than any single company claim or catalog listing.

Cloud product pages and pricing were used to establish what customers can access today. Google Cloud, AWS and Lambda provided evidence on the generations available, reservation structures, hourly pricing and the price gap between B200, H100 and A100 capacity.

Financial results were used to track what suppliers are actually shipping and whether greater output is weakening pricing power. NVIDIA’s fiscal 2027 first-quarter results were especially important because they showed data-center revenue, margins, compute growth and networking growth after the Blackwell ramp.

We treated revenue growth carefully because higher system prices, networking content and rack-level sales can increase revenue faster than the number of GPUs shipped. Revenue was therefore read alongside margins, guidance, cloud prices and evidence of capacity constraints.

Upstream bottlenecks were examined through Micron and SK hynix disclosures on HBM demand and supply, and TSMC’s comments on advanced-packaging capacity. These sources help show whether more GPU dies can actually become finished accelerators.

Data-center readiness was assessed through Microsoft’s capacity guidance and the results of Eaton and Vertiv. Their disclosures add evidence on GPU activation times, electrical-equipment orders, cooling demand, manufacturing expansion and the gap between infrastructure orders and delivered revenue.

Hyperscaler capital-expenditure plans were used as a forward-looking demand check, not as a direct measure of GPU purchases. Microsoft, Amazon, Alphabet and Meta also spend on buildings, networking, logistics, satellites and other infrastructure, but the combined scale of their plans shows that AI capacity remains a central investment priority.

We also examined whether alternative accelerators are loosening NVIDIA-specific scarcity. AMD deployment agreements, Amazon and Anthropic’s Project Rainier, Google TPUs and Microsoft Maia show that supply is broadening, although much of the new capacity is reserved before it arrives and supplements NVIDIA systems rather than replacing them one for one.

Key sources used for this analysis include: NVIDIA’s first-quarter fiscal 2027 results, NVIDIA’s first-quarter fiscal 2026 results, NVIDIA’s fourth-quarter and fiscal 2026 results, Google Cloud’s GPU machine documentation, AWS on EC2 P6-B200 availability, AWS Capacity Blocks pricing, AWS on NVIDIA GPU instance price reductions, Lambda GPU Cloud pricing, Microsoft’s fiscal 2026 third-quarter earnings call, Microsoft’s fiscal 2026 third-quarter results, Micron’s quarterly results, TSMC’s first-quarter 2026 earnings materials, SK hynix investor-relations results, Eaton’s first-quarter 2026 results, Vertiv’s first-quarter 2026 results, AMD and Meta’s accelerator agreement, AWS and Anthropic on Project Rainier, NVIDIA’s Dynamo inference platform, and NVIDIA’s RTX 5090 product page.

Table scoring and prioritizing the main pain points faced by companies in the AI chip market

In our AI chip market deck, we identify pain points entrepreneurs should prioritize

Who is the author of this content?

NEW MARKET PITCH TEAM

We track new markets so founders and investors can move faster

We build living "market pitch" documents for emerging markets: AI, synthetic biology, new proteins, and more. Instead of outdated PDFs or hallucinated LLM answers, our clients get a clean, visual, always-updated view of what's really happening: key players, deals, regulations, and signals that matter. Learn more about us.

Back to blog