Is storage becoming the bottleneck for AI?

Last updated: 31 July 2026
market research pitch 2026 statistics AI infrastructure market

In our AI infrastructure market deck, you will find everything you need to understand the market

SUMMARY

Yes, storage is becoming a major bottleneck for AI data centers, although it does not replace GPUs, networking, or HBM as the main constraint in every workload.

The problem is usually not a lack of total capacity. It is a lack of fast, well-placed storage that can deliver the right data beside the accelerators without long metadata lookups, cross-region copies, or slow recovery paths.

Faster GPUs have made small storage delays much more expensive. A few slow reads can stall synchronized training jobs, so tail latency now matters almost as much as headline bandwidth.

Text pretraining is less storage-hungry than the size of the dataset suggests. Once text has been cleaned and tokenized, steady-state delivery may require surprisingly little bandwidth; checkpoints and data preparation are often the harder parts.

Checkpointing is the clearest frontier-scale storage problem. Multi-terabyte saves must happen more frequently as clusters grow, creating short bursts that can demand several terabytes per second even when average storage traffic looks manageable.

Multimodal AI changes the equation. Images, video, audio, lidar, radar, and robotics logs can require gigabytes per second per GPU and produce many derivative files that teams still want to keep.

Inference has a different storage bottleneck: cold starts, model loading, cache recovery, and document retrieval. Once weights and active KV cache are in HBM, storage matters less, but getting them there quickly can decide whether a fleet scales cleanly.

AI agents are turning storage into a live application problem. The hard part is no longer keeping months of history; it is retrieving the correct, current, permitted fact without flooding the model with everything that came before.

Hard drives are not disappearing. AI data centers are settling on a hierarchy in which hard drives hold the durable source of truth, flash keeps active datasets and model files close to compute, DRAM absorbs working data, and HBM handles the immediate loop.

Our conclusion is that storage is now a first-order AI infrastructure constraint at the system level. Buying more accelerators without fixing data placement, checkpointing, model loading, and retrieval is an increasingly reliable way to leave expensive GPUs underused.

When do we know storage is slowing AI down?

Storage becomes an AI bottleneck when GPUs, researchers, or applications spend meaningful time waiting for persistent data.

That definition keeps us from mixing three different things together. HBM on a GPU holds the weights, activations, and KV cache needed immediately. Server DRAM provides a larger working area. SSDs, object stores, parallel file systems, and hard drives keep data after the job stops.

We should call storage the bottleneck when one of four things happens. GPUs stall while reading training samples or writing checkpoints. A model takes too long to load when an inference fleet scales up. Researchers wait hours for a dataset to reach the same region as the GPUs. An agent has plenty of stored information but cannot retrieve the right piece quickly or safely.

Capacity alone tells us very little. A company can own petabytes of storage and still have a serious bottleneck because the useful data sits in another region, the metadata lookup is slow, or every worker tries to read the same model file at once.

The clean test is simple: would faster or better-organized persistent storage produce more useful GPU work, faster experiments, cheaper inference, or better agent answers? When the answer is yes, storage is genuinely in the way.

Type of storage bottleneck What the user sees What is usually happening
Training throughput GPUs pause or utilization drops Data reads or checkpoint writes arrive late
Research speed An experiment takes hours to launch Data must be copied, prepared, or re-indexed
Inference scaling New workers take too long to become ready Model files and caches move too slowly
Agent quality The agent forgets, retrieves old facts, or misses context Stored information is poorly indexed or governed
Capacity economics Storage bills, power use, and rack needs climb Too much data sits on an expensive tier

Why has AI storage become urgent now?

AI storage has become urgent because accelerators improved faster than the systems feeding them, and every minute of delay now wastes much more expensive hardware.

Meta’s latest storage redesign gives us the clearest evidence. The company said its older object-storage stack was built for globally replicated consumer-app data and cheap hard-drive capacity. AI changed the priorities: much higher IOPS, flash close to GPUs, flatter metadata lookups, fewer software hops, and regional storage deployed beside each AI cluster.

Meta has effectively rewritten a large part of its storage architecture. It now uses memory and flash on GPU hosts as early cache layers, regional flash as another tier, and hard-drive-backed global storage as the durable source of truth. It also prefetches data several minutes ahead, allowing researchers to start before a full regional copy finishes.

Google reached a similar point from the cloud side. Its current Cloud Storage Rapid service offers more than 15 terabytes per second from one zonal bucket, sub-millisecond latency, and up to 20 million queries per second. Google built a new high-performance object-storage tier because conventional object storage had reached what it called a performance tipping point for AI.

The commercial market is moving in the same direction. In its latest fiscal results, Micron said data-center SSD revenue exceeded $5 billion and more than doubled from the previous quarter. It expects DRAM and NAND demand to remain above supply beyond 2027.

The pattern is hard to miss: one hyperscaler rebuilt its stack, another introduced a much faster storage tier, and a major supplier reported a sudden jump in SSD demand. Storage is now important enough to change architecture, products, and purchasing.

If you want more recent data on this point, please see our latest AI infrastructure market report.

Market map chart showing top companies and startups in the AI infrastructure market

This market map, featured in our AI infrastructure market deck, highlights top companies and startups in the AI infrastructure market

Are GPUs really waiting for storage?

Yes, GPUs are already waiting for storage, and the losses can be large enough to change the economics of an AI cluster.

Meta explains the failure mode clearly. A data loader normally fetches the next batch while the GPU processes the current one. Most reads can therefore happen in the background. Trouble begins when a slow request outlasts the prefetched buffer. One GPU runs out of work, and a synchronized training job may then wait for that straggler.

Average bandwidth can look healthy while the cluster still underperforms. A few unusually slow metadata lookups, cache misses, or overloaded storage nodes can hurt more than a modest reduction in average speed.

Google says its Rapid Bucket cut blocked GPU time by 50% and made multimodal data loading up to 2.5 times faster in its tests. The same service produced checkpoint restores up to five times faster and writes 3.2 times faster than traditional object storage. Those are vendor measurements, so we treat them as product evidence rather than universal benchmarks. Even with that caveat, the gains are too large to dismiss as cosmetic.

Look at what both companies chose to optimize. Meta removed metadata layers, eliminated a data-plane proxy, added direct streaming, and placed flash beside GPUs. Google built a zonal object store with extreme throughput and low latency. That is a lot of engineering for a supposedly minor problem.

Storage can now consume a meaningful slice of GPU time when the architecture was designed for ordinary cloud workloads rather than synchronized AI jobs.

Do giant language models actually need huge storage bandwidth?

Giant language models usually need far less bandwidth for the training text itself than people assume.

DeepSeek-V3 gives us a useful reality check. The model was pretrained on 14.8 trillion tokens using 2,048 H800 GPUs and about 2.788 million GPU-hours. That works out to roughly 57 days of cluster time.

If each prepared token occupies two to four bytes, the final token stream represents about 30 to 60 terabytes. Spread across 57 days, the average sequential read rate is only around 6 to 12 megabytes per second.

Real pipelines are larger. Teams keep raw documents, cleaned versions, deduplicated copies, metadata, indexes, mixtures, and replicas. They may shuffle examples or process data during training. Even after allowing for that overhead, the order of magnitude remains striking: feeding a text token stream is often easy compared with moving activations across GPUs or reading and writing model checkpoints.

Large-model engineering reports consequently spend far more time discussing HBM, network fabrics, expert routing, and communication overlap. Text storage causes operational headaches during collection, cleaning, versioning, and placement, while steady-state token delivery rarely needs the enormous bandwidth people imagine.

“The dataset is huge” and “the dataset is hard to stream” are two different claims. For text LLMs, they often describe two very different problems.

Google Trends chart showing rising interest in AI infrastructure

As this chart shows, and as featured in our AI infrastructure market deck, search interest in AI infrastructure has risen sharply

Are checkpoints the storage problem nobody can ignore?

Checkpointing is currently the hardest storage problem in frontier AI training because it combines multi-terabyte files, synchronized writes, frequent failures, and almost no tolerance for delay.

A checkpoint saves enough model and optimizer state to restart training after a failure. The optimizer usually accounts for most of the data. MLCommons estimates a full checkpoint at 105 gigabytes for an 8-billion-parameter model, 912 gigabytes for 70 billion parameters, 5.29 terabytes for 405 billion, and 15 terabytes for one trillion parameters.

Large clusters make the timing brutal. Meta’s Llama 3 training run used 16,000 accelerators for 54 days and encountered 419 interruptions, close to eight per day. Based on that failure pattern, MLCommons estimates that a 16,000-accelerator cluster may need a checkpoint about every 9.3 minutes to keep lost work under control. At 100,000 accelerators, the estimate falls to roughly every 1.5 minutes.

A one-trillion-parameter job saving 15 terabytes every 1.5 minutes would write more than 14 petabytes per day. Keeping checkpoint overhead below 5% would leave about 4.4 seconds for each synchronous save, which implies roughly 3.4 terabytes per second across the cluster.

Software can hide part of the pause. PyTorch’s asynchronous checkpointing reduced one seven-billion-parameter example from 148.8 seconds of effective downtime to 6.3 seconds. A later test on 1,856 H200 GPUs cut background checkpoint processing from about 436 seconds to 67 seconds while blocking GPU training for less than a second during staging.

The GPUs feel much less pain with these techniques, although the bytes still need to reach persistent media. At frontier scale, checkpointing remains the storage event that cluster designers cannot bluff their way around.

Model size MLCommons checkpoint size What that means in practice
8 billion parameters 105 GB A single save can already strain a small shared system
70 billion parameters 912 GB Each checkpoint approaches one terabyte
405 billion parameters 5.29 TB Saves become major distributed events
1 trillion parameters 15 TB Frontier clusters need multi-terabyte-per-second bursts

If you want more recent data on this point, please see our latest AI infrastructure market report.

Does video AI push storage much harder than text AI?

Video, robotics, and computer-vision AI push storage much harder than text models because every useful example contains far more bytes.

NVIDIA’s current SuperPOD guidance says high-resolution computer-vision datasets can easily exceed 30 terabytes and may require around four gigabytes per second of read performance for each GPU when images are uncompressed. At 1,000 GPUs, that theoretical requirement reaches four terabytes per second. At 10,000 GPUs, it reaches 40 terabytes per second.

Compression, caching, sharding, and lower-resolution sampling reduce the real load. The reference figure still shows why multimodal storage sits in a different category from text-token delivery. DeepSeek’s prepared text stream averaged megabytes per second across the entire cluster, while a vision workload may ask for gigabytes per second per GPU.

Video adds more layers. One source clip can produce frames, audio tracks, captions, object labels, embeddings, safety classifications, and several resized or augmented copies. Robotics and autonomous-driving systems add lidar, radar, telemetry, maps, and control logs.

Teams also have stronger reasons to keep the raw data. A public webpage can often be downloaded again. A rare robot failure, unusual road event, or factory accident may never repeat in exactly the same way.

Google’s latest results fit this pattern. Its largest reported gain from faster data loading came from multimodal training, where blocked GPU time fell by half. Storage pressure is moving toward the workloads with the richest inputs: video generation, world models, autonomous systems, scientific imaging, and physical AI.

Chart showing annual VC investment in AI infrastructure startups

This chart, included in our AI infrastructure market deck, shows annual VC investment in AI infrastructure startups

Is AI genuinely running short of storage?

AI has plenty of places to put bytes; the squeeze is in fast storage close to compute, where supply, cost, and power now matter much more.

Micron’s latest numbers are the strongest current warning. The company said data-center SSD revenue more than doubled sequentially to above $5 billion, while DRAM and NAND demand continued to exceed supply. It expects those tight conditions to persist beyond 2027 and has signed 16 strategic customer agreements covering future supply relationships.

Manufacturers are also pushing capacity much faster. Micron has announced a 122-terabyte enterprise SSD and plans a 245-terabyte model. Seagate has qualified hard drives up to 44 terabytes in production-scale hyperscale environments and is working toward drives as large as 100 terabytes.

Those roadmaps make a literal capacity crisis unlikely. Density is rising, and older or colder data can move to cheaper media. The harder issue is getting the right capacity in the right place with enough bandwidth and an acceptable power budget.

Meta now says every kilowatt spent on storage is a kilowatt unavailable to GPUs in a power-constrained data center. That changes the buying decision. A storage system can be cheap per terabyte and still be a poor AI system if it uses too much power, causes too many network trips, or sits far from the accelerators.

Where shortages appear, they will be local and specific: too little flash in an AI region, insufficient bandwidth during checkpoint bursts, too little power for another storage rack, or too much valuable data trapped on a slow tier.

Will AI data centers replace hard drives with SSDs?

AI data centers will keep buying both SSDs and hard drives because flash wins on speed while hard drives still win on the cost of keeping enormous datasets.

Meta’s current design makes the division easy to see. Memory and on-host flash handle the hottest data. Regional flash keeps active training material close to GPUs. Global hard-drive-backed storage remains the durable source of truth.

Putting everything on SSDs would simplify performance planning, although the bill would rise quickly once a company stores exabytes of raw video, synthetic outputs, logs, old checkpoints, and sensor data. Hard drives remain attractive for data that must be kept but rarely needs millisecond access.

The manufacturers are adapting each tier to AI. Micron is increasing SSD speed and density, including PCIe Gen6 drives and capacities above 100 terabytes. Seagate is pushing 44-terabyte hard drives into hyperscale qualification and targeting 100-terabyte products over time.

We are likely to see more aggressive movement between tiers. A dataset may begin on hard drives, hydrate into regional flash before training, enter host memory through prefetching, and disappear from the fast tiers after the experiment ends. Model files and popular checkpoints may remain on SSDs because many workers reuse them.

The hierarchy is the winning design. AI needs flash for urgency and hard drives for memory at scale.

Chart showing why CoreWeave is winning in the AI infrastructure market

This chart, included in our AI infrastructure market deck, shows why CoreWeave is winning in AI infrastructure

Is storage now a bigger problem than networking?

Networking remains the broader bottleneck in frontier training, while storage creates sharper bursts around data loading, checkpoints, and recovery.

During a distributed training step, GPUs constantly exchange gradients, activations, parameters, and routed tokens. That traffic runs through NVLink, InfiniBand, or high-performance Ethernet. A weak network can slow almost every iteration.

Storage behaves differently. Training samples can often be prefetched and cached. Once an inference worker loads its model, the weights may stay in memory. The severe storage moments arrive when thousands of workers read the same files, a checkpoint must finish, a failed job restarts, or a new dataset enters the cluster.

DeepSeek’s engineering work supports this distinction. Its reports focus heavily on communication overlap, expert routing, and network topology because those constraints sit directly in the training loop. Meta has also published extensive work on RoCE networks for tens of thousands of GPUs.

Storage becomes the binding constraint after the network is good enough and the job hits a burst. Meta’s latest roadmap includes scaling storage to network limits and checkpointing at higher scale without stalling GPUs. Google’s 15-terabyte-per-second bucket aims at the same gap.

A frontier cluster can be network-limited during computation and storage-limited a few minutes later during a checkpoint. Asking which one wins misses how quickly the bottleneck moves.

If you want more recent data on this point, please see our latest AI infrastructure market report.

Is storage becoming more important than GPU memory?

GPU memory still matters more for immediate AI performance, although storage increasingly decides how efficiently companies use that scarce memory.

Weights, activations, and active KV caches need HBM’s extreme bandwidth. A seven-billion-parameter model at 16-bit precision needs roughly 14 gigabytes for weights before we count the KV cache or runtime overhead. Larger models, longer contexts, and higher concurrency quickly fill the available memory.

An SSD can hold those bytes cheaply, but moving active weights or KV blocks down to storage adds a large latency penalty. Offloading helps a system fit more users or a larger model. It cannot make cold storage behave like HBM.

The change is happening between tiers. Model files sit in object storage, move to local NVMe before deployment, and then load into GPU memory. Active KV blocks stay in HBM, cooler blocks may move to host memory, and reusable or low-priority blocks can spill to SSDs. The quality of those transfers affects startup time, concurrency, and cost.

Storage works as a support system for memory. HBM sets the speed ceiling. Storage helps the fleet approach that ceiling without keeping every model, prompt, and cache permanently in the most expensive tier.

Chart showing the projected CAGR of the AI infrastructure market

This chart, included in our AI infrastructure market deck, shows annual funding in AI infrastructure startups

Does AI inference have a storage bottleneck today?

AI inference currently hits storage bottlenecks during model loading, fleet expansion, cache recovery, and document retrieval, while token generation itself usually depends more on compute and HBM bandwidth.

A large model may occupy hundreds of gigabytes. When a serving platform starts hundreds of workers, restarts failed nodes, deploys a new version, or opens capacity in another region, all those workers need the weights. A central object store can become congested if every server downloads the same files independently.

Google says Rapid Cache can deliver 2.5 terabytes per second from existing buckets and has produced model-load speeds up to 2.1 times faster in its tests. Thinking Machines Lab reported read-throughput peaks above 1.8 terabytes per second on the service for data preparation, pretraining, training, and model loading.

The same issue appears in KV-cache management. Reusing the cache for a shared system prompt or long document can avoid repeating expensive computation. The most active blocks belong in GPU memory, while less urgent blocks may sit in DRAM or SSDs. Once those blocks move, storage latency starts shaping time to first token.

Retrieval-augmented generation adds another path. The model may be ready, yet the application still waits for documents, embeddings, permission checks, and reranking.

Inference storage is primarily a deployment and context problem. After the model and active cache reach HBM, the GPU takes over. Before that point, storage can decide how quickly the service becomes useful.

Are AI agents turning storage into a live production problem?

AI agents are turning storage into a live production problem because useful agents must remember months of decisions without rereading every past interaction.

A chatbot can forget most of a conversation when the session ends. A project agent, coding agent, or research agent cannot. It accumulates tool results, failed attempts, user preferences, files, plans, permissions, and decisions that may matter much later.

Simply pasting the full history into every prompt becomes slower and more expensive as the task grows. Crude summaries create the opposite problem: they save tokens but lose names, dates, exceptions, and reasoning steps.

Microsoft’s newly published Memora research shows how serious this has become. On two long-memory benchmarks, Memora beat full-context inference and several retrieval systems while using up to 98% fewer context tokens. On LoCoMo, where conversations average 600 turns, it reached 86.3% judged accuracy. On LongMemEval, built around 115,000-token contexts, it reached 87.4%.

The architecture is the interesting bit. Memora keeps rich memories but retrieves them through short abstractions and links between related events. The model reads much less history while retaining details that a flat summary would erase.

Enterprise agents make the challenge harder because their memory lives across email, documents, databases, analytics tools, and customer systems. The storage problem now includes freshness, permissions, provenance, and knowing which record is authoritative.

For long-running agents, retrieval quality will matter more than raw byte capacity. Companies already have plenty of stored information. Their agents still struggle to find the right memory at the right moment.

If you want more recent data on this point, please see our latest AI infrastructure market report.

Chart comparing business model options for AI cloud infrastructure providers

This chart, included in our AI infrastructure market deck, compares the main business model options for AI cloud infrastructure providers

Can vector databases fix AI retrieval on their own?

Vector databases cannot fix AI retrieval on their own because semantic similarity covers only one part of the problem.

Embeddings are useful when a user asks for an idea expressed in different words. They work less reliably when the answer depends on an exact amount, the newest version of a document, a relationship between several records, or whether the user has permission to see the result.

Chunking creates another weakness. A vector search may retrieve the paragraph approving a contract while missing the later amendment that cancelled it. Adding more chunks increases coverage, yet it can also add noise and contradictory versions.

Freshness requires an entire pipeline. A changed file must be detected, parsed, split, embedded, indexed, and connected to its access rules. If any step lags, the agent can retrieve a convincing but obsolete answer.

Google’s new Smart Storage features point toward a broader design. The company is adding automatic object annotations, structured context, compliance labels, and agent access through MCP. Microsoft’s Memora research also moves beyond flat vector fragments by linking memories and separating what is stored from how it is found.

A useful retrieval stack mixes several methods. Vectors find semantically related candidates. Keyword search handles exact phrases. Structured databases enforce dates and filters. Graphs preserve relationships. Rerankers decide what deserves prompt space. Source systems settle conflicts.

The bottleneck is increasingly the coordination of those methods. Another vector index alone leaves an enterprise’s information just as stale, fragmented, and difficult to govern.

Which AI workloads are most likely to hit a storage wall?

Multimodal training, frontier checkpointing, autonomous systems, and persistent agents face the highest storage risk today.

The reason changes by workload. Video training needs sustained read bandwidth. Frontier models produce huge checkpoint bursts. Robotics generates raw sensor data that may be impossible to recreate. Agent systems need accurate retrieval from a growing history.

Standard text inference has lower exposure after the weights and active KV cache reach GPU memory. Small fine-tuning jobs also tend to hit compute or memory limits first, unless the data pipeline is unusually poor.

We should avoid one universal storage metric. Peak bandwidth says a lot about video training but little about an enterprise agent. Total petabytes reveal very little about checkpoint speed. Average latency can hide the slowest request that stalls an entire synchronized job.

AI workload Storage pressure today Most likely failure
Text pretraining Moderate Checkpoint delays and dataset placement
Multimodal training Very high GPUs wait for images, audio, or video
Frontier checkpointing Very high Multi-terabyte saves exceed the time budget
Standard LLM inference Low to moderate Slow model loading and cold starts
Long-context inference High and rising KV-cache movement adds latency
Enterprise RAG High Stale, incomplete, or unauthorized retrieval
Persistent AI agents High and rising Memory grows faster than useful recall
Robotics and autonomous driving Very high Sensor data overwhelms fast local tiers
Chart showing the share of revenue generated by each customer segment in the AI infrastructure market

This chart, featured in our AI infrastructure market deck, shows the share of revenue generated by each customer segment in the AI infrastructure market

Is storage becoming the bottleneck for AI?

Storage has become a major AI bottleneck, although compute, networking, and HBM still dominate many models and stages.

The claim is strongest in four places. Checkpoints already require multi-terabyte bursts and ever shorter save intervals as clusters grow. Multimodal workloads can demand gigabytes per second for each GPU. Inference fleets need faster model loading and multi-tier cache movement. Persistent agents increasingly depend on accurate long-term retrieval.

The claim weakens inside the hottest computational loops. Distributed training still leans heavily on GPU networking. Token generation still depends on HBM bandwidth, available memory, and accelerator throughput. Prepared text tokens can be surprisingly easy to stream.

The freshest industry behavior leaves little doubt that storage has moved up the priority list. Meta rebuilt its object-storage architecture around AI. Google created a zonal object-store tier above 15 terabytes per second. Micron reported a sequential doubling in data-center SSD revenue and expects supply tightness to continue beyond 2027.

Our judgment is direct: the claim is mostly true at the AI-system level and exaggerated when applied to every workload. Storage now decides whether costly GPUs stay busy during data loading and checkpoints, whether models deploy quickly, and whether agents can use the information companies already possess.

The next generation of AI infrastructure will be judged by the whole data path, from hard drives and object stores through SSDs and DRAM to HBM. Buying more GPUs while leaving that path unchanged is now a reliable way to waste part of the investment.

If you want more recent data on this point, please see our latest AI infrastructure market report.

OUR METHODOLOGY

This analysis tests whether storage is becoming a real bottleneck for AI data centers. We separated training throughput, checkpointing, multimodal data loading, inference deployment, capacity, power, retrieval, and persistent agent memory because each one creates a different storage problem.

We compared storage with networking, DRAM, and HBM so that a broader data-movement constraint would not automatically be labelled a storage bottleneck. The conclusion therefore applies to the full AI system, not to every model or every moment of execution.

We prioritized production architectures, engineering reports, technical documentation, research papers, supplier disclosures, and infrastructure roadmaps. Evidence carried more weight when an organization had changed what it built, bought, or operated, or when it disclosed measurable GPU waiting time, checkpoint performance, model-loading speed, or retrieval accuracy.

Vendor benchmarks were used to show the direction and possible scale of a problem, not as universal performance expectations. Derived figures in the article, including average token-stream bandwidth and checkpoint write requirements, were calculated from disclosed model sizes, training durations, failure rates, and timing assumptions.

Key infrastructure sources include Meta’s AI Storage Blueprint at Scale, Meta’s GenAI infrastructure report, Google Cloud Storage Rapid documentation, and Google’s Rapid Bucket and Rapid Cache deployment results.

Checkpointing and workload comparisons rely on MLCommons’ checkpointing workload, Meta’s Llama 3 technical report, PyTorch’s asynchronous checkpointing work, PyTorch’s large-scale H200 checkpointing test, the DeepSeek-V3 technical report, and NVIDIA’s DGX SuperPOD storage guidance.

Capacity, supply, and media-tier evidence comes from Micron’s fiscal Q3 2026 results, Micron’s 245TB data-center SSD announcement, Micron’s G9 NAND data-center portfolio, and Seagate’s Mozaic 4+ hyperscale roadmap.

For retrieval and agent memory, we used Microsoft Research’s Memora paper, Google Cloud’s Smart Storage work, and Google Cloud Storage’s MCP server documentation.

The final judgment aggregates the evidence across these dimensions. A dramatic checkpoint benchmark or a fast object-storage product was not allowed to stand in for the whole market; the conclusion reflects where storage is already binding, where it appears in bursts, and where networking, compute, or HBM still matters more.

Chart showing how GPU cloud infrastructure technology has evolved over time

This chart, included in our AI infrastructure market deck, shows how GPU cloud infrastructure technology has evolved over time