Do AI agents need different GPUs?

Last updated: 31 July 2026
market research pitch 2026 statistics AI chip market

In our AI chip market deck, you will find everything you need to understand the market

SUMMARY

AI agents do not require an entirely new species of GPU, but they increasingly need a different kind of GPU system: more memory, faster data movement, stronger CPU support and better control over persistent context.

The workload has changed from isolated prompts to chains of dependent model calls. A coding or research agent may return to the model dozens of times, carrying forward tool results, errors, files and unfinished work.

This makes task-completion time more useful than raw tokens per second. A server can look efficient on a throughput benchmark while repeatedly evicting context and making the full agent workflow slower.

HBM capacity becomes a concurrency limit, not just a model-size limit. More memory lets the system keep additional agents and cached prefixes warm instead of rebuilding their state after every pause.

Prefix caching makes mature agent workloads surprisingly decode-heavy. Once most of the old prompt is reused, memory bandwidth and cache availability can matter more than the arithmetic needed to process the original context.

Raw GPU speed still matters during prefill, reasoning and large batches. But adding FLOPs without enough bandwidth, HBM or cache-aware software can leave those compute units waiting for data.

The GPU is not always the slowest component. Tool execution, tokenization, retrieval, browser control and kernel scheduling can leave an expensive accelerator idle, and faster GPUs sometimes make those CPU bottlenecks more obvious.

Ordinary request schedulers also miss the structure of agent work. Related calls arrive in bursts, disappear during tool use and return with reusable state, so maximizing batch size can be worse than keeping one workflow close to its cached context.

Agent context is becoming a storage hierarchy of its own. Active KV blocks stay in HBM, paused sessions move into DRAM or shared memory, and colder context may eventually sit on networked flash until the agent returns.

For most companies, the first upgrade should be software rather than silicon: prefix caching, workflow-aware routing, adequate CPUs and better context placement. New GPUs become necessary when measurements show that memory capacity or bandwidth is the hard limit, not simply because the workload has been labeled “agentic.”

Market map chart showing top companies and startups in the AI chip market

This market map, featured in our AI chip market deck, highlights top companies and startups in the AI chip market

Why are AI agents changing the GPU debate now?

AI agents are changing the GPU debate because they keep returning to the model with more history, more tools and more unfinished work.

A normal chatbot request has a fairly clean beginning and end. The user asks something, the model generates an answer, and the server can release most of the temporary state. A coding or research agent may call the model dozens of times, run external tools, wait for results, create subagents and then continue from the same growing history.

This shift has become much clearer lately. Nvidia’s latest Rubin GPU pairs 288 GB of HBM4 with 22 TB per second of memory bandwidth and is explicitly presented as hardware for agentic inference. AMD and Moonshot AI have rebuilt a coding-agent serving stack around KV-cache placement, memory tiers and workflow scheduling. Google has split its newest TPU family into a training chip and a separate inference chip with much larger on-chip memory for active context.

Those companies could have responded by adding more arithmetic power alone. Instead, they are spending a remarkable amount of silicon and system design on memory, data movement, CPUs, networking and storage.

Here, “different GPUs” does not have to mean a completely new chip category. The practical question is whether agents work best with the same hardware balance built for training and ordinary chatbot serving. Increasingly, they do not.

Are AI agents really harder to run than chatbots?

AI agents are harder to serve than chatbots because one user task can trigger dozens of linked model calls instead of one clean request.

A recent study called Agentic AI Workload Characteristics traced reasoning and non-reasoning agents across coding, terminal and general-assistant benchmarks. Most configurations averaged roughly 12 to 62 turns per completed task. One extreme run reached 786 turns.

Each turn can include the full conversation history, tool instructions, files, previous code changes, error messages and observations from the environment. The model repeatedly returns to work it started several seconds or minutes earlier.

That creates a different traffic pattern. Chatbot requests from separate users are mostly independent, so the server can mix them into large batches. An agent produces bursts of related requests separated by unpredictable pauses. Preserving the connection between those requests can matter more than fitting one extra unrelated prompt into the batch.

The best performance metric changes as well. Tokens per second remains useful, but users care about how long the whole agent takes to finish a task. A server can generate tokens very efficiently while repeatedly losing context and making the agent noticeably slower.

Workload feature Typical chatbot AI agent
Model calls per task Usually one or a few Often dozens and occasionally hundreds
Context Mostly tied to one request Grows and returns across many turns
External work Limited Search, code, files, databases and APIs
Traffic Relatively independent requests Bursts separated by tool pauses
Main user concern Fast response Fast and successful task completion
Google Trends chart showing rising interest in AI chips

As this chart shows, and as featured in our AI chip market deck, search interest in AI chips has grown significantly

Do AI agents mainly need more compute or more memory?

For well-optimized AI-agent systems, memory usually becomes the harder scaling problem.

Agents certainly consume a lot of compute. A coding agent that calls a large model 40 times will perform far more arithmetic than a chatbot producing one answer. Yet agent workloads have a second characteristic that ordinary token totals hide: much of the earlier work remains useful during later turns.

The Agentic AI Workload Characteristics study found average contexts of roughly 68,700 to 80,100 tokens in several SWE-bench Pro configurations. Maximum contexts reached about 146,000 to 166,000 tokens. Terminal and assistant tasks also regularly carried more than 50,000 tokens.

Failed attempts make the problem worse. An agent can add a broken command, a stack trace, a failed patch and a new explanation of the failure before trying again. In the same study, unsuccessful agents sometimes carried up to 1.8 times the average context of successful agents.

All of that context creates key-value data, usually called KV cache, which lets the model avoid recalculating attention for every previous token. The server must keep that state in GPU memory, move it to another memory tier or throw it away and rebuild it later.

Running one agent call is rarely the difficult part. Almost every modern inference GPU can do it. The expensive part is keeping many agents alive without repeatedly discarding work they have already completed.

If you want more recent data on this point, please see our latest AI chip market report.

Does long context make HBM capacity the first GPU limit?

Long-context AI agents can hit the HBM ceiling long before they exhaust a modern GPU’s matrix compute.

Nvidia has estimated that the KV cache for one 128,000-token Llama 3 70B session can occupy about 40 GB. The model’s FP16 weights require roughly another 140 GB. A single full-length session would already exceed the 80 GB of an H100 before allowing for temporary tensors, memory fragmentation or other users.

Production systems rarely use that exact configuration. They can quantize the model, distribute it across several GPUs, shorten contexts or store the KV cache at lower precision. Even so, the order of magnitude explains why agent concurrency becomes difficult.

Ten long-running sessions could require hundreds of gigabytes of cache. A multi-agent coding task can add several temporary workers at once. Another agent may disappear during a tool call and return later expecting its context to remain available.

The current hardware direction reflects this pressure. Nvidia moved from 80 GB on the H100 to 288 GB on Blackwell Ultra and Rubin. AMD’s MI355X also carries 288 GB. Google’s TPU 8i combines 288 GB of HBM with 384 MB of on-chip SRAM to keep more of the active working set close to its compute units.

Larger HBM therefore increases more than the maximum context length. It raises the number of agents that can remain warm at the same time, directly reducing waiting, cache eviction and expensive recomputation.

Chart showing annual VC investment in AI chip startups

This chart, featured in our AI chip market deck, shows annual VC investment in AI chip startups

Is memory bandwidth now more important than GPU speed?

Memory bandwidth often sets AI-agent decoding speed today, while raw FLOPs still dominate the initial prompt pass.

When a model first reads a large prompt, it can process many tokens in parallel. That stage, called prefill, makes heavy use of the GPU’s matrix units. Once the model begins generating tokens, each new token depends on the previous one. The workload becomes more sequential and repeatedly reads model weights and KV-cache data from memory.

Agents spend a large share of their model time in that second phase when earlier context is successfully cached. This makes memory bandwidth especially valuable.

The hardware progression is striking. Nvidia’s H100 provides around 3 TB per second of HBM bandwidth. Blackwell Ultra reaches about 8 TB per second. Rubin raises the figure to 22 TB per second. Bandwidth has increased by more than seven times from H100 to Rubin, while memory capacity has grown by about 3.6 times.

That gap shows where Nvidia expects pressure to build. More tensor performance still helps with long prompts, reasoning models and large batches, but agent decoding needs data to reach the compute units quickly enough.

Peak bandwidth should not be treated as a complete speed score. Kernel-launch overhead, cache placement and serving software can leave much of that bandwidth unused. As GPUs become faster, previously minor delays suddenly become visible. Buying the widest memory pipe only helps when the rest of the serving stack can keep it busy.

Does prefix caching change which GPU AI agents need?

Prefix caching pushes well-behaved AI-agent workloads toward decode-heavy, memory-heavy inference.

During every new agent turn, much of the prompt may be identical to the previous one. The system instructions, original user request, conversation history and earlier tool results have already appeared. Prefix caching stores the model’s processed representation of those tokens and reuses it.

The Agentic AI Workload Characteristics study measured cache-hit ratios ranging from 84.6% to 99.5%. Once cached tokens were separated from genuinely new input, decoding represented between 91% and 98.6% of model-execution time in the tested workloads.

Nvidia found a similar pattern while serving coding-agent harnesses. After the first call, later Claude Code-style requests showed cache-hit rates between 85% and 97%. A four-agent coding team reached an aggregate hit rate of 97.2%.

These results dramatically change how the workload looks. A raw prompt may contain 100,000 tokens, yet the model could receive only a short new tool result before generating its next action. Reprocessing all 100,000 tokens would waste a huge amount of compute.

High cache reuse favors GPUs with enough memory to keep useful prefixes available and enough bandwidth to read them quickly. The advantage disappears when memory pressure forces the server to evict those prefixes. Then the agent has to reload or recompute context, and the expensive prefill stage comes back.

If you want more recent data on this point, please see our latest AI chip market report.

Chart showing how Nvidia is leading in the AI chip market

This chart, featured in our AI chip market deck, shows how Nvidia is leading in AI chips

Why do AI tool calls confuse ordinary GPU schedulers?

Tool-calling AI agents make request-by-request GPU scheduling increasingly wasteful.

Most inference schedulers see individual prompts. They have limited awareness that one request belongs to a coding agent that will return after running a test, or that four incoming requests are subagents working on the same parent task.

The SAGA research project tested a scheduler that treats the whole agent workflow as the main unit of work. It keeps related calls close together, predicts which KV-cache blocks will be reused and moves work only when load balancing makes the move worthwhile.

On a 64-GPU cluster running coding and browser agents, SAGA reduced task-completion time by 1.64 times compared with vLLM using prefix caching and affinity routing. It also improved GPU-memory utilization by 1.22 times and reached 99.2% of its service-level targets under multi-tenant load.

The system accepted roughly 30% lower peak token throughput than a scheduler built to maximize batching. That trade-off sounds unattractive until we remember what users are waiting for. A coding agent that finishes in ten minutes is more useful than one that helps the cluster win a tokens-per-second benchmark while finishing in sixteen.

Agent infrastructure needs a clearer view of workflows. The scheduler should know which requests belong together, which agent is likely to return soon and which cached context is worth protecting.

Can CPUs slow down AI agents more than GPUs?

CPU tools can dominate AI-agent latency even when expensive GPUs are waiting nearby.

Many agent capabilities happen outside the language model. Search, code execution, data parsing, database queries, browser control and scientific software may run on CPUs or depend on CPU orchestration.

A recent CPU-focused agent study examined coding, retrieval, web, chemistry and tool-calling workloads across two CPU-GPU systems. Tool processing consumed as much as 88% of end-to-end latency in the most tool-heavy cases. Exact-nearest-neighbor retrieval accounted for 81% to 89% of latency in one RAG setup, while molecular conformer generation represented 85% to 88% in a chemistry workflow.

Faster GPUs sometimes made the imbalance more obvious. In two coding tasks, Bash and Python execution represented 38% and 25% of latency on the weaker-GPU system. Their shares rose as high as 65% on the system equipped with an H200 because model inference finished sooner while the CPU work barely changed.

A separate 2026 study of multi-GPU inference found that CPU shortages could delay kernel launches, tokenization and communication. Adding enough CPU resources reduced time to first token by between 1.36 and 5.40 times in its tested configurations without adding another GPU.

An agent server has to be sized as a complete machine. A top-end GPU paired with too few CPU cores, insufficient memory or a slow tool environment can be worse value than a more balanced system. It happens more often than people admit.

Chart showing the projected CAGR of the AI chip market

This chart, featured in our AI chip market deck, shows annual funding in AI chip startups

Should AI agents use different chips for prefill and decoding?

At large scale, separating prefill from decoding already makes technical and economic sense for AI agents.

Prefill and decoding want different things from the hardware. Reading a large prompt rewards parallel matrix compute. Generating one token after another rewards memory bandwidth, low latency and quick access to persistent context.

Research systems such as DistServe demonstrated the value of assigning those phases to separate resources before agent workloads became mainstream. The commercial market is now moving in the same direction.

Google’s newest TPU family includes TPU 8t for large-scale training and TPU 8i for inference and reinforcement learning. The TPU 8i carries more SRAM and HBM so that large KV caches can stay on the chip. Google also pairs the system with Arm-based Axion CPUs to handle preprocessing and orchestration without starving the accelerator.

Nvidia is taking a broader rack-level approach. Its latest platform combines Rubin GPUs, Vera CPUs, BlueField storage processors and Groq 3 LPUs. The Rubin GPUs provide large HBM capacity, while the LPUs use SRAM and static scheduling for low-latency inference.

Disaggregation adds complexity. The system must transfer KV cache between stages, route requests correctly and keep both sides busy. Small deployments may gain little from that machinery.

Large providers face a different equation. When millions of agent turns repeat every day, using one expensive accelerator type for every stage leaves large parts of the chip underused. At that scale, specialization is hard to avoid.

If you want more recent data on this point, please see our latest AI chip market report.

Will AI-agent context need its own storage layer?

Persistent AI-agent context is becoming a storage workload in its own right.

GPU HBM is fast and expensive. Keeping every inactive agent’s entire KV cache there would sharply reduce the number of active users a server can support. Throwing the cache away also carries a cost because the model must process the old context again.

AMD recently introduced Infinity Context, a shared KV-cache tier for distributed inference. AMD says a single one-million-token request can generate more than 600 GB of KV data in some model configurations, exceeding the HBM capacity of an entire MI455X node. Its design moves colder context out of HBM while keeping it available to several GPUs.

AMD and Moonshot AI have also described a hierarchy that spans GPU HBM, host DRAM, shared memory and SSD storage. Their scheduler checks where a cached prefix currently lives and whether moving it back will save time.

Nvidia is following the same path through BlueField-4-based context storage. The idea is to attach fast, networked flash directly to the inference system so that inactive KV blocks can return without passing through a slow traditional storage path.

These products are still new, and vendor performance claims need independent testing. Still, the architectural direction is obvious: agent context will move through several tiers according to how soon the system expects to use it again.

Context tier Best use Main limitation
GPU HBM Active turns and the hottest KV blocks Very fast but scarce and expensive
On-chip SRAM Tiny, extremely hot working sets Fastest tier with limited capacity
Host DRAM Paused sessions likely to return soon Slower than HBM
Shared memory tier Context reused across several GPUs Requires fast networking and coordination
NVMe or flash Large quantities of colder context Restore time can become the bottleneck
Chart comparing business model options for AI accelerator chip companies

This chart, featured in our AI chip market deck, compares the main business model options for AI accelerator chip companies

Do multi-agent systems need a different GPU setup?

Multi-agent systems need better memory sharing and burst handling more urgently than a special “multi-agent GPU.”

A single agent usually creates a chain of calls. A multi-agent system can suddenly create several parallel branches. A coordinator may launch workers for research, coding, testing and review, then wait for all of them to return.

That pattern produces sharp memory peaks. AMD describes coding subagents as short-lived bursts that quickly expand the active KV-cache working set. Average GPU utilization can look comfortable even though one burst forces the server to evict valuable context.

Multi-agent workflows may also reuse a large amount of shared information. Several workers can receive the same repository map, user request, system prompt or research material. Placing them on unrelated servers can cause each worker to rebuild almost identical cache blocks.

The GPU requirement therefore depends on coordination. Large shared memory domains and fast interconnects help when workers exchange or reuse context. Separate GPUs help when several models genuinely need to execute at once. A scheduler may keep the parent agent’s state warm while its children run elsewhere.

More agents do not automatically improve the result. Weak orchestration can produce duplicate searches, contradictory changes and several copies of the same prompt. In that case, the cheapest infrastructure improvement is simply making fewer unnecessary agent calls.

Useful multi-agent systems still create a distinctive serving problem. Their demand arrives in correlated waves, which makes peak memory and placement more important than average token volume.

Can KV-cache compression delay the need for new GPUs?

KV-cache compression can postpone expensive GPU upgrades, although the gains depend heavily on the agent workload.

Storing each key and value at fewer bits allows more context to fit in the same HBM. That can support longer sessions, more concurrent agents or fewer cache evictions.

AMD’s recent UltraQuant research focused specifically on long-context, multi-turn agents. Its four-bit KV-cache approach reduced median time to first token by 3.47 times during cache-pressured later rounds and by 2.3 times across all rounds. Output throughput improved by 1.63 times compared with an FP8 KV-cache baseline on the tested AMD system.

Those gains came from more than shrinking the data. AMD built specialized attention kernels and adapted the format to its CDNA4 hardware. Simply switching a configuration from eight bits to four bits would not automatically reproduce the result.

Compression can also affect quality. Keys and values play different roles in attention, and some models tolerate aggressive quantization better than others. A broader 2026 benchmark comparing quantization, pruning and cache-merging methods found that the highest compression ratio was a poor predictor of overall serving performance. Different methods won on different tasks and models.

KV compression is a useful pressure valve, not a permanent escape from memory limits. It can extend the life of current GPUs, especially when the workload is stable and carefully tested. Continued growth in context length and concurrent agents will eventually consume the saved space.

Chart showing how revenue is split across customer segments in the AI chip market

This chart, featured in our AI chip market deck, shows how revenue is split across customer segments in the AI chip market

Can software fix AI-agent inference before new GPUs arrive?

Better serving software can deliver larger near-term gains than replacing a recent GPU fleet.

Many agent deployments still treat every turn as an unrelated API request. They discard useful KV cache, route follow-up calls to different workers and optimize token throughput without tracking how long the full task takes.

Current serving frameworks are closing those gaps. vLLM can reuse KV-cache blocks across requests with matching prefixes. Nvidia Dynamo adds cache-aware routing, distributed memory management and separate prefill and decode workers. AMD’s newer serving stack makes routing decisions based on whether context sits in HBM, DRAM or shared storage.

Workflow-level scheduling offers another large opportunity. SAGA improved agent task-completion time without changing the GPUs. MARS, a separate GPU-CPU co-scheduling system, reported up to a 5.94-times reduction in end-to-end latency across its tests. When integrated with the OpenHands coding agent, it accelerated task completion by as much as 1.87 times.

These are research and vendor results rather than universal guarantees. Even so, they show how much performance can remain hidden inside an existing cluster.

New hardware solves physical limits such as HBM capacity and bandwidth. Software determines how quickly a deployment reaches those limits. A company replacing its GPUs before fixing cache reuse, CPU scheduling and request routing may spend millions of dollars to preserve a badly designed system.

If you want more recent data on this point, please see our latest AI chip market report.

Can TPUs, LPUs and inference ASICs replace GPUs for AI agents?

Specialized accelerators can beat GPUs on narrow AI-agent workloads, while GPUs still cover the widest mix of models and tools.

Token generation has several properties that custom chips can exploit. It is sequential, memory-intensive and repeated at enormous scale. Groq’s LPU uses large amounts of SRAM and static scheduling. Google’s TPU 8i dedicates more on-chip memory to inference. Cerebras keeps compute and memory close together across a wafer-scale processor.

Those designs can provide very low and predictable latency. That is valuable for agents because every model call can delay the next tool action. Saving a fraction of a second once is modest. Saving it across 40 dependent calls can noticeably shorten the task.

Agent workloads also contain messy parts. A computer-use agent may process screenshots, run OCR, embed documents, decode text, transcribe audio and call several external programs. A coding agent may switch models or execute custom kernels. GPUs handle that diversity well because their software stack supports a huge range of models and operations.

Nvidia’s latest platform quietly acknowledges both sides of the debate. The company now combines Rubin GPUs with Groq 3 LPUs instead of expecting one processor to handle every phase equally well. Google similarly offers different TPUs for training and inference.

Specialized chips will capture portions of agent inference where the models and traffic are predictable. GPUs should remain the central flexible resource for mixed workloads, frontier models and multimodal tools. The agent data center is heading toward several accelerator types rather than one universal winner.

Chart showing how AI accelerator chip technology has evolved over time

This chart, featured in our AI chip market deck, shows how AI accelerator chip technology has evolved over time

Which hardware setup fits each type of AI agent?

AI agents need different hardware mixes today because coding, research, voice and computer-use agents stress different parts of the system.

A small internal agent that makes a few API calls may run perfectly well on an existing inference GPU. Adding prefix caching and enough CPU capacity could provide all the improvement the company needs.

Long-context coding and research agents put much more pressure on HBM. They benefit from large-memory GPUs, cache compression and a way to offload inactive context. High concurrency strengthens the case for fast interconnects and shared storage.

Voice agents care heavily about latency. A specialized decoding accelerator may be valuable when the model stack is stable and supported. Computer-use and multimodal agents need a broader mix of GPU, CPU and media-processing capabilities.

The right purchasing metric also changes by workload. Training FLOPs say little about how quickly a browser agent finishes a task. Teams should measure end-to-end completion time, success rate, peak memory, cache-hit ratio, CPU utilization and cost per successful task.

Agent workload Likely bottleneck Best current setup
Lightweight tool assistant Orchestration and API latency Existing inference GPU with adequate CPUs
Long coding or research agent KV-cache capacity and reuse Large-HBM GPU with cache-aware serving
High-volume voice agent Sequential decoding latency Inference GPU, TPU or low-latency ASIC
Multi-agent coding system Burst memory and placement Shared memory domain with fast interconnects
Computer-use agent Mixed vision, language and CPU tools Flexible GPU-CPU system
Million-token workflows Context movement and storage Large-HBM accelerators with DRAM or flash offload

Do AI agents need different GPUs?

AI agents increasingly need a different kind of GPU system, although most companies can keep using their existing GPUs for now.

The mathematical core has not changed. Agents still run transformer models, and modern Nvidia, AMD and cloud accelerators can execute them. A team launching a moderate number of short agents does not need to replace a recent cluster.

The hardware balance changes once agents become long-running and heavily used. They repeatedly call the model, preserve large KV caches, pause for tools, create bursts of subagents and return to earlier state. Memory capacity, memory bandwidth, CPUs, interconnects and context storage begin to shape performance as much as raw matrix compute.

The latest hardware releases already reflect that change. Nvidia has paired far more bandwidth with a rack containing GPUs, CPUs, LPUs and storage processors. Google has separated its training and inference TPUs. AMD is building shared context tiers because HBM alone cannot economically hold every active and paused session.

For most deployments today, the first move should be improving prefix caching, workflow scheduling, CPU provisioning and context placement. Replacing the GPU comes later, when measurements show that HBM capacity or bandwidth has become the hard limit.

So the answer is fairly clear. AI agents do not require an entirely new species of GPU, but they are pushing the market away from training-shaped hardware and toward memory-heavy, inference-focused, heterogeneous systems. The GPU remains at the center, just surrounded by more specialized hardware than before.

If you want more recent data on this point, please see our latest AI chip market report.

Table scoring and prioritizing the main pain points faced by companies in the AI chip market

In our AI chip market deck, we identify pain points entrepreneurs should prioritize

OUR METHODOLOGY

We approached “Do AI agents need different GPUs?” as an infrastructure question rather than a prediction about which chip company will win. We assessed the workload across model-call patterns, context growth, KV-cache pressure, memory bandwidth, CPU and tool latency, scheduling, multi-agent bursts, context storage and prefill-decoding separation.

We prioritized first-hand workload measurements and systems research showing how agents behave in practice. Agentic AI Workload Characteristics provided evidence on turn counts, context sizes, cache reuse and decoding time, while SAGA and MARS helped us evaluate whether workflow-aware scheduling and CPU-GPU coordination could improve full task-completion time without replacing the accelerators.

We also reviewed current hardware and serving architectures from Nvidia, AMD and Google. We focused on where those companies are adding resources, including HBM capacity, memory bandwidth, SRAM, CPUs, interconnects and dedicated context-storage tiers, rather than relying on broad product claims about “agentic AI.”

Performance figures are used only for the models, systems and configurations in which they were tested. Research-cluster results and vendor benchmarks can identify a bottleneck or demonstrate a mechanism, but we do not assume that the same improvement will transfer unchanged to every deployment.

We separated the current purchasing decision from the longer-term direction of the market. Evidence from very large agent deployments can show where infrastructure is heading without implying that every company should replace a recent GPU fleet today.

Key sources used for this analysis include Agentic AI Workload Characteristics, SAGA, MARS, UltraQuant, DistServe, Nvidia’s Rubin architecture documentation, Nvidia’s H100 specifications, Nvidia Dynamo’s KV-cache-aware routing documentation, Google’s TPU 8i infrastructure announcement, AMD’s MI355X specifications, and AMD Infinity Context.

Chart showing how revenue is split by region across Europe, Asia, North America, Africa, and South America in the AI chip market

This chart, featured in our AI chip market deck, shows how revenue is split by region across Europe, Asia, North America, Africa, and South America in the AI chip market