What business models actually work in the AI chip market?

Last updated: 25 August 2026
market research pitch 2026 statistics AI chip market

In our AI chip market deck, you will find everything you need to understand the market

SUMMARY

The AI chip business models that actually work today are full-stack merchant platforms, custom silicon, networking and interconnect, captive cloud chips, IP licensing, and foundry and advanced packaging. Proprietary hardware sold through contracted cloud capacity is also becoming credible, while the standalone accelerator startup remains the hardest model to make work.

The market increasingly splits between companies that own a computing platform and companies that remove a bottleneck for everyone else. NVIDIA and AMD belong to the first group; Broadcom, Marvell, Arm and TSMC largely benefit from supplying infrastructure that can survive changes in whichever accelerator wins.

Merchant AI accelerators are clearly a real business, but the evidence increasingly suggests that this is a very small club. NVIDIA operates at extraordinary scale, AMD has crossed into a serious second platform, and any new entrant now has to compete on software, networking, systems, availability and roadmap rather than silicon performance alone.

Custom silicon solves one of the worst semiconductor problems before development even begins: demand. Broadcom and Marvell can build around hyperscalers that already know the workloads they need to run, which removes much of the commercial uncertainty faced by a startup launching a general-purpose accelerator into an open market.

Networking may be an even cleaner way to participate in the AI boom. Every additional accelerator creates more demand for switches, optical links, network interfaces, retimers and scale-up connectivity, and suppliers can sell into NVIDIA GPUs, custom hyperscaler chips and competing accelerator architectures at the same time.

AWS and Google show why captive silicon is unusually powerful. The company designing the chip already owns workloads, data centers, software and customer distribution; Google is now going one step further by selling complete TPU systems externally after more than a decade of internal and cloud deployment.

Arm and TSMC show two very different ways to avoid picking a single AI architecture. Arm licenses reusable IP and collects royalties with software-like gross margins, while TSMC gets paid to manufacture and package chips across competing platforms. One is capital-light; the other is brutally capital-intensive but extremely difficult to displace.

Cerebras is providing some of the strongest evidence yet that a specialized chip company can turn itself into a compute provider instead of relying mainly on hardware sales. The important part is not the cloud API itself but the contracted demand behind it; the unresolved question is whether the infrastructure required to serve that demand can produce durable operating profits.

Groq and Graphcore illustrate a less comfortable pattern. Valuable chip technology can attract customers, strategic buyers and even multibillion-dollar transactions without producing a durable independent semiconductor platform. Good IP and a good chip company are not the same thing.

The practical lesson for a new AI chip company is to remove as much market risk as possible before funding the next generation. A committed customer, reusable IP, an architecture-neutral bottleneck or workloads that can already fill capacity currently looks much stronger than building a better general-purpose accelerator first and hoping the ecosystem follows.

Market map chart showing top companies and startups in the AI chip market

This market map, featured in our AI chip market deck, highlights top companies and startups in the AI chip market

When can we say an AI chip business model actually works?

An AI chip business model works today when customers keep buying across chip generations and the cash from those sales can help fund the next generation.

That sounds obvious, but it eliminates a surprising number of companies. A startup can tape out a good accelerator, win a famous customer and still have a weak business. AI chips require repeated spending on design teams, software, advanced packaging, manufacturing capacity and customer support. The test comes two or three years later: does the first generation make the second one easier to finance and sell?

We therefore care more about repeat purchases than benchmark wins. We also care about who brings the workload. NVIDIA can launch a new GPU into an existing CUDA ecosystem. AWS can launch Trainium into its own cloud. Broadcom can develop an ASIC with a hyperscaler that already knows what it wants to run. Arm can license the same architecture across many customers. Each model removes a different part of the commercial risk.

A chip can benchmark brilliantly and still be a lousy business. The models that work consistently remove uncertainty before billions of dollars have to be committed to the next product cycle.

If AI chips are booming, why do so many startups still get stuck?

AI chip demand is booming now, yet the market has become harder for standalone startups because customers increasingly expect software, systems, networking and a credible multigeneration roadmap along with the silicon.

The scale of the incumbents explains part of the problem. NVIDIA now generates tens of billions of dollars of Data Center revenue every quarter. AMD has turned Data Center into its largest segment. TSMC is spending more than $15 billion of capital expenditure in a single quarter. These companies can spread software development, packaging commitments and engineering costs across volumes that a startup cannot approach in its first few years.

Customers have also changed. A large AI lab buying accelerators today cares about rack design, memory, interconnect, compiler support, availability and the next chip on the roadmap. A 20% advantage on one benchmark becomes much less persuasive if deploying the product creates months of engineering work.

Graphcore is a useful historical example. The company once reached a multibillion-dollar valuation around its IPU architecture, yet eventually became a SoftBank subsidiary after struggling to build enough commercial scale independently. Graphcore is growing engineering operations again these days, including new campuses in India and Taiwan, but it is doing so inside SoftBank's much larger AI-computing group alongside Arm and Ampere.

The AI boom has made the prize much bigger while also raising the minimum scale needed to compete for it.

If you want more recent data on this point, please see our latest AI chip market report.

Google Trends chart showing rising interest in AI chips

As this chart shows, and as featured in our AI chip market deck, search interest in AI chips has grown significantly

Can anyone besides NVIDIA make the merchant AI accelerator model work?

Yes. AMD now shows that a second large merchant AI-compute platform can work, although NVIDIA remains in a completely different league by revenue.

AMD's latest results make that clear. Data Center revenue reached $6.7 billion in the quarter, up 107% from a year earlier, as EPYC CPUs and Instinct MI350 GPUs both grew. Data Center now represents 58% of AMD's total revenue. The segment also produced $2.1 billion of operating income.

NVIDIA's latest reported Data Center quarter reached $75.2 billion, up 92% year over year. Even allowing for the fact that NVIDIA includes networking while AMD includes server CPUs, the gap is enormous. We are looking at roughly an order of magnitude between the two businesses.

Still, AMD has crossed the threshold that matters. Large customers are deploying Instinct, ROCm continues to mature, and AMD is moving into complete rack-scale systems with Helios. Meta has outlined plans for up to 6 GW of AMD Instinct deployments, while companies including Anthropic, Microsoft, OpenAI, Oracle and others are attached to AMD's expanding AI infrastructure ecosystem.

So the merchant accelerator business remains viable today, but the evidence supports a very small number of serious platforms. A new entrant would be competing with companies that already sell CPUs, GPUs, networking, systems and software at massive scale.

Company Latest Data Center revenue YoY growth What it tells us
NVIDIA $75.2B quarterly 92% The full-stack merchant model can produce extraordinary scale
AMD $6.7B quarterly 107% A second large merchant platform is commercially real
New accelerator startup Usually far below this scale Varies Breaking into this model requires far more than a faster chip

Is custom AI silicon a better business than trying to beat NVIDIA?

For most semiconductor suppliers, custom AI silicon currently looks like a cleaner business than launching another general-purpose accelerator and fighting NVIDIA for developers.

Broadcom is the clearest example. Its latest quarter produced $10.8 billion of AI semiconductor revenue, up 143% from a year earlier, driven by custom accelerators and AI networking. Broadcom then guided the following quarter to around $16 billion, which would represent growth above 200%. Broadcom does not disclose custom accelerators separately from AI networking, but the order of magnitude is already huge.

Marvell shows the same model from a smaller base. The company says custom silicon has grown from almost nothing to roughly 25% of its Data Center revenue within a few years. Its latest reported Data Center quarter reached $1.83 billion, up 27% year over year, and Marvell says AI bookings are running at exceptional levels.

A very recent expansion of Marvell's relationship with Google gives us a rare look at how large these programs can become. The agreement covers inference accelerators, network interfaces, storage controllers, memory controllers and near-memory compute around Google's TPU ecosystem. Google also received warrants whose performance vesting is divided into 240 tranches, with one tranche tied to every additional $500 million of qualifying custom-product revenue. We calculate that the entire ladder spans $120 billion of cumulative revenue. The $120 billion figure describes the warrant's full vesting range rather than a guaranteed Google order, but it shows the scale contemplated when a chip supplier becomes embedded in a hyperscaler's roadmap.

Custom silicon removes one of the hardest problems in semiconductors: finding a workload after the chip exists. Google, Meta, Amazon or another hyperscaler arrives with enormous workloads already in hand. Broadcom or Marvell can focus on designing and delivering the silicon.

That trade gives away some of the upside available to a platform owner like NVIDIA. For most chip companies, it is still a much more realistic route to a multibillion-dollar AI business.

If you want more recent data on this point, please see our latest AI chip market report.

Chart showing annual VC investment in AI chip startups

This chart, featured in our AI chip market deck, shows annual VC investment in AI chip startups

Is AI networking a better business than AI accelerators?

For a supplier without a CUDA-sized software ecosystem, AI networking currently looks like one of the safest places to make serious money from the AI buildout.

NVIDIA itself shows how large the opportunity has become. Its latest Data Center quarter included roughly $14.8 billion of networking revenue, up 199% year over year. Networking is now large enough inside NVIDIA to be a major semiconductor business on its own.

Marvell provides an even cleaner view because connectivity sits at the center of its strategy. The company says optical interconnect has compounded at roughly 50% annually for five straight years and now represents about half of Data Center revenue. Marvell has lately added Celestial AI for optical scale-up technology and XConn for advanced interconnect, while launching 102.4 Tbps switching products aimed directly at AI clusters.

The reason is simple. More accelerators create more traffic between accelerators. Clusters built around NVIDIA GPUs, custom Google TPUs, AWS Trainium or future accelerators still need switches, optical DSPs, network interfaces, retimers and increasingly sophisticated scale-up connections.

That gives networking suppliers several ways to win at once. They can sell into competing compute architectures without making customers choose their accelerator as the new software standard.

Is AWS's best AI chip business the one where Amazon keeps the chips inside AWS?

Yes. Amazon's captive-chip model is currently one of the strongest AI semiconductor businesses in the world because AWS already owns the customer, the workload and the cloud where the chips run.

Amazon's latest earnings put the annual revenue run rate of its chips business above $25 billion, growing at a triple-digit rate. That figure includes Trainium, Graviton and Nitro rather than AI accelerators alone, but the growth of Trainium has become especially hard to dismiss. Amazon says it now has more than $225 billion of Trainium revenue commitments.

The demand is moving several generations ahead. Trainium2 has largely sold out. Trainium3, which only recently started shipping, is already close to fully subscribed. Customers have also reserved meaningful future Trainium4 capacity before broad availability. Anthropic has committed to as much as 5 GW of current and future Trainium capacity, while OpenAI has agreed to consume roughly 2 GW beginning with future deployments. Amazon Bedrock now runs most of its inference on Trainium.

AWS gets two benefits from the same chip. Customers pay AWS to consume the compute, while Amazon reduces how much external hardware it needs to buy. Andy Jassy has said that Trainium at scale could save tens of billions of dollars of annual capital expenditure and add several hundred basis points to AWS operating margins compared with relying entirely on third-party chips. Those are Amazon's estimates rather than separately reported Trainium profits, but they explain the logic behind the model.

The captive-cloud model works so well because the semiconductor product never has to create its own distribution system. Amazon already has hundreds of thousands of AI customers, data centers, power contracts, software services and a billing relationship with the customer.

A startup cannot recreate those conditions easily. For AWS, they make custom silicon unusually powerful.

If you want more recent data on this point, please see our latest AI chip market report.

Chart showing how Nvidia is leading in the AI chip market

This chart, featured in our AI chip market deck, shows how Nvidia is leading in AI chips

Is Google now turning TPUs into a real hardware business?

Yes. Google has quietly started selling TPU systems directly to external customers, which makes its current business model much more interesting than the old description of TPUs as chips used inside Google.

Alphabet's latest 10-Q says Google Cloud has begun recognizing revenue from TPU system sales. These systems combine TPU hardware with software, installation, support and extended warranties. Google says it has signed a limited number of agreements with customers that need specialized, very large on-premise infrastructure, with most of the related revenue expected to arrive later.

That changes the model. Google spent more than a decade developing TPUs around its own workloads and Google Cloud before asking outside customers to buy complete systems. That sequence gave Google internal demand, mature software and many generations of operating experience before entering the merchant hardware market.

The timing is also interesting. Google Cloud revenue recently jumped 82% to $24.8 billion, while operating income reached $8.8 billion. Those figures cannot be attributed to TPUs because Alphabet does not break TPU revenue or profit out separately. The latest filing simply tells us that TPU system sales have now become a real Google Cloud revenue stream.

We are therefore seeing a new hybrid model emerge: build silicon for internal use, rent it through the cloud, then sell complete systems once outside demand becomes large enough. Google can take that route with much less risk than a startup launching its first accelerator directly into the open market.

Should AI chip startups sell compute instead of selling chips?

For specialized AI chip companies, selling compute can work very well when real customer commitments arrive before the company builds enormous amounts of capacity.

Cerebras gives us the freshest public test of this model. In its latest quarter, GAAP cloud and other services revenue reached $126 million, up 281% from a year earlier. Cloud had represented about 43% of Cerebras revenue one quarter earlier; it reached about 70% in the latest quarter. Hardware revenue meanwhile fell to $54 million.

The company has also accumulated $25.4 billion of remaining performance obligations and says more than 600 MW of data-center capacity is live or under contract for delivery. Customers now include AI coding companies such as Cognition and Lovable as well as Block, Figma, AlphaSense, GSK and CrowdStrike. AWS has also started working with Cerebras on fast inference.

This is much stronger evidence than simply launching a pay-per-token API and hoping usage appears. Cerebras already has contracted future business against the infrastructure it is building.

The financial picture still deserves caution. Core revenue doubled year over year to about $210 million in the latest quarter, yet core operating margin remained negative at 16%. Cerebras expects full-year core operating margin around negative 17% to negative 19%, even after raising its revenue outlook. Selling compute solves the customer's procurement problem while shifting more capital and utilization risk onto Cerebras.

The model is clearly working commercially now. We still need more time to see whether it becomes consistently self-financing.

Cerebras metric Previous quarter Latest quarter
GAAP total revenue $193.4M $180.1M
GAAP cloud and other services revenue $82.8M $126.0M
Cloud share of GAAP revenue ~43% ~70%
Cloud revenue YoY growth 178% 281%
Core operating margin About -2% -16%

If you want more recent data on this point, please see our latest AI chip market report.

Chart showing the projected CAGR of the AI chip market

This chart, featured in our AI chip market deck, shows annual funding in AI chip startups

What does Groq's latest pivot say about standalone inference chips?

Groq's latest moves make the risk of the standalone inference-chip model unusually clear: excellent technology can create huge value while the independent business ends up moving somewhere else.

Groq spent years building its LPU architecture around extremely fast inference. NVIDIA eventually licensed the technology, and NVIDIA's annual report puts total consideration for the deal at roughly $17 billion, including amounts paid at closing and payable later. Groq founder Jonathan Ross, president Sunny Madra and other members of the team joined NVIDIA as part of the arrangement.

What happened afterward is even more revealing. Groq continued as an independent company, raised $650 million to expand its inference cloud, then recently announced another $350 million financing round. The company says it now operates 13 data centers, serves more than five million developers and processes trillions of tokens each week.

Groq has also become an NVIDIA Cloud Partner and is building infrastructure that can run NVIDIA accelerated computing. Groq's business has moved away from being only a proprietary-silicon story and deeper into operating AI inference infrastructure.

We should not call the underlying LPU technology unsuccessful. NVIDIA's willingness to pay billions for access to the technology and team points in the opposite direction. The harder part was turning that technology into a durable independent semiconductor platform.

For AI chip startups, that distinction is uncomfortable and useful. Creating valuable chip IP and creating a lasting chip company are two different achievements.

Is AI chip IP licensing the most capital-efficient business model?

AI chip IP licensing is still the most capital-efficient proven model in the sector, and Arm's latest results make the case unusually clearly.

Arm's latest quarter produced $1.29 billion of revenue, up 22% year over year, with a 97.2% GAAP gross margin. Royalty revenue reached $715 million, up 22%, while licensing revenue reached $574 million, up 23%. Data Center royalties more than doubled.

The model is powerful because Arm can get paid at two different moments. A customer licenses the architecture or compute subsystem during chip development, then Arm collects royalties as products containing that technology ship. One design can therefore generate revenue for years without Arm financing every wafer itself.

AI is making that model more valuable. Arm architectures now sit inside data-center CPUs, cloud processors, edge AI devices and increasingly custom silicon. Higher-end architectures such as Armv9 and Compute Subsystems also carry higher royalty rates than older products.

Arm itself is starting to push beyond pure licensing. The company recently introduced the Arm AGI CPU, its first Arm-designed production silicon for the data center. That move suggests Arm sees room to capture more of the value where the opportunity is large enough, even while licensing and royalties remain the financial engine.

Startups such as Tenstorrent are testing a similar hybrid idea by licensing AI cores and RISC-V technology while also building their own chips and systems. The approach is interesting, but public financial evidence is still too thin for us to put Tenstorrent in the same proven category as Arm.

Chart comparing business model options for AI accelerator chip companies

This chart, featured in our AI chip market deck, compares the main business model options for AI accelerator chip companies

Is TSMC the safest way to make money from the AI chip race?

TSMC currently has the cleanest architecture-neutral AI chip business because almost every serious winner still needs leading-edge manufacturing and advanced packaging.

TSMC's latest quarter reached $40.2 billion of revenue, up 34% year over year in US-dollar terms. Gross margin reached 67.7% and operating margin 60.3%. High-performance computing accounted for 66% of revenue after growing 20% sequentially, so we calculate that HPC contributed roughly $26.5 billion in one quarter. The category includes processors beyond AI accelerators, but AI is now one of its main growth engines.

Advanced manufacturing is doing most of the work. Nodes at 7 nanometers and below accounted for 77% of wafer revenue. TSMC has raised its annual capital budget to $60 billion to $64 billion, with around 70% to 80% going into advanced process technology and another 10% to 20% going toward areas that include advanced packaging.

The architecture neutrality is extremely valuable these days. NVIDIA GPUs, AMD accelerators, hyperscaler ASICs, Arm-based CPUs and many emerging designs can fight each other while sending fabrication and packaging demand toward the same supplier.

There is a catch for anyone looking for a startup blueprint. TSMC's current model works partly because replicating it requires decades of manufacturing knowledge and tens of billions of dollars every year. It is an excellent AI chip business and an almost impossible one to copy from scratch.

Which AI chip business models actually work today?

Today, six AI chip business models are clearly working, one newer model is becoming credible, and the standalone accelerator startup remains the hardest route.

The full-stack merchant platform works spectacularly for NVIDIA and increasingly well for AMD. Custom silicon works for Broadcom and Marvell because hyperscalers bring massive workloads before the chip is built. Networking and interconnect suppliers benefit from nearly every architecture. AWS and Google show how powerful custom chips become when the same company owns the cloud. Arm shows that reusable IP can produce software-like gross margins. TSMC captures demand from almost everyone through manufacturing and packaging.

Cerebras now gives us credible evidence for another model: proprietary hardware combined with contracted cloud capacity. Its cloud revenue is growing very quickly and its backlog is huge, although the company still has to prove that the model can generate durable operating profits.

Groq points the other way. The company's technology became valuable enough for NVIDIA to license, while Groq itself has lately moved deeper into the neocloud business. That makes the classic startup pitch of “we built a much better accelerator, now we need the ecosystem to follow” the weakest business model in the group.

If we were building an AI chip company now, we would start with a committed customer, reusable IP, a networking or memory bottleneck, or a workload that can already fill the capacity. Fighting NVIDIA head-on for a general-purpose accelerator market would come last.

AI chip business model Does it work today? Proven examples Our judgment
Full-stack merchant accelerators Yes, strongly NVIDIA, AMD Huge upside, brutally difficult for new entrants
Custom AI silicon Yes, strongly Broadcom, Marvell One of the best supplier models today
AI networking and interconnect Yes, strongly NVIDIA, Marvell, Broadcom Attractive because several compute architectures can win at once
Captive chips inside a cloud Yes, strongly AWS, Google Exceptional model when the company already owns workloads and distribution
IP licensing and royalties Yes, strongly Arm The most capital-efficient proven model
Foundry and advanced packaging Yes, exceptionally TSMC Architecture-neutral and highly defensible, but almost impossible to replicate
Proprietary systems + contracted AI cloud Increasingly Cerebras Commercially credible now; profitability still needs proving
Standalone accelerator startup Weakest evidence Groq and Graphcore illustrate the difficulty Great technology alone rarely creates enough distribution and ecosystem power

If you want more recent data on this point, please see our latest AI chip market report.

Chart showing how revenue is split across customer segments in the AI chip market

This chart, featured in our AI chip market deck, shows how revenue is split across customer segments in the AI chip market

OUR METHODOLOGY

This analysis tests which AI chip business models actually work today by separating the market into distinct models rather than treating every semiconductor company as if it faced the same economics. We look separately at merchant accelerators, custom silicon, networking and interconnect, captive cloud chips, IP licensing, foundry and advanced packaging, proprietary systems sold through cloud capacity, and standalone accelerator startups.

For each model, we prioritize the evidence that most directly reveals whether the business can survive successive product cycles. That includes revenue and growth, repeat deployments, multigeneration customer commitments, operating economics, capacity investment, backlog and evidence that customers are expanding beyond experiments.

The standard throughout the analysis is straightforward: there needs to be enough evidence of repeatable demand, meaningful commercial scale and a credible path toward funding or supporting the next generation of products. A benchmark win, one famous customer or a large financing round is not enough on its own.

Freshness carries extra weight because the AI infrastructure market is changing quickly. We therefore give greater weight to the latest reported quarters, regulatory filings, customer commitments, product roadmaps and capacity decisions available at the time of writing. Historical examples such as Graphcore are used mainly to show how a business model played out over a longer period.

The conclusions are not generated from a mechanical score. The evidence that matters for an IP licensing company is different from the evidence that matters for a cloud operator or a foundry, so we assess each model on its own economics and then judge the combined direction and strength of the evidence.

Where companies disclose broader categories, we keep those boundaries visible. Broadcom's AI semiconductor revenue includes custom accelerators and AI networking; Amazon's chip-business figure includes Trainium, Graviton and Nitro; TSMC's high-performance-computing category includes processors beyond AI accelerators; and Google Cloud's overall revenue and profit cannot be attributed specifically to TPUs.

We also distinguish reported figures from calculations made from those disclosures. The estimated $120 billion full vesting range in Marvell's Google warrant comes from 240 performance tranches multiplied by $500 million of qualifying revenue, while the roughly $26.5 billion estimate for TSMC's quarterly HPC revenue comes from applying its reported 66% revenue share to $40.2 billion of total quarterly revenue.

The analysis prioritizes company filings, investor-relations disclosures and other first-hand sources. Key sources include NVIDIA's Q1 FY2027 results, NVIDIA's FY2026 Form 10-K, AMD's Q2 2026 results, Broadcom's Q2 FY2026 results, Marvell's filing on its expanded Google relationship, Amazon's disclosure on Trainium commitments and multigeneration demand, Alphabet's Q2 2026 Form 10-Q, Cerebras's Q2 2026 results, Groq's June 2026 financing and cloud update, Arm's Q1 FYE2027 filing, TSMC's Q2 2026 results, and SoftBank's discussion of the Graphcore acquisition.

Chart showing how AI accelerator chip technology has evolved over time

This chart, featured in our AI chip market deck, shows how AI accelerator chip technology has evolved over time

Who is the author of this content?

NEW MARKET PITCH TEAM

We track new markets so founders and investors can move faster

We build living "market pitch" documents for emerging markets: AI, synthetic biology, new proteins, and more. Instead of outdated PDFs or hallucinated LLM answers, our clients get a clean, visual, always-updated view of what's really happening: key players, deals, regulations, and signals that matter. Learn more about us.

Back to blog