Where can a new company still break in AI chips?

Last updated: 31 August 2026
market research pitch 2026 statistics AI chip market

In our AI chip market deck, you will find everything you need to understand the market

SUMMARY

A new company can still break into AI chips, but the best openings are now in disaggregated inference, optical interconnect, memory-centric computing and specialized edge systems—not in building another general-purpose GPU.

The market is still growing fast enough to support new entrants, and young chip companies are already shipping hardware, signing contracts and reaching production. What has narrowed is the kind of technical advantage customers will tolerate the risk of adopting.

Nvidia does not need to lose its core GPU position for startups to win. In fact, the stronger clue is that Nvidia itself is adding specialized processors around the GPU, which makes heterogeneous AI infrastructure look more like the direction of travel than an exception.

Hyperscaler custom silicon closes one door and opens another. Google, Amazon and Meta make generic accelerators harder to sell to the very largest buyers, but they also prove that specialized chips can take a large share of real workloads when the economics are compelling.

The most attractive customer sits below the hyperscaler: a frontier lab, neocloud, government, financial firm or large enterprise with a very large compute bill but no reason to build a five-generation internal semiconductor program.

Inference is the clearest startup wedge because it breaks into more specific bottlenecks than training. Prefill, decode, memory movement, latency, context handling and token generation do not all want exactly the same machine.

Optical connectivity may be the strongest risk-adjusted opportunity because it benefits from growth across GPUs, TPUs and custom ASICs instead of requiring one accelerator architecture to displace another. Recent financing and acquisitions show strategic buyers are already paying heavily for that bottleneck.

Memory is becoming part of the product experience. Once token speed and inference cost are constrained by moving model weights and context rather than by theoretical FLOPS, architectures that reduce data movement can create a much larger advantage than a modestly faster compute engine.

Manufacturing strategy matters almost as much as architecture. A startup that can avoid the newest node, HBM and the most constrained packaging stack may have a better business even if its chip looks less impressive on a specification sheet.

Software remains the adoption tax. PyTorch, Triton, MLIR and open interconnect standards reduce the amount of rewriting required, but a chip that forces customers into a painful proprietary toolchain can still lose even with excellent hardware.

The common pattern among the more credible entrants is narrow technical focus and broad product delivery: solve one expensive workload unusually well, then ship enough of the rack, networking, software or cloud layer that the customer can actually use the advantage. That is where the room still is.

Market map chart showing top companies and startups in the AI chip market

This market map, featured in our AI chip market deck, highlights top companies and startups in the AI chip market

Is there actually room for a new AI chip startup today?

Yes. There is still room for new AI chip startups today, but the open ground has moved away from general-purpose GPUs and toward narrower problems in inference, memory, interconnect and edge computing.

The market itself is still expanding fast enough to create openings. TrendForce recently raised its forecast for AI server shipment growth this year from 28% to nearly 31%, after estimating that spending by the nine largest cloud companies could rise by roughly 90%. Stanford's latest AI Index puts global AI compute capacity at 17.1 million H100-equivalents, growing about 3.3 times per year since 2022. Nvidia still supplies more than 60% of that capacity, which leaves a large gap between “Nvidia dominates” and “there is nowhere else to build.”

Young semiconductor companies are also getting past the presentation stage. Etched has now delivered its first rack to Jane Street after Jane Street tested the hardware, and says it has signed more than $1 billion of customer contracts. Cerebras reported $193 million of quarterly revenue and signed a multiyear OpenAI compute agreement valued above $20 billion. d-Matrix has moved Corsair into full production. Axelera AI says it has shipped to more than 500 customers.

Capital is following those deployments. If we add just the recent disclosed rounds from Etched, Ayar Labs, Rebellions, d-Matrix, Positron and Axelera AI, they total more than $2.3 billion. That excludes Cerebras' public-market financing and earlier rounds at several of those companies.

What has changed is the kind of company that can plausibly win. Investors and customers are still backing new silicon, but increasingly when the chip solves a specific problem several times better rather than trying to become a smaller version of Nvidia.

Is Nvidia leaving any real opening for AI chip startups?

Nvidia is leaving openings around its platform, but there is currently very little evidence that its core GPU position is weakening.

Nvidia's latest reported quarter makes that hard to argue. Data-center revenue reached $75.2 billion, up 92% year over year. Compute revenue alone was $60.4 billion, while networking reached $14.8 billion and grew 199%. Gross margin stayed around 75%. A company under serious competitive pressure usually does not nearly double its largest business while preserving that level of pricing power.

The software lead is just as uncomfortable for challengers. Nvidia says CUDA now has roughly six million developers after twenty years of development, alongside libraries, compilers, profilers and frameworks that customers already use in production. Nvidia also won every benchmark in the latest MLPerf Training release and scaled some submissions across thousands of Blackwell GPUs. Frontier training remains particularly hard to attack because the workload changes quickly and buyers value flexibility.

Yet Nvidia itself is revealing where specialized chips can fit. Its newest Vera Rubin platform uses multiple types of processors for different jobs, including the Groq-derived LPX accelerator for very fast token generation. Nvidia recently put Groq 3 LPX into full production, with Nebius as the first announced cloud adopter. Artificial Analysis measured roughly 3,400 output tokens per second on a long-context Gemma 4 workload, around four times the fastest public endpoint in Nvidia's comparison.

The useful clue is that AI systems are becoming more heterogeneous. A startup does not need Nvidia to become weak. It needs one increasingly expensive workload where a GPU is no longer the best machine for every step.

If you want more recent data on this point, please see our latest AI chip market report.

Google Trends chart showing rising interest in AI chips

As this chart shows, and as featured in our AI chip market deck, search interest in AI chips has grown significantly

Are Google, Amazon and Meta making AI chips harder for startups to enter?

Google, Amazon and Meta are making generic AI accelerators harder to sell, while proving that specialized AI chips can take a surprisingly large share of real workloads.

Google is already deep into this transition. TrendForce estimates that TPUs could represent nearly 78% of Google's own AI-server shipments this year. Across the broader market, TrendForce expects ASIC-based AI servers to reach 27.8% of shipments and approach 40% by 2030. We are well past the stage where custom AI silicon is a research project.

Amazon provides an even clearer commercial example. Andy Jassy recently said AWS has more than $225 billion of Trainium revenue commitments. Trainium2 has largely sold out, Trainium3 is nearly fully subscribed, and customers have already reserved a meaningful share of Trainium4 despite broad availability still being some distance away. Amazon says Trainium2 offers roughly 30% better price-performance than comparable GPUs, while Trainium3 improves on Trainium2 by another 30% to 40%.

Meta says it has deployed hundreds of thousands of MTIA chips in production and has four additional generations either deployed or scheduled through 2027. Broadcom, which helps large customers build custom accelerators and networking chips, reported $10.8 billion of quarterly AI semiconductor revenue, up 143% year over year, and expects $16 billion in the following quarter.

Those numbers narrow the startup opportunity considerably. A new company selling a generic accelerator to Google or Amazon is competing against an architecture designed around that customer's exact workloads. The more interesting market sits one level below: customers with hyperscaler-sized AI bills but without hyperscaler-sized chip-design teams.

If hyperscalers build their own chips, who can an AI chip startup sell to?

The best customers for an AI chip startup today are probably frontier labs, neoclouds, governments, financial firms and large enterprises that spend heavily on compute but cannot justify designing a processor themselves.

Building an internal TPU or Trainium program requires more than hiring chip designers. The buyer needs predictable multiyear volume, access to advanced manufacturing, packaging expertise, compiler teams, server engineering, networking, model optimization and enough internal workloads to keep improving the silicon generation after generation. Very few companies clear that bar.

There is a much larger group sitting underneath it. Jane Street can run enough AI infrastructure to care deeply about latency and cost, but it does not need its own semiconductor organization. Neoclouds such as Nebius and specialized inference providers face the same problem. Frontier AI labs can easily consume billions of dollars of compute without wanting to own every part of the silicon supply chain.

AWS itself shows how large this middle market can become. Bedrock now serves more than 125,000 customers, according to Amazon, and most Bedrock inference runs on Trainium. The underlying demand for cheaper specialized compute therefore extends far beyond the handful of companies capable of designing chips.

For a startup, the sweet spot is a customer for whom a 30%, 50% or 3x improvement in compute economics is worth tens of millions of dollars, but building five generations of proprietary silicon would cost even more.

Chart showing annual VC investment in AI chip startups

This chart, featured in our AI chip market deck, shows annual VC investment in AI chip startups

Is AI inference the best place for a new chip company to attack?

Inference is currently the strongest part of the AI chip market for a new entrant because it contains far more specific bottlenecks than frontier training.

Training rewards flexibility. Model architectures change, numerical formats change and frontier labs constantly invent new techniques. A GPU's ability to handle many kinds of parallel computation becomes valuable insurance. Nvidia has also spent years optimizing the complete training system, from CUDA kernels to NVLink.

Inference looks very different once a model reaches production. Operators suddenly care about output tokens per second, time to first token, memory bandwidth, batch size, context length, power and cost per request. One architecture can be excellent for processing a huge prompt and mediocre at generating one token after another.

We can see the industry adapting around that difference. Nvidia's new Groq 3 LPX is aimed specifically at low-latency token generation rather than replacing Rubin GPUs. Cerebras sells extremely fast inference based on wafer-scale processors. d-Matrix focuses on low-latency inference using in-memory compute. Positron is attacking memory-heavy inference. Etched has built its entire company around frontier-scale inference clusters.

TrendForce expects inference computing power deployed this year to grow roughly 122%, faster than training. Faster demand growth plus greater workload specialization gives new architectures somewhere concrete to enter.

If you want more recent data on this point, please see our latest AI chip market report.

Can a startup still win by building a better GPU?

A startup building another general-purpose GPU today has probably chosen the hardest possible entry point in AI chips.

The competition already includes Nvidia, AMD, hyperscaler ASIC teams and increasingly Intel alternatives around parts of the stack. AMD illustrates the scale of the challenge. Its latest accelerators can compete closely with Nvidia on selected inference benchmarks, yet AMD needed an established semiconductor organization, an existing server business, years of ROCm development and deep relationships with cloud providers to get there.

A new entrant would have to build comparable silicon while simultaneously convincing developers to move workloads away from CUDA. Even a 20% or 30% hardware advantage can disappear quickly when Nvidia ships another software optimization. In its latest MLPerf Training work, Nvidia improved DeepSeek-V3 throughput by about 30% in three months without changing the underlying Blackwell hardware.

Small benchmark advantages are fragile. We would want to see an architectural gap measured in multiples, or a workload where the GPU's flexibility itself creates unnecessary cost.

“Faster GPU” is easy to pitch because everyone understands the market. It is much harder to turn into a durable company.

Chart showing how Nvidia is leading in the AI chip market

This chart, featured in our AI chip market deck, shows how Nvidia is leading in AI chips

Is disaggregated inference becoming a real AI chip market?

Yes. Disaggregated inference is moving into production now, and it gives specialized AI chips one of the cleanest ways into existing GPU clusters.

Large-model inference contains several different jobs. Prefill processes the prompt and benefits from heavy parallel compute. Decode generates tokens sequentially and becomes much more sensitive to memory access and latency. Attention, feed-forward layers and speculative decoding can create further differences. Using the same processor for every phase is convenient, but increasingly expensive.

d-Matrix and Parasail recently announced one of the clearest commercial deployments. Parasail is putting d-Matrix Corsair accelerators beside Nvidia Hopper and Blackwell GPUs across its inference infrastructure. Independent testing cited by d-Matrix reduced a 24-second response to under two seconds when GPUs and Corsair were combined for the tested workload. The companies are now exploring deployment across Parasail's fleet of more than 40 data centers.

AWS and Cerebras are following the same broad idea from another direction. Their planned infrastructure pairs AWS silicon for parts of inference with Cerebras systems where very fast generation is useful. Nvidia's own Vera Rubin architecture can also split work between Rubin GPUs and LPX accelerators.

Three different ecosystems have reached a similar conclusion: specialized hardware becomes much easier to adopt when it takes over one expensive phase instead of demanding the whole cluster.

Approach Specialized job Commercial evidence Why the opening is interesting
d-Matrix + Nvidia GPUs Low-latency inference stages Parasail production deployment Startup hardware can enter an existing GPU fleet
Cerebras + AWS Fast inference alongside AWS silicon Multiyear AWS collaboration Different processors can own different inference phases
Nvidia Rubin + Groq 3 LPX High-speed token generation LPX now in production, Nebius adopting Nvidia itself is embracing heterogeneous inference
Specialized inference clouds Route workloads to the best hardware Growing Cerebras and other specialist capacity Customers can adopt new silicon without owning it

Is memory now a bigger AI chip bottleneck than raw compute?

For many inference workloads, memory movement is already as important as raw compute, which makes memory-centric AI hardware one of the strongest areas for a new company.

Token generation repeatedly moves huge model weights and intermediate data while performing comparatively limited computation on each byte. If arithmetic units spend their time waiting for data, adding more theoretical FLOPS does surprisingly little.

Several startups are being built around exactly that imbalance. d-Matrix puts computation close to SRAM through digital in-memory compute and is developing 3D DRAM architectures for larger models. Positron raised $230 million at a $1 billion valuation to expand a memory-heavy inference architecture. Its investors include Jump Trading, a customer profile that cares intensely about latency rather than benchmark marketing.

The incumbents are redesigning around memory too. Each Trainium generation has increased memory capacity and bandwidth. Nvidia's LPX rack uses 128 GB of SRAM with 40 petabytes per second of aggregate SRAM bandwidth specifically to push token generation harder. Nvidia has also been expanding dedicated infrastructure for storing and moving inference context.

A new memory architecture still has to translate bandwidth into lower cost or visibly faster applications. But the underlying problem is real enough that startups, hyperscalers and Nvidia are all spending heavily on it at the same time.

Chart showing the projected CAGR of the AI chip market

This chart, featured in our AI chip market deck, shows annual funding in AI chip startups

Is optical interconnect a better startup opportunity than another AI accelerator?

Optical interconnect currently looks like one of the best risk-adjusted places to build in AI chips because the opportunity grows whichever accelerator vendor wins.

AI systems are becoming enormous distributed machines. Once hundreds or thousands of accelerators have to work together, the speed and energy cost of moving data between them starts limiting the useful compute of the cluster. Copper becomes especially awkward as bandwidth and distance increase.

The financing tells us how seriously the semiconductor industry takes that problem. Ayar Labs raised $500 million at a $3.75 billion valuation to move co-packaged optical I/O into volume production. Nvidia and AMD both back the company, alongside ASIC partners such as Alchip and MediaTek. A startup that can sell into several competing compute ecosystems has a much healthier strategic position than one betting everything on stealing GPU share.

Marvell went further and bought Celestial AI. Its latest regulatory filing puts total purchase consideration at roughly $3.5 billion for a company focused on photonic scale-up connectivity. Marvell subsequently said unusually strong AI bookings were pushing it to raise its revenue expectations, naming 800G and 1.6T optics, scale-up optical products, switching and custom AI silicon among the drivers.

TrendForce estimates that 800G-and-faster optical modules could rise from less than 20% of the market in 2024 to more than 60% this year. The bottleneck is turning into a large market quickly.

Company What it attacks Recent evidence Our read
Ayar Labs Co-packaged optical I/O $500M round, $3.75B valuation One of the cleanest independent startup wedges
Celestial AI Photonic scale-up fabric Acquired by Marvell for roughly $3.5B purchase consideration Strategic buyers are paying billions for the bottleneck
Marvell Optics, switches and custom silicon AI bookings drove a higher revenue outlook Connectivity spending is broadening beyond GPUs
Google Optical cluster networking Millions of TPUs drive huge 800G+ demand Hyperscalers already need optics at massive scale

If you want more recent data on this point, please see our latest AI chip market report.

Can an AI chip startup win without the newest process node or HBM?

Yes. A startup can sometimes create a better AI chip business by avoiding the most crowded manufacturing technologies instead of fighting Nvidia for exactly the same supply.

d-Matrix is a useful case. Corsair is built on TSMC's mature N6 process, uses organic substrates and LPDDR5 memory, and avoids an HBM-based CoWoS design. The company says this was deliberate: it wanted manufacturing capacity that could scale without depending on the same constrained packaging stack used by the highest-end GPUs.

That sounds counterintuitive in a market obsessed with smaller process nodes. Yet the useful metric is the performance of the full system for a particular workload, not the transistor density printed on the chip's specification sheet.

For memory-bound inference, edge systems, networking chips and some specialized accelerators, moving less data can save more energy than adding another generation of arithmetic density. Mature manufacturing can also reduce cost, shorten supply-chain risk and make volume easier to secure.

The trade-off is obvious for frontier training, where maximum density and efficiency remain extremely valuable. But a startup that needs TSMC's most constrained node, HBM and advanced packaging from day one is voluntarily entering the same procurement queue as the largest semiconductor customers on earth.

Chart comparing business model options for AI accelerator chip companies

This chart, featured in our AI chip market deck, compares the main business model options for AI accelerator chip companies

Does CUDA still make alternative AI chips too hard to adopt?

CUDA remains the biggest adoption barrier for alternative AI chips, although frameworks such as PyTorch and Triton are slowly making that barrier less absolute.

Nvidia says roughly six million developers now use CUDA after twenty years of ecosystem development. The advantage goes well beyond programming syntax. Customers inherit tuned kernels, libraries, debuggers, profilers, TensorRT, networking software and years of operational knowledge.

A startup asking every customer to rewrite models for a proprietary toolchain has a serious problem. Even strong hardware can sit unused if the migration cost wipes out the infrastructure savings.

The escape route is appearing higher in the software stack. AWS recently released tooling specifically designed to move PyTorch and Triton kernels onto Trainium with far less manual porting. d-Matrix builds around PyTorch, MLIR and Triton. UALink is creating an open scale-up interconnect backed by AMD, AWS, Google, Meta, Microsoft, Apple, Intel and others, with its newest specification designed to support multi-vendor accelerator systems.

These efforts do not reproduce everything CUDA offers. They do make it more realistic for customers to route selected workloads to another processor without changing how every developer works.

For a startup today, software compatibility should be treated as part of the chip architecture from the first day. A great processor with a painful migration path is still a bad product.

Can sovereign AI create real customers for chip startups?

Yes. Sovereign AI can give a chip startup a valuable first large customer, especially in countries that do not want their entire AI infrastructure tied to American GPU suppliers.

South Korea shows how aggressive this can become. Rebellions recently raised $400 million in a pre-IPO round led partly by the Korea National Growth Fund, following a $250 million round only six months earlier. The company says those two rounds brought in $650 million, more than three quarters of all the capital it had raised up to that point. It is now selling complete RebelRack and RebelPOD systems rather than just accelerator chips.

Europe is pushing in a similar direction. Public European money participated in Axelera AI's latest financing, with energy efficiency and local AI infrastructure forming part of the rationale. Governments in the Middle East are also funding compute capacity, while countries facing export restrictions have much stronger incentives to create domestic semiconductor alternatives.

The risk is building a company whose economics work only when a government is paying. Sovereign procurement can hide weak software, poor utilization or expensive manufacturing for longer than a normal commercial buyer would tolerate.

We would use sovereignty to secure an anchor deployment, then judge the product by whether customers in other countries still want it.

Chart showing how revenue is split across customer segments in the AI chip market

This chart, featured in our AI chip market deck, shows how revenue is split across customer segments in the AI chip market

Is edge AI still open to new chip companies?

Edge AI is still open, especially in robotics, defense, industrial vision and other applications where power, heat and latency make data-center hardware awkward.

Axelera AI is one of the better current examples. The company recently raised more than $250 million and says it has shipped to its 500th customer across areas including robotics, manufacturing, defense, retail and security. Those are real environments where watts, physical size and local processing can matter as much as raw model speed.

The market is also becoming easier to define. A robot does not need the world's fastest general-purpose accelerator. It may need a processor that runs perception and action models inside a strict power budget, survives an industrial environment and responds without sending every frame to a remote cloud.

Hailo shows why hardware alone still does not guarantee a great standalone business. The edge-AI chipmaker built a serious product portfolio and won designs across cameras, robots and embedded systems, but Microchip is now acquiring the company. Microchip's CEO later said Hailo had run into financial trouble and argued that putting those products through Microchip's much larger sales channel could accelerate them substantially.

That tells us where the edge opportunity is strongest. A new chip company should probably attach itself to a specific growing hardware market with clear distribution, such as drones, industrial robots, autonomous machines or smart cameras. “Edge AI” by itself is too vague to be a strategy.

If you want more recent data on this point, please see our latest AI chip market report.

What do the latest AI chip winners and failures have in common?

The AI chip companies gaining real traction lately are unusually narrow at the beginning and unusually broad by the time they reach the customer.

Etched focuses tightly on inference, yet it delivers racks rather than asking Jane Street to integrate a loose accelerator. Rebellions began with inference silicon and now sells RebelRack and RebelPOD systems. d-Matrix has expanded from accelerators into networking, rack-scale infrastructure and deployment software. Cerebras turned an unusual wafer-scale chip into systems and cloud inference that customers can consume directly.

The pattern solves a practical problem. A novel processor usually gets its advantage from the way memory, networking, compilation and scheduling work around it. Selling only the chip hands much of that advantage back to whoever integrates the server.

Recent acquisitions point in the same direction. Nvidia has now put Groq-derived LPX hardware inside Vera Rubin rather than treating low-latency inference as an isolated accelerator category. Marvell bought Celestial AI because optical fabric fits its broader data-center connectivity stack. Microchip wants Hailo because edge acceleration becomes more valuable inside a large embedded portfolio.

There is also a warning here. Building complete racks, clouds and software consumes much more capital than being a fabless chip supplier. Cerebras, Etched and other ambitious entrants have needed billions of dollars of financing or committed infrastructure.

The attractive company is narrow in the technical problem it chooses and broad enough in the product it delivers that the customer can actually use the advantage.

Chart showing how AI accelerator chip technology has evolved over time

This chart, featured in our AI chip market deck, shows how AI accelerator chip technology has evolved over time

Do falling AI inference prices make chip startups less attractive?

Cheaper AI inference makes average chip startups less attractive, but it makes genuinely differentiated hardware more valuable because customers have to keep cutting the cost of an exploding amount of computation.

Stanford's AI Index previously measured a more than 280-fold drop in the price of reaching roughly GPT-3.5-level performance in less than two years. Hardware performance per dollar has also continued improving quickly. A startup whose entire advantage is “our inference costs 20% less” can easily wake up one product generation later with no advantage left.

Demand has been moving in the opposite direction. Reasoning systems spend extra compute before answering. Coding agents can make hundreds of sequential model calls. Voice agents need continuous low-latency generation. Long-context systems move enormous amounts of memory. Stanford's latest AI Index also points out that inference can consume more energy than training within months once a model reaches large-scale use.

The economics resemble earlier computing markets: each unit gets cheaper while customers consume vastly more units. Amazon is selling out generations of Trainium, Nvidia's data-center business is still expanding at extraordinary speed, and TrendForce recently raised its AI server forecast again.

We would therefore set a much higher bar for a new architecture. A 20% cost advantage is vulnerable. A processor that changes the economics by 3x, removes a power constraint, or makes an application usable at a latency that GPUs cannot reach has a chance to survive several product cycles.

So where can a new company still break in AI chips?

A new company can still break into AI chips today, and our strongest bets are disaggregated inference, optical interconnect, memory-centric computing and specialized edge hardware rather than another general-purpose GPU.

Disaggregated inference ranks first because the evidence has moved beyond theory. Nvidia now combines Rubin GPUs with LPX accelerators. Parasail is deploying specialist inference hardware beside Nvidia GPUs. AWS and Cerebras are building heterogeneous inference infrastructure. Different companies are independently splitting AI inference into pieces and assigning those pieces to different processors.

Optical connectivity is almost as attractive. Every additional GPU, TPU or custom ASIC creates more demand for bandwidth, and the industry's recent behavior is unusually strong: a $500 million financing for Ayar Labs, a multibillion-dollar acquisition of Celestial AI, rapidly rising demand for 800G and 1.6T optics, and huge networking growth at established semiconductor companies.

Memory-centric compute comes next. Modern inference increasingly makes memory capacity, bandwidth and data movement visible to users through token speed and cost. Several architectures are attacking that problem from different directions, and Nvidia's own newest inference systems devote enormous silicon area and system design to SRAM and memory movement.

Edge and physical AI can also support new companies when the workload is narrow enough. Robots, drones, industrial machines and cameras create power, thermal, latency and privacy constraints that cloud-oriented hardware cannot simply ignore. Distribution will decide which of those chip companies survive.

Generic accelerators are much less convincing. Custom silicon is spreading through the hyperscalers, AMD is finally becoming a credible second GPU supplier, and Nvidia continues to strengthen the GPU, networking and software layers at once. A new company does not need to volunteer for that fight.

The best startup idea in AI chips now starts by finding an expensive part of AI computation that the existing machine handles badly. If that problem is growing quickly, repeats across many customers and can be solved several times better with a different architecture, there is still plenty of room to build a large company.

Opportunity Our conviction now Why there is still room What could kill it
Disaggregated inference Very high Production systems are already mixing specialized processors with GPUs Nvidia can absorb the best ideas into its own platform
Optical scale-up interconnect Very high Bandwidth demand rises across every accelerator ecosystem Long qualification cycles and difficult manufacturing
Memory-centric inference High Token generation increasingly exposes memory bottlenecks Advantages must survive rapid GPU improvements
Merchant inference systems High Large buyers want custom-chip economics without designing chips Heavy capital requirements and concentrated customers
Physical and edge AI Selectively high Robots and machines have hard power and latency limits Fragmented markets and difficult distribution
Sovereign AI silicon Medium Governments can finance large first deployments Dependence on subsidies or protected procurement
Generic inference accelerator Low Large market, but too many credible alternatives Price compression and weak differentiation
General-purpose GPU challenger Very low The theoretical market is enormous Nvidia, AMD, CUDA and hyperscaler ASICs all compete for the same workload

If you want more recent data on this point, please see our latest AI chip market report.

Table scoring and prioritizing the main pain points faced by companies in the AI chip market

In our AI chip market deck, we identify pain points entrepreneurs should prioritize

OUR METHODOLOGY

This analysis tests where a new AI chip company can still build a defensible business today. We break the market into the dimensions that most directly determine whether an opening is real: incumbent strength, customer structure, workload specialization, technical bottlenecks, software adoption, manufacturing constraints, commercial deployment and the direction of infrastructure spending.

We prioritized recent evidence from 2025 and 2026 because AI hardware is changing too quickly for older competitive assumptions to carry the same weight. Earlier data is used mainly when it helps establish a longer-term trend rather than describe the current market.

We did not treat every type of evidence equally. Production deployments, customer commitments, disclosed revenue, real infrastructure decisions and repeat commercial adoption carry more weight here than financing announcements or theoretical benchmark performance.

Benchmarks, funding rounds and acquisitions are used as supporting evidence when they show technical advantage, customer demand and strategic capital converging around the same bottleneck. A large round by itself does not prove product-market fit, and a strong benchmark by itself does not prove adoption.

We also avoided letting one datapoint determine the answer. A fast-growing market can still be structurally difficult for startups, while an incumbent with an overwhelming core position can still leave profitable openings around memory, interconnect, inference phases or edge constraints.

The final opportunity ranking is therefore a judgment based on clusters of evidence rather than a mechanical score. We looked for areas where the technical problem is expensive, demand is growing, multiple customers could share the same need, and a specialist architecture can plausibly create an advantage large enough to survive several product cycles.

Key sources used for market growth and infrastructure direction include TrendForce on 2026 AI server shipments and hyperscaler capex, Stanford HAI's 2026 AI Index, Nvidia's Q1 fiscal 2027 results, MLCommons' MLPerf Training v6.0 results, Amazon on Trainium demand and economics, and Broadcom's Q2 fiscal 2026 results.

For startup deployment and specialization, we relied especially on Etched's Jane Street deployment, Cerebras' Q1 2026 results, d-Matrix and Parasail's heterogeneous inference deployment, Nvidia's Groq 3 LPX production announcement, Ayar Labs on co-packaged optical I/O, Marvell's SEC filing on Celestial AI, Rebellions' financing and rack-scale systems, Axelera AI's financing and commercial deployment, UALink's open interconnect specifications, and Microchip's agreement to acquire Hailo.

Chart showing how revenue is split by region across Europe, Asia, North America, Africa, and South America in the AI chip market

This chart, featured in our AI chip market deck, shows how revenue is split by region across Europe, Asia, North America, Africa, and South America in the AI chip market

Who is the author of this content?

NEW MARKET PITCH TEAM

We track new markets so founders and investors can move faster

We build living "market pitch" documents for emerging markets: AI, synthetic biology, new proteins, and more. Instead of outdated PDFs or hallucinated LLM answers, our clients get a clean, visual, always-updated view of what's really happening: key players, deals, regulations, and signals that matter. Learn more about us.

Back to blog