AI inference chips: which startup is ahead?

Last updated: 31 July 2026
market research pitch 2026 statistics AI chip market

In our AI chip market deck, you will find everything you need to understand the market

SUMMARY

Cerebras is clearly ahead in AI inference chips today, while Groq and SambaNova lead the private-company chase.

The field looks crowded, but it is already split into two groups. Cerebras, Groq and SambaNova operate mature commercial platforms; the other six contenders are still proving that production milestones can become repeated customer deployments.

Revenue is the cleanest separator. Cerebras reported $510 million in 2025 revenue and another 94% year-over-year increase in the first quarter of 2026, while every private rival still relies on bookings, usage, customer logos or unofficial estimates.

Cerebras is no longer technically a startup after its IPO, but excluding it would make the comparison less useful. It is the clearest example of what a venture-backed inference-chip company can become, and it now sits in a tier of its own.

Groq has built the widest private operating footprint, with five million developers and 13 data centers. That distribution is real, though the five-million figure is not a paying-customer count, and Nvidia’s licensing deal and hiring of key Groq leaders make the long-term hardware moat less secure.

SambaNova is the strongest intact private challenger. Its enterprise and sovereign deployments, new financing and SN50 rollout give it a credible route past Groq, but only broad shipments and clearer revenue disclosure will settle that argument.

Cerebras also leads the cleanest public speed comparison. On GPT-OSS-120B, Artificial Analysis measured roughly 1,955 output tokens per second for Cerebras, compared with about 709 for SambaNova and 478 for Groq. The gap is hard to explain away.

Groq currently offers the cheaper leading startup API, while Cerebras sells a speed premium. That creates a sensible split: background workloads can chase lower token prices, while coding, voice and customer-service applications may pay more to remove visible latency.

The most practical challengers may not try to replace Nvidia outright. d-Matrix can sit beside existing GPUs and accelerate decoding, which lowers adoption friction and gives customers a way to improve inference without rebuilding their whole stack.

Etched has the biggest expectations gap. Working silicon, $800 million in funding and more than $1 billion in contracts make it a serious contender, but manufacturing and customer validation still have to catch up with the headline numbers.

Power availability may shape the market as much as raw chip speed. Air-cooled systems, lower rack power and better performance per watt could become decisive for enterprises and sovereign clouds that cannot secure more electricity quickly.

Our ranking is Cerebras first, Groq second and SambaNova third. Groq wins on private infrastructure already operating, SambaNova has the better chance of moving up, and the rest of the field needs more evidence that shipped chips are turning into durable revenue.

Market map chart showing top companies and startups in the AI chip market

This market map, featured in our AI chip market deck, highlights top companies and startups in the AI chip market

Which AI inference chip startups are we actually comparing?

The serious AI inference chip field currently contains nine startup-origin contenders, but only Cerebras, Groq and SambaNova have mature commercial platforms.

We include companies designing data-center chips, systems or cloud infrastructure mainly for generative AI inference. That covers businesses selling complete servers, operating their own inference clouds or supplying accelerators to other infrastructure providers.

We exclude Nvidia, AMD, Intel, Google and Amazon because they are established chip or cloud companies rather than startups. We also leave out edge-chip specialists such as Hailo and Axelera AI, training-first companies such as MatX, and inference clouds that do not design their own silicon.

Cerebras needs a special note. It recently completed an initial public offering, so it is no longer technically a startup. We still include it because it came from the same venture-backed generation, competes directly with these companies and provides the clearest benchmark for what an inference-chip startup can become.

Funding figures remain approximate. Companies do not always separate equity, debt, secondary transactions and strategic investments in the same way.

Company What it does Current maturity Approximate capital raised
Cerebras Wafer-scale systems and cloud infrastructure for training and fast inference Large commercial deployments; now publicly traded About $2.9 billion privately, followed by roughly $6.4 billion through its IPO
SambaNova Reconfigurable dataflow chips, racks and managed enterprise inference Commercial SN40L systems; SN50 rollout beginning About $2.5 billion
Groq Purpose-built LPUs and a globally distributed inference cloud Commercial cloud operating across 13 data centers About $2.4 billion
Rebellions Korean NPUs, servers and rack-scale inference systems Commercial products available; international expansion beginning About $850 million
Etched Transformer-focused chips and complete inference clusters Working silicon undergoing customer validation $800 million
d-Matrix Digital in-memory accelerators focused on fast AI decoding Full production and first commercial-scale deployments $450 million
Positron Energy-efficient transformer inference servers and custom silicon Atlas systems shipping; next-generation chip in development More than $300 million
FuriosaAI Power-efficient data-center NPUs for language, multimodal and vision models RNGD in mass production $246 million
Taalas Chips that hardwire specific AI models into silicon Working model-specific chip with limited deployment evidence $219 million

Is there already a clear leader in AI inference chips?

Cerebras is ahead today, and the gap is too large to describe this market as a three-way tie.

Cerebras is the only company in the group reporting audited revenue at substantial scale. It generated $193.4 million in the first quarter of 2026, almost twice as much as one year earlier. Its latest filing also disclosed a 750-megawatt agreement with OpenAI valued at more than $20 billion and a partnership that will bring Cerebras inference to Amazon Web Services.

Its public performance is equally difficult to ignore. Artificial Analysis recently measured Cerebras at around 1,955 output tokens per second on GPT-OSS-120B. SambaNova reached about 709, while Groq produced 478. On that identical model, Cerebras was almost three times faster than SambaNova and more than four times faster than Groq.

The private-company race is much tighter. Groq has the largest visible operating footprint, with 13 data centers and more than five million developers. SambaNova has fewer disclosed users, but it recently raised the first $1 billion of a new financing round at an $11 billion valuation and added JPMorgan Chase as an on-premise customer.

Cerebras now occupies its own tier. Groq and SambaNova form the main private chasing group, while the remaining startups are still turning production milestones into repeated commercial deployments.

If you want more recent data on this point, please see our latest AI chip market report.

Google Trends chart showing rising interest in AI chips

As this chart shows, and as featured in our AI chip market deck, search interest in AI chips has grown significantly

Who is making real money from AI inference chips?

Cerebras is currently the only company we can confidently describe as a large and fast-growing AI inference business.

Its annual revenue rose from $290.3 million in 2024 to $510 million in 2025, an increase of roughly 76%. First-quarter revenue then grew another 94% year over year. Cerebras now expects approximately $855 million to $865 million in core revenue for the full year.

Cloud and services revenue reached $82.8 million in the latest quarter, up 178% from one year earlier. Cerebras is beginning to earn more from recurring access to its infrastructure rather than depending entirely on occasional sales of expensive systems.

Groq probably operates the second-largest business, but its financial evidence is much thinner. Reporting from The Information indicated that Groq had reduced its 2025 revenue target from more than $2 billion to around $500 million. Even the lower figure would represent substantial sales, yet Groq has never confirmed an audited result or disclosed its gross margin.

SambaNova said it finished 2025 with record bookings and revenue but did not provide the actual numbers. Its new customers and financing suggest demand is improving, although we cannot tell whether it is running a $50 million business, a $300 million business or something larger.

Etched has disclosed more than $1 billion in signed customer contracts, but its first systems are still being validated. d-Matrix, FuriosaAI and Rebellions have entered production without reporting meaningful revenue figures. Positron is shipping Atlas systems but has not revealed unit volume or sales.

Cerebras is already operating at a different commercial scale. The rest of the field still asks us to infer revenue from usage, customer logos and production announcements.

Who has the strongest customers and can actually deliver at scale?

Cerebras has the strongest contract book and the financial capacity to support the largest deployments, while Groq has built the widest private-startup infrastructure footprint.

The OpenAI agreement is the largest disclosed contract in the category. Cerebras will provide 750 megawatts of computing capacity over several years under a deal valued at more than $20 billion. Its remaining contracted obligations reached approximately $25 billion in the latest quarter, almost 49 times its entire 2025 revenue.

That figure is not revenue already earned. Cerebras must build the infrastructure, deliver the capacity and meet the contract’s conditions before recognizing most of the money. Even so, the agreement is far more substantial than a pilot, reservation or memorandum of understanding.

Amazon gives Cerebras another route to customers. AWS can place Cerebras behind a cloud interface already used by large companies, reducing the procurement and integration burden. CrowdStrike has also selected Cerebras for latency-sensitive security applications.

Cerebras ended its latest quarter with about $3.3 billion in cash and investments before raising roughly $6.4 billion through its initial public offering. It has not completed the 750-megawatt build, but it now has both the contract and the capital required to attempt it.

Groq’s strength is its infrastructure already in operation. It runs 13 data centers across North America, Europe, the Middle East and Asia-Pacific and plans to move toward 200 megawatts by the end of 2027.

Its customer relationships cover several valuable markets. Saudi Arabia announced a $1.5 billion commitment connected to Groq infrastructure. Bell Canada chose Groq as the exclusive inference provider for its sovereign AI network, and Meta incorporated Groq into the official Llama API.

The Saudi commitment is not the same as recognized Groq revenue, and neither the Bell nor Meta relationship includes a public contract value. Still, no other private inference-chip startup has documented a comparable combination of operating locations and strategic distribution.

SambaNova recently added JPMorgan Chase as an on-premise RDU customer. SoftBank plans to use its next-generation SN50 platform in Japan, while sovereign infrastructure providers in Australia, Europe and the United Kingdom have selected SambaNova systems.

SambaNova also has a practical deployment advantage. Its systems are designed to operate in existing enterprise and sovereign facilities, often with air cooling and lower rack power. One British deployment described racks using around 10 kilowatts, far below the power requirements associated with the densest new GPU systems.

Etched’s reported $1 billion contract book is important but remains less proven. The company has not named the buyers or disclosed how much depends on successful product validation. Its next challenge is manufacturing reliable chips and complete racks at the scale those orders require.

If you want more recent data on this point, please see our latest AI chip market report.

Chart showing annual VC investment in AI chip startups

This chart, featured in our AI chip market deck, shows annual VC investment in AI chip startups

Which AI inference chip company is growing fastest now?

Cerebras is growing fastest in verified revenue, Groq in visible usage, and Etched in investor and customer expectations.

Cerebras increased annual revenue by 76% in 2025 and almost doubled its first-quarter sales. Its cloud business grew even faster, and the midpoint of its current forecast would add about $350 million of revenue in one year.

Groq’s developer count rose from slightly more than two million to more than five million in less than a year. Its infrastructure now spans 13 data centers across four regions. The company also raised another $650 million recently, after securing $750 million in its previous major round.

That usage growth is impressive, although a registered developer is much easier to acquire than a large paying customer. Groq does not disclose how many of its five million developers pay, how much capacity they consume or whether larger users stay after their initial experiments.

SambaNova has experienced the sharpest change in investor confidence. It raised more than $350 million when announcing SN50, then completed the first $1 billion close of another round shortly afterward. Its valuation reached $11 billion, supported by JPMorgan Chase, SoftBank and several sovereign deployments.

Etched has moved even faster from obscurity to serious contender. It emerged with working silicon, $800 million in financing and more than $1 billion in contracts. The Wall Street Journal has since reported discussions around financings that could value the company as high as $20 billion, although those talks had not closed.

Most of Etched’s growth currently exists in capital raised and future orders. Cerebras still has the strongest expansion visible in delivered revenue.

Who has the most mature product and strongest route to customers?

Cerebras and Groq currently have the most mature AI inference products, while Groq leads with developers and SambaNova stands out in sovereign and on-premise deployments.

Cerebras sells complete CS-3 systems, operates a public inference cloud and provides private infrastructure for large customers. Its API follows familiar OpenAI conventions, and enterprises can use custom model weights, reserved capacity and dedicated endpoints.

Groq has taken the low-friction approach further with developers. More than five million accounts now have access to GroqCloud, and the platform supports popular models, speech recognition, tool use, web search and code execution. A developer can usually test Groq by changing a few lines of an existing application.

The five-million figure is not a paying-customer count. It still gives Groq a valuable lead in awareness, integrations and developer familiarity. Its international data centers also allow governments and telecom groups to keep workloads inside their own regions.

SambaNova’s existing SN40L platform is available through cloud and on-premise systems. Its new SN50 generation is more complicated. Early benchmark systems are running, and SoftBank has agreed to deploy the platform, but broad customer shipments are scheduled for later in 2026.

SambaNova has the clearest enterprise on-premise story. JPMorgan Chase selected its RDUs for secure internal inference, while sovereign providers in Australia, Europe and the United Kingdom are building services around its systems.

d-Matrix recently moved Corsair into full production and announced a commercial deployment with inference provider Parasail. Its approach allows customers to add decoding accelerators beside Nvidia GPUs rather than replacing their existing infrastructure.

FuriosaAI has begun mass production of RNGD and is working with Samsung SDS, Equinix and several Korean enterprise groups. Rebellions now offers cards, servers and rack-scale products with native support for PyTorch and vLLM.

Positron says Atlas is shipping today. Etched has produced working first-pass silicon and is validating full systems with customers. Taalas has demonstrated a working chip and public inference service, but its first processor is tied to a specific model.

Company Product position now Strongest maturity evidence Main remaining question
Cerebras Fully commercial Systems, cloud revenue and large capacity contracts Can it build capacity quickly enough for its backlog?
Groq Fully commercial 13 data centers and more than five million developers Can it maintain its chip roadmap after losing key leaders to Nvidia?
SambaNova Commercial, with a new generation arriving Existing systems, JPMorgan Chase and early SN50 deployments How quickly will SN50 reach broad production?
d-Matrix Entering commercial-scale production Corsair production and Parasail deployment Which large customers will adopt it repeatedly?
FuriosaAI Mass production RNGD shipments and named cloud and enterprise partners How many chips are actually being deployed?
Rebellions Commercial products available Cards, servers and rack-scale systems Can it expand beyond the Korean ecosystem?
Positron Early commercial shipping Atlas systems available today Can it scale before its next-generation chip arrives?
Etched Customer validation Working silicon and signed orders Will its production ramp match the size of its order book?
Taalas Working specialized silicon Public model-specific inference Will customers accept such limited model flexibility?
Chart showing how Nvidia is leading in the AI chip market

This chart, featured in our AI chip market deck, shows how Nvidia is leading in AI chips

Which AI inference chip is actually fastest?

Cerebras is currently the fastest broadly accessible platform in the cleanest public comparison, although the winner changes with the model and workload.

Artificial Analysis recently tested GPT-OSS-120B across 22 providers. Cerebras produced about 1,955 output tokens per second, compared with 709 for SambaNova and 478 for Groq. Its first answer token also arrived in 1.52 seconds, versus 3.78 seconds for SambaNova and 4.90 seconds for Groq.

These gaps are large enough for users to notice. Cerebras generated tokens almost three times faster than SambaNova and more than four times faster than Groq. For coding assistants, voice applications and research agents, that can turn a long pause into something closer to a normal conversation.

SambaNova has since produced another important result with its SN50 platform. Artificial Analysis measured a hybrid system using Nvidia H200 GPUs for prefill and SambaNova chips for decoding as the fastest MiniMax M2.7 endpoint. The result shows that SambaNova can lead on some newer workloads, although it does not represent a standalone SN50 comparison.

d-Matrix recently announced a Parasail deployment that combines Nvidia GPUs with Corsair accelerators. The companies say the system can generate tokens up to ten times faster on selected workloads by letting GPUs handle prefill and d-Matrix handle the latency-sensitive decoding phase. Earlier testing by Gimlet Labs reduced one response from 24 seconds to less than two.

Taalas reports around 17,000 output tokens per second for Llama 3.1 8B. Its chip hardwires that specific model into silicon, so the result cannot be compared fairly with a flexible service running a 120-billion-parameter model.

Etched claims that its Sohu system can deliver roughly 20 times the throughput of an Nvidia H100 on transformer workloads. The company has working silicon, but independent customer-scale benchmarks are still limited.

We place the most confidence in tests that run the same model through accessible services. Vendor claims using different models, quantization levels, batch sizes or latency targets cannot support a clean ranking by themselves.

Provider GPT-OSS-120B output speed First answer token Position in this comparison
Cerebras About 1,955 tokens per second 1.52 seconds Clear leader
SambaNova About 709 tokens per second 3.78 seconds Roughly 36% of Cerebras’s speed
Groq About 478 tokens per second 4.90 seconds Roughly 24% of Cerebras’s speed

If you want more recent data on this point, please see our latest AI chip market report.

Which startup offers the best inference economics?

Groq currently offers the cheaper leading startup API, while Cerebras asks customers to pay more for much faster responses.

On current public pricing for GPT-OSS-120B, Groq charges approximately $0.15 per million input tokens and $0.60 per million output tokens. Cerebras charges about $0.35 and $0.75 respectively.

Using a workload containing three input tokens for every output token, Groq costs roughly $0.26 per million blended tokens. Cerebras costs around $0.45. Groq is therefore about 40% cheaper under that simple mix.

Cerebras delivers more than four times Groq’s output speed on the same Artificial Analysis test. Paying an extra $0.19 per million tokens could be an easy decision for an interactive coding agent, voice assistant or customer-service application where every second affects usage.

A background summarization system may reach the opposite conclusion. It can tolerate slower generation and send work to whichever provider offers the lowest token price. Several conventional GPU clouds already serve GPT-OSS-120B more cheaply than either Groq or Cerebras.

On-premise comparisons remain harder. Positron says Atlas provides more than four times the performance per watt and more than three times the performance per dollar of an Nvidia H200 system on its chosen Llama workload. d-Matrix promotes similar savings from pairing Corsair with existing GPUs, and FuriosaAI emphasizes air-cooled servers with lower power demand.

Those claims could become decisive for data centers that cannot obtain more electricity. Buyers still need public hardware prices, utilization rates and independent total-cost measurements before one company can claim the economics lead across every deployment.

Chart showing the projected CAGR of the AI chip market

This chart, featured in our AI chip market deck, shows annual funding in AI chip startups

What can Nvidia copy or crush?

Cerebras has built the hardest architecture to copy, but Nvidia can still squeeze every startup through its software, distribution and ability to absorb useful technology.

Cerebras’s WSE-3 places approximately four trillion transistors and 900,000 computing cores on one wafer-scale processor. Replicating it would require expertise in chip design, manufacturing yields, packaging, cooling, compilers and distributed systems.

Its advantage now extends beyond the chip. Cerebras operates a cloud, sells complete systems, supports custom models and is building large amounts of dedicated capacity. A rival would have to reproduce that whole stack.

Groq’s recent history shows how Nvidia can respond to a serious threat. Nvidia licensed Groq’s inference technology, while Groq founder Jonathan Ross, president Sunny Madra and other engineers joined Nvidia. Groq remains independent and continues to operate its cloud, but much of the team that built its original advantage now works for the incumbent.

The agreement validates Groq’s technology while making its future hardware differentiation harder to defend. Nvidia can combine ideas from the LPU with its own software ecosystem, networking, customer base and manufacturing scale.

SambaNova has retained more strategic independence. Its reconfigurable dataflow architecture, three-level memory system and complete software platform are difficult to reproduce quickly. Collaboration with Intel gives it manufacturing and distribution support without handing the whole architecture to Nvidia.

Etched and Taalas make a more concentrated bet. They remove general-purpose features and hardwire a transformer architecture or individual model into silicon. That can produce enormous speed and efficiency, provided customers keep using compatible models.

The risk is the specialization itself. A major shift in model architecture could reduce the value of chips already manufactured. Etched says Sohu supports several transformer-like designs, but its performance advantage still depends on a narrower range of workloads than a GPU.

d-Matrix may have found a safer route by working beside Nvidia GPUs. Customers can use GPUs for the computation-heavy prefill phase and Corsair for faster decoding without abandoning CUDA.

Capturing the latency-sensitive, power-constrained or sovereign parts of inference would already give these startups a large market. None needs to replace Nvidia across training, networking and general-purpose computing.

If you want more recent data on this point, please see our latest AI chip market report.

Is the best-funded startup using its money well?

Cerebras has turned capital into products and revenue most convincingly, while SambaNova still needs to show what its $2.5 billion has produced financially.

Cerebras raised approximately $2.9 billion privately before going public. As noted earlier, it generated $510 million of revenue in 2025. Its annual sales were therefore equal to roughly 18% of all the private capital it had accumulated.

That ratio is not a profit margin and ignores the timing of each financing. It still gives us a rough view of how much commercial activity emerged from the money invested.

Groq has raised approximately $2.4 billion across its major rounds. The result is a large global cloud, millions of developers and a strategically valuable architecture. Without audited revenue, we cannot tell whether it has converted capital more efficiently than Cerebras.

SambaNova has raised about $2.5 billion after its latest financing. It has produced several generations of chips, enterprise systems and a growing cloud but does not disclose annual revenue, gross margin or backlog.

Several smaller companies have reached difficult milestones with much less funding. FuriosaAI brought RNGD into mass production after raising $246 million. d-Matrix reached full production with $450 million, and Positron began shipping Atlas before closing its $230 million Series B.

Taalas created a working model-specific chip with $219 million. Its approach may be too narrow for many customers, although the speed of its development deserves more attention than its position near the bottom of the funding table suggests.

Etched has raised $800 million before proving large-scale production. Working first silicon and signed contracts support the investment case, but the capital will only look efficient once customers are running dependable systems.

Chart comparing business model options for AI accelerator chip companies

This chart, featured in our AI chip market deck, compares the main business model options for AI accelerator chip companies

Who has the strongest momentum lately?

Cerebras still has the strongest overall momentum, but SambaNova, d-Matrix and Etched have made the biggest recent jumps.

Cerebras has added several reinforcing advantages within a short period. Revenue almost doubled in its latest quarter, AWS agreed to distribute its inference capacity, OpenAI committed to a major buildout, and the company completed an unusually large semiconductor IPO.

The IPO finances capacity, the capacity supports the OpenAI agreement, and AWS gives more customers access to the resulting infrastructure. Cerebras is turning financial, commercial and operating momentum into one connected expansion.

SambaNova has become the fastest-rising private challenger. It introduced SN50, raised $350 million, added JPMorgan Chase and then completed the first $1 billion close of another round. Its recent MiniMax result also gives the new hardware more substance than a launch presentation alone would provide.

d-Matrix has crossed a practical threshold. Corsair entered full production, and Parasail is now deploying the accelerator alongside Nvidia infrastructure. That customer installation tests whether heterogeneous inference can work inside a commercial cloud.

Etched emerged with working silicon and a large contract book. The company has since attracted discussions around much higher valuations, although expectations have risen faster than delivered systems.

FuriosaAI has quietly strengthened its position. RNGD is in mass production, Samsung SDS is offering it through a domestic NPU service, and Broadcom will help develop Furiosa’s next multi-die accelerator.

Rebellions recently raised $400 million before a planned public offering and launched production-scale rack products. Its momentum is strongest inside Korea, where financing, telecom customers and semiconductor partners support one another.

Groq presents the strangest picture. It added three million developers, expanded to 13 data centers and raised another $650 million. At the same time, Nvidia gained access to its technology and hired several of its most important leaders. GroqCloud is moving forward quickly, while Groq’s independent chip position looks less secure.

Which AI inference chip startup is actually ahead?

Cerebras is the clear overall leader today; among private companies, Groq remains second and SambaNova is closing the gap.

Cerebras leads on the evidence we trust most: audited revenue, revenue growth, contracted demand, independently measured speed, commercial availability and financial capacity. None of its competitors currently combines all six.

Its position still has weaknesses. The OpenAI agreement creates customer concentration, the required infrastructure build is enormous, and Nvidia has enough capital to respond aggressively. Cerebras must now prove that its backlog can become delivered capacity without damaging margins.

Groq keeps second place because its cloud already operates at a scale no other private rival has matched. Five million developers, 13 data centers and major sovereign relationships represent real distribution. The Nvidia licensing deal prevents us from treating its technology and team as strongly independent as before.

SambaNova ranks third and is now the most likely company to overtake Groq. It has kept control of its architecture, raised enough money to fund SN50, won important enterprise customers and shown competitive performance. More transparent revenue and broad SN50 shipments would move it higher.

d-Matrix takes fourth because it has reached full production and found a practical way into existing Nvidia data centers. Its first commercial deployment could become a repeatable model if customers get faster inference without replacing their GPU fleets.

FuriosaAI follows with mass-produced hardware, major Korean customers and a next-generation partnership with Broadcom. Rebellions has deeper funding and strong national support, but FuriosaAI currently shows slightly better evidence of chips entering real cloud and enterprise services.

Etched has more upside than either company. Its $1 billion contract book and working silicon justify the attention, but validation and manufacturing remain unfinished. We rank what customers can run today above what they have ordered for tomorrow.

Positron already ships an efficient product and has raised enough money for a second generation. Its current scale remains modest. Taalas offers the most radical architecture and some remarkable performance, but hardwiring individual models sharply limits its addressable market.

SambaNova could pass Groq by showing strong revenue and large SN50 deployments. d-Matrix could enter the top three with several hyperscale customers. Etched could jump into the leading group once customers confirm that its contract book is becoming dependable production infrastructure.

Rank Company Why it ranks here today
1 Cerebras The clear leader in audited revenue, growth, major contracts, public inference speed, product maturity and available capital; now public rather than technically a startup
2 Groq The largest private operating footprint, with five million developers, 13 data centers and strong sovereign distribution, although the Nvidia agreement weakened its independence
3 SambaNova The strongest intact private challenger, backed by fifth-generation silicon, major new financing and credible enterprise customers
4 d-Matrix Full production, a real commercial deployment and a practical strategy that improves existing Nvidia infrastructure
5 FuriosaAI Mass-produced hardware, strong Korean customers, European expansion and a credible next-generation partnership with Broadcom
6 Rebellions Deep financial and industrial support, complete rack-scale products and a strong home market, with less proof of international adoption
7 Etched Working silicon and unusually strong contracted demand, held back by unfinished customer validation and manufacturing risk
8 Positron A shipping, energy-efficient system with credible financing, but a much smaller installed footprint
9 Taalas Extraordinary model-specific performance, limited today by narrow model support and little evidence of large commercial deployments

If you want more recent data on this point, please see our latest AI chip market report.

Chart showing how revenue is split across customer segments in the AI chip market

This chart, featured in our AI chip market deck, shows how revenue is split across customer segments in the AI chip market

OUR METHODOLOGY

This analysis asks which AI inference chip company is ahead based on the evidence available today. We compare commercial traction, customer quality, deployment capacity, product maturity, measured performance, economics, capital efficiency and recent momentum.

We gave the most weight to audited financial results, independent same-model benchmarks, products already available, named deployments and infrastructure operating today. Funding, bookings, vendor benchmarks and future commitments were included, but they carried less weight when delivery or revenue remained unproven.

Cerebras is included as a startup-origin benchmark even though it recently completed an initial public offering. The private-company ranking is therefore assessed separately, with Groq and SambaNova compared most closely on operating footprint, customers, product maturity and evidence of commercial scale.

Funding totals are approximate because companies do not consistently separate equity, debt, grants, secondary transactions and strategic investments. Contract values and remaining obligations are not treated as revenue already earned, and registered developer counts are not treated as paying-customer counts.

Performance comparisons prioritize accessible services running the same model. The GPT-OSS-120B results from Artificial Analysis provide the cleanest direct comparison among Cerebras, SambaNova and Groq. Vendor claims based on different models, batch sizes, quantization levels or hybrid systems are used as supporting evidence rather than as a universal speed ranking.

The API economics comparison uses public GPT-OSS-120B prices and a simple workload containing three input tokens for every output token. On-premise economics remain less certain because public hardware prices, utilization rates and independently measured total cost of ownership are still limited.

Key sources used for this analysis include: Cerebras Q1 2026 financial results, Cerebras’s Q1 2026 Form 10-Q, the Cerebras IPO registration statement, the OpenAI and Cerebras agreement, the AWS and Cerebras collaboration, Artificial Analysis’s GPT-OSS-120B provider benchmark, and Cerebras inference pricing.

Additional company evidence comes from Groq’s June 2026 financing and infrastructure update, Groq’s GPT-OSS-120B documentation, SambaNova’s $1 billion financing close, the SambaNova SN50 announcement, SambaNova’s SN50 MiniMax benchmark, d-Matrix’s Corsair production announcement, the d-Matrix and Parasail deployment, FuriosaAI’s Samsung SDS launch, the FuriosaAI and Broadcom partnership, Rebellions’ $400 million pre-IPO round, and Taalas HC1 product documentation.

Chart showing how AI accelerator chip technology has evolved over time

This chart, featured in our AI chip market deck, shows how AI accelerator chip technology has evolved over time

Who is the author of this content?

NEW MARKET PITCH TEAM

We track new markets so founders and investors can move faster

We build living "market pitch" documents for emerging markets: AI, synthetic biology, new proteins, and more. Instead of outdated PDFs or hallucinated LLM answers, our clients get a clean, visual, always-updated view of what's really happening: key players, deals, regulations, and signals that matter. Learn more about us.

Back to blog