GPU efficiency: which startup is ahead?

Last updated: 23 July 2026
market research pitch 2026 statistics AI infrastructure market

In our AI infrastructure market deck, you will find everything you need to understand the market

SUMMARY

Cast AI is the GPU-efficiency startup ahead overall today, although ScaleOps is closing quickly and Decart leads the separate race to make AI workloads run efficiently across different chip families.

The market does not have one universal performance winner because the companies attack different layers of waste. Cast AI and ScaleOps manage infrastructure, FriendliAI optimizes inference, Decart works closer to the runtime, Zymtrace finds code bottlenecks, and Rapt.AI focuses on model-aware GPU allocation.

Cast AI’s lead rests on breadth rather than one spectacular benchmark. It has the largest disclosed user footprint, a mature Kubernetes platform, major enterprise references and a product that now spans GPU sharing, workload scaling and cross-cloud capacity.

ScaleOps is the most credible threat to that lead. Its enterprise customer list is unusually strong, its reported growth is far faster than the rest of the field, and its workload-level GPU controls look more specialized than Cast AI’s broader infrastructure proposition.

Decart is the most strategically important outlier. It has raised more than $450 million and attracted both Amazon and Nvidia, but its public commercial evidence is concentrated around a few major relationships rather than a broad, visible customer base.

The biggest GPU-utilization claims are also the least comparable. Rapt.AI says it can push true utilization toward 90% to 98%, while Cast AI and ScaleOps show how low ordinary cluster utilization can be, but none has published a neutral, workload-matched comparison that settles the argument.

FriendliAI has the best-supported case in inference efficiency. Its product connects defined benchmarks, production deployment options and named customers more convincingly than the larger but less transparent claims from Decart or Rapt.AI.

Cross-chip portability may become the most valuable technical moat as companies try to reduce dependence on Nvidia. Decart currently leads that layer through support for Nvidia GPUs, Amazon Trainium and Google TPUs, while FlexAI offers the broader commercial deployment model across Nvidia and AMD.

The acquisition pattern confirms that this software layer is strategically valuable. Nvidia, Red Hat and Qualcomm have all moved to acquire companies that sit between AI applications and expensive compute, shrinking the independent field while raising the value of the remaining specialists.

The ranking can still change quickly. ScaleOps can pass Cast AI by sustaining its growth and disclosing deployment scale; Decart can move to first by turning DOS into a repeatable product across several hyperscalers; and FriendliAI can enter the top tier if it reveals meaningful volume, revenue or fleet data.

Market map chart showing top companies and startups in the AI infrastructure market

This market map, featured in our AI infrastructure market deck, highlights top companies and startups in the AI infrastructure market

GPU efficiency: which startup is actually ahead?

Which GPU efficiency startups belong in this comparison?

We currently see seven independent startups with enough product and market evidence to compare seriously: Cast AI, ScaleOps, Decart, FriendliAI, FlexAI, Zymtrace and Rapt.AI.

The category includes software that helps companies get more useful work from GPUs they already own, rent or manage. That covers cluster allocation, GPU sharing, workload scheduling, inference optimization, cross-chip portability and code-level performance analysis.

The boundaries need to stay fairly strict. We exclude chipmakers such as Groq, Cerebras and Etched because they sell alternative hardware. We also exclude GPU clouds and managed inference providers when customers mainly buy compute from them rather than optimization software they can use across their own infrastructure.

Several former leaders have already disappeared into larger companies. Nvidia acquired Run:ai, Deci, CentML and SchedMD, while Red Hat acquired Neural Magic. Qualcomm has agreed to acquire Modular, so we use Modular as an important technical reference but leave it out of the independent-startup ranking.

The remaining field still mixes several product types. Cast AI and ScaleOps manage infrastructure. Decart optimizes execution across different chips. FriendliAI specializes in model serving. FlexAI combines managed compute with infrastructure software. Zymtrace finds waste inside applications, while Rapt.AI dynamically allocates GPUs around the needs of each model.

Startup Main GPU-efficiency product Publicly reported funding Included because
Cast AI Automated Kubernetes, GPU allocation and cross-cloud capacity management More than $180M Broad enterprise infrastructure platform
ScaleOps Autonomous Kubernetes and fractional GPU management More than $210M Closest direct challenger to Cast AI
Decart Training and inference optimization across several chip families More than $450M Strongest independent cross-chip contender
FriendliAI High-throughput inference engine, APIs and private deployment software About $26.7M Mature specialist in inference efficiency
FlexAI Managed inference, training and private AI infrastructure About $30M Hardware-flexible infrastructure platform
Zymtrace Whole-system CPU and GPU performance analysis $12.2M Promising code-level optimization specialist
Rapt.AI Model-aware GPU scheduling and workload packing Undisclosed Focused pure-play GPU utilization startup

Is there a clear GPU efficiency leader today?

Cast AI is ahead overall today, with ScaleOps close behind and Decart leading a separate cross-chip race.

Cast AI has the broadest disclosed commercial footprint. Its Series C announcement said more than 2,100 organizations used the platform, including Akamai, BMW, FICO, Hugging Face, NielsenIQ and Swisscom. The company has since expanded from general Kubernetes optimization into GPU sharing, workload scaling and cross-cloud capacity through OMNI Compute.

ScaleOps sits one tier behind on current scale, although the gap looks much smaller on momentum. Its disclosed customers include Adobe, Wiz, DocuSign, Salesforce and Coupa, and its latest financing valued the company above $800 million. ScaleOps has yet to publish a customer total comparable with Cast AI’s, so we cannot measure the distance precisely.

Decart has gone further than either company at the execution layer. Its DOS software works across Nvidia GPUs, Amazon Trainium and Google TPUs, while Amazon has deployed Decart technology across several business units. The commercial evidence remains concentrated around a small number of strategic relationships, which keeps Decart below the two infrastructure platforms in our overall ranking.

The market structure is fairly clear. Cast AI leads through breadth and maturity. ScaleOps is the only direct rival close enough to threaten that position soon. Decart has the strongest chance of changing how the category works, but its wider customer base is still largely hidden.

If you want more recent data on this point, please see our latest AI infrastructure market report.

Google Trends chart showing rising interest in AI infrastructure

As this chart shows, and as featured in our AI infrastructure market deck, search interest in AI infrastructure has risen sharply

Which startup has the strongest customers and commercial traction?

Cast AI has the strongest commercial proof, while Decart has landed the most valuable single customer relationship.

Cast AI’s advantage comes from the combination of customer count and customer variety. Its named users span automotive, telecommunications, financial software, AI development and internet infrastructure. That range gives us more confidence that the product works across different workload patterns rather than one narrow technical setup.

ScaleOps has fewer disclosed data points, but the quality of its references is excellent. Adobe, Salesforce, DocuSign, Wiz and Coupa all operate software products where infrastructure failures immediately affect customers. ScaleOps says its platform makes live resource decisions inside production environments, creating a much deeper relationship than a dashboard or occasional consulting project.

Decart’s Amazon deployment carries more strategic weight than a normal startup pilot. The Wall Street Journal reported that Amazon is Decart’s largest customer and uses the technology across Twitch, retail and entertainment operations. One customer spread across several large business units can reveal more product depth than a long list of small trials. Still, one customer is one customer.

FriendliAI leads the smaller inference specialists. It has reported roughly 25 to 30 large clients, with public examples including SK Telecom, Upstage and LG AI Research. Its recent Samsung Cloud Platform alliance around Nvidia B300 infrastructure also gives it a stronger route into enterprise deployments.

FlexAI has started publishing useful customer evidence as well. LegML reported training a 32-billion-parameter legal model for about €22,500, roughly 75% below its comparison cost. DragonLLM reported more than 99.9% uptime on sovereign infrastructure, while Pixelcut used FlexAI for pay-per-use image-model fine-tuning. These are real deployments, although they still form a much smaller public portfolio than Cast AI’s.

Commercially, Cast AI remains comfortably first. ScaleOps has the next-best enterprise footprint, Decart owns the deepest strategic account, and FriendliAI has built the strongest specialist customer base.

Which GPU efficiency startup is growing fastest now?

ScaleOps is growing fastest on the clearest recent operating evidence.

TechCrunch reported that ScaleOps said its business had grown by more than 450% year over year and that headcount had tripled over twelve months. A 450% increase means the underlying metric reached roughly 5.5 times its previous level. Even allowing for the usual startup optimism, that pace stands well above the disclosed growth of the other contenders.

There is a small reporting inconsistency worth noticing. The Next Web described ScaleOps’ growth as more than 350% on the same funding announcement, while TechCrunch used more than 450%. Neither report clearly identified whether the figure referred to revenue, annual recurring revenue, bookings or another internal metric. We therefore treat ScaleOps as the obvious hypergrowth company without pretending that 450% is an audited revenue figure.

Its financing trajectory supports the broader conclusion. ScaleOps raised a $58 million Series B, followed about sixteen months later by a $130 million Series C. Total funding reached more than $210 million, and the company included a sizeable employee secondary sale in the latest round. Investors usually reserve that structure for a company that has moved beyond early product validation.

Decart has the fastest strategic acceleration. It raised $153 million at a $3.1 billion valuation, then added $300 million at nearly $4 billion while bringing Nvidia into the investor group and Amazon deeper into the business. That tells us powerful companies value Decart’s technology, although it reveals little about recurring revenue growth.

FriendliAI previously projected revenue growth of as much as 600%. We have yet to see a later disclosure confirming that outcome, so it carries less weight than ScaleOps’ completed customer, hiring and financing expansion.

Cast AI remains larger, while ScaleOps is closing ground faster. Decart could eventually grow past both, but its public operating numbers currently lag far behind its fundraising story.

If you want more recent data on this point, please see our latest AI infrastructure market report.

Chart showing annual VC investment in AI infrastructure startups

This chart, included in our AI infrastructure market deck, shows annual VC investment in AI infrastructure startups

Which GPU efficiency products are genuinely ready for large customers?

Cast AI and ScaleOps have the most mature enterprise products, while FriendliAI is the strongest production-ready inference specialist.

Cast AI already handles automated rightsizing, GPU sharing, workload scaling, cost visibility and access to capacity across cloud providers. OMNI Compute adds another useful layer by letting customers find GPUs outside their main cloud environment, including Oracle capacity, without rebuilding the whole application stack.

ScaleOps also operates continuously inside live Kubernetes environments. Its AI platform measures GPU demand at the individual workload level, shares devices through fractional allocation, optimizes GPU memory and adjusts inference replicas. The platform is self-hosted, which suits companies that need tighter control over security and data.

FriendliAI offers three mature deployment paths: shared model APIs, dedicated GPU endpoints and containers that run inside a customer’s own environment. Its current platform advertises a 99.99% uptime service-level agreement and access to hundreds of thousands of open and custom models. That is a much fuller production product than a benchmark repository or bespoke optimization project.

FlexAI has also moved past the prototype stage. Customers can buy token-priced inference, dedicated endpoints or an AI Factory deployment inside their own cloud, data center or air-gapped environment. The company now covers both Nvidia and AMD hardware, although its public evidence still comes mostly from smaller AI companies.

Decart has proved that DOS can run inside a hyperscaler, yet the product remains oriented toward strategic customers and pilots. Zymtrace has published a strong customer result and is expanding enterprise deployments, but its recent seed round shows how early the company remains. Rapt.AI has working software and a neocloud integration, though its public rollout evidence is still limited.

Startup Product maturity today Ability to support large deployments Main limitation
Cast AI Mature production platform Strong across clouds and industries GPU-specific adoption remains undisclosed
ScaleOps Mature production platform Strong in Kubernetes environments No public total for customers or managed GPUs
FriendliAI Mature inference platform Strong across API, dedicated and private deployments Smaller disclosed customer base
FlexAI Commercial managed and private platform Growing, with several published case studies Limited evidence from very large fleets
Decart Production use with strategic customers Proven inside Amazon Wider repeatability remains unclear
Zymtrace Early production product Promising customer results Seed-stage company with few public deployments
Rapt.AI Early commercial product Neocloud integration and pilots Most results remain company-reported

Which startup improves GPU utilization the most?

Rapt.AI makes the biggest utilization claim, but nobody has proved the category-wide performance lead.

Rapt.AI says customers can move from roughly 20% to 35% true utilization into the 90% to 98% range. It also claims three to five times more parallel jobs and as much as ten times higher inference throughput on the same hardware. Those numbers would put Rapt far ahead, but they come from the company’s own pilots and production rollouts without a published workload-by-workload benchmark.

Cast AI supplies the strongest evidence about the size of the underlying problem. Its latest Kubernetes optimization report examined tens of thousands of clusters across AWS, Azure and Google Cloud and found average GPU utilization around 5% in the non-optimized environments it observed. Google Cloud clusters reached about 6%, AWS about 5% and Azure about 2%.

That dataset is valuable because of its scale, but it measures the market before Cast AI optimization. It shows an enormous amount of waste. It does not show how far the same clusters improved after adopting the product.

ScaleOps takes a more targeted approach. Its platform measures consumption per pod, then changes GPU shares, memory allocation and replica counts around actual demand. The company says ordinary production-inference workloads often use only 5% to 20% of available capacity, with even well-managed Kubernetes clusters struggling to pass 20% to 30%.

An independent Gartner playbook published earlier this year reached a similar broad conclusion, describing many enterprise GPU clusters as operating below 15% to 20% efficiency. That agreement across several sources makes widespread underutilization credible, even though each company measures it differently.

Rapt.AI currently owns the highest headline result. Cast AI owns the largest observational dataset, while ScaleOps has the most convincing workload-level mechanism. We need comparable before-and-after results from the same models and hardware before naming a proven utilization winner.

Chart showing why CoreWeave is winning in the AI infrastructure market

This chart, included in our AI infrastructure market deck, shows why CoreWeave is winning in AI infrastructure

Which startup makes AI inference fastest and cheapest?

FriendliAI is currently the best-supported inference-efficiency leader among independent startups.

FriendliAI’s current platform combines custom GPU kernels, caching, continuous batching, speculative decoding and parallel inference. The company advertises more than twice the performance of common serving alternatives, with some model-specific tests reaching two to five times faster output speed and 50% to 90% lower GPU costs.

The benchmark details are better than the usual homepage claim. One current comparison uses GLM-5 on four Nvidia B200 GPUs with 10,000 input tokens and 500 output tokens, measured against vLLM and SGLang. The result still comes from FriendliAI, but at least buyers can see the model, hardware and request shape behind the number.

Its commercial packaging strengthens the case. The same engine powers shared APIs, dedicated endpoints and customer-hosted containers. FriendliAI recently launched InferenceSense, which lets GPU cloud operators fill spare capacity with paid inference traffic while preserving priority for their main workloads. That directly links utilization with revenue rather than treating efficiency as a laboratory exercise.

Zymtrace has published the clearest individual customer improvement. Anam said Zymtrace helped its Cara3 model achieve 2.5 times lower inference latency and 90% greater throughput without replacing the hardware. The result comes from one workload, but it connects a named product with a measurable operational gain.

Decart claims roughly 1,600 tokens per second for agentic inference, compared with an industry reference near 200. An eightfold gap sounds decisive until we ask about the model, chip, batch size, context length, latency target and precision. The public material provides too little detail for a fair comparison with FriendliAI.

Rapt.AI claims throughput gains of up to ten times, while FlexAI’s strongest economics evidence comes from complete customer projects. LegML’s reported 75% lower compute cost is impressive, though part of the saving came from model design and training choices rather than GPU orchestration alone.

FriendliAI wins this subcategory because its performance story connects technical methods, defined benchmarks, production deployment options and real customers. Zymtrace has the strongest early customer proof. Decart and Rapt may deliver larger gains on selected workloads, but their public comparisons remain too thin to support the overall lead.

If you want more recent data on this point, please see our latest AI infrastructure market report.

Which startup leads across Nvidia, AMD and alternative AI chips?

Decart is ahead on cross-chip optimization now, with FlexAI the closest independent challenger.

DOS targets Nvidia GPUs, Amazon Trainium and Google TPUs. That goes deeper than choosing the cheapest available cloud instance. Decart adjusts how training and inference workloads execute on each hardware architecture, reducing the engineering work required when a customer changes chip families.

Amazon’s deployment gives this claim unusual credibility. Amazon has every reason to make Trainium easier to use, while Nvidia benefits when software extracts more work from its GPUs. Their willingness to support the same startup suggests that DOS solves a real portability problem rather than merely promising vendor independence.

FlexAI supports Nvidia and AMD hardware across managed and private deployments. Its Workload Co-Pilot evaluates models, token volumes, latency requirements and request rates before recommending a deployment configuration. FlexAI has also published an open benchmarking framework and lets customers place its stack inside their own cloud or data center.

Cast AI and ScaleOps offer hardware flexibility at the orchestration layer. They can place, share and scale workloads across different resources, but the model still needs a compatible runtime for each chip. Their strength lies in allocation rather than deep execution portability.

Modular would have ranked beside Decart because its MAX platform and Mojo language were designed to separate AI software from the underlying hardware. Qualcomm’s agreement to acquire the company removes it from the independent field and reinforces the strategic value of this software layer.

Decart has the clearest technical lead in cross-chip execution. FlexAI offers the broader commercial delivery model, while Cast AI and ScaleOps remain stronger higher up the infrastructure stack.

Chart showing the projected CAGR of the AI infrastructure market

This chart, included in our AI infrastructure market deck, shows annual funding in AI infrastructure startups

Which GPU efficiency startup has the strongest moat?

Cast AI has the strongest commercial moat, while Decart has the hardest technical moat to copy.

Cast AI becomes harder to replace as customers let it make continuous infrastructure decisions. The platform observes workload history, selects resources, changes scaling policies and interacts with several cloud providers. Removing it can force a customer to rebuild operational processes, accept more manual work and relearn how its applications behave.

The company’s dataset adds another layer. Observing tens of thousands of clusters gives Cast AI a broad view of workload patterns, pricing and infrastructure waste. That information should improve recommendations over time, although outsiders cannot test how much of the product’s performance actually comes from proprietary data.

Decart’s moat sits closer to the hardware. Optimizing the same workload across CUDA, Trainium and TPU environments requires compiler knowledge, model expertise, custom kernels and close work with chip providers. A normal cloud-cost startup would struggle to recreate that stack quickly.

FriendliAI also owns meaningful technical assets. Its team helped develop continuous batching, now a standard technique in LLM serving, and the company says key parts of its iteration-batching technology are patented in the United States, South Korea and China. Its challenge comes from aggressive open-source projects and Nvidia’s own inference stack, which keep narrowing any static performance lead.

ScaleOps relies more heavily on operational depth than unique algorithms. Its autonomy, production integrations and workload-level decision engine can create strong switching costs, though several larger infrastructure companies could build similar features.

Recent acquisitions show where the durable value lies. Nvidia bought orchestration, inference and workload-management companies, Red Hat bought Neural Magic, and Qualcomm moved for Modular. The market keeps rewarding software that sits deeply between applications and expensive hardware.

Cast AI currently owns the strongest installed commercial position. Decart’s cross-chip engineering may prove more defensible over the long run, especially as companies spread workloads across Nvidia and alternative accelerators.

If you want more recent data on this point, please see our latest AI infrastructure market report.

Is the best-funded startup using its money most effectively?

Cast AI and ScaleOps are turning capital into visible adoption more convincingly than Decart.

Decart has raised more than $450 million, roughly two and a half times Cast AI’s disclosed total and more than twice ScaleOps’ funding. Its valuation and strategic backers confirm that the technology is highly prized. The public evidence still tells us very little about customer count, recurring revenue or how much of the company’s activity comes from DOS rather than its world models.

Cast AI has raised more than $180 million and disclosed thousands of customers, global offices, a broad product suite and relationships with several large enterprises. The company also reached a valuation above $1 billion after its latest strategic investment. Among the well-funded contenders, it has produced the clearest link between capital and market penetration.

ScaleOps has raised slightly more than Cast AI but remains below it in disclosed scale. The difference is shrinking quickly. Its enterprise customers, reported hypergrowth and rapid hiring suggest that the latest capital is supporting a business already expanding rather than financing a search for product-market fit.

FriendliAI looks efficient at a smaller scale. It built a full production inference stack, signed several major Korean technology customers and launched new commercial products on about $26.7 million of disclosed funding. The company has less room for expensive expansion, but the output per dollar appears strong.

FlexAI also started with about $30 million and now offers shared inference, dedicated endpoints, private infrastructure and published customer cases. Zymtrace reached a named 2.5-times latency improvement with $12.2 million in total funding. Both have achieved credible technical progress, although neither has yet shown broad enterprise penetration.

Funding changes the order less than the headline numbers suggest. Cast AI ranks first on conversion of capital into adoption. ScaleOps follows closely and may be improving faster. Decart has converted funding into strategic influence, while its broader commercial efficiency remains impossible to judge.

Chart comparing business model options for AI cloud infrastructure providers

This chart, included in our AI infrastructure market deck, compares the main business model options for AI cloud infrastructure providers

Is one startup ahead across the whole GPU efficiency market?

Cast AI leads the overall category, while every major technical subsegment has its own winner.

Cast AI ranks first because infrastructure allocation affects almost every GPU workload. A company can gain substantial savings before changing its model or inference engine simply by buying fewer idle resources, sharing devices and moving workloads to available capacity. Cast AI has also shown the broadest ability to sell that proposition repeatedly.

ScaleOps competes in the same control-plane layer and currently has the strongest chance of taking the lead. Its GPU product appears more deeply focused on workload-level demand, and its recent growth is much faster. Cast AI still has the larger visible footprint.

Decart leads the runtime-portability segment. Its product becomes especially valuable when companies want to reduce dependence on Nvidia or place different workloads on the chip that offers the best economics.

FriendliAI leads optimized inference serving. It helps customers generate more tokens from each GPU and packages the technology across cloud, dedicated and private environments. FlexAI follows with a wider infrastructure proposition that includes training, inference and private AI factories.

Zymtrace leads the emerging code-analysis niche. Its product finds expensive lines of code across an application, including third-party libraries and mixed CPU-GPU activity. That can uncover bottlenecks an infrastructure scheduler would never see.

Rapt.AI is the specialist in model-aware allocation. Its product has the most aggressive utilization targets, although the evidence base remains small.

The overall leader wins through breadth rather than domination of every technical metric. Cast AI touches the largest part of the customer problem today. Decart and FriendliAI may deliver bigger gains inside narrower workloads.

How much can we trust the startups’ GPU efficiency numbers?

We trust the customer and product evidence far more than the headline efficiency percentages.

Funding rounds, acquisitions, named customers and commercially available products are usually straightforward to verify. Cast AI’s customer announcement, ScaleOps’ financing, Qualcomm’s agreement to buy Modular and Amazon’s use of Decart all belong in this higher-confidence group.

Named customer outcomes come next. Anam’s 2.5-times latency gain with Zymtrace, LegML’s training cost with FlexAI and FriendliAI’s published enterprise case studies are useful because they tie the claim to a real workload. They still represent selected success stories rather than an average across all customers.

Large company-controlled datasets require more care. Cast AI’s analysis of tens of thousands of clusters provides strong evidence that GPU waste is widespread. The sample comes from environments the company can observe and focuses on clusters before optimization, so it cannot prove the savings every new customer will achieve.

Maximum performance claims sit at the bottom of the confidence ladder. Rapt.AI’s 98% utilization, Decart’s 1,600 tokens per second and FriendliAI’s 90% cost reduction may all occur under the right conditions. Buyers need the model, hardware, precision, batch size, context length, traffic pattern and latency target before comparing them.

Even ordinary growth figures can become slippery. ScaleOps’ latest announcement produced reports of both 350% and 450% year-over-year growth. The direction is obvious, but the precise number carries less authority when the underlying metric remains unnamed.

Our strongest conclusions rely on patterns rather than one spectacular claim. Cast AI repeatedly shows breadth. ScaleOps repeatedly shows acceleration. Decart repeatedly shows strategic technical value. FriendliAI repeatedly shows inference maturity. The evidence becomes weaker as we move further down the ranking.

Chart showing the share of revenue generated by each customer segment in the AI infrastructure market

This chart, featured in our AI infrastructure market deck, shows the share of revenue generated by each customer segment in the AI infrastructure market

Which GPU efficiency startups are actually ahead?

Cast AI is the best overall GPU-efficiency startup today, with ScaleOps close enough to overtake it and Decart capable of changing the market from a different angle.

Cast AI earns first place through the broadest commercial footprint, a mature product, major enterprise users and coverage across several sources of infrastructure waste. Its lead over ScaleOps is real but fairly narrow. We see a difference of one competitive tier rather than an order-of-magnitude gap.

ScaleOps ranks second because it combines strong customers with the fastest recent operating growth. Its main missing piece is a clear measure of deployed scale. A customer total, GPU count, managed workload count or audited revenue disclosure could quickly strengthen its case for first place.

Decart ranks third. It has raised more money than the top two combined, attracted Nvidia and Amazon, and built the strongest independent cross-chip software. Its ranking stays below them because so much of the commercial picture remains private.

FriendliAI ranks fourth and clearly wins the inference-specialist group. Its current product is mature, its benchmark disclosures are more useful than most competitors’, and its combination of cloud and private deployment creates several ways to sell the same engine.

FlexAI ranks fifth after a productive recent expansion. It now has customer case studies, transparent product availability, private infrastructure deployments and support for several GPU families. Its position will improve if those early customer examples turn into a much larger recurring business.

Zymtrace ranks sixth because its technical approach and Anam result are unusually promising for a seed-stage company. It needs more customers before we can separate a repeatable platform from one excellent optimization engagement.

Rapt.AI ranks seventh. Its claimed utilization and throughput gains would justify a much higher position if independently reproduced across several customers. For now, commercial scale and benchmark transparency remain too limited.

Three developments could change the order. ScaleOps can pass Cast AI by sustaining its growth while disclosing broader deployment scale. Decart can move to first by turning DOS into a standardized product used by several hyperscalers and major AI labs. FriendliAI can enter the top three if its inference engine gains enough volume to make tokens served, revenue or GPU fleet size visible.

Rank Startup Why it ranks here
1 Cast AI Broadest commercial footprint, mature automation and the strongest overall evidence
2 ScaleOps Fastest recent growth and excellent enterprise traction, with less disclosed scale
3 Decart Clear cross-chip leader and major strategic backing, but limited commercial transparency
4 FriendliAI Strongest independent inference-efficiency specialist with a mature production platform
5 FlexAI Broad hardware-flexible platform with improving customer and product evidence
6 Zymtrace Compelling early customer result and differentiated code-level optimization
7 Rapt.AI Exceptional claimed performance, with too little independently verifiable adoption

If you want more recent data on this point, please see our latest AI infrastructure market report.

OUR METHODOLOGY

This analysis asks which independent GPU-efficiency startup is ahead based on the evidence available today. We compare the companies across commercial traction, customer quality, growth, product maturity, GPU utilization, inference performance, cross-chip support, technical defensibility, capital efficiency and the quality of the underlying evidence.

We define GPU efficiency as software that helps customers get more useful work from GPUs they already own, rent or manage. The comparison includes infrastructure allocation, GPU sharing, workload scheduling, inference optimization, cross-chip portability and code-level performance analysis. We exclude alternative chipmakers, general GPU clouds and managed inference providers whose main product is compute rather than portable optimization software.

We prioritized recent first-hand evidence, including company product documentation, funding announcements, named customer deployments, engineering material, benchmark methodologies, partnerships and acquisition announcements. Tier-one reporting was used when it added details that companies did not publish directly, particularly around financing, growth and strategic customer relationships.

Not every data point proves the same thing. Funding reflects investor conviction rather than adoption. Benchmarks show technical potential but often depend heavily on model, hardware, batch size, context length, traffic pattern and latency target. Named production customers and documented deployments carry more weight because they show that a product has survived real operational use.

Company-reported performance claims were included when the workload and setup were clear enough to interpret. Named customer outcomes carried more weight than maximum homepage claims, while broad observational datasets were used mainly to show the scale of GPU waste rather than to prove the savings delivered after optimization.

The final ranking does not come from a fixed mathematical score. It is an editorial assessment based on the accumulation of evidence across all dimensions, with more weight given to companies that repeatedly show strength through different forms of proof rather than one exceptional benchmark, funding round or customer announcement.

Key sources include Cast AI’s official website, Cast AI’s Series C announcement, ScaleOps’ official website, TechCrunch reporting on ScaleOps’ financing and growth, Decart’s official website, FriendliAI’s official website, FriendliAI’s documentation and benchmark material, FlexAI’s official website, Zymtrace’s official website, Rapt.AI’s official website, Qualcomm’s acquisition announcements, Nvidia’s acquisition announcements, Red Hat’s acquisition announcements, The Wall Street Journal’s reporting on Amazon and Decart, and Gartner research on enterprise GPU utilization.

Chart showing how GPU cloud infrastructure technology has evolved over time

This chart, included in our AI infrastructure market deck, shows how GPU cloud infrastructure technology has evolved over time

Who is the author of this content?

NEW MARKET PITCH TEAM

We track new markets so founders and investors can move faster

We build living "market pitch" documents for emerging markets: AI, synthetic biology, new proteins, and more. Instead of outdated PDFs or hallucinated LLM answers, our clients get a clean, visual, always-updated view of what's really happening: key players, deals, regulations, and signals that matter. Learn more about us.

Back to blog