AI Infrastructure: what are startups building now?

In our AI infrastructure market deck, you will find everything you need to understand the market
SUMMARY
AI Infrastructure: what are startups building now? Startups are building the full machinery around AI compute: inference platforms, specialized chips, optical interconnects, runtimes, serverless execution, agent sandboxes, routing, evaluation and observability.
The strongest buildout is happening around inference. Training remains huge, but inference is where successful AI products create recurring demand every time a user, application or agent does useful work.
The category is also spreading in two directions at once. Startups are moving downward into silicon, memory and networking, while others move upward into deployment, context management, sandboxes and AI-behavior monitoring.
That makes AI infrastructure less dependent on predicting which foundation model wins. A good serving, routing or execution layer can benefit from OpenAI, Anthropic, Google, Meta, DeepSeek or whichever model becomes popular next.
Raw GPU access is already the weakest part of the stack. The more providers expose similar accelerators, the harder it becomes to defend a business built mainly around renting the same H100, H200 or B200 by the hour.
The more durable infrastructure companies are trying to turn hardware efficiency into software economics. Better batching, lower cold-start times, higher utilization, faster runtimes and smarter scheduling can create hardware-equivalent savings without manufacturing a new chip.
Agents are creating a genuinely different infrastructure problem. Once software can browse, call tools, run code, keep state and act for minutes or hours, isolated execution environments and detailed traces become core production infrastructure.
Memory movement and networking are becoming as important as arithmetic. Etched, Fractile and d-Matrix attack inference from the silicon side, while Ayar Labs, Lightmatter and HyperLight are trying to keep huge clusters fed with data fast enough to behave like one machine.
The cloud opportunity is still real, but simple neocloud economics look less attractive than they did during the GPU shortage. CoreWeave shows how large the revenue opportunity can become, and also how heavy the financing and infrastructure burden can be.
Evaluation, observability and model control are becoming more valuable as AI systems get less deterministic. A system can return a technically successful response while still choosing the wrong tool, retrieving bad context, wasting money or producing an incorrect answer.
The clearest pattern across the market is that investors are funding bottlenecks, not generic middleware. The biggest checks are going to inference, specialized hardware and connectivity where better performance can directly reduce cost or unlock more capacity.
The practical conclusion is simple: the next phase of AI infrastructure is about making expensive compute produce more useful work per dollar, per watt and per second, while giving applications and agents a safer, more reliable way to use it.

This market map, featured in our AI infrastructure market deck, highlights top companies and startups in the AI infrastructure market
What does “AI infrastructure” actually mean now?
AI infrastructure now means far more than renting GPUs: we found startups building almost every layer between electricity entering a data center and an AI application returning a useful answer. The category increasingly covers specialized chips, GPU clouds, optical networking, distributed training, inference engines, serverless compute, model deployment, routing, sandboxes, retrieval, evaluation and observability.
During the first generative-AI investment wave, “infrastructure” often meant securing enough NVIDIA GPUs to train larger models. Today, much more startup activity is concentrated on what happens after a model exists: keeping accelerators busy, moving data between them, serving millions of requests economically, switching between models, running agents safely and finding out why production AI fails.
Fireworks AI, Together AI, Baseten and Modal, for example, all sit above raw hardware but sell different abstractions over it. RunPod sells flexible GPU infrastructure to more than one million developers. Etched and Fractile are attempting to replace general-purpose GPUs with inference-specific silicon. Ayar Labs and Lightmatter attack the increasingly expensive problem of moving data between chips. Braintrust focuses on evaluating AI applications rather than executing them. Inferact, founded by creators and maintainers of vLLM, works even lower in the software stack by optimizing how models are served.
AI infrastructure is becoming an industrial stack whose bottlenecks are large enough to support separate companies.
| Layer | What startups are building | Examples |
|---|---|---|
| Physical compute | AI clouds, GPU clusters, capacity orchestration | Crusoe, Lambda, RunPod |
| Silicon | Inference accelerators and AI systems | Etched, Fractile, d-Matrix |
| Connectivity | Optical I/O and rack-scale interconnects | Ayar Labs, Lightmatter, HyperLight |
| Runtime | Faster execution and model serving | Inferact, Modular, ZML |
| Deployment | Serverless GPU and inference platforms | Fireworks AI, Together AI, Baseten, Modal |
| Control layer | Routing, gateways, monitoring and evals | OpenRouter, Braintrust, LangSmith |
Why are so many startups building AI infrastructure instead of another AI model?
AI infrastructure has become attractive precisely because the model layer is getting harder to attack directly. Startups are increasingly betting that they can make money from every successful model rather than trying to build the winning model themselves.
Frontier-model development requires enormous training budgets, privileged access to compute, large research teams and continuous spending simply to remain near the frontier. Infrastructure companies can instead benefit when OpenAI, Anthropic, Google, Meta, DeepSeek or the next open model improves, because better models generally create more inference, more data movement and more production workloads.
The scale reached by independent infrastructure companies shows that this has moved well beyond a simple picks-and-shovels story. Fireworks AI said when announcing its latest financing that it had exceeded $1 billion in annualized revenue and was processing more than 40 trillion tokens per day. Together AI has said its platform processes more than 400 trillion tokens per month. Modal disclosed more than $300 million in annualized revenue after growing roughly fivefold between financing rounds. RunPod says more than one million developers use its platform.
These companies do not need to predict which foundation model wins. Fireworks can serve different open models. Together supports training, fine-tuning and inference across multiple model families. Modal sells programmable computing rather than intelligence itself.
The harder question is whether they own something more defensible than temporarily scarce compute.
If you want more recent data on this point, please see our latest AI infrastructure market report.

As this chart shows, and as featured in our AI infrastructure market deck, search interest in AI infrastructure has risen sharply
Is AI infrastructure becoming mostly an inference business?
Yes. The startup buildout is increasingly organized around inference rather than frontier training, because inference is where successful AI products turn adoption into recurring infrastructure consumption.
Fireworks AI is the clearest example. The company disclosed more than $1 billion in annualized revenue and over 40 trillion tokens served per day when it announced a $1.5 billion financing at a $17.5 billion valuation. It also said roughly 95% of its token volume came from customized models rather than generic off-the-shelf deployments.
Together AI has followed a similar path. It originally combined training and inference infrastructure, but its current product surface increasingly emphasizes production inference, dedicated deployments and traffic management. Its infrastructure was handling more than 400 trillion tokens per month according to its own July update.
Baseten is even more explicitly an inference company. Its platform revolves around dedicated model deployments, autoscaling, model APIs, runtimes and enterprise reliability. It raised $300 million at a $5 billion valuation earlier this year and subsequently raised substantially more capital, according to its financing announcement.
The hardware startups tell the same story from underneath. Etched is building accelerator systems specifically around transformer inference. Fractile is designing silicon around the memory-bandwidth problem that slows token generation. d-Matrix focuses on digital in-memory computing for inference. Groq built its entire LPU architecture around fast deterministic inference.
Training is not disappearing; frontier laboratories still consume extraordinary amounts of training compute. But only a small number of companies need frontier-scale training, while every successful AI product creates inference demand repeatedly.
What are inference startups actually selling besides cheaper tokens?
The strongest inference startups are selling an operating system for putting models into production. Raw model execution is becoming the entry product rather than the whole business.
Fireworks illustrates the change. Its pitch has expanded from fast model APIs into fine-tuning, custom deployments, optimization and production infrastructure for specialized models. Together AI offers serverless inference but also dedicated model endpoints, training, fine-tuning and mechanisms for gradually shifting traffic between deployments. Baseten combines deployment tooling, autoscaling, observability, model APIs and dedicated infrastructure.
Modal approaches the problem from the other direction. Rather than turning every workload into a model API, it lets developers execute arbitrary Python workloads on elastic CPU and GPU infrastructure. That has made it especially useful for AI workloads that do not fit neatly into a request-response API: reinforcement learning, batch generation, scientific computing and agent sandboxes.
Tokens per second still matter, but customers also care about cold-start time, GPU utilization, availability, model switching, geographic deployment, observability, caching, batching and the cost of unused capacity.
A consumer application can suddenly generate a hundred times normal demand. An agent can execute for minutes rather than milliseconds. A newly released open model can become popular overnight. A production team may need to route traffic away from a failing model without modifying its application.
The product increasingly looks like this: make heterogeneous AI compute behave like reliable cloud software. The H100 itself is only one piece of it.

This chart, included in our AI infrastructure market deck, shows annual VC investment in AI infrastructure startups
Are startups still trying to build new AI clouds?
Yes, but the successful AI-cloud startups are moving away from simple GPU rental. The market increasingly rewards providers that control enough infrastructure to offer predictable capacity while adding software that makes that capacity easier to use.
RunPod shows the developer end of the market. It has grown around self-service GPU access, serverless workloads and per-second pricing rather than giant multiyear infrastructure contracts. When Summit Partners invested $100 million in the company in June, RunPod said it had surpassed one million developers and was valued at $1 billion.
At the opposite end sit capital-intensive providers such as Crusoe and Lambda, which compete for large clusters and long-duration workloads. Public-company CoreWeave provides the clearest benchmark for how far that model can scale. CoreWeave reported $2.575 billion of quarterly revenue, approximately $104 billion of backlog and 1.5 GW of active power in its latest quarter. Yet the same results showed $640 million of quarterly net interest expense and enormous infrastructure obligations.
RunPod emphasizes flexibility. Together AI increasingly combines its own infrastructure with capacity sourced elsewhere. Modal abstracts multiple underlying clouds. Specialized serving companies can purchase capacity and resell useful inference rather than competing to finance entire data centers.
The pure neocloud opportunity is real, but it looks narrower than it first appeared. Owning GPUs is not enough; the durable advantage has to come from utilization, software, financing efficiency or reliability.
If you want more recent data on this point, please see our latest AI infrastructure market report.
Can AI infrastructure startups really compete with AWS, Microsoft and Google?
AI infrastructure startups can compete with hyperscalers, mostly where specialization beats breadth. Startups are not replacing AWS, Azure or Google Cloud as general-purpose clouds; they are carving out workloads where hyperscaler abstractions are too expensive, rigid or complicated.
A company such as Baseten can optimize nearly every part of its product around production inference. Modal can design the developer experience around launching ephemeral GPU workloads from Python. RunPod can offer AI developers immediate access to a wide GPU menu without requiring the rest of an enterprise cloud relationship.
General-purpose clouds were built to serve databases, websites, storage, enterprise applications and thousands of other workloads. AI infrastructure companies can redesign scheduling, networking, storage and billing around expensive accelerators whose economics deteriorate rapidly when they sit idle.
But hyperscalers control data centers, network backbones, enterprise procurement relationships and enormous capital budgets. AWS, Microsoft and Google can also bundle model serving with storage, databases, security and committed-cloud-spend agreements.
The specialist case is still strong. Fireworks has reached more than $1 billion in annualized revenue, Together is handling more than 400 trillion tokens per month, Modal surpassed $300 million in annualized revenue, and RunPod passed one million developers.
Startups are winning narrow workloads where specialization is valuable enough to justify another vendor.

This chart, included in our AI infrastructure market deck, shows why CoreWeave is winning in AI infrastructure
Are serverless GPUs and AI agents creating a new infrastructure layer?
Yes. Serverless compute and AI agents are converging into a new infrastructure layer built around short-lived, unpredictable and isolated computational tasks rather than permanently rented machines.
A dedicated accelerator can cost dollars per hour whether useful work is happening or not. That makes sense for continuously loaded models but becomes painful for experimentation, batch jobs and applications with sharp traffic peaks.
Modal, RunPod and similar platforms attack exactly that waste. Developers send work to the platform, infrastructure starts when needed and the customer pays closer to actual usage. Modal has invested heavily in shortening container and GPU startup times; RunPod similarly markets rapid serverless worker startup.
Modal's growth offers unusually strong evidence that the abstraction is useful. The company disclosed more than $300 million in annualized revenue when raising $355 million at a $4.65 billion valuation, saying revenue had increased roughly fivefold since its previous round.
More revealingly, Modal says Sandboxes already generate more than one-third of its revenue. These are isolated execution environments rather than conventional model-serving endpoints.
Agents need infrastructure conventional model APIs did not. A chatbot largely sends text to a model and returns text. An agent may generate code, launch processes, access files, call external applications, browse websites, retry failed actions and continue working for minutes or hours.
Cloud providers and AI infrastructure companies are consequently adding sandbox products, durable execution, agent tracing and fine-grained permission systems. Gateways are becoming control points that determine which models and tools agents may access.
Observability becomes harder too. A failed agent can result from the model, prompt, retrieval, memory, tool API, authentication or a sequence of earlier decisions. Companies such as Braintrust, LangSmith and Arize therefore trace entire execution paths instead of only model latency.
The unit being scheduled is increasingly one computational task: somewhere safe for software to act, retain state and be inspected when it fails.
Are startups really building alternatives to NVIDIA GPUs, and why are memory and networking suddenly so important?
Yes. AI infrastructure startups are building serious alternatives to NVIDIA GPUs, and the broader shift is just as important: they are attacking the memory and networking bottlenecks that determine whether accelerators can actually be used efficiently at scale.
Etched is the most aggressive example. Its architecture is designed around transformer inference rather than general-purpose accelerated computing. The company shipped its first rack to Jane Street and subsequently announced a $700 million financing at a $21 billion valuation led by Jane Street. That followed a $300 million financing only weeks earlier. Etched has also said it has more than $1 billion of contracted demand.
Fractile is taking a different route. Its design attempts to solve the memory bottleneck of inference by bringing computation and model weights much closer together rather than repeatedly moving weights between external high-bandwidth memory and compute units. The company raised $220 million in May.
d-Matrix similarly attacks inference through in-memory compute. SambaNova continues to develop its own accelerator systems. Groq built its LPU architecture around fast deterministic inference.
The same bottleneck appears at cluster scale. Once thousands of accelerators are connected, keeping them supplied with model weights, activations and intermediate results becomes a networking problem as much as a compute problem.
That explains the capital moving into optical interconnects. Ayar Labs raised $500 million at a reported $3.75 billion valuation to expand production of optical I/O technology. Its TeraPHY chiplets are intended to bring optical communication much closer to processors. Lightmatter's Passage architecture similarly uses photonics to move data between AI chips and racks. HyperLight raised $80 million to scale thin-film lithium-niobate photonics for AI infrastructure.
NVIDIA still sells far more than a GPU: CUDA, networking, libraries and complete rack-scale systems. New entrants therefore need to win at the system level.
The infrastructure frontier is increasingly about keeping processors fed with data and connected closely enough to behave as one machine.
| Startup | Main bet | Recent evidence |
|---|---|---|
| Etched | Transformer-specific inference systems | First rack delivered; $700M round at $21B |
| Fractile | In-memory inference architecture | $220M financing |
| d-Matrix | Digital in-memory inference | Moving Corsair systems into production |
| SambaNova | Full-stack AI accelerator systems | Continued large-scale financing and deployments |
| Groq | Deterministic LPU inference | Demonstrated extremely high inference throughput |
If you want more recent data on this point, please see our latest AI infrastructure market report.

This chart, included in our AI infrastructure market deck, shows annual funding in AI infrastructure startups
Is raw GPU access already becoming a commodity?
Basic GPU rental is moving toward commoditization, but reliable access to well-utilized AI compute still has real value. That distinction increasingly determines whether an infrastructure startup has a durable business or merely temporary scarcity economics.
Customers can already compare H100, H200, B200 and other accelerator prices across an expanding list of clouds. RunPod, hyperscalers, neoclouds and marketplaces all expose variations of GPU-hour pricing. As supply expands and older generations remain useful, raw hourly prices face natural pressure.
Inference providers show how companies are escaping that commodity trap. Fireworks charges customers for useful model execution rather than simply passing through GPU hours. Baseten sells production deployments and reliability. Modal charges for precisely scheduled computational work. Together combines model APIs with dedicated infrastructure.
CoreWeave provides an instructive benchmark. Its latest results showed that even a huge, rapidly growing AI cloud can carry massive depreciation and financing costs. Quarterly revenue reached $2.575 billion, but net interest expense alone was $640 million.
Software-heavy infrastructure startups avoid some of that balance-sheet burden, but suppliers and hyperscalers can squeeze their margins instead.
Owning an H100 is not a moat. Making one H100 perform meaningfully more useful work per dollar can be.
Are inference runtimes becoming companies rather than open-source projects?
Yes. Some of the newest AI infrastructure companies are commercializing what used to look like a low-level open-source engineering problem: squeezing more useful inference out of the same hardware.
The clearest sign is Inferact. The company was founded by people behind vLLM, one of the most important open-source model-serving systems, and announced a $150 million seed financing. That is an unusually large amount of capital for infrastructure whose basic job can be described as “run models more efficiently.”
Modular is pursuing the same problem more broadly through its MAX platform and Mojo language. It is attempting to create a portable compute layer that can optimize AI execution across heterogeneous hardware rather than requiring developers to code directly against one accelerator ecosystem.
ZML is building an inference stack intended to run efficiently across different chips. Wafer is applying AI itself to low-level GPU optimization, including kernel generation and benchmarking. Kog is focusing on low-latency inference for agent workloads.
If a runtime improves effective GPU utilization or token throughput by 20%, that gain applies repeatedly to expensive infrastructure. At large scale, software optimization can therefore create hardware-equivalent savings without manufacturing a chip.
It also makes a heterogeneous accelerator market more plausible. Customers are much less likely to adopt alternative chips if each requires a separate developer stack.
That is why low-level runtimes have become investable companies rather than merely engineering projects.

This chart, included in our AI infrastructure market deck, compares the main business model options for AI cloud infrastructure providers
Do AI companies need a new control and context layer between applications and models?
Increasingly, yes. As companies use several models and increasingly complex retrieval systems, model gateways, routing and context management are becoming production infrastructure rather than developer conveniences.
An AI application may use one frontier model for difficult reasoning, a small model for classification, an open model for high-volume work and a specialist model for coding or images. Prices and performance can change faster than the application itself.
A control layer lets developers send those requests through one interface and decide dynamically where they should go. OpenRouter has built one of the most visible versions of this model, aggregating a large catalog of models behind a common API. Production gateways increasingly add fallbacks, caching, rate limits, data-governance policies, cost controls and observability.
Agents strengthen the case further. An organization may want to constrain which models an agent may call, what information can leave a region, how much an agent may spend or what happens when one provider fails.
At the same time, vector databases are becoming less central as standalone infrastructure. During the first RAG boom, companies frequently assembled a stack consisting of embeddings, a dedicated vector database and a language model. Pinecone, Weaviate, Qdrant and others became synonymous with that architecture.
Traditional databases and cloud platforms have since added vector search, while longer-context models reduce the need to retrieve tiny fragments for some workloads. Retrieval remains important, especially for fresh enterprise information, but the harder problem is increasingly deciding what context an AI system should receive, enforcing permissions and combining structured and unstructured information.
The risk for both gateways and vector databases is commoditization. Basic routing is technically reproducible, and vector search is increasingly a standard database feature.
The defensible layer is moving toward governance, policy and context orchestration rather than simply “one API for many models” or “we store embeddings.”
If you want more recent data on this point, please see our latest AI infrastructure market report.
Are AI evaluation and observability becoming real infrastructure?
Yes. Evaluation and observability are becoming one of the clearest software layers in AI infrastructure because production AI behaves too unpredictably for conventional software monitoring alone.
Traditional application monitoring can tell engineers whether an API returned an error or exceeded a latency threshold. An AI application can return HTTP 200 and still be completely wrong. Agents make the problem worse because a superficially plausible final answer may hide a failed retrieval, incorrect tool selection or unnecessary sequence of expensive model calls.
Braintrust has built around this problem by making evaluations, datasets, experiments and production traces part of one engineering workflow. The company raised $80 million in a Series B earlier this year.
Arize expanded from classical machine-learning observability into LLM and agent evaluation, tracing and automated diagnosis. The strategic value of that layer became unusually visible when Dynatrace agreed to acquire Arize for approximately $915 million.
LangChain has followed a parallel path with LangSmith, which gives teams tracing, evaluation and deployment tooling around LLM applications and agents. Open-source systems such as Arize Phoenix and Langfuse further show that tracing has become a standard development primitive.
The underlying shift is from monitoring infrastructure health to monitoring AI behavior: hallucination rates, tool-selection accuracy, retrieval relevance, token expenditure, judge scores, task completion and failure paths.
As AI systems take actions rather than merely produce text, those metrics become production controls rather than optional analytics.

This chart, featured in our AI infrastructure market deck, shows the share of revenue generated by each customer segment in the AI infrastructure market
Where is the biggest startup money going inside AI infrastructure?
Capital is concentrating around bottlenecks whose value scales directly with AI usage: inference, specialized silicon and high-bandwidth connectivity. Financing is much less enthusiastic about generic middleware than technologies that can materially change compute economics.
Fireworks AI's $1.5 billion round is especially striking because the company does not train a frontier foundation model. Investors are valuing the infrastructure that executes and customizes other companies' models at $17.5 billion.
Together AI raised $800 million in its latest round at an $8.3 billion valuation while expanding both its software platform and physical compute footprint. Modal raised $355 million after disclosing more than $300 million in annualized revenue.
At the hardware level, Etched raised $700 million after shipping its first rack, while Ayar Labs raised $500 million to move optical I/O toward high-volume production. Fractile raised $220 million before commercial-scale deployment of its inference hardware.
The financings show where investors expect value to accumulate: around the cost and capacity constraints created by successful AI applications.
| Startup | Infrastructure layer | Recent disclosed financing |
|---|---|---|
| Fireworks AI | Inference platform | $1.5B |
| Together AI | AI cloud / inference | $800M |
| Etched | Inference hardware | $700M |
| Ayar Labs | Optical interconnect | $500M |
| Modal | Serverless AI compute | $355M |
| Baseten | Inference platform | $300M Series E, followed by a larger later round |
| Fractile | Inference hardware | $220M |
| RunPod | AI developer cloud | $100M |
| Braintrust | AI evaluation | $80M |
| HyperLight | AI photonics | $80M |
Which AI infrastructure businesses look easiest to commoditize?
The most vulnerable AI infrastructure startups are those selling something customers can compare primarily on price: generic GPU hours, undifferentiated model APIs, simple routing and basic vector search.
GPU rental is the obvious example. A customer that simply needs eight H100s can increasingly compare multiple vendors. Unless a provider offers better availability, networking, software or financing terms, the workload can move.
Basic inference APIs face similar pressure. Popular open models can appear simultaneously on several providers. If outputs are identical, price and latency become dominant purchasing criteria.
Generic gateways face another problem: routing requests between APIs is not intrinsically difficult. Vector search has already demonstrated how quickly a specialist category can become a platform feature as PostgreSQL extensions, cloud databases and data platforms add similar capabilities.
The stronger businesses control a harder bottleneck. Etched cannot be duplicated by adding a software feature because it manufactures specialized silicon. Ayar Labs requires photonics technology, packaging and production expertise. Inferact and other runtime companies can defend themselves through deep performance engineering. Modal becomes harder to replace when applications depend on its execution model, autoscaling and sandboxes rather than merely its GPU price.
Infrastructure does not automatically mean defensibility. The strongest positions are where a company can make something measurably cheaper, faster, safer or more reliable.
If you want more recent data on this point, please see our latest AI infrastructure market report.

This chart, included in our AI infrastructure market deck, shows how GPU cloud infrastructure technology has evolved over time
What are AI infrastructure startups building now?
AI infrastructure startups are now building the machinery required to turn increasingly abundant AI intelligence into reliable, economical computing systems. The category has moved decisively beyond “more GPUs” and is spreading both downward into chips, networking and power constraints and upward into serverless compute, inference runtimes, agent sandboxes, routing and evaluation.
The strongest current buildout is around inference. Fireworks AI, Together AI and Baseten are turning model serving into large independent businesses. Modal and RunPod are redesigning cloud computing around short-lived and irregular AI workloads. Inferact, Modular, ZML and others are trying to extract more performance from every accelerator through software.
Below them, Etched, Fractile and d-Matrix are betting that inference eventually deserves purpose-built silicon rather than permanent dependence on general-purpose GPUs. Ayar Labs, Lightmatter and HyperLight are working on the bottleneck created when thousands of those chips need to exchange enormous amounts of data.
Above compute, another layer is appearing because agents behave differently from ordinary applications. Sandboxes, model gateways, tracing systems and eval platforms are becoming basic components for software that can execute actions autonomously.
The clearest conclusion is that the most important AI infrastructure startups today are building around the inefficiencies created by AI at scale.
The first phase of the AI boom rewarded whoever could obtain GPUs. The next phase is rewarding companies that can make those GPUs—and increasingly alternative chips—produce more useful work per dollar, per watt and per second.
The infrastructure opportunity is shifting toward removing the bottlenecks around compute: inference cost, utilization, memory bandwidth, networking, deployment complexity, agent execution and reliability.
OUR METHODOLOGY
This analysis examines what AI infrastructure startups are building now by breaking the market into the layers where new companies are concentrating: physical compute, silicon, connectivity, runtimes, deployment, model control, agent execution, evaluation and observability.
We prioritized recent evidence that showed what was happening in practice: product launches, production deployments, disclosed revenue and usage, developer adoption, financing, technical architecture changes and, where available, public financial results. We gave the most weight to signals that revealed actual operating scale, customer demand or technical commitment.
The companies featured here were selected because they made a particular shift or bottleneck unusually visible. This was not designed as a directory of AI infrastructure startups. A company was more useful to the analysis when there was enough concrete evidence to understand what it was building, how the product was being used and whether the underlying market was gaining real economic weight.
We treated large financings as one signal among several. They helped show where investors were placing unusually large bets, especially across inference, specialized hardware and connectivity, but we compared those capital flows with product development, usage and commercial evidence before drawing broader conclusions.
For private companies, we prioritized direct company announcements, technical documentation and product releases. Public-company filings and investor disclosures gave us a clearer benchmark for the economics of capital-intensive AI clouds.
The main conclusions came from aggregation. We looked for the same direction to appear across several independent types of evidence before treating it as a broader market shift, including the growing weight of inference, the emergence of agent execution infrastructure, the importance of memory and networking, and the parts of the stack most exposed to commoditization.
For defensibility, we focused on substitutability: how easily a customer could compare providers, reproduce the functionality elsewhere or move a workload without giving up meaningful performance, reliability or integration.
Key sources used for this analysis include: Fireworks AI on its Series D, annualized revenue and token volume, Together AI on its $800 million financing, Together AI on provisioned throughput and token scale, Baseten on its $300 million Series E, Baseten on its later Series F, Modal on its Series C and annualized revenue, Modal on large-scale agent sandboxes, RunPod on reaching one million developers, CoreWeave on revenue, backlog, power and interest expense, Etched on its first rack and financing, Fractile on its $220 million financing and inference architecture, d-Matrix on Corsair entering production, Ayar Labs on optical I/O and its $500 million Series E, Lightmatter on its Passage photonic interconnect roadmap, QIA on HyperLight's $80 million Series C, Inferact on its inference-runtime strategy, OpenRouter on model routing, fallbacks and provider selection, Braintrust on its Series B and production AI evaluation, Dynatrace on its agreement to acquire Arize, and NVIDIA documentation on rack-scale AI systems, NVLink and networking.

In our AI infrastructure market deck, we identify pain points entrepreneurs should prioritize
Related blog posts
- What are the latest funding developments in AI infrastructure?
- AI Infrastructure: what are the top startups now?
- The main fundraising trends in AI infrastructure
- How strong is fundraising in the AI infrastructure market right now?
Who is the author of this content?
NEW MARKET PITCH TEAM
We track new markets so founders and investors can move fasterWe build living "market pitch" documents for emerging markets: AI, synthetic biology, new proteins, and more. Instead of outdated PDFs or hallucinated LLM answers, our clients get a clean, visual, always-updated view of what's really happening: key players, deals, regulations, and signals that matter. Learn more about us.