Voice AI: what is getting real adoption now?

In our conversational AI market deck, you will find everything you need to understand the market
SUMMARY
Voice AI is already getting real adoption now, especially in repetitive phone workflows, healthcare documentation, customer support, scheduling, restaurant ordering and other jobs where the conversation leads to a small number of defined actions.
The clearest evidence comes from usage rather than demos. Retell reports more than 100 million calls a month, Vapi has passed one billion cumulative calls, and Bland reports more than 650 million calls resolved.
Voice AI has also become a meaningful software business. ElevenLabs has passed $500 million in ARR, while voice-agent platforms such as Retell and Vapi have built substantial recurring-revenue businesses around production workloads.
The strongest deployments share a surprisingly simple characteristic: the language can be complicated while the job itself stays narrow. Booking an appointment, taking an order, collecting a payment or changing an account record gives the agent a clear destination even when callers phrase things unpredictably.
Customer service is already moving beyond experimental deployments. Klarna, Revolut, Sunshine Loans and SafeRide Health show Voice AI handling large call volumes, shortening resolution times, reducing abandonment and taking routine workload away from human teams.
Healthcare may be the deepest enterprise deployment so far. Ambient documentation systems are expanding from hundreds of clinicians to thousands, while independent studies are beginning to measure lower documentation time rather than relying only on vendor claims.
Restaurant ordering and outbound calling show where the boundary currently sits. Phone ordering, collections, reminders and warm follow-up already have credible production evidence; open-ended drive-thru automation and mass autonomous cold calling remain much less consistent.
Consumer Voice AI is much larger by user count than enterprise voice agents. Gemini and Alexa+ are bringing conversational voice to tens or hundreds of millions of people, although those numbers mix voice with broader multimodal assistant usage.
Automotive voice shows a similar split between enormous distribution and newer generative capability. Traditional assistants already sit in hundreds of millions of vehicles, while modern LLM-powered systems are only beginning to spread through that installed base.
The market is therefore developing around bounded authority rather than unlimited autonomy. Voice AI works best today when companies know what the agent is allowed to do, can measure whether it succeeded and can hand unusual cases to a person without much friction.

This market map, featured in our conversational AI market deck, highlights top companies and startups in the conversational AI market
What counts as real Voice AI adoption, and is Voice AI there yet?
Real Voice AI adoption now means production usage that survives the pilot, and Voice AI clearly meets that bar in several parts of the market.
A funding round, a polished demo or an announced partnership tells us very little on its own. We give much more weight to recurring call volume, customer expansion, money collected, appointments booked, staff hours removed from routine work, or deployments that grow from hundreds of users to thousands.
Retell's current website says more than 4,000 businesses now generate over 100 million calls a month on its platform. In April 2026, the company was publicly reporting more than 50 million monthly calls, so the disclosed run rate has roughly doubled within a few months.
Vapi shows that this scale extends beyond one vendor. TechCrunch reported in May 2026 that Vapi had processed more than one billion calls cumulatively and was running between one million and five million calls per day, with enterprises responsible for most of that traffic. Bland's own site currently displays more than 650 million calls resolved to date.
The revenue is meaningful too. ElevenLabs passed $500 million in annual recurring revenue during the first four months of 2026 after ending 2025 around $350 million. Retell publicly reported about $50 million in ARR earlier in 2026; Sacra now estimates roughly $80 million in annualized revenue by August. TechCrunch reported that Vapi was running at a “healthy” eight-figure ARR when it raised its latest round.
Large contact-center vendors are seeing the same budget shift. Genesys's latest quarterly disclosure put AI ARR above $400 million, while NICE reported 66% year-over-year growth in AI ARR in its first quarter. Those numbers include AI products beyond voice, so we use them as evidence of broader contact-center AI spending.
Company-reported call counts can cover very different call lengths and outcomes. Even with that caveat, recurring usage and recurring revenue are already far beyond what we would reasonably call experimental adoption.
| Voice AI company | Latest scale | What it shows |
|---|---|---|
| Retell | 100M+ calls per month; roughly $80M annualized revenue estimated by Sacra | Large recurring usage and a substantial voice-agent-native business |
| Vapi | 1M–5M calls per day; 1B+ cumulative; eight-figure ARR | Another independent platform reaching serious enterprise scale |
| Bland | 650M+ calls resolved cumulatively | Large lifetime usage outside the two biggest disclosed platforms |
| ElevenLabs | $500M+ ARR | Enterprise audio and voice software has become a major software category |
| Genesys | $400M+ AI ARR | Large contact-center customers are expanding AI budgets |
If you want more recent data on this point, please see our latest conversational AI market report.
Why is Voice AI suddenly useful enough for real work?
Voice AI has become useful enough for real work because live conversations are faster, cheaper and much easier to connect to the software that completes the task.
Latency used to kill the experience. Retell currently advertises end-to-end responses of roughly 600 milliseconds, while ElevenLabs says its Flash speech models can generate audio with model latency around 75 milliseconds. Those figures measure different parts of the stack, so a direct comparison would be misleading. Together, they show how far the industry has moved away from the awkward multi-second pauses that made earlier voicebots feel broken.
Cost has moved in the same direction. Bland currently charges self-serve customers around $0.12 to $0.14 per connected minute, depending on the plan. Retell starts much lower at the infrastructure layer, with the final price changing according to the language model, voice, telephony and other components selected. At high call volumes, even a few dimes per minute can be attractive beside a fully staffed phone queue.
The bigger change is what the agent can do during the call. Voice platforms now connect to CRMs, calendars, payment systems, ticketing tools, electronic health records and contact-center software. An agent that can check an account, move an appointment, collect a payment or create a ticket has a much easier business case than one that simply talks well.

As this chart shows, and as featured in our conversational AI market deck, search interest in conversational AI has increased sharply
Are customer service and AI receptionists the clearest Voice AI winners today?
Customer service, reception and scheduling are currently among the clearest Voice AI winners because the calls repeat, the desired actions are usually known and human escalation is easy to keep.
Klarna has put an ElevenLabs agent at the front of US phone support for a customer base of 35 million people. Klarna says queries completed by the AI can reach resolution up to ten times faster. Revolut is doing something similar across the UK and Europe for more than four million customers; its first rollout reported over an eightfold drop in time to resolution, a 99.7% successful-call rate and support across more than 30 languages.
The smaller-company examples make the labor impact easier to see. Sunshine Loans says AI now handles 75% to 80% of its calls, while abandonment fell from peaks of 20%–30% to around 5%–6%. The lender also says the system replaced coverage previously provided by more than 100 offshore agents.
Reception and scheduling show the same pattern. Pine Park Health says its Retell agents achieve a 55.7% booking success rate when the scheduling agent reaches a patient, improved scheduling-related NPS by 38% and recovered about 2.1 full-time-equivalent medical-assistant roles for other work.
SafeRide Health shows how large that workflow can become. PolyAI says its voice agent handles more than one million calls a month for non-emergency medical transportation, has saved 47,000 staff hours on authentication alone and now handles the majority of calls by itself.
Consumers still want a human escape route. Five9 surveyed 3,000 consumers and 600 CX decision-makers across the US, UK and Germany in April 2026 and found that 80% of consumers were willing to use AI-powered customer service, while about two-thirds still preferred speaking with a human. Gartner later found that 87% of 3,566 surveyed customers considered access to a human agent essential when generative AI is used for support.
The practical setup is pretty clear: Voice AI handles the repetitive first layer, and unusual or sensitive cases move to a person.
If you want more recent data on this point, please see our latest conversational AI market report.
Is healthcare ambient Voice AI already a major market?
Healthcare ambient Voice AI is currently one of the deepest real deployments of voice technology anywhere.
Abridge now says its platform is live across more than 300 health systems and will support more than 100 million patient-clinician conversations over a year. Its latest expansion with BJC Health and WashU Medicine shows what happened after the pilot: access grew from 450 clinicians in 2025 to roughly 4,000 clinicians, while measured documentation-time reduction improved from 8% early in the rollout to 15% by day 150. After-hours documentation fell by as much as 20% by that point.
Independent research points in the same direction. A prospective Hawaii Pacific Health study published in 2026 followed 79 providers using Abridge across more than 25,000 notes. Heavy users spent 21% less time on notes per day, equivalent to 13.6 minutes, and the share of providers reporting at least eight hours of weekly after-hours documentation fell sharply during the pilot.
The safety problem remains real. A separate UK survey of 1,003 general practitioners found that 14% were already using ambient AI scribes and another 39% planned to adopt them soon. Among current users, 32% said errors appeared often or always, and 14% reported errors with potentially significant to critical implications. Heidi Health accounted for 86% of users in that survey, showing that ambient adoption extends beyond Abridge.
| Healthcare Voice AI evidence | Scale | What changed |
|---|---|---|
| Abridge overall | 300+ health systems; 100M+ annual conversations expected | Ambient documentation has reached large-system production scale |
| BJC Health / WashU Medicine | 450 clinicians → about 4,000 | A pilot expanded by almost 9x |
| Hawaii Pacific Health study | 79 providers; 25,000+ notes | Heavy users spent 21% less daily time on notes |
| UK GP survey | 1,003 GPs | 14% already using ambient scribes; 39% planned near-term adoption |

This chart, included in our conversational AI market deck, shows annual VC investment in conversational AI startups
Are restaurants really automating orders with Voice AI?
Restaurant Voice AI has clearly reached production, especially for phone ordering, although drive-thru automation is still uneven.
Casey's is one of the strongest examples we found. SoundHound says its ordering agents are deployed in more than 2,600 of Casey's roughly 2,900 locations, close to 90% of the chain, and have already handled more than 21 million guest interactions. SoundHound currently says its restaurant and retail voice technology reaches more than 15,000 locations overall.
Phone ordering gives Voice AI a fairly contained job: menu information, modifications, store information and orders that can be sent straight into restaurant software.
Drive-thru ordering is harder. McDonald's removed its IBM automated ordering system from more than 100 restaurants after a two-year test in 2024. Yum has continued pushing in the other direction, using its own Byte platform and NVIDIA technology to expand voice ordering across brands including Taco Bell and Pizza Hut.
Thousands of stores and tens of millions of interactions are enough to call restaurant Voice AI real adoption. Drive-thru results are still messier.
Is outbound Voice AI actually working for sales, collections and reminders?
Outbound Voice AI works today when the company already knows why it is calling; broad cold prospecting has a much weaker case.
Medical Data Systems gives us one of the clearest economic examples. The healthcare collections company says Retell agents now handle roughly 30,000 inbound and outbound calls per month, collect about $280,000 monthly and complete around 70% of conversations without a live agent.
BrightChamps uses outbound Voice AI for trial reminders, lead qualification and re-engagement. The education company reports a roughly 15% lift in trial joining, a 5% lift in trial completion and a re-engagement campaign contributing 10%–15% of monthly revenue.
Regulation makes cold calling harder. The FCC has ruled that AI-generated voices fall under Telephone Consumer Protection Act restrictions governing artificial or prerecorded voices. The FTC's Telemarketing Sales Rule also requires prior written agreement for most prerecorded telemarketing calls and imposes opt-out and do-not-call obligations.
So far, the better evidence sits with collections, reminders, renewals, lead follow-up and existing-customer outreach. Mass autonomous prospecting still looks much less proven.
If you want more recent data on this point, please see our latest conversational AI market report.

This chart, included in our conversational AI market deck, breaks down Cognigy's playbook in conversational AI
Is consumer Voice AI finally mainstream?
Consumer Voice AI is now mainstream inside general AI assistants, especially when voice sits beside screens, cameras and actions.
Google's August 2026 Gemini update is the clearest scale data we have. Gemini passed one billion monthly active users, and Google says 63% of users now talk directly to Gemini. That figure includes people who mix typing and speaking, so it cannot be read as 630 million voice-only users.
Google says one in five Gemini Live interactions now extends beyond speech into camera or screen sharing. A user can show an object or screen, ask a question aloud and keep talking while the model reasons over what it sees.
Alexa+ adds another large-scale example. Amazon says the upgraded assistant reached tens of millions of customers within its first nine months. Those users were having roughly twice as many conversations as before, making about three times as many purchases and requesting recipes five times as often.
Generative assistants are pushing consumer voice beyond the old smart-speaker habit of timers, weather and music into longer conversations, planning and shopping.
Is in-car Voice AI already real adoption or mostly future hype?
Automotive Voice AI already has huge distribution, while the newest LLM-powered assistants are only starting to spread through that installed base.
Cerence says its technology has shipped in more than 525 million vehicles over time and appeared in more than 25 million new vehicles during fiscal 2025, around 52% of global vehicle production for that year. Much of that installed base comes from earlier generations of speech and assistant technology, so it gives us distribution scale rather than proof that hundreds of millions of cars already have modern generative agents.
The upgrade cycle is now visible in production. On Cerence's latest quarterly earnings call, management said roughly 100,000 xUI-powered cars were already on the road and that programs with Stellantis, BYD, Geely, a Volkswagen Group brand and others were entering production or approaching it. Volvo has also started rolling Gemini out to cars with Google built in, including existing models dating back several years.
Cars give voice an obvious interface advantage because drivers already have a reason to avoid touching a screen. Conventional in-car voice is enormous today; generative assistants are still early beside that installed base.
If you want more recent data on this point, please see our latest conversational AI market report.

This chart, included in our conversational AI market deck, shows annual funding in conversational AI startups
Is AI dubbing and narration getting real adoption too?
AI dubbing and synthetic narration are already real production markets, with localization showing the clearest adoption.
Vimeo's latest deployment data is unusually concrete. Since integrating ElevenLabs dubbing at the start of 2026, customers have generated about 1.4 million dubbed minutes across more than 25,000 videos and 134,000 dubbing jobs. More than 27% of Vimeo enterprise accounts have tried the feature, and each dubbed video is translated into 3.3 target languages on average.
Some customers already dub a single video into more than 30 languages, which explains why corporate training, product updates and marketing content fit the technology so well.
The creator economy around synthetic voice is becoming meaningful too. ElevenLabs says creators in its voice marketplace have earned more than $22 million, twice the amount reported six months earlier, across more than 10,000 earning creators. ElevenReader now offers more than 200,000 premium audiobooks and ebooks, although that catalog mixes human-narrated and AI-narrated material.
Dubbing has the cleaner adoption case today because companies can lower localization costs without changing the original performance.
Why are narrow Voice AI agents winning faster than general voice agents?
Narrow Voice AI agents are winning faster because the company can define the job, the allowed actions and the handoff point before the conversation begins.
A clinic agent books or moves an appointment. A restaurant agent takes an order. A collections agent identifies an account and works toward payment. A customer-service agent answers recurring questions, changes a known record or transfers the caller.
The language can be messy while the action space stays small. The system can be tested against real calls, given access to a limited set of tools and measured through booking rate, payment collection, containment or handle time.
Open-ended agents face a much harder problem. The caller can switch topics, ask for an exception, introduce sensitive information or request an action the system has never been allowed to take.
Recent customer research supports keeping a human exit. Gartner found that 87% of surveyed customers want access to a person when generative AI is used for service. Five9 found that 83% of consumers still have to repeat themselves at least sometimes after an AI-to-human transfer.
Companies are comfortable giving Voice AI bounded authority over repetitive jobs today. Trust drops quickly once the conversation becomes unusual, emotional or hard to reverse.

This chart, included in our conversational AI market deck, compares the main business model options for conversational AI enterprise platforms
So what Voice AI is getting real adoption now?
Real Voice AI adoption today is strongest in healthcare documentation and high-volume phone workflows where the conversation repeats and the AI can complete a limited set of actions.
Inbound customer service, reception and scheduling are already firmly real. Large financial companies are putting millions of customers behind AI first-line support, while smaller businesses are automating most routine calls, recovering staff time and cutting abandonment.
Ambient clinical documentation may be the deepest enterprise case. Hospitals are moving from hundreds of clinicians in pilots to thousands in systemwide rollouts, and independent studies are finding measurable reductions in documentation time. Restaurant ordering, collections and warm outbound follow-up sit just behind those leaders.
Consumer voice has reached much larger user numbers through Gemini and Alexa+, while automotive voice has enormous legacy distribution and has begun a generative upgrade cycle. Dubbing is also crossing into normal production work.
Broad cold calling and agents expected to handle almost any high-stakes conversation from beginning to end remain weaker.
Our final judgment is clear: Voice AI has already found product-market fit in repetitive, bounded workflows with defined actions and clean human escalation. The further a deployment moves away from those conditions, the faster the evidence thins out.
If you want more recent data on this point, please see our latest conversational AI market report.
OUR METHODOLOGY
This analysis tests where Voice AI has genuinely reached real adoption. We split the market into platform-scale voice agents, customer service and reception, healthcare documentation, restaurant ordering, outbound calling, consumer assistants, automotive voice and AI dubbing, then looked for production evidence inside each one.
We gave the most weight to recurring call volume, recurring revenue, completed transactions, appointments booked, staff time recovered, measurable workflow improvements and deployments that expanded materially after an initial pilot. Funding rounds, demonstrations and announced partnerships were useful context, but they were not treated as proof of adoption on their own.
Company disclosures were used for concrete deployment figures where those companies were the direct source of the data. Broader conclusions were made only when similar patterns appeared across several vendors, customers or independent datasets. We also kept the measurements separate: monthly calls, ARR, booking rates, collections, documentation time and customer-service outcomes describe different parts of adoption and were not forced into a single score.
Healthcare evidence received additional weight where independent research was available. The Hawaii Pacific Health study gave us prospective evidence on documentation time across more than 25,000 notes, while the UK GP survey added a separate view of ambient-scribe adoption and error frequency. Regulatory limits on outbound calling were checked against the FCC's interpretation of AI-generated voices under the TCPA and the FTC's Telemarketing Sales Rule.
We also kept contrary evidence in the analysis. McDonald's withdrawal of its earlier automated drive-thru ordering system, reported ambient-scribe errors, consumers' continued preference for access to human agents and the much smaller installed base of new generative in-car assistants all help define where Voice AI remains less mature.
Key sources used for this analysis include Retell on production call scale, TechCrunch on Vapi's call volume and enterprise usage, Bland on cumulative call volume, ElevenLabs on $500 million-plus ARR, Genesys on AI ARR, NICE on AI ARR growth, Abridge on the BJC Health and WashU Medicine expansion, the Hawaii Pacific Health prospective study, the UK GP ambient-scribe survey, SoundHound on Casey's restaurant deployment, Retell on Medical Data Systems, the FCC on AI-generated voices under the TCPA, the FTC's Telemarketing Sales Rule guidance, Google on Gemini usage, Amazon on Alexa+ adoption, Cerence's vehicle-distribution disclosures, Volvo on Gemini's production rollout, and ElevenLabs and Vimeo on AI dubbing usage.

This chart, featured in our conversational AI market deck, illustrates revenue distribution by customer segment in the conversational AI market
Related blog posts
- Voice AI: what is actually working now?
- What are the latest funding developments in conversational AI?
- Conversational AI: what are the top startups?
- The main fundraising trends in conversational AI
- How strong is fundraising in the conversational AI market right now?
Who is the author of this content?
NEW MARKET PITCH TEAM
We track new markets so founders and investors can move fasterWe build living "market pitch" documents for emerging markets: AI, synthetic biology, new proteins, and more. Instead of outdated PDFs or hallucinated LLM answers, our clients get a clean, visual, always-updated view of what's really happening: key players, deals, regulations, and signals that matter. Learn more about us.