Voice AI: what is actually working now?

In our conversational AI market deck, you will find everything you need to understand the market
SUMMARY
Voice AI is already working now, especially when it handles bounded, high-volume jobs where people can speak naturally but the business actions behind the conversation stay controlled.
The biggest change is that voice quality is no longer the main test. Current systems can sound natural, respond quickly and survive interruptions; the harder question is whether they can authenticate callers, use tools correctly, complete transactions and recover when a conversation gets messy.
Production deployments show that meaningful call automation is already possible, but there is no normal containment rate. Some companies remove only a fifth or a quarter of calls from human queues, while narrow reservation-heavy workflows can automate well above half.
The safest automation pattern is surprisingly consistent: flexible language on the front end, rigid business rules underneath. The AI interprets what the caller means, while conventional software still controls what can actually be booked, changed, disclosed or approved.
Voice AI becomes much less dependable when the call requires judgment rather than workflow execution. Complaints, policy exceptions, emotional situations, disputed charges and high-stakes financial or regulatory decisions remain much harder than scheduling, status checks or routine account servicing.
AI assistance for human agents may currently be the strongest economic use case. Transcription, summaries, knowledge retrieval and after-call documentation can produce measurable productivity gains without asking the model to own the whole customer interaction.
Cost comparisons favor voice AI only when the conversation actually gets resolved. Cheap automated minutes are not useful if a failed call later creates a human escalation, a repeat call and an annoyed customer who has to explain everything again.
The market has already reached serious scale. Voice companies are reporting tens of millions of calls a month and hundreds of millions of dollars in recurring revenue, which shows that enterprises are buying the technology even though financial returns remain uneven across customer-service AI more broadly.
Natural conversation is getting close to infrastructure. Native speech models, lower latency, interruption handling, realtime translation and expressive synthetic voices are improving quickly, while workflow reliability, escalation and tool use increasingly determine whether a deployment feels good or falls apart.
The clearest boundary today is simple: voice AI is strong when it has to understand messy human speech and map it onto a known process. General-purpose artificial phone workers that can safely handle unusual situations and decide what a business should do next still have a lot to prove.

This market map, featured in our conversational AI market deck, highlights top companies and startups in the conversational AI market
What does “working” actually mean for voice AI today?
Voice AI is working today when it can finish a real job reliably enough that companies keep sending customers through it and can show a useful business result.
That is a higher bar than sounding human. Voice AI can now transcribe speech well, answer quickly and produce remarkably natural voices. Those improvements are real, but they tell us little about whether an agent can authenticate a caller, understand what the person wants, retrieve the right information, make a change in another system and recover gracefully when the conversation goes off script.
For this article, we looked for three levels of proof. First, can the technology hold a usable spoken conversation? That part is increasingly solved. Second, can it repeatedly complete a defined workflow in production? Several voice AI products can now do this. Third, can companies hand over large parts of human phone work without creating enough errors, escalations or unhappy customers to wipe out the savings? That third level is where the picture becomes much more uneven.
Fresh spending data shows why the distinction matters. In Gartner's recent survey of 199 customer-service leaders, AI spending had risen 38% while overall service budgets increased only 2%. Yet a separate Gartner survey of 1,303 senior leaders found that only 24% of customer-service and support leaders could demonstrate positive financial returns across their AI use cases.
Companies are clearly spending. The harder question is where they are already getting enough back.
Why has voice AI suddenly become much better?
Voice AI became genuinely more useful once speech recognition, language models, speech generation and latency all improved enough to work together during a live conversation.
Older phone automation forced people to speak like machines. Callers navigated menus, chose from a handful of intents or repeated themselves until the system recognized the expected phrase.
Generative models changed that interaction. Someone can now say, “I booked Friday, but my flight changed, so could you move everything to Saturday morning?” and the system can usually understand the underlying request without somebody programming that exact wording.
The newer improvement is speed. OpenAI's current realtime models, for example, can process spoken conversation, reason, call tools and respond while keeping delays short enough for normal back-and-forth. Its newer realtime generation also improved silence handling, recognition of letters and numbers, background-noise performance and interruption behavior. Those are very practical problems for phone agents.
The architecture is changing too. Earlier systems commonly ran speech through three separate steps: transcription, a language model and then text-to-speech. Native speech models increasingly work directly with audio. OpenAI's GPT-Live can listen while speaking, handle brief acknowledgements and stop when someone interrupts. Google and other model providers are moving in the same direction.
The result is easy to notice in a real call. Voice AI understands much freer language than traditional phone bots, while the awkward pauses and rigid turn-taking that used to expose automation immediately are shrinking.

As this chart shows, and as featured in our conversational AI market deck, search interest in conversational AI has increased sharply
Is voice AI actually replacing call-center agents now?
Voice AI is already taking meaningful call volume away from human agents, although complete replacement of customer-service teams is nowhere close to being proven.
The useful metric here is how many calls reach a successful outcome without a person taking over. Unfortunately, vendors use different terms such as containment, resolution and deflection, so comparisons need some caution.
Even with that caveat, the production numbers have become too large to dismiss.
Simplyhealth says its PolyAI assistant fully contains roughly 25% of calls. Quicken's deployment started around 5% containment and reached 21% after a year as the company expanded what the assistant could safely handle. An Atos deployment reduced human-agent call volume by 30%, with PolyAI reporting workload equivalent to as many as 95 full-time employees.
Other deployments go further. Decagon reports more than 50% voice deflection at mortgage servicer Valon. Landry's has previously reported containment above 80% in a hotel deployment, including one hotel where its assistant handled roughly 12,000 calls and generated more than 3,000 reservations in a month.
The interesting pattern is the spread. Voice AI does not appear to have one normal automation rate. A complicated operation may initially remove only 10% or 20% of calls from human queues, while repetitive reservation or scheduling workloads can go much higher.
| Deployment | Reported result | What it tells us |
|---|---|---|
| Simplyhealth / PolyAI | About 25% call containment | A meaningful chunk of routine insurance calls can already be automated |
| Quicken / PolyAI | About 5% to 21% containment over one year | Voice automation can expand as the company learns which calls are safe to hand over |
| Atos / PolyAI | 30% reduction in human-agent call volume | Automation can remove substantial workload rather than just answer FAQs |
| Valon / Decagon | More than 50% voice deflection | Even regulated financial workflows can support significant automation |
| Landry's / PolyAI | More than 80% in the cited hotel rollout | Narrow reservation-heavy calls can reach much higher automation rates |
If you want more recent data on this point, please see our latest conversational AI market report.
Which calls can voice AI actually handle by itself?
Voice AI works best today when customers can speak freely but the business only allows a limited number of possible actions.
Appointment scheduling is a good example. A patient can explain the request in dozens of ways, change their mind halfway through or ask whether Tuesday afternoon is possible. Behind that messy conversation, the system is still solving a fairly structured problem: identify the patient, inspect available slots, book one and confirm it.
Hotels have a similar shape. Callers ask about parking, breakfast, check-in time, room types, pets or availability in unpredictable language. The final action usually comes from a small set of known workflows.
The same pattern shows up in order tracking, basic account servicing, reservation changes, qualification of inbound leads, simple insurance questions, collections reminders and routing.
Modern voice AI is especially useful here because the conversation itself no longer has to follow a rigid script. The business rules still can.
A language model may interpret “Can you push my appointment until after lunch next Wednesday?” while a separate scheduling system decides what times actually exist. The model handles language; conventional software keeps control of the transaction.
Gartner's recent customer survey supports the broader shift toward this kind of action. Among people who already use generative AI, 58% said they had used it to complete a task for them. In B2B settings, the figure reached 74%.

This chart, included in our conversational AI market deck, shows annual VC investment in conversational AI startups
Where does voice AI still fail badly?
Voice AI still struggles with the conversations where the caller's situation is unusual, emotionally loaded or expensive to get wrong.
A request such as “move my booking to Saturday” maps neatly onto a known workflow. “Your cancellation made me miss my daughter's wedding and you charged me twice” requires a very different kind of judgment. The system may need to investigate what happened, understand policy, decide whether an exception is justified, handle an angry customer and possibly authorize compensation.
That is where current agents become much less dependable.
VoiceAgentBench gives us a useful technical view of the same problem. The benchmark contains more than 5,500 spoken tasks involving single and multiple tools, multi-turn conversations, multilingual requests and adversarial cases. Its researchers found clear weaknesses in contextual tool orchestration, multilingual generalization and robustness when spoken requests became more complicated.
Actual phone calls make things harder again. People mumble. Calls drop. A child shouts in the background. A customer changes languages halfway through. Names have unusual spellings. Someone says “fifteen” and the system hears “fifty.” A tiny speech-recognition error can become a serious problem when the value being captured is a bank amount, a flight date or an account number.
The downside is also uneven. Ninety-five routine calls can go perfectly and five failures can still make the workflow unacceptable if those five involve unauthorized payments or incorrect regulatory disclosures.
If you want more recent data on this point, please see our latest conversational AI market report.
Is AI more useful helping call-center agents than replacing them?
AI assistance for human call-center workers is currently one of the strongest and best-proven voice AI use cases.
The economics are easier because the AI does not have to become good enough to own the entire conversation. Transcription, summaries, knowledge retrieval and automatic note-taking can save time even when a human remains responsible for the customer.
We have unusually strong evidence here. A large study covering more than 5,000 customer-support agents found that access to a generative AI assistant increased productivity by about 14%, measured through issues resolved per hour. Less experienced and lower-performing workers gained much more than the strongest agents.
Chime offers a more specific voice example. According to an AWS case study, its AI system now summarizes customer calls automatically and saves more than 250,000 hours a year. Average handling time fell by 18 seconds per call, producing an estimated $700,000 in annual efficiency gains.
Similar gains show up elsewhere. Japanese retailer deployments using automatic transcription and summaries have reported removing minutes of administrative work after calls. JR West Customer Relations, for example, reported cutting average after-call work from 12 minutes 47 seconds to 6 minutes 23 seconds after adopting AI-generated summaries more broadly.

This chart, included in our conversational AI market deck, breaks down Cognigy's playbook in conversational AI
Are voice AI agents already cheaper than human phone agents?
Voice AI can already cost far less than a human agent on routine calls, provided the AI actually finishes the job.
The raw cost structure favors software. Human phone support carries wages, training, management, idle time, turnover and after-call work. An AI agent adds telephony, model inference, speech processing and platform fees, then handles many conversations at once.
Peak demand makes the difference even larger. A call center needs enough people to survive busy periods even if some of that capacity sits idle later. Software can add simultaneous conversations much more easily.
Underlying model costs are also falling while capability improves. Providers now offer smaller realtime models designed specifically for cheaper voice interactions, and speech transcription itself has become inexpensive enough that it is rarely the main cost of a call.
Production results show why companies care. The Atos deployment mentioned earlier reported that automation performed work equivalent to as many as 95 full-time employees at roughly half the cost. Other platforms report large reductions in support costs when voice is combined with chat and other automated channels.
But the correct unit to measure is a successfully resolved conversation.
If the AI spends four minutes talking to somebody, misunderstands the issue and eventually transfers the customer to a human who has to start again, the automation may have increased the cost. Poor containment can also increase repeat calls and customer frustration.
| Cost issue | Human-heavy support | Voice AI |
|---|---|---|
| Handling more calls at once | Usually needs more staffing | Concurrency can increase quickly |
| Nights and weekends | Requires staffing or outsourcing | Similar software economics throughout the day |
| Call summaries | Uses paid agent time | Can be generated automatically |
| Repetitive requests | Human labor repeats every time | Strong candidate for automation |
| Complex exceptions | Human can use judgment immediately | Usually needs escalation |
| Failed interaction | Human is already handling it | Can create both AI cost and human cost |
Are companies actually using voice AI at serious scale?
Voice AI has clearly moved beyond pilots: several suppliers now process tens of millions of conversations or generate hundreds of millions of dollars in recurring revenue.
ElevenLabs is the clearest financial example. The company finished 2025 above $330 million in annual recurring revenue and said it passed $500 million ARR during the first four months of 2026. ElevenLabs specifically said enterprise deployments of voice agents across support, sales, recruitment and marketing were helping drive the growth.
Dedicated phone-agent companies are smaller but already handling enormous traffic. Retell says it now powers more than 50 million realtime AI phone calls every month after reaching roughly $50 million ARR during 2025.
Parloa passed $50 million ARR during 2025 after roughly doubling from the previous year, according to company and investor disclosures. Its reported 150% net revenue retention is especially interesting because it suggests existing enterprise customers were spending considerably more after adopting the product.
Conversational-agent companies such as Decagon and Sierra are also moving heavily into voice alongside chat and other channels. Decagon added more than 100 enterprise customers during 2025 and subsequently raised money at a $4.5 billion valuation. Sierra entered 2026 with more than $150 million in ARR across its broader customer-service agent business.
Revenue does not tell us whether every deployment works. Still, companies do not build businesses of this size from weekend demos.

This chart, included in our conversational AI market deck, shows annual funding in conversational AI startups
Does voice AI really solve customer problems, or just keep people away from humans?
Some voice AI systems are clearly resolving real customer problems, but “deflection” numbers should still make us suspicious until we know what happened after the call.
A call that never reaches a person could mean the AI solved it. The caller could also have given up, switched channels or called again later.
The stronger case studies connect automation to an actual result.
Hunter Douglas says conversations fully handled by Decagon have generated more than $1 million in revenue. The company also found that customers interacting with its AI agents had 85% higher average order values than those who did not, although that comparison does not prove the AI caused the difference.
Quicken provides another useful pattern. Its automation rate rose from roughly 5% to 21% over a year. A slow increase like that tells us more than an extraordinary launch-day containment number. It suggests the company kept adding workflows after seeing what the agent could handle safely.
Customer research gives us the boundary. In Gartner's recent survey of 3,566 customers, half said generative AI made customer-service interactions easier. At the same time, 87% said access to a human remained essential when companies used AI.
Customers seem perfectly willing to use automation when it gets them to an answer or an action faster. Their patience disappears quickly when automation blocks the way to somebody who can fix a complicated problem.
If you want more recent data on this point, please see our latest conversational AI market report.
Has voice AI basically solved natural conversation now?
Voice AI has become dramatically more natural, and native speech-to-speech models are making conversations smoother, but the hard part is increasingly what the agent does after understanding the caller.
The improvement is easy to hear. Current systems can vary pacing and emotion, respond quickly, survive interruptions and avoid the exaggerated robotic cadence that made older assistants painful to use.
Older voice systems often converted speech into text, sent that text through a language model and generated audio from the answer. That pipeline works, but every handoff adds delay and can throw away information contained in the original audio.
Tone is a simple example. “Yeah, great” can communicate enthusiasm, irritation or sarcasm depending on how somebody says it. A transcript keeps the words and loses much of the rest.
Current native voice models can use more of that audio directly. OpenAI's GPT-Live architecture can listen and speak simultaneously, handle brief acknowledgements, interruptions and faster exchanges instead of forcing strict alternating turns. Current API models can also reason and call tools directly during voice sessions.
Latency has improved too. OpenAI said newer realtime models reduced p95 latency by at least 25% through caching improvements. A few hundred milliseconds genuinely change how a spoken interaction feels because humans notice strange pauses very quickly.
ElevenLabs has pushed expressive speech quality aggressively, while companies such as Deepgram and Cartesia compete around latency, transcription and speech generation. Convincing synthetic speech is becoming widely available infrastructure.
The bottleneck is increasingly workflow reliability. A natural-sounding agent still needs to access the right account, use tools correctly, obey business rules and know when to escalate. Fluent speech can even make a wrong answer sound more convincing.

This chart, included in our conversational AI market deck, compares the main business model options for conversational AI enterprise platforms
Are multilingual voice agents actually useful now?
Multilingual voice AI is already useful for translation, localization and selected support workflows, although quality still varies too much across languages to treat “supports 70 languages” as one uniform capability.
The underlying models have expanded quickly. OpenAI's current realtime translation model accepts speech in more than 70 input languages and translates into 13 output languages while the speaker talks. Google has also pushed realtime speech translation across a broad language set.
Dubbing provides some of the clearest commercial proof because the task is controlled and easy to measure. Companies can now turn one recording into many localized versions without hiring separate voice actors and studios for every market. That completely changes the economics for corporate training, product videos, education and creator content.
Live customer calls are tougher. Recognition has to survive local accents, names, noisy connections, slang and code-switching while the agent simultaneously understands the customer and performs a task.
Research still finds obvious gaps here. As discussed earlier, VoiceAgentBench found weaker performance when current speech agents had to generalize across several Indic languages while using tools and handling adversarial requests.
So multilingual voice AI has reached different levels of maturity depending on the job. Dubbing and controlled translation are already strong. Realtime translation is becoming genuinely useful. Autonomous customer service across many languages remains much more dependent on the specific language, model and workflow.
Is outbound voice AI useful, or will it mostly create more spam?
Outbound voice AI works for reminders, requested callbacks and existing-customer workflows, while mass automated cold calling has a much uglier long-term outlook.
The attraction is obvious. Once one software system can run hundreds of simultaneous calls, outbound calling becomes dramatically cheaper.
Some legitimate use cases fit that economics very well. Appointment reminders, delivery coordination, collections, renewals, requested callbacks and qualification after someone has already submitted a lead form all involve a pre-existing reason to contact the person.
Companies are already reporting meaningful results. Bland has published examples of customers using AI for lead qualification, mandatory disclosures and follow-up. MyPlanAdvocate, for instance, uses AI agents for inbound qualification and required disclosures, with Bland reporting $40 million in additional revenue associated with the broader workflow over five months.
Cold outreach creates a different equation. Making calls becomes cheaper precisely when the recipient's attention remains scarce. If thousands of companies use AI to multiply outbound volume, the result can quickly become intolerable.
Regulators have already reacted to that risk. The U.S. Federal Communications Commission has ruled that AI-generated voices fall under Telephone Consumer Protection Act restrictions covering artificial or prerecorded voices. Telemarketing uses therefore remain subject to the applicable consent requirements. Regulators are also paying close attention to voice cloning and impersonation scams.
If you want more recent data on this point, please see our latest conversational AI market report.

This chart, featured in our conversational AI market deck, illustrates revenue distribution by customer segment in the conversational AI market
Are consumer voice assistants finally becoming useful?
Consumer voice assistants are much smarter today than the old Alexa-and-Siri generation, but voice still looks like one interface for general AI rather than the interface that will replace screens and keyboards.
Amazon shows how dramatic the reset has been. The original Alexa worked well for timers, music, weather and smart-home commands while struggling with open-ended requests. Alexa+ now uses generative AI for broader conversations and multi-step requests and has moved beyond early access in the U.S.
Google is following a similar path by building voice into Gemini rather than maintaining a completely separate low-intelligence assistant. Gemini can use personal context from Google's ecosystem and lets users move between spoken and visual interaction. OpenAI has taken the same basic direction with ChatGPT Voice and GPT-Live.
That fixes the biggest weakness of the old assistants: the intelligence behind the microphone.
Usage behavior still gives us reasons to avoid a bigger conclusion. Voice is extremely convenient while driving, cooking, walking or working hands-free. Reading is faster for many dense information tasks. Screens are much easier for comparing ten products, checking a table or scanning several possible answers. Text also works quietly in offices and public places.
Where is voice AI creating value outside phone calls?
Some of today's most useful voice AI systems simply listen to human conversations and turn them into usable work.
Healthcare may be the strongest example. Ambient clinical documentation records the doctor-patient conversation and drafts the clinical note automatically. The doctor reviews the output rather than typing the visit from scratch.
Companies such as Abridge, Microsoft and others have pushed this workflow rapidly into large health systems. Microsoft has reported that its clinical documentation products can save several minutes per patient encounter in some deployments. Abridge has expanded across hundreds of enterprise health systems and is moving beyond transcription into richer clinical context and workflow support.
Meeting software follows the same idea. Otter, Fireflies, Zoom, Microsoft Teams, Google Meet and specialized hardware such as Plaud can already transcribe discussions, summarize them and pull out action items. None of this requires an AI agent to replace the people in the meeting.
Dubbing and localization belong in this category too. Here the speech becomes the raw material for a different voice output, allowing one recording to be reused cheaply across several languages.
| Voice AI use case | How mature it looks today | Why it works |
|---|---|---|
| Call transcription and summaries | High | The task is repetitive and easy to check |
| Clinical documentation | High | Most of the information already exists in the spoken conversation |
| Meeting notes and action items | High | Humans remain responsible for the actual decisions |
| Dubbing and localization | High | AI dramatically lowers the cost of producing additional language versions |
| Live translation | Improving quickly | New realtime models reduce delay and preserve more spoken information |
| Open-ended autonomous decision-making | Uneven | Errors become much more expensive once the AI controls the outcome |

This chart, included in our conversational AI market deck, shows how AI chatbot platform technology has evolved over time
Can companies trust voice AI with healthcare, banking and other sensitive calls?
Voice AI is already operating in regulated industries, but companies are keeping tight controls around what the model can access, say and change.
We can see this in real deployments. Valon uses voice automation in mortgage servicing. Chime uses AI across financial customer service. Healthcare voice products operate inside environments covered by strict privacy requirements. Simplyhealth designed its voice assistant to recognize situations where vulnerable callers should be moved to an appropriate human team.
The model is only one piece of these systems.
A serious regulated deployment needs identity checks, access controls, logs, approved disclosures, clear transaction rules, monitoring and escalation. Teams increasingly run agents through simulated conversations before changing production behavior so they can catch failures without testing them on customers.
Voice cloning creates another issue. The sound of somebody's voice can no longer carry the same weight as proof of identity. A short recording can be enough to imitate a relative, an executive or another trusted person convincingly enough to fool some listeners.
Providers and regulators are responding. OpenAI now embeds SynthID watermarking in supported GPT-Live audio and offers tools for checking provenance. The FTC has repeatedly warned about voice-cloning scams, while also pointing out that no single detection technology can solve the problem.
Will people actually put up with talking to AI?
People will talk to AI when it solves something quickly, but they still want an obvious way to reach a human when the conversation goes wrong.
Gartner's latest customer research captures the trade-off unusually well. Half of surveyed customers said generative AI had made their service interactions easier. Yet 87% said companies using generative AI should provide access to a human agent.
There is no real contradiction there.
Someone checking whether a package arrives tomorrow may happily take an AI answer in 20 seconds instead of waiting on hold. The reaction changes when a customer disputes a large charge, has already explained the problem twice and cannot escape the automated system.
Poor experiences can also poison future adoption. In another recent Gartner finding, only 27% of customers said they would be willing to try a company chatbot again after a negative experience. That study covered customer-service chatbots rather than voice alone, but the implication for automated conversations is hard to miss.
This is why successful voice AI deployments increasingly include clean transfers rather than treating escalation as a failure metric.
If you want more recent data on this point, please see our latest conversational AI market report.

In our conversational AI market deck, we identify pain points entrepreneurs should prioritize
So what is actually working in voice AI right now?
Voice AI is genuinely working today in bounded, high-volume jobs where the conversation can be messy but the possible business actions remain fairly controlled.
Routine inbound calls are already one clear winner. Scheduling, reservations, status checks, routing, FAQs, simple account servicing and qualification can remove meaningful volumes from human queues. Depending on the workflow, production automation rates can range from the low double digits to well above half of calls.
Agent assistance looks even stronger. Call transcription, summaries, knowledge retrieval and automatic documentation already save measurable time without asking the model to own every customer decision.
Voice as an input layer is also mature enough to matter. Healthcare documentation, meeting capture, dubbing and translation are processing real workloads because speech contains information that previously had to be typed, summarized or recorded again manually.
The industry around these products now has real scale. ElevenLabs has passed $500 million in ARR. Dedicated platforms such as Retell process tens of millions of AI calls each month. Enterprise customer-service companies are pushing voice into increasingly large production deployments. At the same time, the recent Gartner data showing positive AI returns at only 24% of customer-service organizations is a useful warning against treating adoption as proof of success.
The boundary we keep finding is remarkably consistent.
Voice AI works very well when it needs to understand what somebody says and map that request onto a known workflow. Performance becomes much shakier when the AI has to interpret an unusual situation, exercise judgment and decide what the business should do next.
That is why the companies getting the clearest value today are automating specific calls, removing administrative work around human calls and gradually giving the AI more authority as the workflow proves itself.
General-purpose artificial phone workers still have a lot to prove. The underlying voice layer is already useful enough to be a real product category.
OUR METHODOLOGY
This analysis tests what “working” actually means for voice AI today by separating conversational quality from production reliability, business outcomes, commercial scale, customer acceptance and the ability to operate safely in higher-stakes workflows.
We use three levels of proof throughout the article: whether a system can hold a usable spoken conversation, whether it can repeatedly complete a defined workflow in production, and whether companies can hand over meaningful work without errors, escalations or customer frustration wiping out the gains.
We gave the most weight to repeated production evidence. Named deployments with containment, deflection, reduced call volume, productivity, handling-time or customer-outcome data tell us more about practical usefulness than a demo. Benchmarks are used to expose technical weaknesses, while ARR and call volume are used to establish market scale rather than prove that every interaction works.
We did not force different metrics into one score. Containment, deflection, resolution, productivity gains, customer outcomes, recurring revenue, monthly call volume and benchmark performance answer different questions, so each is used only where it adds direct evidence.
Freshness matters especially for realtime model capability, enterprise adoption, commercial scale and customer behavior. Older research is kept where it still provides unusually strong measured evidence, particularly the large study of more than 5,000 customer-support agents.
Key sources used for this analysis include OpenAI on current realtime voice models, OpenAI on GPT-Live, Google DeepMind on Gemini Audio, Gartner on customer-service AI spending, Gartner on customer acceptance and human access, PolyAI's Quicken deployment, PolyAI's Atos deployment, Decagon's Valon deployment, AWS on Chime's call-summary deployment, the NBER study on generative AI and customer-support productivity, VoiceAgentBench, ElevenLabs on commercial scale, Retell AI on monthly call volume, the FCC ruling on AI-generated voices and the TCPA, and the FTC on voice-cloning risks.

This chart, included in our conversational AI market deck, illustrates revenue distribution by region across Europe, Asia, North America, Africa, and South America in the conversational AI market
Related blog posts
- Voice AI: what is getting real adoption now?
- What are the latest funding developments in conversational AI?
- Conversational AI: what are the top startups?
- The main fundraising trends in conversational AI
- How strong is fundraising in the conversational AI market right now?
Who is the author of this content?
NEW MARKET PITCH TEAM
We track new markets so founders and investors can move fasterWe build living "market pitch" documents for emerging markets: AI, synthetic biology, new proteins, and more. Instead of outdated PDFs or hallucinated LLM answers, our clients get a clean, visual, always-updated view of what's really happening: key players, deals, regulations, and signals that matter. Learn more about us.