Healthcare AI: what is actually working now?

Last updated: 11 September 2026
market research pitch 2026 statistics healthcare AI market

In our healthcare AI market deck, you will find everything you need to understand the market

SUMMARY

Healthcare AI is actually working now in ambient documentation, AI-assisted mammography and colonoscopy, selected radiology and pathology tasks, narrow autonomous retinal screening, and parts of medical coding; broad clinical judgment remains much less proven.

The clearest dividing line is job design. Healthcare AI looks strongest when it has a bounded task, a recognizable input, a measurable output and a clinician who can quickly check what it did.

Ambient scribes have become the clearest generative-AI success because hospitals are no longer measuring only pilot satisfaction. Thousands of clinicians are using them repeatedly, some health systems have generated millions of notes, and independent studies now show real reductions in documentation burden.

The clinical winners are also unusually specific. AI-assisted mammography has increased cancer detection while cutting reading workload, and AI colonoscopy has improved adenoma detection across dozens of randomized trials involving tens of thousands of patients.

Speed and outcomes need to be separated. Stroke AI can get urgent cases in front of treatment teams faster, yet pooled evidence still has not shown a convincing improvement in mortality or functional independence.

Autonomous diagnosis is real, but only inside narrow lanes. Diabetic-retinopathy screening works because the image, disease threshold, referral rule and next step can all be tightly defined inside primary care.

Some of healthcare AI's biggest commercial successes are happening away from the bedside. Medical coding has reached large deployments and strong customer retention even though independent standards for comparing vendor accuracy and ROI still lag behind the market.

Text generation by itself is a weaker advantage than it first appears. AI-drafted patient messages can reduce mental effort, but several real-world studies found little or no improvement in reply time, and clinicians often still spend time checking the draft.

General clinical LLMs remain the awkward middle. They can perform impressively on packaged cases, yet randomized physician studies are mixed and a pragmatic trial involving almost 10,000 patients did not improve the main treatment-failure outcome.

The pattern across healthcare AI is now fairly sharp: tools that remove repetitive work or improve detection inside a controlled workflow are crossing into routine use, while claims about broad medical judgment still run ahead of the evidence. AI drug discovery sits between those worlds, with credible candidates reaching randomized human trials but approved, repeatedly successful medicines still the harder test.

Market map chart showing top companies and startups in the healthcare AI market

This market map, featured in our healthcare AI market deck, highlights top companies and startups in the healthcare AI market

Why does healthcare AI suddenly feel more real?

Healthcare AI feels more real today because several tools have moved from impressive pilots into boring, repeated, everyday use.

The clearest shift has happened inside hospitals rather than in consumer-facing “AI doctor” products. AI is drafting clinical notes, reading scans alongside specialists, flagging polyps during colonoscopies, screening diabetic eyes and automating parts of medical coding. Those jobs already existed, so hospitals can see fairly quickly whether the software saves time, catches something useful or reduces costs.

Ambient documentation shows how far things have moved. UCHealth expanded Abridge after testing it with 250 providers, and more than 2,300 clinicians were using it by early 2026. Geisinger passed 1,000 clinicians after roughly ten months of scaling. BJC Health and WashU Medicine have since expanded access from an initial group of around 450 clinicians to roughly 4,000.

Sutter Health offers an even better measure of repeated use. The health system says its clinicians had generated five million notes with Abridge by mid-2026. Five million completed notes tell us much more about product maturity than another hospital announcing a pilot.

Clinical AI has been building quietly for longer. The FDA's public database now contains well over 1,000 AI-enabled medical-device authorizations, with radiology still dominating the list. Generative AI arrived on top of that existing base and opened another large category around clinical language and paperwork.

So when we ask what is working in healthcare AI today, the useful question has changed. The interesting areas are the ones where repeated use is already producing a measurable benefit.

What should count as healthcare AI actually “working”?

Healthcare AI is working when it repeatedly makes a real healthcare job better enough that clinicians or hospitals keep using it.

The standard changes with the job.

For an AI scribe, saving documentation time and reducing after-hours work is a meaningful result. Asking whether an ambient scribe lowers mortality would make little sense because that is not the job being bought.

Diagnostic AI needs a tougher test. If an algorithm helps radiologists find more cancers, helps endoscopists miss fewer adenomas or allows diabetic-eye screening to happen reliably in primary care, we have direct evidence that the clinical process improved.

The bar rises again for general clinical decision-making. A model scoring highly on medical questions tells us that the model knows a lot of medicine. Once doctors use the model with real patients, we want to know whether decisions improve, mistakes fall and patients actually do better.

That distinction clears up a lot of the hype. Healthcare AI currently looks very strong under some definitions of “working” and surprisingly weak under others.

Healthcare AI job What would convince us it works? Where the evidence is today
Clinical documentation Less documentation work and sustained use Strong
Medical imaging More useful findings, fewer misses or faster care Strong in selected uses
Procedure assistance Better detection during real procedures Strong for colonoscopy
Workflow triage Faster action on urgent cases Strong for speed, weaker for outcomes
Clinical decision support Better decisions and patient outcomes Mixed
Autonomous diagnosis Reliable care without routine specialist interpretation Proven only in narrow tasks

If you want more recent data on this point, please see our latest healthcare AI market report.

Google Trends chart showing rising interest in AI for healthcare

As this chart shows, and as featured in our healthcare AI market deck, search interest in healthcare AI has grown rapidly

Are AI scribes actually saving doctors time now?

Ambient AI scribes are currently the most convincing generative-AI product in healthcare.

We now have enough independent and real-world evidence to move beyond vendor anecdotes.

A multicenter JAMA Network Open study followed 263 clinicians using Abridge across six health systems. The share reporting burnout fell from 51.9% before adoption to 38.8% after 30 days. Clinicians also reported less after-hours documentation and lower cognitive load.

A much larger Providence study published in 2026 looked at 1,547 active users of Microsoft's DAX across a health system serving more than two million patients. Median time spent in notes fell from 7.1 minutes per appointment before AI use to 6.1 minutes afterward. More importantly, the researchers found a sustained decline in after-hours documentation over time.

The productivity story is more modest. Providence saw no meaningful increase in appointments per day, and the initial increase in relative-value units did not continue growing. Scribes currently look better at giving clinicians time back than at turning doctors into dramatically higher-throughput workers.

Randomized studies point in the same general direction while showing that product quality matters. In a trial involving 238 outpatient physicians across 14 specialties, Nabla reduced note-writing time by 9.5% relative to usual care. DAX Copilot produced a much smaller change that did not reach statistical significance.

The latest hospital behavior is perhaps the strongest commercial evidence. BJC Health and WashU Medicine expanded from roughly 450 clinicians to around 4,000 after evaluating documentation time and physician experience. Sutter moved to enterprise-wide access. HonorHealth even skipped a traditional pilot after several years of testing the category.

Hospitals rarely expand unwanted software across thousands of doctors for fun. Ambient documentation has crossed the line into a real healthcare product category.

Is radiology AI genuinely useful, or are there simply hundreds of FDA-cleared tools?

Radiology AI is genuinely useful today, although the huge number of cleared products makes the market look more clinically proven than it really is.

Radiology dominates medical AI for a practical reason. Scans are already digital, many useful tasks can be narrowly defined, and a radiologist remains in the workflow to check what the algorithm finds.

The FDA's continuously updated AI-device list still shows radiology taking the largest share of new authorizations. Products cover tasks such as detecting suspected brain bleeds, pulmonary embolisms and fractures, measuring anatomy, segmenting lesions and helping prioritize urgent studies.

Regulatory authorization answers a fairly narrow question. Most AI devices historically entered the U.S. market through the 510(k) pathway, where the manufacturer generally shows that the product is substantially equivalent to an existing legally marketed device. That tells us far less about whether a hospital subsequently saves time or patients get better care.

We get a better picture by looking at individual uses. Breast-cancer screening now has large randomized data. Stroke imaging shows repeatable reductions in treatment delays. Other radiology products mainly have retrospective validation, single-hospital studies or workflow evidence.

Radiology AI as a whole has clearly arrived. The strength of the evidence still varies enormously from one algorithm to another.

If you want more recent data on this point, please see our latest healthcare AI market report.

Chart showing annual VC investment in healthcare AI startups

This chart, featured in our healthcare AI market deck, shows annual VC investment in healthcare AI startups

Is AI finding breast cancers that radiologists would otherwise miss?

AI-assisted mammography is one of the strongest clinical AI use cases we can point to today.

The Swedish MASAI trial randomized roughly 106,000 women between AI-supported screening and standard double reading. AI-supported screening detected 338 cancers among 53,043 women, compared with 262 among 52,872 women receiving standard screening.

That works out to 6.4 cancers detected per 1,000 women versus 5.0 per 1,000, an increase of roughly 29%.

The workload result was almost as striking. The AI group required 61,248 screen readings, compared with 109,692 in the control group, cutting radiologist reading workload by about 44%.

Follow-up analysis strengthened the clinical case. Sensitivity reached 80.5% with AI-supported screening versus 73.8% with standard double reading, while specificity remained 98.5% in both groups. Researchers also found no concerning increase in interval cancers.

Those results answer a much harder question than whether an algorithm can classify mammograms accurately. AI was inserted into a real national screening workflow, more cancers were found and radiologists performed tens of thousands fewer readings.

We still need longer follow-up to measure things such as mortality and lifetime cost-effectiveness. For the immediate job of helping mammography programs find cancer efficiently, the evidence is already unusually convincing.

MASAI screening result AI-supported screening Standard screening
Women analysed 53,043 52,872
Cancers detected 338 262
Cancers per 1,000 screens 6.4 5.0
Sensitivity 80.5% 73.8%
Specificity 98.5% 98.5%
Mammogram readings 61,248 109,692

Does AI actually help doctors find more precancerous polyps during colonoscopies?

AI-assisted colonoscopy clearly helps endoscopists find more adenomas, and the result now appears across dozens of randomized trials.

The technology is simple from the doctor's point of view. Computer vision watches the live colonoscopy feed and highlights areas that could contain a polyp. The endoscopist decides whether to inspect or remove it.

An updated 2026 meta-analysis pooled 46 randomized controlled trials covering 37,206 people. AI-assisted colonoscopy increased the adenoma detection rate by 22% and cut the adenoma miss rate by roughly 47%. Detection of sessile serrated lesions rose by around 25%.

Those figures are especially useful because the result survives across many trials rather than depending on one hospital or one unusually strong research team.

The harder question is whether the extra lesions found eventually translate into fewer colorectal cancers and deaths. That takes much longer to measure. Adenoma detection still matters today because it is one of the established quality indicators for colonoscopy and is associated with future interval-cancer risk.

Colonoscopy gives us another case where healthcare AI has moved beyond a benchmark. Tens of thousands of randomized procedures show that doctors using AI find abnormalities they would otherwise miss.

Chart showing Tempus AI’s strategy in the healthcare AI market

This chart, featured in our healthcare AI market deck, looks at Tempus AI’s strategy in healthcare AI

Is stroke AI actually helping patients, or just speeding up the hospital?

Stroke AI is clearly making emergency stroke workflows faster, while the evidence for better final patient outcomes remains unconvincing.

That gap matters because stroke is one of the most heavily marketed clinical AI categories.

Systems such as Viz.ai analyse brain imaging for suspected large-vessel occlusions and automatically alert stroke teams. Faster alerts can matter a lot when patients may need mechanical thrombectomy.

A systematic review covering 12 studies and 15,595 patients found shorter CT-to-treatment time, shorter door-to-groin-puncture time, faster recanalization and shorter transfer times after Viz.ai implementation.

But when researchers looked at outcomes such as mortality and functional independence, the pooled results did not show a statistically significant improvement.

So stroke AI deserves credit for what it has actually proved: hospitals can identify likely urgent cases and assemble treatment teams faster.

Whether those saved minutes consistently translate into more patients walking out of hospital independently is still open. Stroke AI is useful operational technology today, with a clinical-outcome claim that remains ahead of the evidence.

If you want more recent data on this point, please see our latest healthcare AI market report.

Can healthcare AI diagnose anything on its own yet?

Autonomous healthcare AI already handles a few narrow diagnostic jobs, with diabetic-retinopathy screening providing the clearest example.

Systems such as LumineticsCore and EyeArt can analyse retinal photographs and return a screening result without requiring an ophthalmologist to interpret every image.

That changes something practical. A person with diabetes can be screened during a primary-care visit instead of completing one appointment and then successfully navigating a separate specialist visit just to find out whether referral is needed.

The diagnostic performance has been good enough for FDA authorization and clinical deployment. Recent research has also started looking beyond algorithm accuracy at what happens after primary-care implementation.

A 2026 Johns Hopkins study examined 3,745 adults with diabetes who entered eye care either through conventional primary-care referral or after autonomous AI screening. After researchers adjusted for clinical and social differences between patients, AI-assisted screening was associated with higher presentation to specialist eye care among Black patients, a group that historically has lower annual screening rates and a higher risk of presenting with advanced diabetic retinopathy.

That is a useful evolution in the evidence. We are beginning to see whether autonomous diagnosis closes gaps in care rather than simply returning an accurate result on a camera screen.

The scope is still very narrow. Eye screening works because the input, disease threshold, referral rule and next step can all be tightly defined. General diagnosis across hundreds of possible conditions is a completely different problem.

Chart showing the projected CAGR of the healthcare AI market

This chart, featured in our healthcare AI market deck, shows annual funding in healthcare AI startups

Is pathology AI doing useful work in real labs?

Pathology AI is starting to earn its place in real laboratories by helping pathologists reduce unnecessary tests and feel more confident about difficult slides.

The prospective CONFIDENT-P study tested Paige Prostate Detect during actual prostate-biopsy assessment. Researchers included 82 patients and 237 whole-slide images.

When pathologists had AI assistance, the amount of immunohistochemistry required per detected cancer fell substantially. The relative risk was 0.55 at the patient level and 0.41 at the slide level.

Pathologists also reported being confident or highly confident in 80% of AI-assisted diagnoses, compared with 56% under the conventional workflow.

One result keeps us from overselling it: reading speed did not improve. Median assessment took 139 seconds with AI and 112 seconds in the control workflow, a difference that was not statistically significant.

The useful outcome here is more specific. AI helped pathologists avoid extra staining and increased diagnostic confidence. In the trial alone, investigators calculated €1,700 in reduced immunohistochemistry costs.

Pathology AI still faces a practical bottleneck that radiology passed years ago. Many pathology departments first need to digitize glass slides at scale. Once that infrastructure exists, the evidence suggests that selected second-reader tools can already pay their way.

Is AI medical coding quietly becoming one of healthcare AI's biggest successes?

AI medical coding is already a serious commercial healthcare application, although we have much better evidence for adoption than for the exact performance numbers vendors advertise.

Coding fits automation unusually well. Every hospital produces large volumes of documentation, each encounter eventually needs billing codes, trained coders are expensive and coding delays can slow reimbursement.

CodaMetrix gives us a useful view of how large the category has become. The company says customers using its coding technology represent about $180 billion in annual net patient revenue, equivalent to roughly 12% of U.S. hospital net patient revenue.

Customer behavior also looks sticky. In KLAS research cited when CodaMetrix received the inaugural 2026 Best in KLAS award for autonomous coding, 93% of interviewed customers said they would buy the product again and 93% said it formed part of their long-term plans.

We should be more cautious with headline accuracy claims. CodaMetrix advertises figures such as 98% average coding accuracy, a 70% reduction in manual coding work and coding-related denial reductions of up to 60%. Those numbers largely come from the company and its customer programs. The company itself recently created a health-system council partly because the industry still lacks a consistent definition for claims such as “95% coding accuracy.”

That tells us something about the maturity of the market. Hospitals are clearly buying autonomous coding and keeping it. Independent measurement standards are still catching up with commercial adoption.

For now, we can confidently call coding AI an operational success. We should avoid pretending that every vendor ROI percentage has been independently established.

Chart comparing business model options for ambient AI companies

This chart, featured in our healthcare AI market deck, compares the main business model options for ambient AI companies

Are AI-written patient messages actually clearing doctors' inboxes faster?

AI-written patient messages currently help with drafting and mental effort more than they help clinicians empty the inbox faster.

Several real-world studies have reached versions of the same conclusion.

At Stanford, 162 clinicians received AI-generated draft replies inside the electronic health record. Doctors and other clinicians reported improvements in workload and work exhaustion, but researchers did not find meaningful reductions in reading, writing or total reply time.

A randomized study at UC San Diego was even less flattering on pure speed. AI drafts failed to shorten physician reply time and increased the amount of time doctors spent reading messages. Clinicians still liked starting from a draft and often felt that the generated replies sounded more empathetic.

At UCHealth, AI generated 21,323 draft messages across nine clinics, yet staff ultimately used only 2,596 of them, around 12%. Nurses were substantially more enthusiastic than physicians and advanced-practice clinicians.

Patients seem reasonably comfortable with the idea when a clinician remains responsible for the final response. Studies of patient attitudes generally find decent acceptance, with some decline in satisfaction when people are explicitly told that AI drafted the message.

The pattern is pretty clear now. Generating text is easy. Producing the exact response a clinician is comfortable sending without much checking is harder.

That puts inbox AI below ambient documentation in our ranking. It is useful today, though the big time-saving story has yet to appear.

If you want more recent data on this point, please see our latest healthcare AI market report.

Does giving doctors ChatGPT-style AI actually improve medical decisions?

General-purpose clinical LLMs can sometimes make doctors better, but current randomized evidence is too mixed to call broad AI decision support a proven clinical success.

This is where the gap between impressive model demos and real healthcare becomes especially obvious.

One randomized trial gave 50 physicians access to GPT-4 while they worked through diagnostic cases. Doctors using GPT-4 scored 76% on the diagnostic-reasoning assessment, compared with 74% among physicians using ordinary resources. The difference was not statistically significant, even though GPT-4 performing the cases by itself scored very well.

Another randomized trial involving 92 physicians produced a stronger result. Doctors with GPT-4 access scored 6.5 percentage points higher on simulated patient-management cases, although they spent roughly two extra minutes on each case.

The hardest test so far came from real primary care rather than simulated cases. A pragmatic randomized trial across 16 facilities in Kenya included 9,691 patients and 103 clinical officers. Treatment failure occurred in 2.2% of patients receiving LLM-assisted care and 2.0% under usual care. After adjustment, researchers found no statistically significant difference.

The Kenya study also found no serious safety signal linked to the AI, which is reassuring. “Safe and no better on the main outcome” still leaves us a long way from a general AI copilot that reliably improves medicine.

These studies expose a recurring problem. A model can solve a case well when the whole problem is neatly packaged for it. A busy clinician has to decide when to ask the AI, what information to provide, which answer to trust and how much time to spend checking it.

The next breakthrough in clinical LLMs may depend as much on workflow design as on a smarter underlying model.

Clinical LLM study Setting Result
50 physicians Diagnostic cases 76% with GPT-4 vs 74% with usual resources; no significant gain
92 physicians Simulated patient management AI-assisted doctors scored 6.5 points higher
9,691 patients in Kenya Real primary care No significant reduction in treatment failure
Chart illustrating how revenue is distributed across customer segments in the healthcare AI market

This chart, featured in our healthcare AI market deck, illustrates how revenue is distributed across customer segments in the healthcare AI market

Is AI drug discovery producing real drugs yet?

AI drug discovery has reached real patients and randomized trials, which is genuine progress, but the industry still has to prove that AI can produce successful medicines repeatedly.

Rentosertib is the most useful example.

Insilico Medicine used AI in both target identification and molecule generation for the drug, which is being developed for idiopathic pulmonary fibrosis. The program has already moved through preclinical development, Phase 1 and a randomized Phase 2a trial.

The Phase 2a study included 71 patients receiving three rentosertib dosing regimens or placebo over 12 weeks. The primary aim was safety, and treatment-emergent adverse events were broadly similar across treatment and placebo groups. Exploratory lung-function results also showed encouraging dose-related changes.

That is a meaningful milestone for AI drug discovery. A molecule created through an AI-heavy discovery process has survived long enough to produce randomized human data.

Pharma sets a much tougher finish line, though. Phase 2 success does not guarantee Phase 3 success, approval, reimbursement or meaningful use in patients. Most experimental drugs fail somewhere along that road no matter how cleverly the molecule was discovered.

AI drug discovery has now proved that it can generate credible clinical candidates. We still need approved drugs and a broader portfolio before we can say that AI has materially improved pharmaceutical R&D productivity.

Why does healthcare AI work so much better on narrow jobs?

Healthcare AI works best today when the software has one clear job, one recognizable input and an easy way for a human to check the output.

The strongest examples in this article share that shape.

Mammography AI studies one type of image. Colonoscopy AI highlights possible lesions in a live video. Stroke software looks for specific imaging patterns and alerts the right team. Pathology AI points a pathologist toward suspicious tissue. Ambient AI listens to a consultation and drafts a note for review.

Broad clinical reasoning creates a much messier problem. The system may need symptoms, medical history, laboratory values, scans, previous treatments, guidelines, patient preferences and information that was never recorded properly in the first place. A plausible mistake can then change treatment rather than simply create a sentence that a doctor edits.

The regulatory data reflect that difference. Medical imaging has dominated FDA-authorized AI devices for years because images provide a relatively structured input and specialists already have established review workflows around them.

Generative AI is gradually pushing into less structured territory. Abridge, for example, now says more than 300 partner health systems have adopted its newer context-aware clinical decision-support technology since its launch earlier in 2026. That is a fresh sign that vendors want to move from recording what happened during a visit toward helping clinicians decide what to do next.

We should watch that transition closely because it moves AI into a much harder category. Millions of successfully generated notes tell us that documentation works. Clinical recommendations need their own evidence.

Chart showing how symptom checker app technology has evolved over time

This chart, featured in our healthcare AI market deck, shows how symptom checker app technology has evolved over time

So what healthcare AI is actually working now?

Healthcare AI is working today in several important areas, and the winners are much easier to name than they were a few years ago.

Ambient documentation is the clearest generative-AI success. We now have large deployments, millions of notes produced at individual health systems and studies involving hundreds or more than a thousand clinicians showing less documentation burden. The improvement in raw physician productivity looks smaller, but clinicians clearly value getting some of their documentation time back.

AI-assisted screening has the strongest clinical evidence. Mammography AI increased cancer detection by roughly 29% in a randomized trial of around 106,000 women while cutting reading workload sharply. AI colonoscopy has improved adenoma detection across 46 randomized trials and more than 37,000 participants.

Radiology triage and pathology assistance also work for selected jobs. Stroke software reliably speeds up emergency workflows, although better patient outcomes have yet to show up clearly in pooled data. Prostate-pathology AI has reduced the need for additional staining while increasing pathologist confidence.

Autonomous AI has made real progress where the task can be tightly contained. Diabetic-retinopathy screening shows that an algorithm can already make a defined screening decision in primary care without an ophthalmologist interpreting every image.

Behind the scenes, medical coding has become one of AI's bigger operational businesses in healthcare. The deployment evidence is strong, while independent standards for measuring accuracy and ROI remain immature.

The weaker areas are also becoming easier to identify. AI-generated inbox replies have yet to produce large time savings. General LLM decision support gives doctors mixed results and failed to improve the main patient outcome in a recent trial involving almost 10,000 people. AI drug discovery has reached randomized Phase 2 testing, though successful approved medicines remain the next real test.

Our final judgment is quite sharp. Healthcare AI already works when we give it a bounded, repetitive job with measurable output and sensible human supervision. Documentation, selected medical-image tasks, colonoscopy detection, narrow autonomous screening and parts of revenue-cycle work have crossed that threshold.

The evidence drops quickly as the job expands toward broad medical judgment. An all-purpose AI doctor remains far ahead of what today's clinical evidence supports.

Healthcare AI use case Our judgment now What has actually been proved
Ambient clinical documentation Working at scale High usage plus measurable reductions in documentation burden
AI mammography Strongly working More cancers detected with much lower reading workload
AI colonoscopy Strongly working More adenomas found across dozens of randomized trials
Radiology detection and triage Working in selected tasks Detection and workflow gains; patient outcomes vary
Pathology assistance Working in selected tasks Less ancillary testing and higher diagnostic confidence
Autonomous retinal screening Working for a narrow diagnosis Independent screening can function inside primary care
Medical coding Commercially working Large deployments and strong customer retention
AI patient messaging Useful but limited Easier drafting without convincing time savings
General LLM decision support Still unproven Mixed physician results and no clear patient-outcome gain
AI drug discovery Real but early AI-designed candidates can reach randomized human trials
General autonomous AI doctor Unsupported for routine care Broad independent diagnosis and treatment still lack comparable clinical proof

If you want more recent data on this point, please see our latest healthcare AI market report.

OUR METHODOLOGY

This analysis tests Healthcare AI: what is actually working now? by separating healthcare AI into the main jobs already being used or seriously tested, then assessing each one with the kind of evidence that fits the job. We gave the most weight to recent randomized and prospective clinical studies, large real-world implementations, repeated hospital use, regulatory records and credible adoption data.

Clinical claims were held to a higher bar than workflow claims. For documentation, coding and other operational tools, sustained use and measurable reductions in work can be enough to show that a product is working. For diagnosis, treatment or clinical decision support, we looked for evidence that clinicians detect more useful findings, make better decisions, act faster when that speed matters, or improve patient outcomes.

We avoided letting one strong study or one large deployment define an entire category. The final judgments come from aggregating the strongest evidence within each area, checking whether results hold across larger populations, multiple sites or several studies, and separating what has been directly demonstrated from what is still being inferred.

Regulatory authorization was used to establish that technologies have genuinely entered clinical use, especially in imaging and autonomous screening, but it was not treated as proof of real-world benefit by itself. Company data were useful for adoption and commercial scale, while independent clinical evidence carried more weight for claims about accuracy, effectiveness and patient outcomes.

Key sources include the FDA's AI-enabled medical-device database, the JAMA Network Open multicenter study of ambient AI scribes, the Providence real-world evaluation of ambient documentation, the randomized trial of Nabla and DAX ambient scribes, the MASAI randomized mammography trial, the MASAI follow-up on sensitivity, specificity and interval cancers, the meta-analysis of 46 randomized AI-colonoscopy trials, and the systematic review and meta-analysis of Viz.ai stroke implementations.

Other important sources include the FDA De Novo record for autonomous diabetic-retinopathy screening, the Johns Hopkins real-world study of autonomous retinal screening, the CONFIDENT-P prospective pathology trial, KLAS research on autonomous coding, the UC San Diego study of AI-drafted patient messages, the randomized GPT-4 diagnostic-reasoning trial, the pragmatic Kenya trial of LLM-assisted primary care, and the randomized Phase 2a study of rentosertib.

Table scoring and prioritizing the main pain points faced by companies in the healthcare AI market

In our healthcare AI market deck, we identify pain points entrepreneurs should prioritize

Who is the author of this content?

NEW MARKET PITCH TEAM

We track new markets so founders and investors can move faster

We build living "market pitch" documents for emerging markets: AI, synthetic biology, new proteins, and more. Instead of outdated PDFs or hallucinated LLM answers, our clients get a clean, visual, always-updated view of what's really happening: key players, deals, regulations, and signals that matter. Learn more about us.

Back to blog