Healthcare AI: what are the biggest unsolved problems?

In our healthcare AI market deck, you will find everything you need to understand the market
SUMMARY
Healthcare AI’s biggest unsolved problems are real-world clinical reliability, uncertainty under changing conditions, and human-AI interaction. The technology is already widely used, but the hardest question is whether its performance survives real patients, messy data, changing hospitals, and imperfect human use.
Healthcare AI adoption is now well ahead of the clinical evidence base. Physician and hospital use has moved quickly, while most published medical-AI studies still stop at preclinical evaluation and only a small fraction reach randomized testing.
The strongest deployments cluster around tasks where a clinician can inspect the result quickly. Documentation, summarization, coding, discharge instructions, and research support have spread much faster than autonomous diagnosis or triage.
Benchmark performance can be misleading once a real person enters the loop. Models that identify the right condition from a clean case description can perform far worse when patients have to describe symptoms, answer follow-up questions, and decide what to do next.
Healthcare AI is also bad at knowing when it is wrong. Models can sound nearly certain even when accuracy is much lower, which makes abstention, escalation, and uncertainty calibration central safety problems rather than side features.
Performance is fragile across hospitals and patient populations. Data drift, changing laboratory assays, missing fields, different workflows, and local patient mixes can materially alter model accuracy after deployment.
Average accuracy can hide unequal performance. Evidence across demographic labels, languages, hospital types, and non-English encounters shows that healthcare AI can create very different outcomes for patients who are supposedly receiving the same system.
Putting a doctor in the loop helps only when the doctor can challenge the model effectively. Randomized studies show both meaningful gains and clear automation-bias failures, so the quality of the human-AI team matters as much as standalone model strength.
The clearest economic win so far is clinician time, especially around documentation. Large savings in hospital costs or major increases in patient throughput are much less established once integration, monitoring, cybersecurity, validation, training, and review are counted.
Security and accountability become harder after deployment because prompts, data, model updates, and vendor dependencies can all change the behavior of a clinical system. Hospitals increasingly need to know exactly which model ran, what data it saw, what changed, and what recommendation reached the clinician.
The overall pattern is fairly sharp: healthcare AI already works well enough in narrow settings to matter, but trust at scale is still unresolved. Progress now depends less on producing another impressive medical benchmark and more on proving durable benefit across real hospitals, real users, and changing conditions.

This market map, featured in our healthcare AI market deck, highlights top companies and startups in the healthcare AI market
How much healthcare AI is actually being used today?
Healthcare AI is already mainstream inside hospitals and medical practices, although most of the successful use today sits in documentation, summarization and other tightly bounded tasks.
The American Medical Association’s latest physician survey makes the shift hard to miss. In 2026, 81% of the 1,692 physicians surveyed said they used AI professionally, up from 38% in 2023. The average physician reported 2.3 use cases, more than double the 1.1 reported three years earlier. Research summaries were the most common use at 39%, followed by discharge instructions or care plans at 30% and documentation or coding at 28%. Assistive diagnosis was much lower at 17%.
Hospitals are moving in the same direction. A JAMA Network Open study covering 2,174 nonfederal US hospitals found that 31.5% were already using generative AI integrated with their electronic health record in 2024, while another 24.7% planned to adopt it within a year. Among large health systems, ambient documentation has become one of the clearest early winners.
Healthcare AI is spreading fastest where clinicians can inspect the output quickly and keep final authority over high-stakes decisions.
| Healthcare AI signal | Latest measured level | What it says |
|---|---|---|
| Physicians using AI professionally | 81% | AI use has become normal in the surveyed physician population |
| US hospitals using generative AI in the EHR | 31.5% | Hospital deployment has moved well beyond pilots |
| Physicians using AI for research summaries | 39% | Information work is a major early use |
| Physicians using AI for assistive diagnosis | 17% | Higher-stakes clinical use still trails workflow use |
Why is healthcare AI spreading faster than we can prove it works?
Healthcare AI has an evidence gap today: deployment is accelerating far faster than rigorous clinical testing.
A 2026 npj Digital Medicine review gives us the clearest aggregate. Researchers pulled 4,667 primary studies from 218 systematic reviews. They classified 4,114 studies, or 88.2%, as preclinical. Only 113 studies, 2.4%, were randomized controlled trials, and 76 of those 113 RCTs were single-center.
FDA-authorized medical AI shows a similar imbalance. A JAMA Network Open analysis of 903 AI-enabled devices found publicly reported clinical performance studies for 505. Among those studies, retrospective designs were far more common than prospective or randomized testing.
There are areas where the picture is getting stronger. A 2026 systematic review of 32 randomized cardiovascular AI trials found improvements across workflow, patient engagement and clinical outcomes. In 12 studies covering 33,698 patients, AI-supported interventions were associated with lower all-cause mortality, with a pooled risk ratio of 0.84.
The evidence is improving, but it remains much thinner than the pace of deployment suggests.
If you want more recent data on this point, please see our latest healthcare AI market report.

As this chart shows, and as featured in our healthcare AI market deck, search interest in healthcare AI has grown rapidly
Can healthcare AI that aces medical benchmarks safely handle real patients?
Healthcare AI can look excellent on medical exams and still perform much worse once real people have to describe symptoms, answer questions and decide what to do next.
A preregistered Nature Medicine study tested that gap with 1,298 people across ten medical scenarios. When GPT-4o, Llama 3 and Command R+ received the cases directly, the models identified relevant conditions in 94.9% of cases. When ordinary people used those same models, relevant conditions were identified in fewer than 34.5% of cases, and correct decisions about what to do next stayed below 44.2%. The AI-assisted groups did no better than people using their usual information sources.
A separate 2026 JAMA Network Open evaluation pushed 21 frontier models, including GPT-5, Claude 4.5 Opus, Gemini 3 and Grok 4, through sequential clinical cases. Overall scores ranged from 0.64 to 0.78 on the study’s balanced clinical-reasoning measure. Differential-diagnosis failure rates were between 90% and 100% across the models, while final-diagnosis failure rates were below 40%.
A newer real-world benchmark points in the same direction. BRIDGE, published in 2026, tested 95 language models on 87 tasks drawn from 59 real clinical data sources across nine languages. Performance varied sharply by task, language, specialty and model size.
Patient triage makes the gap more serious. Over-triage can send someone to urgent care unnecessarily. Under-triage can delay treatment for a stroke, heart attack, sepsis or psychiatric emergency. A 2026 Nature Medicine evaluation of ChatGPT Health found missed high-risk emergencies and inconsistent activation of crisis safeguards across structured triage cases.
There is still a useful role for patient-facing AI in explaining lab terms, translating discharge instructions, preparing questions for a doctor or summarizing a report. Autonomous diagnosis and triage remain much harder to justify.
Can healthcare AI tell when it is about to give a bad answer?
Healthcare AI still struggles to judge its own uncertainty, and fluent language can make a weak answer sound far more certain than it is.
A Journal of Medical Internet Research study tested nine language models on 2,522 US medical licensing questions. Model accuracy ranged from 56.5% to 89%. Their self-reported confidence stayed between roughly 90% and 100%. That confidence did a poor job of separating correct answers from incorrect ones: the area under the receiver operating characteristic curve ranged from 0.52 to 0.68.
The model’s internal token probabilities did better, reaching 0.71 to 0.87, but even those scores were imperfectly calibrated. We already have better ways to estimate uncertainty than simply asking a chatbot “how sure are you?”, yet the threshold for safe clinical action remains unclear.
The acceptable uncertainty also changes with the task. A doubtful draft note can be reviewed. A doubtful recommendation about chest pain or sepsis carries a very different risk.
Healthcare AI still needs reliable rules for when to answer, when to ask for more information and when to escalate.
If you want more recent data on this point, please see our latest healthcare AI market report.

This chart, featured in our healthcare AI market deck, shows annual VC investment in healthcare AI startups
Will healthcare AI still work in a different hospital with messier data?
Healthcare AI can lose accuracy when the hospital, patient mix, workflow or data source changes, and current health data infrastructure makes those shifts difficult to manage.
A JAMA Network Open study of 143,049 patients across seven Toronto hospitals showed how large the effect can become. Researchers detected meaningful shifts tied to demographics, hospital type, admission source and changes in common laboratory assays. During the COVID-19 period, a drift-triggered continual-learning approach improved the mortality model’s AUROC by 0.44.
The data feeding these systems are getting easier to exchange, but exchange and usability are different things. The latest US Office of the National Coordinator for Health IT figures show that 76% of hospitals participated in all four measured domains of electronic health-information exchange by 2025. Sending, receiving and finding records improved. Integration of received information stayed largely flat.
A medication list may arrive but be duplicated. A missing allergy may mean “none known” or “never recorded.” Two labs can use different units. A scanned PDF can exist in the chart while remaining difficult for structured reasoning.
Healthcare AI therefore needs local validation and continuous monitoring wherever the data or patient mix can change.
Can healthcare AI avoid making care less fair?
Healthcare AI can amplify unfair differences in care, and current testing still leaves too many ways for those differences to slip through.
One of the strongest experiments came from Nature Medicine. Researchers ran more than 1.7 million outputs from nine language models across 1,000 emergency-department cases while changing only the sociodemographic label attached to the patient. Cases labeled as Black, unhoused or LGBTQIA+ were more often pushed toward urgent care, invasive intervention or mental-health evaluation. Some LGBTQIA+ labels triggered mental-health recommendations around six to seven times more often than the clinical facts justified. High-income labels were also more likely to receive advanced imaging recommendations.
Language creates another layer of risk. AgentClinic, a 2026 benchmark covering nine specialties and seven languages, found that every tested model performed best in English. Claude 3.5 Sonnet was the strongest multilingual model in that evaluation, yet average diagnostic accuracy across languages was 48.4%. GPT-4 averaged 20.9%, with performance ranging from about 11% in Chinese to 40% in English.
A more recent preprint from a US safety-net health system studied 54,160 outpatient encounters using an ambient AI documentation tool. Non-English encounters were 21% to 25% less likely than English encounters to reach the study’s high-performance threshold.
There is also an access problem. In the national US hospital study on generative AI adoption, smaller hospitals, independent hospitals and critical-access hospitals were less likely to be early adopters. Hospitals with the highest Medicaid discharge share were also less likely to be early adopters or fast followers.
Overall accuracy can hide large differences between demographic groups, languages and hospitals.
If you want more recent data on this point, please see our latest healthcare AI market report.

This chart, featured in our healthcare AI market deck, looks at Tempus AI’s strategy in healthcare AI
Does putting a doctor in the loop make healthcare AI safe?
A doctor in the loop improves healthcare AI only when the doctor can recognize when the model deserves trust and when it deserves pushback.
Randomized evidence now gives us several very different outcomes. In one US study of 50 physicians, GPT-4 plus conventional resources produced a diagnostic-reasoning score of 76%, versus 74% with conventional resources alone. The two-point gap was statistically insignificant. In another randomized trial with 92 practicing physicians, GPT-4 assistance improved management-reasoning scores by 6.5 percentage points.
Training appears to matter. A separate trial in Pakistan gave physicians around 20 hours of AI-literacy training before testing them with GPT-4o. The AI-assisted group averaged 71.4% versus 42.6% for the conventional-resources group.
Then comes the uncomfortable part. A randomized preprint involving 44 AI-trained physicians exposed half the cases to deliberately flawed GPT-4o recommendations. Diagnostic-reasoning accuracy fell from 84.9% with error-free AI suggestions to 73.3% when flawed suggestions were introduced, an adjusted decline of 14 percentage points.
Access to a strong model alone clearly does not create a strong physician-AI team.
| Human–AI study | Participants | Measured effect |
|---|---|---|
| US diagnostic-reasoning trial | 50 physicians | 76% with GPT-4 vs 74% with conventional resources |
| Management-reasoning trial | 92 physicians | +6.5 percentage points with GPT-4 |
| Pakistan trial after AI-literacy training | 58 physicians | 71.4% vs 42.6% |
| Automation-bias preprint | 44 physicians | Accuracy fell from 84.9% to 73.3% when flawed AI advice was introduced |
Is healthcare AI actually saving hospitals time and money?
Healthcare AI is clearly reducing documentation work in some settings, but the evidence that those gains turn into large financial savings is much weaker.
Providence gives us one of the best large real-world datasets. A 2026 JAMA Network Open study followed 16,149 observations from 1,547 clinicians using an ambient AI scribe. Median note time per appointment fell from 7.1 to 6.1 minutes. After-hours documentation also improved. Appointments per day barely changed, from 12.4 to 12.7, and the study found no significant sustained increase in appointment volume.
Clinician experience looks more positive. A multicenter study of 263 ambulatory clinicians found self-reported burnout falling from 51.9% before ambient-AI use to 38.8% after 30 days. Another study across Mass General Brigham and Emory found large drops in reported burnout among participating clinicians after ambient documentation use.
More recent deployments show a similar pattern. Stanford’s MedAgentBrief generated 1,274 hospital-course summaries across 384 discharges. Physicians used AI-generated content in 57% of discharges. Among the 100 summaries that received feedback, 25 had omissions and 20 had inaccuracies, while hallucinations were rare. Objective time savings were modest and varied by physician.
A recent npj Digital Medicine review looked at 48 economic evaluations of AI-enabled precision medicine. In 89% of base-case analyses, the intervention was judged cost-saving or cost-effective. The underlying gains were usually modest: median incremental quality-adjusted life-year improvement was 0.006, with the middle 50% of estimates ranging from 0.001 to 0.019. The review also found signs of systematic optimism and large differences between applications.
A separate systematic review found that infrastructure costs, indirect costs and equity effects were often undercounted. Integration, validation, cybersecurity, clinician training, monitoring and human review can all sit outside the headline price of the software.
So far, the clearest economic benefit is time returned to clinicians. Whether that becomes lower costs, more patients treated or better long-term outcomes depends heavily on the workflow and hospital.
If you want more recent data on this point, please see our latest healthcare AI market report.

This chart, featured in our healthcare AI market deck, shows annual funding in healthcare AI startups
Can hospitals secure healthcare AI and know who is responsible when it fails?
Healthcare AI still has a serious post-deployment safety problem because malicious inputs, model updates and unclear responsibility can all change a clinical output after the system enters routine use.
A JAMA Network Open study tested prompt-injection attacks in 216 simulated patient-model dialogues. The attacks succeeded in 102 of 108 attacked conversations, or 94.4%. In the highest-harm scenarios, including contraindicated drugs during pregnancy, the success rate was 91.7%. The malicious influence also persisted into later turns in 69.4% of the attacked dialogues.
Training data create another attack surface. Nature Medicine researchers poisoned tiny fractions of medical training data with misinformation. Replacing only 0.001% of training tokens increased harmful medical outputs, while the corrupted models continued to perform similarly to clean models on standard medical benchmarks. In a four-billion-parameter model, poisoning one million of 100 billion training tokens increased harmful content by 4.8%.
The accountability chain is just as complicated. The clinician may act on the recommendation. The hospital chose the tool and its place in the workflow. The healthcare AI vendor designed the application. A separate foundation-model company may control the underlying model.
Regulators are trying to make that lifecycle more visible. The FDA now has final guidance for predetermined change-control plans for AI-enabled device software, allowing manufacturers to describe certain future modifications and how those changes will be validated. The agency also finalized updated Clinical Decision Support guidance in 2026 and continues to develop broader lifecycle guidance.
A recent longitudinal analysis of the FDA registry counted 1,430 AI- or machine-learning-enabled devices authorized through the end of 2025, across 576 manufacturers and 17 specialty panels. The FDA’s current list has continued adding 2026 decisions.
Postmarket evidence shows why this remains unresolved. A 2026 JAMA Network Open study followed 903 FDA-authorized AI devices and identified 43 recalls, or 4.8%, after a median of about 15 months.
Hospitals increasingly need an auditable record of which model ran, what data it received, what changed and what recommendation reached the clinician.
So what are the biggest unsolved problems in healthcare AI?
The biggest unsolved problem in healthcare AI today is reliability in the real world: getting systems to keep improving care across messy patients, changing hospitals and imperfect human use without creating new safety failures.
Some healthcare AI systems already improve narrow clinical outcomes. Others save clinicians meaningful time. Frontier models can answer difficult medical questions at a very high level.
The unresolved problems appear once those systems leave controlled tests. We still need to know whether benefits survive multicenter deployment, whether models recognize uncertainty, whether performance holds across hospitals and languages, whether clinicians resist persuasive errors, and whether health systems can catch drift, attacks and harmful updates quickly enough.
When we put the evidence together, three problems stand above the rest. Real-world clinical evidence comes first because most medical AI research still stops before randomized live evaluation. Reliability under uncertainty and change comes next, covering calibration, drift, messy data and safe escalation. Human-AI interaction follows closely because strong standalone model performance can disappear once clinicians or patients enter the loop.
Equity, security, governance and economics sit behind those three and can still block deployment even when the model itself performs well.
Healthcare AI has already solved enough narrow tasks to become a real part of healthcare. The remaining challenge is whether we can trust that performance repeatedly across people, hospitals and changing conditions.
| Priority | Unsolved healthcare AI problem | What would count as real progress |
|---|---|---|
| 1 | Real-world clinical effectiveness | Multicenter prospective trials showing durable patient benefit |
| 2 | Uncertainty, drift and messy real-world data | Reliable abstention, local validation and continuous monitoring |
| 3 | Human-AI interaction | Consistent gains from patient-AI and clinician-AI teams even when the model makes mistakes |
| 4 | Equity across groups, languages and hospitals | Routine subgroup outcome testing plus comparable performance outside well-resourced English-language settings |
| 5 | Security and lifecycle accountability | Auditable model versions, adversarial testing and rapid postmarket detection |
| 6 | Sustainable economics | Full-cost evidence showing that clinical or operational value exceeds integration and monitoring costs |
If you want more recent data on this point, please see our latest healthcare AI market report.

This chart, featured in our healthcare AI market deck, compares the main business model options for ambient AI companies
OUR METHODOLOGY
There is no obvious consensus on what the biggest unsolved problems in healthcare AI actually are. The answer changes depending on whether someone is looking at benchmark scores, clinical trials, hospital deployment, physician experience, patient safety, economics, regulation or technical reliability.
We therefore broke the question into distinct analytical dimensions before forming any overall conclusion. For each one, we prioritized the freshest evidence capable of showing what is happening in practice: large-scale adoption data, clinical and randomized studies, real-world deployments, systematic reviews, independent model evaluations, post-market evidence and regulatory data.
We then assessed those findings together rather than letting one headline result drive the answer. Where possible, we looked for convergence across different methods and settings. We gave particular weight to evidence that survived the move from controlled testing into real patients, clinicians, hospitals and workflows, because that is where many of healthcare AI’s hardest problems become visible.
Only after assessing the evidence dimension by dimension did we compare the problems against one another. The final hierarchy reflects how consistently each issue appeared across the evidence, how directly it can affect real-world clinical performance, and how fundamental it is to making healthcare AI reliably useful at scale.
Key sources used for this analysis include the American Medical Association’s physician AI survey, JAMA Network Open on generative AI adoption across US hospitals, npj Digital Medicine on the preclinical and randomized evidence gap, Nature Medicine on patient use of medical AI, JAMA Network Open on clinical data drift across seven hospitals, Nature Medicine on sociodemographic bias in emergency-care recommendations, NEJM AI on automation bias in physician-AI collaboration, JAMA Network Open on ambient-AI documentation at Providence, JAMA Network Open on prompt-injection attacks in medical dialogues, and the US Food and Drug Administration’s current registry of AI-enabled medical devices.

This chart, featured in our healthcare AI market deck, illustrates how revenue is distributed across customer segments in the healthcare AI market
Related blog posts
- Healthcare AI: what is getting real adoption now?
- Healthcare AI: what is actually working now?
- What is the latest funding news in healthcare AI?
- Healthcare AI: who are the top startups?
- What are the latest fundraising trends in healthcare AI?
- How strong is fundraising in the healthcare AI market right now?
Who is the author of this content?
NEW MARKET PITCH TEAM
We track new markets so founders and investors can move fasterWe build living "market pitch" documents for emerging markets: AI, synthetic biology, new proteins, and more. Instead of outdated PDFs or hallucinated LLM answers, our clients get a clean, visual, always-updated view of what's really happening: key players, deals, regulations, and signals that matter. Learn more about us.