The Hardest Part of AI in Hospitals Isn't the Model—It's Making the Hospital's Ledger Work
Suchi Saria, whose sepsis alert was adopted by 90% of doctors, asserts: the obstacle for AI isn't model performance, but trust and financial accounting under prospective payment systems.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Publication isn't clinical success
She warns this is the step most researchers misjudge. The 2015 Science Translational Medicine sepsis prediction model looked great on retrospective data, but when actually deploying in the ED and wards, they found the data's veracity, completeness, and messiness were different; even the ICD-10 coded ground truth used to measure quality was too noisy and could lead the team astray. To fill in ground truth, debiasing, transparency, outcome evaluation, and monitoring after deployment, she later did the equivalent of about 40 papers' worth of innovation, publishing only 20.
— Suchi SariaHospitals won't buy because of a good paper
Saria says there are many layers between a tool and making AI adoption a no-brainer for the healthcare system. First, prove the tool is 10x or 15x better than current practice and explain why; then turn the point solution into a platform covering multiple conditions; then partner with EMR vendors to compress IT integration to about 80 hours and go live within six weeks; then handle governance, monitoring, tuning, and even plan with the FDA; only then do clinical outcomes and financial returns come into play. Each step is a 'surprise' and a hill to climb.
— Suchi SariaSoftware itself can be a treatment
She holds a conviction: medical innovation shouldn't be limited to drugs and devices, but should also include software-only AI-based interventional protocols. These interventions don't assist with documentation but use the full EMR's multimodal data to precisely answer 'when, for whom, what to do' and smooth the execution. She says this is a big opportunity to improve care quality because most AI today clusters around pre-visit transcription and post-visit coding, leaving the middle ground of actually 'caring for patients' almost untouched.
— Suchi SariaDoctors actively use alerts if signals are accurate
Bayesian's 2022 Nature Medicine report scale: nearly 750,000 patients, two years, 4,400 clinicians, of which 2,200 were physicians. The key internal number is a 90% physician adoption rate. She says this isn't due to pop-up bombardment but signal quality: many old tools alert 95% of the time after the doctor has already initiated treatment, rendering them valueless; hers is designed with an early warning objective that identifies early, so doctors proactively click, enter assessments, and choose agree or disagree.
— Suchi SariaUnder DRG, early rescue equals cost savings
Under DRG bundled payment, a hospital receives a fixed amount for treating a sepsis patient. Whether the stay is 20 days or 4 days, the extra 16 days of cost come out of the hospital's profit margin and occupy beds that other patients could use; if sepsis isn't identified at admission and documented as present-on-admission, Medicare pays even less. Suchi says for a 2,000-bed system, these missed detections accumulate into $30–40 million in financial penalties annually, so AI early identification is like giving money back to the hospital, making it easier for the CFO to say yes.
— Suchi SariaThe hardest part is clinical trust, not technology
Asked which medical task AI will master in five years, Suchi didn't list new algorithms. She said data quality and veracity are hard, but as an AI researcher, that part is most within her control; what she finds truly difficult is the sociotechnical layer: getting doctors to trust the output, and getting EMR vendors like Epic to make integration easier. In her view, the pain point in medical AI isn't that models aren't strong enough, but that the systems determining whether models can be used are never in the data scientist's hands.
— Suchi SariaClinical multimodal data remains a research blue ocean
She says there is still a vast open field in academic research. She cites FDA-funded translational safety research and the work of several PhD students: the EMR is not text plus images, but thousands of modalities, each with different missingness patterns, bias patterns, and interaction patterns; how to design architectures for this and how to validate generalization are unfinished problems. Foundation model companies are pressured to do well on benchmarks, focusing on Q&A over internet text, but real-world applications are not Q&A—that's the entry point for researchers; choose important problems where giants haven't heavily deployed.
— Suchi SariaLow-hanging fruit is the last to pick
Her advice to founders is blunt: many startups today sell big visions but aren't AI-native themselves, and their moats are unclear. No one knows how long the market gap will last; if a small company can build a moat quickly through execution speed, it's worth a shot, but if even speed can't build a moat, they'll be swallowed by large model companies. Suchi chooses the opposite: everyone advised her against sepsis because it's hard, requires long-term investment, and hospital relationships, but she has always done what others think impossible; she says hundreds of companies are turning notes into codes—that low-hanging fruit—but lasting impact lies in problems no one dares to touch.
— Suchi SariaIn their own words · checked verbatim
That's probably the most common mistake most researchers make. That when they get a paper of that format, they think the goal's been achieved, victory's been had, and it's done.
Suchi Saria12:23
I feel like there was gonna be drugs and devices, and there's gonna be a third vertical, which is more AI-based interventions.
Suchi Saria25:37
In sepsis, every hour is associated with five to eight percent increased risk of mortality, so an hour can be the difference between life and death.
Suchi Saria29:46
once you move into startups, it's the Wild Wild West. You know, people make claims all the time. There's no data to back it up. We don't know if it's gonna generalize. People go from hero to zero in one minute.
Suchi Saria34:49
And what was super interesting is that actually, I think, um, quality is the path to financial sustainability for health systems.
Suchi Saria36:56
To me, the hardest part feels the sociotechnical part, the part of convincing people to trust it, the part of convincing electronic health record providers to make an integration easier, you know? So the, it almost feels like those are the harder parts, not the AI.
Suchi Saria46:05
Figures
| Early sepsis identification lead time (Hopkins control) | median 5.5 hours | 29:46 |
| Increase in death risk per hour of sepsis delay | 5%–8% | 29:46 |
| Physician adoption rate of Bayesian's alert | 90% | 31:47 |
| IT effort and time for EMR vendor integration | about 80 hours of IT work, within 6 weeks | 19:34 |
| Cleveland Clinic deployment scale | about 20–21 hospitals, about 25 emergency departments | 40:59 |
| Clinical benefit of sepsis alert in actual deployment | average length of stay reduced by 2 days; absolute mortality reduced by 4%–5% | 40:59 |
| Mayo Clinic advanced illness pilot results | length of stay reduced by nearly 5.5 days; palliative care referrals increased by 44% | 33:49 |
Glossary
- failure to rescue
- Failing to detect clinical deterioration and intervene in time, leading to preventable death or severe illness.
- DRG
- Diagnosis-related group: a bundled payment per diagnosis; hospital savings from reduced resource use directly become profit.
- NTAP
- New Technology Add-on Payment: an additional CMS payment for new technologies, compensating hospitals beyond routine payments.
- AI-based interventional protocols
- Software-only treatment plans that identify timing, target, and action from multimodal data and assist in execution.
How to listen
Medical AI researchers, entrepreneurs aiming to deploy in hospitals, and hospital CFOs and clinical informatics officers deciding on such systems.