The Mess in AI Evaluation Isn't a Technical Problem, It's a Transparency Problem
In that 2019 systematic review, fewer than 5% of medical imaging AI papers met the rigor standards long established in other areas of medicine — and even the best-performing models only matched radiologists.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Fewer than 5% of AI papers in top journals hold up
Around 2019, new papers appeared every week claiming "AI is as good as radiologists" or "AI beats radiologists at pneumonia detection," yet these conclusions contradicted each other. Xiao Liu and her collaborators did something nobody was doing at the time: they took the rigor standards that the diagnostic test evaluation field had used for years and applied them to these AI papers, keeping only those with external validation, head-to-head comparison, and a meaningful control group. The paper count was cut down drastically — fewer than 5% truly met the standards long established in other areas of medicine. And within that small handful, the best models only matched radiologists. This isn't bad news; it means there's too much noise in the published literature.
— Xiao LiuReporting guidelines don't teach you how to do research, they force you to say what you did
CONSORT, SPIRIT, PRISMA, STARD, TRIPOD — these reporting guidelines are essentially a checklist specifying which items a paper must report. They don't tell you how to do the research — they're not methodology guidelines, they're transparency guidelines. Whether your methods are good or bad, you have to write them transparently and clearly. Their function is to put authors, journal editors and peer reviewers on the same information plane, with no smoke bombs and no hiding behind vague details. Xiao Liu adds a side effect: when authors know they must report a given item, it in turn forces methodological rigor upward, though that's hard to measure.
— Xiao LiuGuidelines come out of Delphi voting; items that can't reach consensus get cut
CONSORT AI and SPIRIT AI weren't written by a few people behind closed doors; they went through a Delphi process: first a literature review and expert opinion generated a set of candidate items, then multiple rounds of multi-stakeholder voting, closing with an in-person discussion. That in-person discussion is often the most valuable part, because people learn from each other. Items that can't reach consensus are excluded; only those that do are kept. Xiao Liu says doing this kind of consensus project properly takes at least six months; she tried a rushed version compressed to three months, but participants contribute their time for free, and getting high-quality feedback and multi-round participation means you can't cut corners on time or reach.
— Xiao LiuWhether explainability belongs in reporting guidelines split the room
Andy Beam recalls that one question in the Delphi discussion was whether clinical trials should add explainability requirements, or at least report any explainability evaluation you did. He and Lauren Oakden-Rayner got up on the small stage and said "no, it shouldn't be done at all, and here's why." That debate was later commissioned as a paper by Rupa, an editor at Lancet Digital Health, and became one of Andy's most-cited papers. In the lightning round, Xiao Liu also clearly takes the "no" side: she thinks you should first ask why you want AI to be explainable, and keep asking all the way down — what you ultimately want is just beneficial outcomes.
— Andy BeamAdvising regulators without ever working in industry felt insincere to her
At the time, Xiao Liu was doing academic AI research and clinical ophthalmology while increasingly consulting for regulators. She began to feel she had never stood on the side being regulated, and that consulting was a bit insincere. So she joined Apple as a health scientist, mainly designing validation studies for wearable algorithms, such as Apple Watch sensing algorithms, including a hypertension notification feature released this year. She says Apple has a well-oiled product-shipping machine, which was exactly the experience she wanted. After more than a year there she went back to Birmingham to build a lab full-time, and then joined Microsoft AI's newly formed medical frontier LLM team.
— Xiao LiuShort term computer scientists set the direction, long term clinicians gatekeep
Asked whether medical AI will be driven more by computer scientists or clinicians, Xiao Liu first pushed back on what "driven" means, and after confirming it meant "changed," gave her judgment: short term, computer scientists and engineers; long term, clinicians. The mechanism is that whoever processes the data, chooses the metrics and chooses the benchmarks determines in the short term which direction progress moves; but if the endpoint is real use on patients, clinicians will play the gatekeeper role at a later stage.
— Xiao LiuDoctors' work won't change much in five years, but should in ten
Xiao Liu thinks doctors' work won't look dramatically different within five years, but very likely will in ten years and beyond — and she hopes it will. The path she envisions: if medical intelligence and models keep advancing at the current pace, the medical profession no longer has to be about accumulating knowledge, or even about the kind of medical reasoning taught in medical school; the center of gravity shifts to being "a companion to experience" — accompanying deteriorating health, sudden health events, day-to-day wellbeing fluctuations, life-changing diagnoses, treatment failures and complications, and maybe recovery, maybe not. She thinks people will continually need something deeply interpersonal on that journey: someone who can call on knowledge but doesn't necessarily carry it themselves, instead translating it into human language.
— Xiao LiuThe most optimistic thing is dismantling old medical concepts; the most pessimistic is that we won't dare
Xiao Liu says she's more optimistic than pessimistic. What excites her most is that aggregating all the knowledge in the medical world might reveal that our understanding of biomarkers, disease phenotype clustering, and existing medical knowledge frameworks is off. She gives the example of the inflammatory eye diseases from her PhD research: a pile of diseases loosely called "white dot syndromes" just because they produce white dots in the fundus, but they're very likely driven by completely different mechanisms. She hopes one way the new intelligence shows up is by dismantling these old concepts in medicine. The pessimistic side: the field is inherently cautious and may resist this dismantling out of unfamiliarity, staying rigidly attached during a period of technological change to the evaluation methods and guidelines of the past ten or twenty years, and missing the chance to discover.
— Xiao LiuIn their own words · checked verbatim
Less than 5% of the papers actually really met that very standard bar in other areas of medicine.
Xiao Liu13:23
we found that the performance of those models within that cohort was at best equivalent to radiologists.
Xiao Liu14:23
What reporting guidelines don't do is tell you how you should do your study. So, they're not methodological guidelines, they're transparency guidelines.
Xiao Liu16:28
it supplies enough information that brings the authors, the editors of a journal, and the peer reviewers on a level playing field.
Xiao Liu17:30
I think we should ask why we want AI or anything to be explainable, and I think if we drill down far enough, ultimately, we just want stuff to have beneficial effect.
Xiao Liu30:48
the computer scientists and the people who are handling the data, choosing the metrics, choosing the benchmarks, get to decide in the short term which direction progress happens.
Xiao Liu32:49
there's something deeply interpersonal that we will continue to want as part of that journey, which someone that can leverage the knowledge but not necessarily embody it but translate it into human terms.
Xiao Liu34:53
Figures
| Share of AI papers meeting the rigor standards of other medical fields | Fewer than 5% | 13:23 |
| Combined citations of CONSORT AI and SPIRIT AI | More than 4000 (since publication in 2024) | 18:33 |
| Minimum time needed to do a consensus project properly | At least 6 months (the rushed version was compressed to 3 months) | 22:39 |
| Xiao Liu's tenure at Apple | More than a year | 26:43 |
| Xiao Liu's tenure at Microsoft AI | A year and a half | 27:44 |
Glossary
- Delphi study
- Multiple rounds of anonymous voting plus an in-person discussion, used to reach consensus among experts on which items matter.
- CONSORT
- The international checklist of what a randomized controlled trial paper must report; the AI version is CONSORT AI.
- SPIRIT
- The reporting checklist for clinical trial protocol papers; the AI version is SPIRIT AI.
- white dot syndromes
- A loose umbrella term for a group of inflammatory eye diseases that produce white dots in the fundus, possibly driven by different mechanisms.
How to listen
Engineers and founders working on medical AI products, clinical validation or regulatory submissions, and researchers interested in how evidence standards for medical AI get built.
The personal-history portion of the lightning round (first job, favorite paper) can be skipped.