A review in Cell from MIT and UC San Diego argues that clinical AI is being validated in a way that systematically fails to predict how it behaves once deployed. The authors sort real-world failures into three classes: outputs that are simply wrong or fabricated, performance that varies across patient groups for reasons unrelated to biology, and performance that decays after launch as data, coding standards, and workflows shift. Their central and more contentious claim is that the standard technical remedies do not fix these problems, because each remedy assumes something about clinical data that clinical data does not satisfy. They also identify a measurement trap: once a model influences care, it alters the very outcomes later used to grade it. The prescription is not better models but lifecycle evaluation, prospective testing, continuous monitoring, and institutional accountability.
Artificial intelligence has arrived in hospitals faster than the evidence that it works there. This review argues that the field has been measuring the wrong thing. Models are validated on fixed datasets, cleared on benchmark scores, then deployed into an environment that looks nothing like the one they were tested in.
The authors group what goes wrong into three categories. The first is straightforward error: models producing wrong or invented output. A widely used transcription tool has been documented inventing sentences that were never spoken, and it remains in hospital use. The second is performance that varies across patient groups for reasons that have nothing to do with medicine. When four leading language model families were given ten identical psychiatric cases and the only thing that changed was the patient’s stated race, diagnostic decisions stayed relatively stable but treatment recommendations did not. The third is decay. Systems that work on launch day stop working as coding standards shift, patient populations change, and clinical workflows get rearranged around them.
The central claim is the uncomfortable one. Most fixes currently on offer do not address these problems, because each rests on an assumption that clinical data violates. Cleaning and imputing missing values assumes missingness is random. In hospitals it is not. Data goes missing because of billing incentives, documentation habits, and who gets ordered which test. Reweighting training data to correct group disparities assumes you already know which groups are affected and why, information electronic health records rarely contain. Retrieval augmentation, the main proposed fix for language model fabrication, reduces it but introduces a fresh failure mode when the retrieved source is stale or mismatched.
Then there is the feedback problem, which the authors treat as the deepest. Once a model is embedded in care, it changes clinician behavior, which changes treatment, which changes the recorded outcomes later used to judge the model. A sepsis predictor that successfully prompts early intervention will appear to be getting worse, because the events it was trained to predict stop happening. Monitoring a system that has begun editing its own report card is a hard measurement problem, and it is not solved.
None of this argues the technology does not work. The review cites a Swedish randomized trial of AI-supported mammography covering more than 100,000 women, which found more cancers caught at screening and fewer of the aggressive tumors that surface between screening rounds. That is real benefit, demonstrated prospectively, in the setting where the tool is actually used.
The lesson the authors draw is that the prospective design is the point, not the AI. They call for evaluation across the whole deployment lifecycle: stress testing before launch, silent deployments that observe a model without letting it touch care, continuous auditing afterward, systems that decline to answer when uncertain, and clear institutional ownership of failures. Benchmarks tell you what a model can do. They do not tell you what it will do.
Actionable Insights
This paper is about how medical AI fails, so the take-home concerns how you should weight AI-derived health information.
Discount benchmark claims heavily. In one cited study, model accuracy fell from about 68 percent to about 55 percent purely because the model had to gather patient information itself rather than receive it pre-packaged. That is 13 percentage points lost, a 19 percent relative decline, and a 41 percent increase in error rate. Expressed as a standardized effect size for proportions, Cohen’s h is 0.27, which is a small-to-moderate but consistent gap.
Do not trust your own judgment of AI medical advice. Among 300 participants, people distinguished physician-written from AI-written answers at 50 percent, which is a coin flip. Effect size approximately zero. They rated inaccurate AI answers as trustworthy as accurate physician answers and reported equal willingness to act on harmful advice.
Monitoring beats model selection. A documentation accuracy collapse from 79 percent to 35 percent, a tripling of the error rate and a large effect at Cohen’s h of 0.92, was caught only because someone was systematically watching.
Where AI is prospectively tested it can deliver. AI-supported mammography in roughly 106,000 women produced 27 percent fewer aggressive-subtype cancers, an absolute gain of about one avoided per 3,300 women screened.
Context and Source
- Full title: Avoiding common failures in AI for health and medicine.
- Authors and institutions: Olawale Salaudeen, Haoran Zhang, Eileen Kim, Karandeep Singh, Marzyeh Ghassemi. Massachusetts Institute of Technology (Department of Electrical Engineering and Computer Science; Institute for Medical Engineering and Science), Cambridge, Massachusetts, and University of California San Diego (Division of General Internal Medicine; Joan and Irwin Jacobs Center for Health Innovation, UCSD Health; Division of Biomedical Informatics), La Jolla, California.
- Country: United States.
- Journal: Cell (Cell Press, Elsevier).
- Impact evaluation: The impact score of this journal is 42.5 (most recent reporting year; the five-year figure is 48.9 and the 2023 value was 45.5), evaluated against a typical high-end range of 0 to 60+ for top general science, therefore this is an Elite impact journal.