This website uses cookies

Read our Privacy policy and Terms of use for more information.

Can Clinical AI Be Trusted?

Seven years ago, at my White Coat Ceremony, my classmates and I recited some version of the Hippocratic Oath: first, do no harm. That was before “clinical AI” meant a large language model you could pull up instantly for almost any question.

Now we have a flood of clinical AI tools and paper after paper asking whether they’re “just as good” as physicians. A new Harvard/Stanford preprint, “First, do NOHARM,” asks a different question: what happens when these tools make errors of commission or omission?

In this piece, I’ll break down what the study found and share my take from both the physician and systems perspectives.

First, Do NOHARM: Methods

While there are now innumerable studies assessing LLMs on medical knowledge and diagnostic/management tasks, far fewer have evaluated safety when it comes to errors.

The authors of NOHARM took a unique approach by creating a new benchmark designed to measure the safety of AI-generated clinical recommendations. They tested 45 frontier LLMs (ChatGPT, Perplexity, Claude), 4 clinical RAG tools (AMBOSS LiSA, Doximity Ask, OpenEvidence, Glass Health), and physicians with and without AI assistance using realistic primary-care-to-specialist eConsult cases.

Here’s how they did it:

  • E-consult cases: 100 e-consult cases were curated from Stanford’s e-consult dataset based on recency and case detail. Physicians reviewed the cases and reconstructed them to avoid reidentification. Note, these were not routine primary care questions. They were specialist-level eConsults where a generalist had already decided the case needed outside input.

  • Case variants: for each of the 100 cases, physicians created variants—changing specific details to stress-test whether models would make the same commission and/or omission errors.

  • Clinical Management Actions: for each case, physicians listed possible management actions spanning diagnostic, medication, counseling, follow-up, and procedural decisions. Each “action” was graded by multiple specialists for clinical appropriateness using a standardized scale. Harmful actions were categorized as error of commission (e.g., ordering an unnecessary or dangerous medication) and error of omission (e.g., missing a critical test, referral, or follow-up). Many papers published thus far that assess “safety” focus on hallucinations or obviously wrong answers.

The primary metric was the severity-weighted F1 score, which combined precision and recall:

  • Precision: Did the model avoid recommending inappropriate actions?

  • Recall: Did the model include important appropriate actions?

This is a sensible approach: a model that says very little may avoid dangerous recommendations but miss critical care steps, while a model that says everything may include key actions but also add unnecessary or harmful ones.
The best clinical answer has to be both careful and complete.

First, Do NOHARM: Results

There are three main takeaways from this paper, plus one additional “food for thought” point I’ll get to in Dashevsky’s Dissection.

  1. Severe potential harm varied widely across AI systems: across the models studied, severe potential harm ranged from 2.9% (AMBOSS LiSA) to 24.6% (Llama 4 Maverick). There was no statistically significant difference in severe potential harm between the clinical RAG models. The authors also tested a hypothetical “Do Nothing” control, which carried potential severe harm in 37% of cases—worse than every AI system tested. This is important since, in medicine, the comparison is rarely AI versus perfect care. It’s AI versus the messy baseline of delayed consults, incomplete workups, and no second opinion.

  1. Clinical RAG systems outperformed generalist LLMs: clinical RAG tools had the lowest severe-error rates and the highest severity-weighted F1 scores (higher scores meaning the clinical answer was both careful and complete). The top clinical RAG systems also outperformed the top frontier generalist LLMs across prompting conditions, suggesting the clinical RAG tools were already better tuned for complete clinical responses, whereas generalist models benefited more from explicit prompting.

  1. The biggest safety issue was omission: most “severe errors” were omissions, meaning these models were more likely to leave out key components of management than to recommend something overtly dangerous (errors of commission). A blatantly bad answer makes you skeptical, but a good-yet-incomplete answer can slip by unnoticed (e.g., missing a repeat troponin when it’s uptrending).

Dashevsky’s Dissection

The most important contribution of this paper is that it shifts the AI conversation from performance to safety. Over and over again, we’ve asked whether models can pass medical exams, solve challenging diagnostic cases, or outperform clinicians on benchmarks. But clinical care is far from “benchmarks” and closer to a series of management decisions where one missing step can be critical.

I recently wrote about the importance of “trust” when it comes to the clinical AI tools we use. One of the best ways to build trust is through external studies like this one (“external” meaning the clinical AI tool isn’t producing a study on itself, which helps minimize bias) that specifically test harm—not just how well a model answers standardized exams and vignettes.

That’s why this paper pairs well with the recent Nature publication from June 2026 showing that frontier models outperformed specialized clinical tools on traditional medical benchmarks. Those benchmarks mostly test knowledge and reasoning. NOHARM tests whether the recommendation is complete and safe enough to act on.

I’m not surprised that clinical RAG tools outperformed frontier generalist LLMs on severe error rates and severity-weighted F1 scores. A model can score well on medical knowledge and still be unsafe when it comes to real clinical management (like a resident who’s a great test-taker but is awful at actual clinical medicine).

A major strength here is the realistic clinical framing: eConsult-style cases feel much closer to day-to-day practice than USMLE-style questions. The “Do Nothing” comparator is also important because it keeps this from becoming a simple AI safety story. In many clinical settings, the alternative to AI assistance is waiting, incomplete workups, delayed referrals, or no second opinion at all, as compared to immediate specialist involvement. That doesn’t make AI automatically safe, but it does make the comparison more honest.

And here’s my food for thought: physicians often didn’t use good AI recommendations. In the randomized physician study (the other part of this study), AI assistance improved performance, but modestly: conventional resources scored 42.2%, AI-assisted care 47.3%, and actual AI use 52.0%. The tool helped, but the benefit depended on whether clinicians meaningfully engaged with it. This suggests the bottleneck may not be access to AI alone. It may be whether the workflow helps us recognize which parts of the AI output are worth incorporating, and which parts need to be challenged.

For clinical AI companies, the lesson is straightforward: trust will not come from cool demos or benchmark screenshots. It will come from rigorous safety testing, transparent limitations, and workflows that make omissions easier for clinicians to catch. For that reason, I still don’t recommend my colleagues use general frontier models like ChatGPT and Claude for clinical questions.

Before closing, a few limitations are important. First, this is a preprint and hasn’t been peer reviewed. The claims still need external scrutiny (and not from me or the LinkedIn crowd). Second, the cases were limited to outpatient eConsults, which makes generalizability tough (I can’t generalize these findings to the ICU environment I’m working in!). Third, the study used single-turn consultation (I, the PCP, send all the info in one eConsult and the AI produces next steps). In real life, the specialist would sift through the EHR, aggregate data and trends, and often ask follow-up questions. A single free-text answer can’t fully represent clinical decision-making.

In summary, this preprint is a reminder that clinical AI should be judged less by whether it can sound right and more by what happens if we act on its recommendations. The most important safety risk may be omission: the missing workup, follow-up, counseling, or referral that turns a polished answer into an incomplete plan. Clinical RAG tools appear to outperform generalist LLMs, and AI assistance can improve physician performance, but human oversight alone is not a complete safety strategy. If AI is going to move from documentation into clinical decision support, we need benchmarks, workflows, and monitoring systems built around patient harm—not just model accuracy.

Huddle+ Members Only

Want to go deeper? Upgrade to Huddle+

Get exclusive courses, expert analysis, and the tools to understand how healthcare really works—from AI to policy to the business of medicine.

Upgrade Now
Premium courses & guides
Community access
Weekly insights

Reply

Avatar

or to participate