Medical AI Benchmark Exposes Severe Consultation Safety Gaps
NOHARM uses specialist-rated consult options across 1,100 cases, finding severe-harm potential up to 24.6% and gains from AI-assisted physicians.
Underlying Paper
First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations
Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain poorly characterized. We present NOHARM (Numerous Options Harm Assessment for Risk in Medicine), a 1,100-task benchmark of primary care-to-specialist consultation cases to measure the frequency and severity of potentially harmful errors from LLM-generated medical consultation recommendations. NOHARM covers 10 specialties, with 12,747 expert annotations for 4,249 clinical management options. Across 20 notable LLMs and 4 widely used retrieval-augmented generation (RAG) clinical AI tools, direct application of recommendations carried potential for severe harm in up to 24.6% of cases, with errors of omission accounting for more than 80% of severe errors. Harm potential was not uniform across systems, with clinical AI tools outperforming generalist LLMs, and multi-agent AI teaming further improving performance in generalist models. In a randomized study of 101 U.S.-licensed generalist physicians, AI assistance improved physician performance compared to conventional resources. However, AI-assisted physicians frequently omitted valuable AI-generated recommendations and still scored lower than many AI systems alone. Had those recommendations been incorporated, combined human-AI responses would have outperformed both the human and AI system as used, suggesting complementary strengths and unrealized potential in human-AI teaming. Collectively, these results show that despite strong performance on medical knowledge benchmarks, widely used AI tools can produce medical consultation advice with the potential for severe harm, and highlight the need for explicit measurement of clinical safety. The benchmark and leaderboard are publicly available to support ongoing evaluation and improvement of AI systems used for clinical care.
Medical LLM evaluation often rewards factual recall or board-style question answering, while clinical consultation safety depends on something narrower and harder: whether the recommendations a clinician might act on omit necessary care, suggest harmful management, or delay escalation. This paper introduces NOHARM, a benchmark built around primary care-to-specialist consultation cases, and pairs it with a randomized study of how physicians use AI assistance in the same decision setting. The core message is uncomfortable: systems that look capable on medical knowledge tasks can still generate advice with severe harm potential.
Core Contribution
NOHARM evaluates clinical management options rather than final-answer accuracy. The authors curate 1,100 consultation tasks across 10 specialties, then ask specialists to rate 4,249 management options using a scale that combines RAND/UCLA appropriateness judgments with harm severity. That design matters because many unsafe medical responses are not obviously wrong statements; they are missing referrals, missing tests, inappropriate reassurance, or recommendations that fail to address high-risk differential diagnoses.
The paper’s main empirical claim is that explicit harm measurement changes the ranking and interpretation of medical AI systems. Across 20 generalist LLMs and 4 retrieval-augmented clinical AI tools, direct use of recommendations had severe-tier harm potential in as many as 24.6% of cases. Errors of omission made up more than 80% of severe errors, which points to a practical weakness in consultation use: the dangerous failure is often what the model fails to mention.
Technical Approach
The benchmark construction is unusually clinician-heavy. Cases are curated, reviewed by specialists, and converted into rubrics over many possible management options rather than a single gold answer. Figure 1 summarizes that pipeline, from case selection through specialist review and rubric creation.
The rating scheme separates appropriateness from possible patient harm. Figure 2 shows the expert scale used to classify recommendations, including the severe-harm tier that drives the paper’s headline safety analysis.
The system evaluation covers both general-purpose LLMs and clinical RAG tools. The authors also test multi-agent AI teaming for generalist models, where multiple AI agents contribute to or critique the recommendation set. Separately, the randomized physician study compares 101 U.S.-licensed generalist physicians using AI assistance against physicians using conventional resources, allowing the paper to distinguish AI-alone performance from the messier human-AI workflow.
Results and Analysis
The strongest result is not that every AI system performs poorly. It is that harm is unevenly distributed by system type and workflow. Clinical RAG systems outperform generalist LLMs on the benchmark, and Figure 3 reports lower severe-tier harm rates for those tools than for generalist models. The finding supports a practical reading: domain retrieval and clinical product constraints appear to reduce some unsafe outputs, but they do not remove the need for case-level safety evaluation.
The randomized physician study adds a second layer. AI assistance improved physician performance relative to conventional resources, so the paper does not argue against clinician use of AI. But AI-assisted physicians frequently left useful AI-generated recommendations out of their final responses, and their submitted answers still scored below many AI systems alone. The counterfactual analysis is the most interesting part: if the omitted useful AI recommendations had been incorporated, the combined human-AI response would have beaten both the physician response as used and the AI system alone.
That result cuts against two simple stories. It is not enough to say that doctors should supervise AI, because supervision can discard correct machine suggestions. It is also not enough to say that the AI should replace the physician, because the authors find complementary strengths that are lost in the actual workflow. The evidence points instead to interface and teaming design as the bottleneck: physicians need help identifying which AI suggestions are clinically valuable, not merely access to a generated consultation note.
Limitations
The evidence is broad for a medical AI safety benchmark, but it is still a benchmark and randomized study rather than measurement of real patient outcomes. Harm potential is judged by expert review of consultation recommendations, not by observed downstream morbidity. The physician study uses U.S.-licensed generalist physicians and a defined consultation task, so results may differ in other health systems, specialties, or longitudinal care settings. The paper is strongest as a safety measurement framework and comparative evaluation; it does not prove that any particular deployment workflow is safe in practice.
Evidence Box
strongKey Claims
- •NOHARM measures clinical consultation safety through specialist-rated management options
- •Clinical RAG systems produce fewer severe-harm recommendations than generalist LLMs
- •AI assistance improves physician performance but leaves human-AI gains partly unrealized
- •Omitted recommendations account for most severe-error risk
Key Results
- •1,100 consultation tasks across 10 specialties
- •12,747 expert annotations for 4,249 clinical management options
- •Severe-harm potential in up to 24.6% of cases across evaluated systems
- •101 U.S.-licensed generalist physicians in the randomized study
Limitations & Caveats
- •Potential harm rated by experts rather than observed patient outcomes
- •Evaluation focused on primary care-to-specialist consultation cases
- •Physician study limited to U.S.-licensed generalist physicians
- •Workflow findings depend on how AI recommendations were presented and incorporated