AcuityBench Exposes Miscalibrated Clinical Triage Uncertainty
A 914-case benchmark tests 12 models across four care levels, showing that conversational triage trades less over-triage for more under-triage.
Underlying Paper
AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment
We introduce AcuityBench, a benchmark for evaluating whether language models identify the appropriate urgency of care from user medical presentations. Existing health benchmarks emphasize medical question answering, broad health interactions, or narrow workflow-specific triage tasks, but they do not offer a unified evaluation of acuity identification across these settings. AcuityBench addresses this gap by harmonizing five public datasets spanning user conversations, online forum posts, clinical vignettes, and patient portal messages under a shared four-level acuity framework ranging from home monitoring to immediate emergency care. The benchmark contains 914 cases, including 697 consensus cases for standard accuracy evaluation and 217 physician-confirmed ambiguous cases for uncertainty-aware evaluation. It supports two complementary task formats: explicit four-way classification in a QA setting, and free-form conversational responses evaluated with a rubric-based judge anchored to the same framework. Across 12 frontier proprietary and open-weight models, we find substantial variation in clear-case acuity accuracy and error direction. Comparing task formats reveals a systematic tradeoff: conversational responses reduce over-triage but increase under-triage relative to QA, especially in higher-acuity cases. In ambiguous cases, no model closely matches the distribution of physician judgments, and model predictions are more concentrated than expert clinical uncertainty. We also compare expert and model adjudication on a subset of maximally ambiguous cases, using those cases to examine the role of clinical uncertainty in label disagreement. Together, these results position acuity identification as a distinct safety-critical capability and show that AcuityBench enables systematic comparison and stress-testing of how well models guide users to the right level of care in real-world health use.
Choosing the appropriate urgency of care is a different problem from answering a medical question correctly. A response can contain clinically plausible information yet still direct a user to the wrong setting, with under-triage at the emergent end carrying particular safety consequences. AcuityBench evaluates that decision directly, separating agreement on clear cases from alignment with clinician disagreement when the case itself admits more than one defensible recommendation.
The benchmark harmonizes five public sources—user conversations, forum posts, clinical vignettes, and patient-portal messages—into a four-level scale: home monitoring, seeing a doctor within weeks, urgent outpatient care within 24–48 hours, and immediate emergency-department care. It contains 914 cases: 697 physician-consensus cases for standard accuracy analysis and 217 physician-confirmed ambiguous cases for uncertainty-aware evaluation. The authors test 12 proprietary and open-weight models in both an explicit four-way QA format and a free-response conversational format scored against the same acuity framework.
Core Contribution
That distinction makes AcuityBench more informative than a benchmark that reports only exact-match accuracy. It asks two separate questions: whether a model recognizes acuity when clinicians agree, and whether its distribution of recommendations reflects the uncertainty clinicians express when the presentation is genuinely borderline. For health-facing assistants, the second question is not cosmetic. An apparently decisive answer can be poorly calibrated even when its modal label is often acceptable.
Figure 1 lays out the pipeline from heterogeneous source data through label normalization and physician-panel annotation to the two evaluation formats.
Technical Approach
For each ambiguous case, the authors query a model five times at temperature 1.0 and convert the samples into an empirical distribution over the four ordered acuity labels. They compare that distribution with the physician-derived soft distribution using Jensen–Shannon divergence and Wasserstein-1 distance. Jensen–Shannon divergence measures distributional mismatch, while Wasserstein-1 preserves the ordinal structure: moving a recommendation from home monitoring to emergency care is treated as a larger error than moving it one adjacent level.
The paper also tests a panel-substitution scenario. A model replaces one physician rater, and the resulting panel is compared with the all-human panel through Krippendorff’s alpha, leave-one-out median changes, and the magnitude of those changes. This does not establish clinical usefulness in deployment; it is a targeted test of whether model outputs behave like one contributor to a clinician panel.
The comparison of QA and free response is equally consequential. A fixed-choice prompt exposes the model’s categorical preference directly, whereas a conversational response permits qualification and explanation. The reported format effect is systematic rather than uniformly beneficial: conversational outputs reduce over-triage but increase under-triage, especially for higher-acuity presentations. A more reassuring tone or broader explanation therefore should not be mistaken for safer triage.
Results and Analysis
On the 217 ambiguous cases, no individual model closely reproduced the physician uncertainty distribution. Table 15 reports Jensen–Shannon divergence from 0.245 to 0.327 and Wasserstein-1 distance from 0.816 to 0.957 across the evaluated models. The narrow spread is itself revealing: the paper’s concern is not confined to one provider or model family. Model samples were consistently more concentrated than physician judgments and under-represented the spread of clinical opinion.
Pooling GPT-5.4, Claude Opus 4.7, and Gemini 2.5 Pro improved distributional alignment: the frontier panel reached Jensen–Shannon divergence 0.191 and Wasserstein-1 0.738, versus individual-model means of 0.270 and 0.882. But the pooled samples selected one acuity label in 91.2% of ambiguous cases, while human raters showed comparably low entropy in only 8.3%. Ensembling makes the aggregate distribution closer to clinicians without solving the central calibration failure: it remains much more certain case by case.
The panel also skewed upward in acuity. Its mean ordinal position exceeded the human distribution by 0.320 levels; 66.4% of cases shifted upward and 28.6% downward. That directionality matters because a benchmark can show reasonable aggregate overlap while concealing a systematic tendency toward more intensive care recommendations.
Limits in Practice
The evidence is broad for a benchmark study, but it remains an evaluation rather than a clinical deployment study. The cases draw on public datasets and may include source-specific artifacts; the physician panel is drawn from emergency medicine; and the four-level framework abstracts care decisions that depend on patient context, resource access, and local practice. The conversational evaluation uses a single judge model, and the adjudication analysis is restricted to GPT-5.4. AcuityBench is therefore a useful stress test for clinical-acuity behavior, not a warrant to use any evaluated model for unsupervised patient triage.
Evidence Box
strongKey Claims
- •Acuity identification is a distinct safety-critical capability
- •Free-response triage changes the direction of model errors
- •Model uncertainty does not match physician disagreement on boundary cases
Key Results
- •914 cases from 5 public datasets, including 697 consensus and 217 ambiguous cases
- •12 models evaluated across 4 ordered acuity levels
- •Ambiguous-case JSD ranged from 0.245 to 0.327 and W-1 from 0.816 to 0.957
- •Frontier panel selected one label in 91.2% of ambiguous cases versus 8.3% for human raters
Limitations & Caveats
- •Public-source cases may retain dataset-specific artifacts
- •Physician panel drawn from emergency medicine
- •Conversational responses evaluated by a single judge model
- •Adjudication analysis restricted to GPT-5.4