AcuityBench Exposes Miscalibrated Clinical Triage Uncertainty

A 914-case benchmark tests 12 models across four care levels, showing that conversational triage trades less over-triage for more under-triage.

Editorial Desk·October 11, 2026·4 min readstrong

Underlying Paper

AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment

Robin Linzmayer (Department of Computer ScienceColumbia UniversityDepartment of Biomedical InformaticsColumbia University)Georgianna Lin (Department of Biomedical InformaticsColumbia University)Di Coneybeare (Department of Emergency MedicineColumbia University Irving Medical Center)Jason Chu (Department of Emergency MedicineColumbia University Irving Medical Center)Trudi Cloyd (Department of Emergency MedicineColumbia University Irving Medical Center)Manish Garg (Department of Emergency MedicineColumbia University Irving Medical Center)Miles Gordon (Department of Emergency MedicineColumbia University Irving Medical Center)Elizabeth Hartofilis (Department of Emergency MedicineColumbia University Irving Medical Center)Benjamin Hong (Department of Emergency MedicineColumbia University Irving Medical Center)Ashraf Hussain (Department of Emergency MedicineColumbia University Irving Medical Center)Eugene Y. Kim (Department of Emergency MedicineColumbia University Irving Medical Center)Oluchi Iheagwara King (Department of Emergency MedicineColumbia University Irving Medical Center)Ross McCormack (Department of Emergency MedicineColumbia University Irving Medical Center)Erica Olsen (Department of Emergency MedicineColumbia University Irving Medical Center)John K. Riggins Jr (Department of Emergency MedicineColumbia University Irving Medical Center)Mustafa N. Rasheed (Department of Emergency MedicineColumbia University Irving Medical Center)Dana L. Sacco (Department of Emergency MedicineColumbia University Irving Medical Center)Vinay Saggar (Department of Emergency MedicineColumbia University Irving Medical Center)Osman R. Sayan (Department of Emergency MedicineColumbia University Irving Medical Center)Amit Shembekar (Department of Emergency MedicineColumbia University Irving Medical Center)Janice Shin-Kim (Department of Emergency MedicineColumbia University Irving Medical Center)Wendy W. Sun (Department of Emergency MedicineColumbia University Irving Medical Center)Bernard P. Chang (Department of Emergency MedicineColumbia University Irving Medical Center)David Kessler (Department of Emergency MedicineColumbia University Irving Medical Center)No\'emie Elhadad (Department of Computer ScienceColumbia UniversityDepartment of Biomedical InformaticsColumbia University)

We introduce AcuityBench, a benchmark for evaluating whether language models identify the appropriate urgency of care from user medical presentations. Existing health benchmarks emphasize medical question answering, broad health interactions, or narrow workflow-specific triage tasks, but they do not offer a unified evaluation of acuity identification across these settings. AcuityBench addresses this gap by harmonizing five public datasets spanning user conversations, online forum posts, clinical vignettes, and patient portal messages under a shared four-level acuity framework ranging from home monitoring to immediate emergency care. The benchmark contains 914 cases, including 697 consensus cases for standard accuracy evaluation and 217 physician-confirmed ambiguous cases for uncertainty-aware evaluation. It supports two complementary task formats: explicit four-way classification in a QA setting, and free-form conversational responses evaluated with a rubric-based judge anchored to the same framework. Across 12 frontier proprietary and open-weight models, we find substantial variation in clear-case acuity accuracy and error direction. Comparing task formats reveals a systematic tradeoff: conversational responses reduce over-triage but increase under-triage relative to QA, especially in higher-acuity cases. In ambiguous cases, no model closely matches the distribution of physician judgments, and model predictions are more concentrated than expert clinical uncertainty. We also compare expert and model adjudication on a subset of maximally ambiguous cases, using those cases to examine the role of clinical uncertainty in label disagreement. Together, these results position acuity identification as a distinct safety-critical capability and show that AcuityBench enables systematic comparison and stress-testing of how well models guide users to the right level of care in real-world health use.

arXiv:2605.11398Submitted: Sep 28, 2026v2

Choosing the appropriate urgency of care is a different problem from answering a medical question correctly. A response can contain clinically plausible information yet still direct a user to the wrong setting, with under-triage at the emergent end carrying particular safety consequences. AcuityBench evaluates that decision directly, separating agreement on clear cases from alignment with clinician disagreement when the case itself admits more than one defensible recommendation.

The benchmark harmonizes five public sources—user conversations, forum posts, clinical vignettes, and patient-portal messages—into a four-level scale: home monitoring, seeing a doctor within weeks, urgent outpatient care within 24–48 hours, and immediate emergency-department care. It contains 914 cases: 697 physician-consensus cases for standard accuracy analysis and 217 physician-confirmed ambiguous cases for uncertainty-aware evaluation. The authors test 12 proprietary and open-weight models in both an explicit four-way QA format and a free-response conversational format scored against the same acuity framework.

Core Contribution

That distinction makes AcuityBench more informative than a benchmark that reports only exact-match accuracy. It asks two separate questions: whether a model recognizes acuity when clinicians agree, and whether its distribution of recommendations reflects the uncertainty clinicians express when the presentation is genuinely borderline. For health-facing assistants, the second question is not cosmetic. An apparently decisive answer can be poorly calibrated even when its modal label is often acceptable.

Figure 1 lays out the pipeline from heterogeneous source data through label normalization and physician-panel annotation to the two evaluation formats.

Figure 1. Overview of AcuityBench construction and evaluation. Heterogeneous data sources were normalized into a four-level acuity framework, labeled through direct mapping or physician-panel annotation, and evaluated in QA and free-response conversation formats, yielding consensus and ambiguous subsets for downstream accuracy, uncertainty, and error analyses.

Technical Approach

For each ambiguous case, the authors query a model five times at temperature 1.0 and convert the samples into an empirical distribution over the four ordered acuity labels. They compare that distribution with the physician-derived soft distribution using Jensen–Shannon divergence and Wasserstein-1 distance. Jensen–Shannon divergence measures distributional mismatch, while Wasserstein-1 preserves the ordinal structure: moving a recommendation from home monitoring to emergency care is treated as a larger error than moving it one adjacent level.

The paper also tests a panel-substitution scenario. A model replaces one physician rater, and the resulting panel is compared with the all-human panel through Krippendorff’s alpha, leave-one-out median changes, and the magnitude of those changes. This does not establish clinical usefulness in deployment; it is a targeted test of whether model outputs behave like one contributor to a clinician panel.

The comparison of QA and free response is equally consequential. A fixed-choice prompt exposes the model’s categorical preference directly, whereas a conversational response permits qualification and explanation. The reported format effect is systematic rather than uniformly beneficial: conversational outputs reduce over-triage but increase under-triage, especially for higher-acuity presentations. A more reassuring tone or broader explanation therefore should not be mistaken for safer triage.

Results and Analysis

On the 217 ambiguous cases, no individual model closely reproduced the physician uncertainty distribution. Table 15 reports Jensen–Shannon divergence from 0.245 to 0.327 and Wasserstein-1 distance from 0.816 to 0.957 across the evaluated models. The narrow spread is itself revealing: the paper’s concern is not confined to one provider or model family. Model samples were consistently more concentrated than physician judgments and under-represented the spread of clinical opinion.

Pooling GPT-5.4, Claude Opus 4.7, and Gemini 2.5 Pro improved distributional alignment: the frontier panel reached Jensen–Shannon divergence 0.191 and Wasserstein-1 0.738, versus individual-model means of 0.270 and 0.882. But the pooled samples selected one acuity label in 91.2% of ambiguous cases, while human raters showed comparably low entropy in only 8.3%. Ensembling makes the aggregate distribution closer to clinicians without solving the central calibration failure: it remains much more certain case by case.

The panel also skewed upward in acuity. Its mean ordinal position exceeded the human distribution by 0.320 levels; 66.4% of cases shifted upward and 28.6% downward. That directionality matters because a benchmark can show reasonable aggregate overlap while concealing a systematic tendency toward more intensive care recommendations.

Figure 4. Prediction distribution on boundary-label cases by model (QA format, mode of five samples; N = 31 A|B, 48 B|C, and 91 C|D cases). Bars decompose each model's predictions into lower constituent, upper constituent, below-set, and above-set outcomes. Dashed lines mark constituent accuracy.

Limits in Practice

The evidence is broad for a benchmark study, but it remains an evaluation rather than a clinical deployment study. The cases draw on public datasets and may include source-specific artifacts; the physician panel is drawn from emergency medicine; and the four-level framework abstracts care decisions that depend on patient context, resource access, and local practice. The conversational evaluation uses a single judge model, and the adjudication analysis is restricted to GPT-5.4. AcuityBench is therefore a useful stress test for clinical-acuity behavior, not a warrant to use any evaluated model for unsupervised patient triage.

Evidence Box

strong

Key Claims

  • •Acuity identification is a distinct safety-critical capability
  • •Free-response triage changes the direction of model errors
  • •Model uncertainty does not match physician disagreement on boundary cases

Key Results

  • •914 cases from 5 public datasets, including 697 consensus and 217 ambiguous cases
  • •12 models evaluated across 4 ordered acuity levels
  • •Ambiguous-case JSD ranged from 0.245 to 0.327 and W-1 from 0.816 to 0.957
  • •Frontier panel selected one label in 91.2% of ambiguous cases versus 8.3% for human raters

Limitations & Caveats

  • •Public-source cases may retain dataset-specific artifacts
  • •Physician panel drawn from emergency medicine
  • •Conversational responses evaluated by a single judge model
  • •Adjudication analysis restricted to GPT-5.4

Related Articles

Readers are encouraged to consult the original arXiv paper for complete details. SOTA Papers does not make claims beyond what is supported by the authors' reported evidence.