Refined Labels Shift Pulmonary Embolism Segmentation Scores

Radiologist-reviewed masks raise measured nnU-Net Dice scores by 0.143–0.188, exceeding score changes from altering training-data composition.

Editorial Desk·August 30, 2026·5 min readstrong

Underlying Paper

Model Effect or Label Effect? Refined Annotations and a Human-Referenced Benchmark for Pulmonary Embolism Segmentation

Purpose: To quantify how evaluation annotations influence measured pulmonary embolism (PE) segmentation performance relative to model training changes, and to establish a human-referenced framework. Materials and Methods: This retrospective study screened 166 voxel-annotated CT pulmonary angiography cases from CADPE (n=91), FUMPE (n=35), and READ (n=40); 149 were included. A primary rater annotated PE by protocol, and a senior thoracic radiologist reviewed and revised all segmentations. Three additional raters at three centers annotated a 15-case subset. The label effect was measured by evaluating two pretrained nnU-Net models (nnU-Net-A, nnU-Net-B) against original and refined annotations. The model effect was measured by comparing the same architecture trained on different dataset combinations with annotations fixed. The benchmark model (nnPE) was trained with leave-one-dataset-out and pooled five-fold cross-validation. Four metric categories were analyzed with case-paired Wilcoxon signed-rank tests, Benjamini-Hochberg correction, and bootstrap 95% CIs. Results: Changing only the annotation increased mean DSC by 0.143 (0.122-0.166) for nnU-Net-A and 0.188 (0.163-0.213) for nnU-Net-B (both P < .001), whereas changing training-dataset composition changed DSC by 0.028. The label effect exceeded the model effect on CADPE and FUMPE and was 0.045 on READ. Within-mask attenuation SD fell in all three datasets after re-annotation (all P < .001). nnPE reached DSC 0.72 +/- 0.22 on pooled cross-validation but scored below all four annotators across 52 paired comparisons (all corrected P < .05). Conclusion: Evaluation annotations affected measured PE segmentation performance at least as much as model training choices. A human-referenced evaluation framework is publicly available for future study.

arXiv:2608.24486Submitted: Aug 26, 2026v1

Pulmonary embolism segmentation is usually judged against a single voxel-level reference mask, even though small boundary decisions, missed emboli, and anatomically implausible regions can change overlap metrics substantially. This study asks a practical question that benchmark reporting often avoids: when a model score changes, how much is a model effect and how much is a label effect? The authors re-annotate three public CT pulmonary angiography datasets under a shared protocol and use the revised masks to build a human-referenced evaluation framework.

Core Contribution

The central contribution is an empirical separation of evaluation-label effects from training-data effects. The authors screened 166 cases from CADPE, FUMPE, and READ, retaining 149 cases: 76 CADPE, 33 FUMPE, and 40 READ. A primary rater produced protocol-guided masks, a senior thoracic radiologist reviewed and revised every segmentation, and three additional raters from three centers annotated a 15-case subset.

That design makes the paper more than a dataset cleanup. It tests pretrained nnU-Net models against both original and refined labels while holding the models fixed, then changes training-dataset combinations while holding the annotations fixed. The resulting comparison targets a recurrent ambiguity in medical-image segmentation: a Dice gain can reflect closer alignment with an imperfect benchmark rather than a better representation of disease.

Figure 1 shows that the three datasets differ in case selection, central pulmonary-artery attenuation, embolus location, and registered voxel-wise embolus density. Those differences matter because training-set composition and annotation conventions are entangled in ordinary cross-dataset comparisons.

Figure 1. Characteristics of the three re-annotated PE datasets. (a) Case-selection flowchart yielding 149 cases (CADPE, n = 76; FUMPE, n = 33; READ, n = 40); excluded case IDs shown. (b) Mean CT attenuation of the central pulmonary arteries per dataset (HU). (c) Per-case distribution of the anatomic location with the maximum cumulative embolus burden. (d) Voxel-wise embolus density after registration to a common lung template, per dataset and pooled (ALL): anterior view of the lung volume (top row), medial (mediastinal) surface view (middle row), and the pulmonary-arterial surface (bottom row); color bars encode local voxel count. R/L = right/left lung.

Technical Approach

The evaluation uses four metric categories and case-paired Wilcoxon signed-rank tests, with Benjamini-Hochberg correction and bootstrap 95% confidence intervals. For the label-effect analysis, two pretrained models, nnU-Net-A and nnU-Net-B, are evaluated against original and refined annotations. For the model-effect analysis, the architecture is kept unchanged while training-dataset combinations vary. The benchmark model, nnPE, is evaluated through leave-one-dataset-out experiments and pooled five-fold cross-validation.

The re-annotation analysis also examines what changed inside the masks. Figure 2 makes the intended distinction visible: original masks are overlaid against an individual re-annotation and a STAPLE consensus from the other three annotators, with axial and coronal close-ups around disputed regions. The comparison is useful because a scalar Dice score alone cannot distinguish a missed small embolus from an over-segmented vessel, lung region, or bronchus.

Figure 2. Visual comparison of original and refined PE masks with STAPLE consensus. For each case, the same PE is displayed in the axial and coronal planes; the leftmost image of each plane shows the whole slice, and the two adjacent panels (Zoom: A, Zoom: B) magnify the two regions of interest outlined by the red dashed boxes labeled A and B on the full image. Overlays: original annotation (red), Annotation 1 (green), and STAPLE consensus of Annotation 2–4 (blue). Ax = axial, Cor = coronal.

Results and Analysis

Changing only the evaluation annotation increased mean Dice similarity coefficient by 0.143 for nnU-Net-A, with a 95% confidence interval of 0.122–0.166, and by 0.188 for nnU-Net-B, with a 95% confidence interval of 0.163–0.213. Both comparisons report P<.001P < .001. By contrast, changing training-dataset composition changed DSC by 0.028. The paper reports that the label effect exceeded the model effect on CADPE and FUMPE; on READ, the reported label effect was 0.045.

This is a consequential result for PE benchmarks. A 0.143–0.188 change in measured DSC is much larger than the 0.028 change attributed to training-set composition in this experiment. It does not imply that model design is unimportant. It shows that headline segmentation comparisons can be dominated by the chosen reference standard when the target lesions are small and annotations contain systematic errors.

Figure 3 connects the score shift to concrete mask changes. It contrasts original and refined lesion-volume distributions, per-case lesion counts, error patterns identified by two reviewers, and within-lesion attenuation. The reported decrease in within-mask attenuation standard deviation across all three datasets after re-annotation, each with P<.001P < .001, is consistent with removing heterogeneous tissue from embolus masks. That supports the claim that refinement changed anatomical validity rather than merely redrawing contours to favor a model.

Figure 3. Per-case embolus volume, count, annotation-error patterns, and mask attenuation before (Original) and after (Refined) re-annotation across the three datasets. Main panel: Kernel density estimates of per-lesion volume (log scale) for original (dashed) and refined (solid) annotations, colored by dataset; triangles mark medians. (a) Per-case lesion counts (box plots; light = Original, dark = Refined). (b) Radar plot of annotation-error types identified by two raters reviewing the original annotations, as a percentage of cases per dataset (FP PA/PV/Lung/Bronchi = false-positive over- segmentation onto pulmonary artery, pulmonary vein, lung parenchyma, or bronchi; FN = missed emboli; Error = anatomically implausible error). (c) Per-lesion mask attenuation (mean and SD, in Hounsfield units) before and after re-annotation, shown as combined box, strip, and half-violin plots.

The benchmark model nnPE reached pooled five-fold cross-validation DSC of 0.72±0.220.72 \pm 0.22. Yet it scored below all four annotators in 52 paired comparisons, all with corrected P<.05P < .05. The result sets an appropriately demanding reference point: a pooled model can produce usable segmentations while still failing to match human agreement under the authors' evaluation. The evidence supports the narrower claim that label quality materially changes measured performance; it does not establish that the revised labels are a final clinical ground truth.

Limits for Deployment

The study is retrospective and covers 149 cases from three datasets, so its numerical effect sizes should not be assumed to transfer unchanged to other scanners, acquisition protocols, embolus prevalence, or institutions. The multi-rater comparison uses 15 cases, limiting how precisely it characterizes inter-rater variation. More fundamentally, the refined masks are radiologist-reviewed references rather than pathology-confirmed truth. The paper strengthens benchmark interpretation, but it does not test whether nnPE changes diagnostic decisions or patient outcomes.

Evidence Box

strong

Key Claims

  • Evaluation-label choice can affect PE segmentation scores as much as or more than training-data composition
  • Protocol-guided refinement reduces annotation errors in public PE datasets
  • Human-referenced evaluation exposes a gap between nnPE and annotator performance

Key Results

  • Refined labels increased mean DSC by 0.143 for nnU-Net-A (95% CI 0.122–0.166; P < .001)
  • Refined labels increased mean DSC by 0.188 for nnU-Net-B (95% CI 0.163–0.213; P < .001)
  • Changing training-dataset composition changed DSC by 0.028
  • nnPE achieved pooled five-fold DSC 0.72 ± 0.22 and trailed all four annotators in 52 paired comparisons

Limitations & Caveats

  • Retrospective evaluation of 149 cases from three datasets
  • Inter-rater subset contains only 15 cases
  • Refined masks are radiologist-reviewed references rather than pathology-confirmed ground truth
  • No evidence that segmentation differences improve clinical decisions or outcomes

Related Articles

Readers are encouraged to consult the original arXiv paper for complete details. SOTA Papers does not make claims beyond what is supported by the authors' reported evidence.