Refined Labels Shift Pulmonary Embolism Segmentation Scores
Radiologist-reviewed masks raise measured nnU-Net Dice scores by 0.143–0.188, exceeding score changes from altering training-data composition.
Underlying Paper
Model Effect or Label Effect? Refined Annotations and a Human-Referenced Benchmark for Pulmonary Embolism Segmentation
Purpose: To quantify how evaluation annotations influence measured pulmonary embolism (PE) segmentation performance relative to model training changes, and to establish a human-referenced framework. Materials and Methods: This retrospective study screened 166 voxel-annotated CT pulmonary angiography cases from CADPE (n=91), FUMPE (n=35), and READ (n=40); 149 were included. A primary rater annotated PE by protocol, and a senior thoracic radiologist reviewed and revised all segmentations. Three additional raters at three centers annotated a 15-case subset. The label effect was measured by evaluating two pretrained nnU-Net models (nnU-Net-A, nnU-Net-B) against original and refined annotations. The model effect was measured by comparing the same architecture trained on different dataset combinations with annotations fixed. The benchmark model (nnPE) was trained with leave-one-dataset-out and pooled five-fold cross-validation. Four metric categories were analyzed with case-paired Wilcoxon signed-rank tests, Benjamini-Hochberg correction, and bootstrap 95% CIs. Results: Changing only the annotation increased mean DSC by 0.143 (0.122-0.166) for nnU-Net-A and 0.188 (0.163-0.213) for nnU-Net-B (both P < .001), whereas changing training-dataset composition changed DSC by 0.028. The label effect exceeded the model effect on CADPE and FUMPE and was 0.045 on READ. Within-mask attenuation SD fell in all three datasets after re-annotation (all P < .001). nnPE reached DSC 0.72 +/- 0.22 on pooled cross-validation but scored below all four annotators across 52 paired comparisons (all corrected P < .05). Conclusion: Evaluation annotations affected measured PE segmentation performance at least as much as model training choices. A human-referenced evaluation framework is publicly available for future study.
Pulmonary embolism segmentation is usually judged against a single voxel-level reference mask, even though small boundary decisions, missed emboli, and anatomically implausible regions can change overlap metrics substantially. This study asks a practical question that benchmark reporting often avoids: when a model score changes, how much is a model effect and how much is a label effect? The authors re-annotate three public CT pulmonary angiography datasets under a shared protocol and use the revised masks to build a human-referenced evaluation framework.
Core Contribution
The central contribution is an empirical separation of evaluation-label effects from training-data effects. The authors screened 166 cases from CADPE, FUMPE, and READ, retaining 149 cases: 76 CADPE, 33 FUMPE, and 40 READ. A primary rater produced protocol-guided masks, a senior thoracic radiologist reviewed and revised every segmentation, and three additional raters from three centers annotated a 15-case subset.
That design makes the paper more than a dataset cleanup. It tests pretrained nnU-Net models against both original and refined labels while holding the models fixed, then changes training-dataset combinations while holding the annotations fixed. The resulting comparison targets a recurrent ambiguity in medical-image segmentation: a Dice gain can reflect closer alignment with an imperfect benchmark rather than a better representation of disease.
Figure 1 shows that the three datasets differ in case selection, central pulmonary-artery attenuation, embolus location, and registered voxel-wise embolus density. Those differences matter because training-set composition and annotation conventions are entangled in ordinary cross-dataset comparisons.
Technical Approach
The evaluation uses four metric categories and case-paired Wilcoxon signed-rank tests, with Benjamini-Hochberg correction and bootstrap 95% confidence intervals. For the label-effect analysis, two pretrained models, nnU-Net-A and nnU-Net-B, are evaluated against original and refined annotations. For the model-effect analysis, the architecture is kept unchanged while training-dataset combinations vary. The benchmark model, nnPE, is evaluated through leave-one-dataset-out experiments and pooled five-fold cross-validation.
The re-annotation analysis also examines what changed inside the masks. Figure 2 makes the intended distinction visible: original masks are overlaid against an individual re-annotation and a STAPLE consensus from the other three annotators, with axial and coronal close-ups around disputed regions. The comparison is useful because a scalar Dice score alone cannot distinguish a missed small embolus from an over-segmented vessel, lung region, or bronchus.
Results and Analysis
Changing only the evaluation annotation increased mean Dice similarity coefficient by 0.143 for nnU-Net-A, with a 95% confidence interval of 0.122–0.166, and by 0.188 for nnU-Net-B, with a 95% confidence interval of 0.163–0.213. Both comparisons report . By contrast, changing training-dataset composition changed DSC by 0.028. The paper reports that the label effect exceeded the model effect on CADPE and FUMPE; on READ, the reported label effect was 0.045.
This is a consequential result for PE benchmarks. A 0.143–0.188 change in measured DSC is much larger than the 0.028 change attributed to training-set composition in this experiment. It does not imply that model design is unimportant. It shows that headline segmentation comparisons can be dominated by the chosen reference standard when the target lesions are small and annotations contain systematic errors.
Figure 3 connects the score shift to concrete mask changes. It contrasts original and refined lesion-volume distributions, per-case lesion counts, error patterns identified by two reviewers, and within-lesion attenuation. The reported decrease in within-mask attenuation standard deviation across all three datasets after re-annotation, each with , is consistent with removing heterogeneous tissue from embolus masks. That supports the claim that refinement changed anatomical validity rather than merely redrawing contours to favor a model.
The benchmark model nnPE reached pooled five-fold cross-validation DSC of . Yet it scored below all four annotators in 52 paired comparisons, all with corrected . The result sets an appropriately demanding reference point: a pooled model can produce usable segmentations while still failing to match human agreement under the authors' evaluation. The evidence supports the narrower claim that label quality materially changes measured performance; it does not establish that the revised labels are a final clinical ground truth.
Limits for Deployment
The study is retrospective and covers 149 cases from three datasets, so its numerical effect sizes should not be assumed to transfer unchanged to other scanners, acquisition protocols, embolus prevalence, or institutions. The multi-rater comparison uses 15 cases, limiting how precisely it characterizes inter-rater variation. More fundamentally, the refined masks are radiologist-reviewed references rather than pathology-confirmed truth. The paper strengthens benchmark interpretation, but it does not test whether nnPE changes diagnostic decisions or patient outcomes.
Evidence Box
strongKey Claims
- •Evaluation-label choice can affect PE segmentation scores as much as or more than training-data composition
- •Protocol-guided refinement reduces annotation errors in public PE datasets
- •Human-referenced evaluation exposes a gap between nnPE and annotator performance
Key Results
- •Refined labels increased mean DSC by 0.143 for nnU-Net-A (95% CI 0.122–0.166; P < .001)
- •Refined labels increased mean DSC by 0.188 for nnU-Net-B (95% CI 0.163–0.213; P < .001)
- •Changing training-dataset composition changed DSC by 0.028
- •nnPE achieved pooled five-fold DSC 0.72 ± 0.22 and trailed all four annotators in 52 paired comparisons
Limitations & Caveats
- •Retrospective evaluation of 149 cases from three datasets
- •Inter-rater subset contains only 15 cases
- •Refined masks are radiologist-reviewed references rather than pathology-confirmed ground truth
- •No evidence that segmentation differences improve clinical decisions or outcomes