Whole-Study MIL Improves Prenatal CHD Screening
Masked-autoencoder pre-training, disease-robust cardiac-frame gating, and transformer pooling produce study-level CHD predictions without clinician-selected frames.
Underlying Paper
Towards Whole-Study Screening for Congenital Heart Disease in Fetal Ultrasound Using Multiple Instance Learning
Congenital heart disease (CHD) is the most common birth defect, yet a large fraction of cases remain undetected on prenatal ultrasound, in part because current artificial-intelligence methods assume that the key diagnostic frames have already been isolated from a study, by a clinician or by a view classifier. We remove that assumption and address CHD screening directly at the level of the whole ultrasound study. We propose a two-stage framework that first learns transferable frame representations by self-supervised masked-autoencoder pre-training on unlabeled fetal ultrasound, then identifies cardiac frames with a disease-robust module and aggregates them with a transformer-based multiple instance learning (MIL) model that produces a case-level diagnosis from study-level labels alone. The model further returns its highest-scoring frames for clinician review, and a hierarchical head separates critical from non-critical CHD. On the internal test set of our multi-source development cohort (FUSE), the proposed cardiac-gated MIL model reaches an area under the curve (AUC) of 0.985 with a specificity of 0.990, outperforming the reproduced NATMED ensemble (AUC 0.861, specificity 0.600) and the FetalCLIP foundation model (AUC 0.867, specificity 0.710). On an independent external cohort, all models initially perform near chance, but label-free CORAL adaptation raises the proposed model from an AUC of 0.513 to 0.944, whereas whole-study and view-dependent baselines do not recover. These results indicate that whole-study MIL with disease-robust cardiac-frame identification is an accurate and deployable route to prenatal CHD screening.
Prenatal ultrasound screening for congenital heart disease (CHD) is often treated as an image-selection problem before it becomes a diagnostic one: a clinician or a view classifier first isolates cardiac images, then a model scores them. That workflow can fail when abnormal anatomy makes a relevant frame look unlike the normal cardiac views used to train the selector. The authors instead treat a complete ultrasound study as the diagnostic unit, using multiple instance learning (MIL) to select and combine informative frames while returning the most influential ones for review.
Core Contribution
The paper's central change is architectural rather than a new image-level CHD classifier. Its cardiac-gated MIL model consumes every frame in a study, identifies frames likely to contain cardiac information, and maps the retained frame features to one subject-level CHD prediction. This removes the requirement for expert frame selection and avoids making standard-view recognition the sole gateway to diagnosis.
Figure 1 contrasts that design with expert-selected-image and two-stage view-classification pipelines. The distinction matters because the proposed gate is intended to identify cardiac content despite disease-related deviations, whereas a conventional view classifier can discard frames precisely because the anatomy is abnormal.
Technical Approach
The framework has two stages. First, the authors pre-train a frame encoder with a masked-autoencoder objective on unlabeled fetal-ultrasound frames. This lets the encoder learn ultrasound representations before it receives study-level CHD labels. In the second stage, the encoder is frozen and produces one feature vector per frame in a preprocessed study.
A cardiac-frame identification module filters these features before a transformer MIL aggregator combines them. A classification token attends across the retained frame sequence, and its pooled representation feeds a CHD-versus-healthy-control diagnosis head. The same attention weights rank the supporting frames returned to the clinician. A separate hierarchical branch further divides predicted CHD cases into critical and non-critical disease, rather than forcing a flat three-class decision.
Figure 2 makes the division of labor clear: pre-training supplies frame features, gating limits the MIL bag to cardiac content, and transformer attention creates both the study prediction and an audit trail of highly weighted frames. That is a more clinically usable output than an unexplained average of per-image scores, although attention-ranked frames are supporting evidence rather than proof of causal reasoning.
Results and Analysis
On the 177-subject internal FUSE test set, the cardiac-gated MIL model reached AUC 0.985 and specificity 0.990. The reproduced NATMED ensemble reached AUC 0.861 and specificity 0.600; FetalCLIP reached AUC 0.867 and specificity 0.710. The specificity gap is consequential for a screening workflow, where false-positive referrals can create downstream burden. These internal results support the claim that whole-study pooling is more effective than the two reported comparison systems in this cohort, though they do not isolate how much of the gain comes from masked-autoencoder pre-training, the cardiac gate, or transformer aggregation.
The external CARDIUM result is more revealing. Without adaptation, the proposed model had AUC 0.513, near chance, as did the comparison models according to the reported screening plot. Label-free CORAL adaptation raised the proposed model to AUC 0.944, while the whole-study and view-dependent baselines did not recover. This is evidence that the learned representation can be adapted across acquisition domains without target labels; it is not evidence of ready-to-deploy, zero-shot generalization. The unadapted collapse is the paper's clearest operational caveat.
Figure 3 places the internal advantage and the external domain shift side by side. The result is persuasive for a multi-source development cohort and promising after label-free adaptation, but the clinical case rests on whether such adaptation can be validated prospectively across scanners, sites, and referral populations.
Interpretation
The study makes a practical argument for aligning model inputs with how ultrasound examinations are acquired: as long, heterogeneous studies rather than curated diagnostic stills. Its strongest contribution is the combination of whole-study supervision with disease-robust cardiac filtering, which addresses a failure mode of normal-anatomy view selection. The reported numbers justify further evaluation, especially because the model also surfaces candidate supporting frames. They do not yet remove the need to manage domain shift before clinical use.
Evidence Box
moderateKey Claims
- •Whole-study MIL can screen CHD without expert-selected frames
- •Disease-robust cardiac-frame identification preserves diagnostically relevant abnormal frames
- •Label-free CORAL adaptation can recover performance across cohorts
Key Results
- •AUC 0.985 and specificity 0.990 on the 177-subject internal FUSE test set
- •AUC 0.861 and specificity 0.600 for reproduced NATMED on internal testing
- •AUC 0.867 and specificity 0.710 for FetalCLIP on internal testing
- •External AUC improved from 0.513 without adaptation to 0.944 after label-free CORAL adaptation
Limitations & Caveats
- •External performance was near chance before adaptation, with AUC 0.513
- •Internal screening evaluation contains 177 test subjects
- •Severity classification is evaluated on the internal test set
- •Reported comparisons do not separate the contributions of pre-training, gating, and MIL aggregation