Whole-Study MIL Improves Prenatal CHD Screening

Masked-autoencoder pre-training, disease-robust cardiac-frame gating, and transformer pooling produce study-level CHD predictions without clinician-selected frames.

Editorial Desk·October 10, 2026·4 min readmoderate

Underlying Paper

Towards Whole-Study Screening for Congenital Heart Disease in Fetal Ultrasound Using Multiple Instance Learning

Congenital heart disease (CHD) is the most common birth defect, yet a large fraction of cases remain undetected on prenatal ultrasound, in part because current artificial-intelligence methods assume that the key diagnostic frames have already been isolated from a study, by a clinician or by a view classifier. We remove that assumption and address CHD screening directly at the level of the whole ultrasound study. We propose a two-stage framework that first learns transferable frame representations by self-supervised masked-autoencoder pre-training on unlabeled fetal ultrasound, then identifies cardiac frames with a disease-robust module and aggregates them with a transformer-based multiple instance learning (MIL) model that produces a case-level diagnosis from study-level labels alone. The model further returns its highest-scoring frames for clinician review, and a hierarchical head separates critical from non-critical CHD. On the internal test set of our multi-source development cohort (FUSE), the proposed cardiac-gated MIL model reaches an area under the curve (AUC) of 0.985 with a specificity of 0.990, outperforming the reproduced NATMED ensemble (AUC 0.861, specificity 0.600) and the FetalCLIP foundation model (AUC 0.867, specificity 0.710). On an independent external cohort, all models initially perform near chance, but label-free CORAL adaptation raises the proposed model from an AUC of 0.513 to 0.944, whereas whole-study and view-dependent baselines do not recover. These results indicate that whole-study MIL with disease-robust cardiac-frame identification is an accurate and deployable route to prenatal CHD screening.

arXiv:2609.31376Submitted: Sep 28, 2026v1

Prenatal ultrasound screening for congenital heart disease (CHD) is often treated as an image-selection problem before it becomes a diagnostic one: a clinician or a view classifier first isolates cardiac images, then a model scores them. That workflow can fail when abnormal anatomy makes a relevant frame look unlike the normal cardiac views used to train the selector. The authors instead treat a complete ultrasound study as the diagnostic unit, using multiple instance learning (MIL) to select and combine informative frames while returning the most influential ones for review.

Core Contribution

The paper's central change is architectural rather than a new image-level CHD classifier. Its cardiac-gated MIL model consumes every frame in a study, identifies frames likely to contain cardiac information, and maps the retained frame features to one subject-level CHD prediction. This removes the requirement for expert frame selection and avoids making standard-view recognition the sole gateway to diagnosis.

Figure 1 contrasts that design with expert-selected-image and two-stage view-classification pipelines. The distinction matters because the proposed gate is intended to identify cardiac content despite disease-related deviations, whereas a conventional view classifier can discard frames precisely because the anatomy is abnormal.

Figure 1. Three designs for AI-based prenatal CHD screening. (a) Expert-selected images: a clinician manually selects representative frames, each is scored by a CHD classifier, and the per-image predictions must still be combined into a diagnosis. (b) Two-stage design: a plane classifier first identifies standard cardiac views before diagnosis. Because these classifiers are designed to recognize standard cardiac structures and are predominantly trained on normal anatomy, they may discard abnormal cardiac frames. Aggregation is again required as the design produces per-image predictions. (c) Proposed whole-study design: the entire study enters at subject level. The frame encoder embeds every frame, a disease-robust cardiac-frame identifier gates which frames enter the bag, and a MIL aggregator produces one subject-level prediction while surfacing the key diagnostic frames that support it.

Technical Approach

The framework has two stages. First, the authors pre-train a frame encoder with a masked-autoencoder objective on unlabeled fetal-ultrasound frames. This lets the encoder learn ultrasound representations before it receives study-level CHD labels. In the second stage, the encoder is frozen and produces one feature vector per frame in a preprocessed study.

A cardiac-frame identification module filters these features before a transformer MIL aggregator combines them. A classification token attends across the retained frame sequence, and its pooled representation feeds a CHD-versus-healthy-control diagnosis head. The same attention weights rank the supporting frames returned to the clinician. A separate hierarchical branch further divides predicted CHD cases into critical and non-critical disease, rather than forcing a flat three-class decision.

Figure 2 makes the division of labor clear: pre-training supplies frame features, gating limits the MIL bag to cardiac content, and transformer attention creates both the study prediction and an audit trail of highly weighted frames. That is a more clinically usable output than an unexplained average of per-image scores, although attention-ranked frames are supporting evidence rather than proof of causal reasoning.

Figure 2. Overview of the proposed two-stage framework. Stage~1: self-supervised pre-training of the frame encoder with a masked-autoencoder objective on unlabeled fetal ultrasound frames. Stage~2: MIL-based CHD detection. A whole-study ultrasound is preprocessed, every frame is embedded by the pre-trained encoder, which is kept frozen (snowflake) and yields one vector per frame; the cardiac-frame identification module passes only the features of the frames it identifies as cardiac, and a transformer aggregator with a classification token pools them into a study representation. A diagnosis head predicts CHD versus HC, a hierarchical severity head, trained as a separate branch, separates critical from non-critical CHD among predicted CHD cases. The attention from the classification token to each frame ranks the key supporting frames returned for clinician review. No expert frame selection is required at any point.

Results and Analysis

On the 177-subject internal FUSE test set, the cardiac-gated MIL model reached AUC 0.985 and specificity 0.990. The reproduced NATMED ensemble reached AUC 0.861 and specificity 0.600; FetalCLIP reached AUC 0.867 and specificity 0.710. The specificity gap is consequential for a screening workflow, where false-positive referrals can create downstream burden. These internal results support the claim that whole-study pooling is more effective than the two reported comparison systems in this cohort, though they do not isolate how much of the gain comes from masked-autoencoder pre-training, the cardiac gate, or transformer aggregation.

The external CARDIUM result is more revealing. Without adaptation, the proposed model had AUC 0.513, near chance, as did the comparison models according to the reported screening plot. Label-free CORAL adaptation raised the proposed model to AUC 0.944, while the whole-study and view-dependent baselines did not recover. This is evidence that the learned representation can be adapted across acquisition domains without target labels; it is not evidence of ready-to-deploy, zero-shot generalization. The unadapted collapse is the paper's clearest operational caveat.

Figure 3 places the internal advantage and the external domain shift side by side. The result is persuasive for a multi-source development cohort and promising after label-free adaptation, but the clinical case rests on whether such adaptation can be validated prospectively across scanners, sites, and referral populations.

Figure 3. CHD screening performance. (A) Internal FUSE test set (177 subjects): AUC, sensitivity, and specificity of NATMED, FetalCLIP, and the proposed cardiac-gated MIL model. (B) External CARDIUM cohort: AUC of NATMED, FetalCLIP, and the proposed model without adaptation, and of the proposed model after label-free CORAL adaptation. The dashed line marks chance.

Interpretation

The study makes a practical argument for aligning model inputs with how ultrasound examinations are acquired: as long, heterogeneous studies rather than curated diagnostic stills. Its strongest contribution is the combination of whole-study supervision with disease-robust cardiac filtering, which addresses a failure mode of normal-anatomy view selection. The reported numbers justify further evaluation, especially because the model also surfaces candidate supporting frames. They do not yet remove the need to manage domain shift before clinical use.

Evidence Box

moderate

Key Claims

  • •Whole-study MIL can screen CHD without expert-selected frames
  • •Disease-robust cardiac-frame identification preserves diagnostically relevant abnormal frames
  • •Label-free CORAL adaptation can recover performance across cohorts

Key Results

  • •AUC 0.985 and specificity 0.990 on the 177-subject internal FUSE test set
  • •AUC 0.861 and specificity 0.600 for reproduced NATMED on internal testing
  • •AUC 0.867 and specificity 0.710 for FetalCLIP on internal testing
  • •External AUC improved from 0.513 without adaptation to 0.944 after label-free CORAL adaptation

Limitations & Caveats

  • •External performance was near chance before adaptation, with AUC 0.513
  • •Internal screening evaluation contains 177 test subjects
  • •Severity classification is evaluated on the internal test set
  • •Reported comparisons do not separate the contributions of pre-training, gating, and MIL aggregation

Related Articles

Readers are encouraged to consult the original arXiv paper for complete details. SOTA Papers does not make claims beyond what is supported by the authors' reported evidence.