Skip to main content

Evaluating deep learning sepsis prediction models in ICUs under distribution shift: a multi-centre retrospective cohort study.

NPJ digital medicine2026-03-04PubMed
Total: 83.0Rigor: 8Innovation: 9Journal: 9Clinical: 7

Summary

Across 216,536 ICU stays from HiRID, MIMIC-IV, and eICU, routine fine-tuning under distribution shift underperformed compared with retraining, fusion-training, and supervised domain adaptation. Retraining/fusion excelled with small/large target data, while domain adaptation delivered the most stable gains for medium target data, improving AUROC and normalized AUPRC.

Key Findings

  • Quantified distribution shifts across HiRID, MIMIC-IV, and eICU (216,536 ICU stays).
  • Compared five deployment strategies across multiple deep learning architectures and four target-data regimes.
  • Routine fine-tuning consistently underperformed versus alternatives.
  • Retraining and fusion-training performed best in small and large target-data regimes.
  • Supervised domain adaptation yielded the most stable gains in medium target-data settings, improving AUROC and normalized AUPRC.

Clinical Implications

ICUs should avoid naive fine-tuning when transferring sepsis predictors; instead, select deployment strategies based on target-data availability to improve reliability, reduce false alarms, and support earlier sepsis recognition.

Why It Matters

This work challenges the field’s default reliance on fine-tuning and provides data-driven guidance on when to use domain adaptation, retraining, or fusion, directly informing real-world deployment of sepsis prediction models.

Limitations

  • Retrospective design without prospective clinical deployment
  • Potential cohort-specific labeling and practice differences not fully controllable

Future Directions

Prospective trials integrating domain adaptation and fusion strategies into clinical workflows, with drift monitoring, cost-effectiveness, and impact on sepsis treatment timing and outcomes.

Study Information

Study Type
Cohort
Research Domain
Diagnosis
Evidence Level
III - Level III: retrospective multi-cohort observational evaluation of predictive models
Study Design
OTHER