Daily Endocrinology Research Analysis
Three studies stand out today in endocrinology. A multicenter AI framework for embryo evaluation showed robust, externally validated performance across clinics. A systematic review/meta-analysis supports Ki-67 (≥5%) as a highly specific marker for adrenocortical carcinoma, while highlighting the need to standardize IGF2 assays. A multicenter cohort found that treating transient hypothyroxinemia of prematurity did not improve neurodevelopment at 2 years.
Summary
Three studies stand out today in endocrinology. A multicenter AI framework for embryo evaluation showed robust, externally validated performance across clinics. A systematic review/meta-analysis supports Ki-67 (≥5%) as a highly specific marker for adrenocortical carcinoma, while highlighting the need to standardize IGF2 assays. A multicenter cohort found that treating transient hypothyroxinemia of prematurity did not improve neurodevelopment at 2 years.
Research Themes
- Trustworthy AI for reproductive endocrinology
- Diagnostic pathology of adrenocortical tumors (Ki-67/IGF2)
- Neonatal thyroid management and neurodevelopment
Selected Articles
1. Application of a methodological framework for the development and multicenter validation of reliable artificial intelligence in embryo evaluation.
Across 10 IVF clinics, a deep learning embryo-ranking model showed consistent, externally validated associations between higher AI score brackets and clinical pregnancy (fetal heartbeat), with top-tier scores yielding OR≈3.8–4.0. The authors provide a four-step methodology emphasizing curated datasets, performance assessment across variable data, and explainability via correlations with morphology.
Impact: Provides a transparent, multicenter framework that demonstrates reliable AI performance with external validation, addressing reproducibility and explainability—key barriers to clinical adoption of AI in IVF.
Clinical Implications: Clinics can consider AI scores as an adjunct for embryo selection, given consistent associations with fetal heartbeat across sites. Prospective trials are still needed to confirm improvements in live birth and to guide clinic-specific calibration and governance.
Key Findings
- AI score brackets showed monotonic increases in fetal heartbeat odds across test and independent datasets (top bracket OR ≈3.84–4.01).
- Performance generalized across clinics and age subgroups; FH-positive embryos had higher average AI scores within each age stratum.
- AI scores correlated with established morphologic quality parameters, supporting interpretability.
- A four-step development/validation framework addressed dataset curation, optimization, performance under data variability, and explainability.
Methodological Strengths
- Multicenter external validation with blind test and independent unseen clinic datasets
- Explainability via correlation of AI scores with morphology; prespecified four-step methodology
Limitations
- Non-randomized, retrospective datasets; no direct evidence of improved live birth rates
- Dependent on time-lapse imaging and specific lab workflows; potential selection biases
Future Directions: Prospective, randomized studies to test impact on live birth; site-level calibration and fairness auditing; integration with clinical decision support and cost-effectiveness analyses.
BACKGROUND: Artificial intelligence (AI) models analyzing embryo time-lapse images have been developed to predict the likelihood of pregnancy following in vitro fertilization (IVF). However, limited research exists on methods ensuring AI consistency and reliability in clinical settings during its development and validation process. We present a methodology for developing and validating an AI model across multiple datasets to demonstrate reliable performance in evaluating blastocyst-stage embryos. METHODS: This multicenter analysis utilizes time-lapse images, pregnancy outcomes, and morphologic annotations from embryos collected at 10 IVF clinics across 9 countries between 2018 and 2022. The four-step methodology for developing and evaluating the AI model include: (I) curating annotated datasets that represent the intended clinical use case; (II) developing and optimizing the AI model; (III) evaluating the AI's performance by assessing its discriminative power and associations with pregnancy probability across variable data; and (IV) ensuring interpretability and explainability by correlating AI scores with relevant morphologic features of embryo quality. Three datasets were used: the training and validation dataset (n = 16,935 embryos), the blind test dataset (n = 1,708 embryos; 3 clinics), and the independent dataset (n = 7,445 embryos; 7 clinics) derived from previously unseen clinic cohorts. RESULTS: The AI was designed as a deep learning classifier ranking embryos by score according to their likelihood of clinical pregnancy. Higher AI score brackets were associated with increased fetal heartbeat (FH) likelihood across all evaluated datasets, showing a trend of increasing odds ratios (OR). The highest OR was observed in the top G4 bracket (test dataset G4 score ≥ 7.5: OR 3.84; independent dataset G4 score ≥ 7.5: OR 4.01), while the lowest was in the G1 bracket (test dataset G1 score < 4.0: OR 0.40; independent dataset G1 score < 4.0: OR 0.45). AI score brackets G2, G3, and G4 displayed OR values above 1.0 (P < 0.05), indicating linear associations with FH likelihood. Average AI scores were consistently higher for FH-positive than for FH-negative embryos within each age subgroup. Positive correlations were also observed between AI scores and key morphologic parameters used to predict embryo quality. CONCLUSIONS: Strong AI performance across multiple datasets demonstrates the value of our four-step methodology in developing and validating the AI as a reliable adjunct to embryo evaluation.
2. The differential diagnosis of adrenocortical tumors: systematic review of Ki-67 and IGF2 and meta-analysis of Ki-67.
This PRISMA-compliant review/meta-analysis supports Ki-67 (≥5% positive cells) as a highly specific marker for adrenocortical carcinoma, with pooled specificity 0.98 and sensitivity 0.82. IGF2 is often positive in carcinomas, but heterogeneity in staining evaluation prevented pooled accuracy estimates, underscoring the need for standardized protocols.
Impact: Provides quantitative diagnostic accuracy for a widely used biomarker (Ki-67) and clarifies gaps for IGF2, informing pathology thresholds and future standardization.
Clinical Implications: In indeterminate ACTs, a Ki-67 index ≥5% strongly favors malignancy (high specificity), while acknowledging limited sensitivity; results should be integrated with Weiss/Helsinki scores and imaging. IGF2 may aid diagnosis once staining methods are standardized.
Key Findings
- Systematic review of 26 studies; meta-analysis feasible for Ki-67 but not for IGF2 due to heterogeneity.
- At a 5% cutoff, Ki-67 showed pooled specificity 0.98 and sensitivity 0.82 for identifying adrenocortical carcinoma.
- IGF2 staining frequently positive in carcinomas versus adenomas, but standardized evaluation is lacking.
Methodological Strengths
- PRISMA-compliant systematic review with meta-analysis for diagnostic accuracy
- Explicit reporting of pooled specificity/sensitivity and diagnostic odds ratio
Limitations
- Heterogeneous IHC protocols and scoring limited comparability, precluding IGF2 meta-analysis
- Potential publication bias and variable reference standards across studies
Future Directions: Standardize IGF2 staining/scoring; evaluate alternative Ki-67 thresholds and continuous measures; prospective multicenter diagnostic studies with uniform reference standards.
Distinguishing benign from malignant adrenocortical tumors (ACT) is not always easy, particularly for tumors with unclear malignant potential based on the histopathological features comprised of the Weiss score. Previous studies reported the potential utility of immunohistochemistry (IHC) markers to recognize malignancy, in particular the Insulin-like growth factor 2 (IGF2) and the proliferation marker, Ki-67. However, this information was not compiled before. Therefore, this review aimed to collect the evidence on the potential diagnosis utility of IGF2 and Ki-67 IHC staining. Additionally, a meta-analysis was performed to assess the Ki-67 accuracy to identify adrenocortical carcinoma. The systematic review and meta-analysis were conducted according to the PRISMA guidelines. From the 26 articles included in the systematic review, 21 articles provided individual data for IGF2 (n = 2) or for Ki-67 (n = 19), while 5 studies assessed both markers. IGF2 staining was positive in most carcinomas, in contrast to adenomas. However, the different immunostaining evaluation methods adopted among the studies impeded to perform a meta-analysis to assess IGF2 diagnostic accuracy. In contrast, for the most commonly used cut-off value of 5% stained cells, Ki-67 showed pooled specificity, sensitivity and log diagnostic odds ratio of 0.98 (95% CI 0.95 to 0.99), 0.82 (95% CI 0.65 to 0.92) and 4.26 (95% CI 3.40 to 5.12), respectively. At the 5% cut-off, Ki-67 demonstrated an excellent specificity to recognize malignant ACT. However. the moderate sensitivity observed indicates the need for further studies exploring alternative threshold values. Additionally, more studies using similar approaches are needed to assess the diagnostic accuracy of IGF2.Registration code in PROSPERO: CRD42022370389.
3. Treatment of Transient Hypothyroxinemia of Prematurity Does Not Improve Neurodevelopment at Two Years of Age.
In 373 very preterm infants (<32 weeks), treating THOP did not improve 2-year neurodevelopment versus no treatment. Both unadjusted and adjusted analyses showed no significant differences, challenging the rationale for routine treatment aimed at neurodevelopmental benefit.
Impact: Addresses a long-standing clinical question with multicenter data, delivering a clinically important negative result that could discourage unnecessary levothyroxine exposure in very preterm infants.
Clinical Implications: Routine levothyroxine treatment for THOP should be reconsidered; clinicians should individualize decisions and prioritize enrollment in randomized trials to assess benefits and harms.
Key Findings
- Treatment of THOP did not improve neurodevelopment at 2 years compared with no treatment (treated vs untreated OR 0.8, 95% CI 0.3–1.9).
- Presence of THOP itself was not associated with worse 2-year neurodevelopment versus no THOP (OR 1.4, 95% CI 0.8–2.3).
- Results persisted after adjustment for confounders, across a multicenter cohort of 373 very preterm infants.
Methodological Strengths
- Multicenter cohort with prospectively collected data and predefined THOP criteria
- Adjusted analyses for confounding; clinically relevant, standardized outcome at 2 years
Limitations
- Non-randomized design introduces potential indication and residual confounding bias
- Variability in treatment timing/dose and in neurodevelopmental assessments across centers
Future Directions: Randomized, placebo-controlled trials to assess levothyroxine for THOP; standardized protocols for treatment timing/dosing; long-term neurocognitive and safety outcomes.
AIM: Transient hypothyroxinemia of prematurity (THOP) has been associated with suboptimal neurodevelopment. We aimed to assess neurodevelopment in very preterm infants with treated and untreated THOP. METHODS: This study was a multicentre, cohort study, based on prospectively collected data in four French level III neonatal intensive care units. Infants born before 32 weeks of gestation between 2009 and 2020 who underwent a thyroid function test were included. THOP was defined as low free thyroxine and unelevated thyroid stimulating hormone. Infants were classified as no THOP, treated THOP, and untreated THOP. The primary outcome was suboptimal neurodevelopment at 2 years of age evaluated by clinical examination. RESULTS: Three hundred and seventy-three infants (54% male) born at a median gestational age of 28 weeks of gestation were included. There was no significant difference in neurodevelopment at 2 years of age when comparing the no THOP to the THOP group (Odds Ratio (OR) 1.4, 95% confident Interval (CI) 0.8-2.3) nor when comparing the treated with the untreated THOP group (OR 0.8, 95% CI 0.3-1.9). Results remained unchanged after adjusting for confounding factors. CONCLUSION: In very preterm infants treated THOP was not associated with improved neurodevelopment compared to untreated THOP. Numerous biases could have limited treatment effect.