Published online Sep 15, 2026. doi: 10.4239/wjd.123276
Revised: June 15, 2026
Accepted: July 8, 2026
Published online: September 15, 2026
Processing time: 114 Days and 11.9 Hours
Gestational diabetes mellitus (GDM) is usually diagnosed at 24-28 weeks of gestation, when opportunities for early prevention may already be limited. First-trimester maternal characteristics and placental biomarkers may provide earlier risk information, but their combined predictive value remains insufficiently de
To develop and evaluate an interpretable machine learning approach for first-trimester risk stratification of GDM.
This retrospective cohort study included singleton pregnancies undergoing first-trimester screening at MacKay Memorial Hospital, a tertiary referral center, be
Among 2756 singleton pregnancies, 352 women developed GDM (12.8%). Women who developed GDM were older and had higher pregestational body mass index and mean arterial pressure than those without GDM. First-trimester placental biomarkers, including pregnancy-associated plasma protein A, placental growth factor, and free β-human chorionic gonadotropin, were significantly lower in women who subsequently developed GDM. Among all evaluated models and resampling strategies, gradient boosting with random oversampling achieved the best overall performance, with an area under the receiver operating characteristic curve of 0.768, sensitivity of 0.647, specificity of 0.777, positive predictive value of 0.289, negative predictive value of 0.940, and F1 score of 0.400. SHAP analysis identified maternal age, mean arterial pressure, pregestational weight, pregnancy-associated plasma protein A, and placental growth factor as the major contributors to model predictions.
An interpretable machine learning approach integrating first-trimester clinical, obstetric, and biochemical para
Core Tip: This retrospective cohort study developed an interpretable machine learning model for early risk stratification of gestational diabetes mellitus using first-trimester clinical and biomarker parameters. In 2756 pregnancies, gradient boosting with random oversampling achieved moderate discrimination with high negative predictive value. Model interpretation using SHapley Additive exPlanations and patient-level heatmaps identified maternal and placental factors as key contributors. This approach may support early identification of at-risk women and facilitate targeted preventive strategies before routine mid-pregnancy screening.
- Citation: Hung SM, Chen CP, Sun FJ, Chen YY, Wang LK, Chen CY. Early risk stratification of gestational diabetes using interpretable machine learning with first-trimester screening parameters. World J Diabetes 2026; 17(9): 123276
- URL: https://www.wjgnet.com/1948-9358/full/v17/i9/123276.htm
- DOI: https://dx.doi.org/10.4239/wjd.123276
Gestational diabetes mellitus (GDM) is one of the most common metabolic complications of pregnancy and is defined as glucose intolerance with onset or first recognition during gestation[1,2]. Its pathophysiology involves progressive insulin resistance driven by placental hormones together with inadequate β-cell compensation, resulting in maternal hyperglyce
GDM is associated with a broad spectrum of adverse maternal and neonatal outcomes. Women with GDM have increased risks of hypertensive disorders of pregnancy, cesarean delivery, and long-term progression to type 2 diabetes mellitus, whereas offspring are at higher risk of macrosomia, neonatal hypoglycemia, and future metabolic disease[6,7]. Large prospective studies have further demonstrated a continuous association between maternal glycemia and adverse pregnancy outcomes, even below traditional diagnostic thresholds[8]. These findings underscore the importance of identifying high-risk women earlier in pregnancy, when preventive strategies may still be clinically meaningful.
Current screening for GDM is typically performed during 24 weeks to 28 weeks of gestation using an oral glucose tolerance test[1]. However, this timing may limit opportunities for early risk modification. Increasing evidence suggests that metabolic and placental alterations related to GDM are already present in the first trimester[9,10]. Accordingly, first-trimester biomarkers, including pregnancy-associated plasma protein A (PAPP-A), placental growth factor (PlGF), and free β-human chorionic gonadotropin (β-hCG), have been investigated for their associations with subsequent GDM, although findings remain heterogeneous across studies[10]. Our previous study showed that lower first-trimester PAPP-A and PlGF levels were associated with subsequent GDM development, and that incorporating these routinely collected biomarkers with maternal clinical characteristics improved predictive performance for early risk assessment[11]. Inter
Machine learning (ML) approaches may further improve early prediction by capturing complex and nonlinear rela
| Ref. | Sample size | Key predictors | Class imbalance | ML algorithms | Best ML performance | ML explainability | External validation | Calibration analysis |
| Wu et al[17] | 31811 | Clinical, biochemical, lipid, thyroid, and obstetric variables; 73-variable model and simplified 7-variable LR model | Not reported | DNN, SVM, KNN, and LR | DNN using 73 variables, AUC: 80%; 7-variable LR, AUC: 77% | Not reported | No | No |
| Xiong et al[18] | 490 | Routine blood tests, hepatic and renal function markers, and coagulation markers, particularly PT and aPTT | Not reported | SVM and LightGBM | SVM using PT and aPTT, sensitivity: 88.3%, specificity: 99.47%, and AUC: 94.2% | Not reported | No | No |
| Kaya et al[19] | 97 | First-visit venous plasma glucose level, maternal BMI, family history of DM, smoking, and obstetric history | Not reported | Extra trees, average blender, LightGBM, XGBoost, LR, and RF | XGBoost, AUC: 55.0%, accuracy: 66.7%, sensitivity: 80.0%, and specificity: 50.0% in nulliparous women; AUC: 73.3%, accuracy: 72.7%, sensitivity: 40.0%, and specificity: 100.0% in primiparous women | SHAP | No | No |
| Li et al[20] | 7594 | Forty-five first-trimester features; top predictors included pre-pregnancy BMI and maternal abdominal circumference at pregnancy initiation, and FPG and HbA1c at the end of the first trimester | Not reported | LR, XGBoost, RF, and other ML algorithms | XGBoost, AUC: 75% at pregnancy initiation and 99% at the end of the first trimester in the XHCM cohort; external validation AUC: 83% in the SPNPH cohort | Feature importance analysis | Yes | No |
| Zorlu et al[21] | 400 | Maternal characteristics, BMI, PAPP-A, and free β-hCG | Not reported | RF, GradBoost, and LR | GradBoost, AUC: 71.5% and accuracy: 71.3% | Not reported | No | No |
| Ni et al[22] | 956 | Common first-trimester clinical and laboratory variables selected by Spearman correlation analysis and Boruta algorithm, including pre-pregnancy BMI, SBP, and HDL-C | Not reported | LR, RF, XGBoost, LightGBM, MLP, KNN, and SVM | LR, AUC: 78.7% (95%CI: 72.3%-85.0%); RF, AUC: 77.6% (95%CI: 71.1%-84.1%) | Not reported | No | Yes |
| Zaky et al[23] | 138 | History of high glucose/diabetes, insulin, HOMA-IR, uric acid, cholesterol, urea, PT, NT-proBNP, thyroid markers, and routine blood markers | Not reported | RF, GradBoost, AdaBoost, DT, LR, SVM, Gaussian NB, KNN, CatBoost, XGBoost, LightGBM, and stacking ensemble | Stacking ensemble, accuracy: 88.8%, precision: 87.3%, sensitivity: 92.1%, and F1 score: 89.6% | SHAP | No | No |
| Bigdeli et al[24] | 106 | Age, BMI, previous abortion history, FPG, demographic variables, medical history, and clinical findings | SMOTE | DT, MLP, KNN, NB, RF, and XGBoost | RF, accuracy: 89%, precision: 86%, sensitivity: 92%, and AUC: 94% | Not reported | No | No |
| Pazaras et al[25] | 797 | Maternal demographics, obstetric history, lifestyle factors, and FFQ-derived dietary micronutrient intake | SMOTE, borderline SMOTE, adaptive synthetic sampling, and SMOTE-Tomek | LR, RF, extra trees, XGBoost, LightGBM, CatBoost, GradBoost, AdaBoost, and MLP | LR without resampling, AUC: 66.4% (95%CI: 54.2%-77.7%), sensitivity: 78.3%, and NPV: 93.2%; reduced 9-feature model, AUC: 71.2% (95%CI: 58.9%-82.5%) | SHAP | No | Yes |
| Prashanthan and Prashanthan[26] | 10000 | Synthetic demographic characteristics, clinical risk factors, and first-trimester laboratory parameters including random blood sugar, post-prandial blood sugar, HbA1c, and OGTT values | SMOTE | Multiple ML algorithms | Best model, accuracy: 71.7% and AUC: 76.9% | SHAP | No | No |
| Zhai et al[27] | 534 | BMI, SAT, VAT, maternal clinical characteristics, and first-trimester ultrasonographic indicators | IPW | XGBoost, ANN, SVM, MLR, and RF | XGBoost with GA-selected features, internal test AUC: 96.2%; external validation AUC: 87.8%, sensitivity: 70.0%, and specificity: 93.5% | GA-based feature selection and heatmaps | Yes | No |
| Louzoun et al[28] | 596 twin pregnancies | WBC count, platelet levels, BMI, and previous GDM | Not reported | LightGBM, XGBoost, and LR | LightGBM, AUC: 72% (95%CI: 69%-75%); detection rates: 28% and 42% at false-positive rates of 10% and 20%, respectively | Not reported | No | No |
| Present study | 2756 | Maternal age, MAP, pregestational weight, PAPP-A, and PlGF | SMOTE, ROS, and RUS | LR, GradBoost, XGBoost, AdaBoost, LightGBM, SVM, RF, and MLP | GradBoost + ROS, AUC: 76.8%, sensitivity: 64.7%, specificity: 77.7%, PPV: 28.9%, NPV: 94.0%, and F1 score: 40.0% | SHAP and patient-level heatmaps | No | Yes |
This retrospective cohort study included singleton pregnant women who underwent routine first-trimester screening at MacKay Memorial Hospital, a tertiary referral center, between January 2019 and July 2024. All participants underwent GDM screening at 24-28 weeks of gestation using either a 75-g or 100-g oral glucose tolerance test. GDM was diagnosed according to either the one-step International Association of Diabetes and Pregnancy Study Groups (IADPSG) criteria or the two-step Carpenter-Coustan or National Diabetes Data Group (NDDG) criteria, according to the clinical protocol used during the study period. Women meeting any of these diagnostic criteria were classified as having GDM for subsequent analyses.
Women with multiple pregnancy, pregestational diabetes mellitus, preeclampsia, major fetal anomalies, autoimmune disease, chronic renal disease, chronic hepatic disease, or maternal age < 18 years or > 50 years were excluded, whereas women with chronic hypertension or cardiovascular disease were retained because these conditions are clinically relevant to GDM risk stratification. This study was approved by the Institutional Review Board of MacKay Memorial Hospital (No. 26MMHIS097e).
Clinical and demographic data were extracted from electronic medical records, including maternal age, pregestational body mass index, previous GDM, previous gestational hypertension, previous preeclampsia, family history of diabetes mellitus, previous macrosomia (fetal birth weight ≥ 4000 g), history of polycystic ovary syndrome (PCOS), conception via in vitro fertilization (IVF), chronic hypertension, and cardiovascular disease. Family history of diabetes mellitus was defined as chronic diabetes in a first- or second-degree relative.
First-trimester parameters included gestational age at scan, crown-rump length (CRL), nuchal translucency (NT), mean arterial pressure (MAP), uterine artery pulsatility index (PI), and biochemical markers, including PAPP-A, PlGF, and free β-hCG. Biochemical markers were measured using a Kryptor analyzer (Brahms GmbH, Germany). Ultrasound examinations were performed using a Voluson E10 ultrasound system (GE Healthcare, Zipf, Austria) equipped with a 3-9 MHz transabdominal probe. CRL, NT thickness, and uterine artery PI were assessed in accordance with the guidelines of the Fetal Medicine Foundation.
The dataset was randomly divided into a development set (90%) and an independent test set (10%) using stratified samp
Eight ML algorithms were evaluated, including logistic regression, gradient boosting (GradBoost), extreme GradBoost, adaptive boosting, light GradBoost machine, support vector machine, random forest, and multilayer perceptron. Con
The primary performance metric was the area under the receiver operating characteristic curve (AUC-ROC). Secondary performance metrics included sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), F1 score, and accuracy. Binary classification metrics were calculated using a fixed probability threshold of 0.50. The final best-performing model was selected according to overall test-set discrimination performance, primarily based on AUC-ROC and F1 score. Ninety-five percent confidence intervals were calculated for the AUC-ROC estimates to improve interpretation of model stability. As a supplementary analysis, calibration curves and Brier scores were generated for the best-performing models under the ROS, RUS, and SMOTE strategies.
Model interpretability was evaluated for the final best-performing model using SHapley Additive exPlanations (SHAP), which decompose model predictions into additive feature-level contributions. SHAP values were calculated on the inde
Statistical analyses were performed using IBM SPSS Statistics for Windows, version 29.0 (IBM Corp., Armonk, NY, United States), and ML analyses were conducted using Python version 3.12.4. The Kolmogorov-Smirnov test was used to assess the normality of continuous variables. Continuous variables were compared using the Student’s t-test or Mann-Whitney U test, whereas categorical variables were analyzed using the χ2 test or Fisher’s exact test depending on expected cell counts. Continuous variables are presented as mean ± SD or median with interquartile range according to data distribution, whereas categorical variables are presented as n (%). All statistical tests were two-tailed, and a P value < 0.05 was considered statistically significant.
A total of 2756 singleton pregnancies were included in the final analysis, of which 352 women (12.8%) developed GDM and 2404 did not (Figure 1). Among the 352 women with GDM, 195 (55.4%) were diagnosed according to the IADPSG criteria, 42 (11.9%) according to the Carpenter-Coustan criteria, and 115 (32.7%) according to the NDDG criteria. Baseline maternal characteristics are summarized in Table 2. Women who developed GDM were significantly older than those without GDM (34.0 years vs 33.0 years, P < 0.001) and had higher pregestational BMI (22.9 kg/m2 vs 21.5 kg/m2, P < 0.001). Previous GDM (8.5% vs 0.8%, P < 0.001), previous macrosomia (1.7% vs 0.3%, P = 0.004), conception via IVF (13.6% vs 9.2%, P = 0.013), chronic hypertension (3.1% vs 0.7%, P < 0.001), and cardiovascular disease (1.1% vs 0.04%, P = 0.001) were also significantly more common in the GDM group. No significant differences were observed for previous gesta
| Variables | GDM (n = 352) | Non-GDM (n = 2404) | P value |
| Age (years) | 34.0 (31.0-37.0) | 33.0 (30.0-35.0) | < 0.001a |
| Pregestational BMI (kg/m2) | 22.9 (20.7-27.1) | 21.5 (19.6-23.9) | < 0.001a |
| Previous GDM | 30 (8.5) | 19 (0.8) | < 0.001a |
| Previous gestational hypertension | 2 (0.6) | 11 (0.5) | 0.671 |
| Previous preeclampsia | 5 (1.4) | 17 (0.7) | 0.181 |
| Family history of diabetes mellitus | 30 (8.5) | 182 (7.6) | 0.561 |
| Previous macrosomia | 6 (1.7) | 7 (0.3) | 0.004a |
| PCOS history | 0 (0) | 9 (0.4) | 0.613 |
| IVF | 48 (13.6) | 221 (9.2) | 0.013a |
| Chronic hypertension | 11 (3.1) | 17 (0.7) | < 0.001a |
| Cardiovascular disease | 4 (1.1) | 1 (< 0.1) | 0.001a |
First-trimester clinical, ultrasound, and biochemical parameters are presented in Table 3. Women who subsequently developed GDM had significantly higher MAP values than women without GDM (84.7 mmHg vs 81.7 mmHg, P < 0.001). First-trimester placental biomarkers were significantly lower in the GDM group, including PAPP-A (5.07 IU/L vs 5.80 IU/L, P < 0.001), PlGF (41.80 pg/mL vs 43.00 pg/mL, P = 0.049), and free β-hCG (38.60 IU/L vs 41.40 IU/L, P = 0.020). No significant differences were observed in gestational age at scan, CRL, NT thickness, or uterine artery PI between the two groups.
| Variables | GDM (n = 352) | Non-GDM (n = 2404) | P value |
| GA at scan (weeks) | 12.7 (12.4-13.0) | 12.7 (12.4-13.0) | 0.277 |
| CRL (mm) | 66.6 (62.4-70.5) | 66.6 (62.9-70.4) | 0.385 |
| NT (mm) | 1.8 (1.7-2.0) | 1.8 (1.7-2.0) | 0.879 |
| MAP (mmHg) | 84.7 (77.3-93.0) | 81.7 (75.7-88.0) | < 0.001a |
| PAPP-A (IU/L) | 5.07 (3.26-6.96) | 5.80 (4.04-8.11) | < 0.001a |
| PlGF (pg/mL) | 41.80 (29.66-53.26) | 43.00 (31.31-57.00) | 0.049a |
| Free β-hCG (IU/L) | 38.60 (25.80-58.80) | 41.40 (28.60-61.00) | 0.020a |
| Uterine artery PI | 1.60 (1.32-1.97) | 1.58 (1.31-1.89) | 0.072 |
The predictive performance of the ML models under different class imbalance handling strategies is summarized in Table 4 and Figure 2. Model performance varied substantially according to the resampling strategy used. Models trained on the original imbalanced dataset generally demonstrated high specificity but low sensitivity, indicating limited ability to identify women who subsequently developed GDM. In contrast, resampling strategies improved sensitivity across multiple models, although this was often accompanied by reduced specificity. Overall, GradBoost with ROS achieved the best overall discrimination and balanced classification performance on the independent test set, with an AUC-ROC of 0.768, sensitivity of 0.647, specificity of 0.777, PPV of 0.289, NPV of 0.940, F1 score of 0.400, and accuracy of 0.761. Cali
| Sampler | Model | AUC-ROC | Sensitivity | Specificity | PPV | NPV | F1 score | Accuracy |
| Original | LogReg | 0.676 (0.565-0.784) | 0.147 | 0.992 | 0.714 | 0.892 | 0.244 | 0.888 |
| XGBoost | 0.720 (0.618-0.808) | 0.118 | 0.988 | 0.571 | 0.888 | 0.195 | 0.880 | |
| AdaBoost | 0.674 (0.548-0.792) | 0.147 | 0.975 | 0.455 | 0.891 | 0.222 | 0.873 | |
| GradBoost | 0.731 (0.621-0.832) | 0.147 | 0.983 | 0.556 | 0.891 | 0.233 | 0.880 | |
| LightGBM | 0.702 (0.600-0.795) | 0.118 | 0.988 | 0.571 | 0.888 | 0.195 | 0.880 | |
| SVM | 0.703 (0.592-0.808) | 0.088 | 1.000 | 1.000 | 0.886 | 0.162 | 0.888 | |
| RandForest | 0.713 (0.605-0.813) | 0.029 | 0.996 | 0.500 | 0.880 | 0.056 | 0.877 | |
| MLP | 0.639 (0.531-0.740) | 0.147 | 0.938 | 0.250 | 0.887 | 0.185 | 0.841 | |
| SMOTE | LogReg | 0.675 (0.562-0.787) | 0.559 | 0.591 | 0.161 | 0.905 | 0.250 | 0.587 |
| XGBoost | 0.694 (0.588-0.795) | 0.235 | 0.979 | 0.615 | 0.901 | 0.340 | 0.888 | |
| AdaBoost | 0.674 (0.564-0.783) | 0.353 | 0.901 | 0.333 | 0.908 | 0.343 | 0.833 | |
| GradBoost | 0.659 (0.553-0.762) | 0.206 | 0.971 | 0.500 | 0.897 | 0.292 | 0.877 | |
| LightGBM | 0.696 (0.597-0.783) | 0.059 | 0.979 | 0.286 | 0.881 | 0.098 | 0.866 | |
| SVM | 0.675 (0.564-0.777) | 0.618 | 0.711 | 0.231 | 0.930 | 0.336 | 0.699 | |
| RandForest | 0.708 (0.611-0.803) | 0.147 | 0.983 | 0.556 | 0.891 | 0.233 | 0.880 | |
| MLP | 0.632 (0.531-0.736) | 0.265 | 0.905 | 0.281 | 0.898 | 0.273 | 0.826 | |
| ROS | LogReg | 0.677 (0.565-0.789) | 0.588 | 0.587 | 0.167 | 0.910 | 0.260 | 0.587 |
| XGBoost | 0.701 (0.605-0.790) | 0.176 | 0.926 | 0.250 | 0.889 | 0.207 | 0.833 | |
| AdaBoost | 0.715 (0.602-0.819) | 0.588 | 0.752 | 0.250 | 0.929 | 0.351 | 0.732 | |
| GradBoost | 0.768 (0.679-0.849) | 0.647 | 0.777 | 0.289 | 0.940 | 0.400 | 0.761 | |
| LightGBM | 0.720 (0.620-0.811) | 0.206 | 0.975 | 0.538 | 0.897 | 0.298 | 0.880 | |
| SVM | 0.707 (0.591-0.813) | 0.618 | 0.723 | 0.239 | 0.931 | 0.344 | 0.710 | |
| RandForest | 0.724 (0.618-0.823) | 0.088 | 0.996 | 0.750 | 0.886 | 0.158 | 0.884 | |
| MLP | 0.610 (0.504-0.723) | 0.235 | 0.909 | 0.267 | 0.894 | 0.250 | 0.826 | |
| RUS | LogReg | 0.680 (0.567-0.789) | 0.559 | 0.624 | 0.173 | 0.910 | 0.264 | 0.616 |
| XGBoost | 0.736 (0.642-0.821) | 0.706 | 0.640 | 0.216 | 0.939 | 0.331 | 0.649 | |
| AdaBoost | 0.662 (0.565-0.761) | 0.647 | 0.612 | 0.190 | 0.925 | 0.293 | 0.616 | |
| GradBoost | 0.717 (0.614-0.816) | 0.706 | 0.628 | 0.211 | 0.938 | 0.324 | 0.638 | |
| LightGBM | 0.737 (0.629-0.826) | 0.735 | 0.570 | 0.194 | 0.939 | 0.307 | 0.591 | |
| SVM | 0.708 (0.594-0.805) | 0.706 | 0.607 | 0.202 | 0.936 | 0.314 | 0.620 | |
| RandForest | 0.740 (0.640-0.828) | 0.706 | 0.616 | 0.205 | 0.937 | 0.318 | 0.627 | |
| MLP | 0.643 (0.524-0.753) | 0.676 | 0.525 | 0.167 | 0.920 | 0.267 | 0.543 |
Model interpretability analysis of the final GradBoost model is presented in Figure 3. Figure 3A shows the SHAP sum
In this retrospective cohort study, we developed and evaluated a prediction model for early GDM risk stratification using routinely available first-trimester maternal characteristics, ultrasound parameters, and placental biomarkers. Among the evaluated models and resampling strategies, GradBoost with ROS achieved the best overall discrimination and balanced classification performance, with an AUC-ROC of 0.768 and a high NPV of 0.940. These results indicate that routinely available first-trimester parameters may help identify women at relatively low risk of subsequent GDM before routine mid-pregnancy screening[1,2]. Because first-trimester screening is already integrated into routine prenatal care, the same clinical visit may also provide an opportunity for earlier metabolic risk assessment before the conventional timing of GDM diagnosis. Maternal age, MAP, pregestational weight, PAPP-A, and PlGF were identified as major contributors to risk prediction. These observations are consistent with previous studies reporting associations between first-trimester placental biomarkers, maternal metabolic characteristics, and subsequent GDM development[10,11].
PAPP-A is involved in the regulation of insulin-like growth factor bioavailability and maternal glucose metabolism, and lower first-trimester PAPP-A levels have been associated with insulin resistance and later development of GDM. PlGF also plays an important role in placental angiogenesis and metabolic adaptation during pregnancy, and abnormal placental vascular function may contribute to impaired glucose homeostasis and beta-cell dysfunction[10,11]. These findings support the concept that GDM is not solely a disorder of maternal glucose metabolism, but may also reflect early abnormalities in placental development and maternal vascular adaptation. Elevated MAP during early pregnancy may likewise reflect underlying insulin resistance and vascular-metabolic dysfunction, and recent studies have demonstrated associations between higher early-pregnancy blood pressure patterns and increased GDM risk[29,30]. The integration of maternal metabolic, hemodynamic, and placental biomarkers within the model therefore appears biologically reasonable. Maternal and placental factors also remained among the most influential variables in the feature importance analysis, supporting the clinical relevance of the prediction results.
Our findings further extend those of our previous study, in which conventional logistic regression analysis demon
Several previous studies have explored the use of ML methods for early GDM prediction; however, study populations, predictor selection, sample size, and reported model performance have varied substantially across studies[16,17]. Some studies have reported favorable predictive performance using extensive biochemical or clinical variables; however, application of these models in routine prenatal care may be limited by variable availability, data complexity, and diffe
Class imbalance handling had a substantial influence on model performance. In the original dataset, the models gene
Interpretability remains an important consideration when prediction models are applied in pregnancy care. Feature importance analysis identified maternal age, MAP, pregestational weight, PAPP-A, and PlGF as major contributors to model predictions, findings that were broadly consistent with established knowledge regarding GDM risk factors[2,3,10]. Patient-level SHAP analysis further illustrated how the relative contributions of individual variables varied across women, providing a more transparent interpretation of model predictions. The importance of interpretable prediction frameworks has also been increasingly emphasized in healthcare applications, particularly when models are intended to support clinical decision-making rather than purely automated systems[31,32].
The present findings may have several implications for early pregnancy care. Because the variables included in the model are routinely obtained during first-trimester prenatal screening, early GDM risk stratification may be performed without substantial additional testing or major modification of existing clinical workflows. Earlier identification of women who may be at increased risk of GDM could potentially facilitate consideration of preventive strategies evaluated in previous studies[12,13]. The relatively high NPV suggests potential utility in identifying women at relatively low risk of GDM. The relatively low PPV may also lead to false-positive classification and unnecessary surveillance during early pregnancy. The model should therefore be regarded as a possible early risk stratification tool rather than a definitive clinical prediction tool. Conventional mid-pregnancy glucose screening remains essential for diagnosis.
This study has several strengths. First, we included a relatively large cohort of singleton pregnancies with comprehensive first-trimester maternal characteristics, ultrasound parameters, and placental biomarker data obtained during routine prenatal screening, supporting potential clinical feasibility and practical applicability without substantial addi
Several limitations should also be acknowledged. First, this was a retrospective single-center study, which may limit generalizability to other ethnic populations or healthcare systems. Differences in ethnicity, prenatal screening protocols, and biomarker assay platforms across healthcare settings may influence model calibration and transportability. Second, external validation was not performed, and the robustness of the model in independent populations remains to be established. Third, some potentially relevant metabolic variables, including fasting glucose, glycated hemoglobin, and longitudinal glycemic measurements during pregnancy, were not included in the analysis. Fourth, different GDM diagnostic criteria were used during the study period, which may have introduced some degree of outcome heterogeneity, and criterion-specific sensitivity analyses were not performed because of the limited number of cases in certain diagnostic subgroups. Future multicenter prospective studies with external validation and standardized diagnostic criteria are warranted to further evaluate the robustness and transportability of the proposed model.
Our interpretable ML model integrating first-trimester maternal characteristics, ultrasound parameters, and placental biomarkers may serve as a possible early risk stratification tool for GDM and demonstrated favorable rule-out perfor
| 1. | American Diabetes Association Professional Practice Committee. 9. Pharmacologic Approaches to Glycemic Treatment: Standards of Care in Diabetes-2024. Diabetes Care. 2024;47:S158-S178. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 534] [Cited by in RCA: 485] [Article Influence: 242.5] [Reference Citation Analysis (0)] |
| 2. | McIntyre HD, Catalano P, Zhang C, Desoye G, Mathiesen ER, Damm P. Gestational diabetes mellitus. Nat Rev Dis Primers. 2019;5:47. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 1412] [Cited by in RCA: 1250] [Article Influence: 178.6] [Reference Citation Analysis (4)] |
| 3. | Plows JF, Stanley JL, Baker PN, Reynolds CM, Vickers MH. The Pathophysiology of Gestational Diabetes Mellitus. Int J Mol Sci. 2018;19:3342. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 1421] [Cited by in RCA: 1234] [Article Influence: 154.3] [Reference Citation Analysis (8)] |
| 4. | Zhu Y, Zhang C. Prevalence of Gestational Diabetes and Risk of Progression to Type 2 Diabetes: a Global Perspective. Curr Diab Rep. 2016;16:7. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 738] [Cited by in RCA: 938] [Article Influence: 93.8] [Reference Citation Analysis (0)] |
| 5. | Guariguata L, Linnenkamp U, Beagley J, Whiting DR, Cho NH. Global estimates of the prevalence of hyperglycaemia in pregnancy. Diabetes Res Clin Pract. 2014;103:176-185. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 536] [Cited by in RCA: 477] [Article Influence: 39.8] [Reference Citation Analysis (1)] |
| 6. | Ye W, Luo C, Huang J, Li C, Liu Z, Liu F. Gestational diabetes mellitus and adverse pregnancy outcomes: systematic review and meta-analysis. BMJ. 2022;377:e067946. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 28] [Cited by in RCA: 637] [Article Influence: 159.3] [Reference Citation Analysis (18)] |
| 7. | Vounzoulaki E, Khunti K, Abner SC, Tan BK, Davies MJ, Gillies CL. Progression to type 2 diabetes in women with a known history of gestational diabetes: systematic review and meta-analysis. BMJ. 2020;369:m1361. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 280] [Cited by in RCA: 807] [Article Influence: 134.5] [Reference Citation Analysis (4)] |
| 8. | HAPO Study Cooperative Research Group; Metzger BE, Lowe LP, Dyer AR, Trimble ER, Chaovarindr U, Coustan DR, Hadden DR, McCance DR, Hod M, McIntyre HD, Oats JJ, Persson B, Rogers MS, Sacks DA. Hyperglycemia and adverse pregnancy outcomes. N Engl J Med. 2008;358:1991-2002. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 4469] [Cited by in RCA: 3905] [Article Influence: 216.9] [Reference Citation Analysis (18)] |
| 9. | Sweeting A, Wong J, Murphy HR, Ross GP. A Clinical Update on Gestational Diabetes Mellitus. Endocr Rev. 2022;43:763-793. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 717] [Cited by in RCA: 639] [Article Influence: 159.8] [Reference Citation Analysis (3)] |
| 10. | Donovan BM, Nidey NL, Jasper EA, Robinson JG, Bao W, Saftlas AF, Ryckman KK. First trimester prenatal screening biomarkers and gestational diabetes mellitus: A systematic review and meta-analysis. PLoS One. 2018;13:e0201319. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 30] [Cited by in RCA: 46] [Article Influence: 5.8] [Reference Citation Analysis (0)] |
| 11. | Lu YT, Chen CP, Sun FJ, Chen YY, Wang LK, Chen CY. Associations between first-trimester screening biomarkers and maternal characteristics with gestational diabetes mellitus in Chinese women. Front Endocrinol (Lausanne). 2024;15:1383706. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 10] [Reference Citation Analysis (0)] |
| 12. | Koivusalo SB, Rönö K, Klemetti MM, Roine RP, Lindström J, Erkkola M, Kaaja RJ, Pöyhönen-Alho M, Tiitinen A, Huvinen E, Andersson S, Laivuori H, Valkama A, Meinilä J, Kautiainen H, Eriksson JG, Stach-Lempinen B. Gestational Diabetes Mellitus Can Be Prevented by Lifestyle Intervention: The Finnish Gestational Diabetes Prevention Study (RADIEL): A Randomized Controlled Trial. Diabetes Care. 2016;39:24-30. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 373] [Cited by in RCA: 325] [Article Influence: 32.5] [Reference Citation Analysis (0)] |
| 13. | Simmons D, Immanuel J, Hague WM, Teede H, Nolan CJ, Peek MJ, Flack JR, McLean M, Wong V, Hibbert E, Kautzky-Willer A, Harreiter J, Backman H, Gianatti E, Sweeting A, Mohan V, Enticott J, Cheung NW; TOBOGM Research Group. Treatment of Gestational Diabetes Mellitus Diagnosed Early in Pregnancy. N Engl J Med. 2023;388:2132-2144. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 285] [Cited by in RCA: 264] [Article Influence: 88.0] [Reference Citation Analysis (1)] |
| 14. | Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. 2019;25:44-56. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 6739] [Cited by in RCA: 4467] [Article Influence: 638.1] [Reference Citation Analysis (9)] |
| 15. | Rajkomar A, Dean J, Kohane I. Machine Learning in Medicine. N Engl J Med. 2019;380:1347-1358. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 3295] [Cited by in RCA: 2161] [Article Influence: 308.7] [Reference Citation Analysis (4)] |
| 16. | Mennickent D, Rodríguez A, Farías-Jofré M, Araya J, Guzmán-Gutiérrez E. Machine learning-based models for gestational diabetes mellitus prediction before 24-28 weeks of pregnancy: A review. Artif Intell Med. 2022;132:102378. [RCA] [PubMed] [DOI] [Full Text] [Cited by in RCA: 26] [Reference Citation Analysis (0)] |
| 17. | Wu YT, Zhang CJ, Mol BW, Kawai A, Li C, Chen L, Wang Y, Sheng JZ, Fan JX, Shi Y, Huang HF. Early Prediction of Gestational Diabetes Mellitus in the Chinese Population via Advanced Machine Learning. J Clin Endocrinol Metab. 2021;106:e1191-e1205. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 158] [Cited by in RCA: 124] [Article Influence: 24.8] [Reference Citation Analysis (0)] |
| 18. | Xiong Y, Lin L, Chen Y, Salerno S, Li Y, Zeng X, Li H. Prediction of gestational diabetes mellitus in the first 19 weeks of pregnancy using machine learning techniques. J Matern Fetal Neonatal Med. 2022;35:2457-2463. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 11] [Cited by in RCA: 26] [Article Influence: 4.3] [Reference Citation Analysis (0)] |
| 19. | Kaya Y, Bütün Z, Çelik Ö, Salik EA, Tahta T, Yavuz AA. The early prediction of gestational diabetes mellitus by machine learning models. BMC Pregnancy Childbirth. 2024;24:574. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 22] [Cited by in RCA: 14] [Article Influence: 7.0] [Reference Citation Analysis (1)] |
| 20. | Li YX, Liu YC, Wang M, Huang YL. Prediction of gestational diabetes mellitus at the first trimester: machine-learning algorithms. Arch Gynecol Obstet. 2024;309:2557-2566. [RCA] [PubMed] [DOI] [Full Text] [Cited by in RCA: 15] [Reference Citation Analysis (0)] |
| 21. | Zorlu U, Elmas B, Ergün GT, Sucu ST, Ozan E, Aydoğdu E, Şahin D, Tekin ÖM. Determining the risk of gestational diabetes using machine learning: A study on first-trimester PAPP-A and β-hCG data. Int J Gynaecol Obstet. 2025;171:1189-1196. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 2] [Cited by in RCA: 6] [Article Influence: 6.0] [Reference Citation Analysis (0)] |
| 22. | Ni H, Miao J, Chen J. Advanced Machine Learning did not Surpass Traditional Logistic Regression in First-Trimester Gestational Diabetes Mellitus Prediction: A Retrospective Single-Center Study From Eastern China. Int J Gen Med. 2025;18:2263-2274. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 4] [Reference Citation Analysis (0)] |
| 23. | Zaky H, Fthenou E, Srour L, Farrell T, Bashir M, El Hajj N, Alam T. Machine learning based model for the early detection of Gestational Diabetes Mellitus. BMC Med Inform Decis Mak. 2025;25:130. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 22] [Cited by in RCA: 11] [Article Influence: 11.0] [Reference Citation Analysis (0)] |
| 24. | Bigdeli SK, Ghazisaedi M, Ayyoubzadeh SM, Hantoushzadeh S, Ahmadi M. Predicting Gestational Diabetes Mellitus in the first trimester using machine learning algorithms: a cross-sectional study at a hospital fertility health center in Iran. BMC Med Inform Decis Mak. 2025;25:3. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 22] [Cited by in RCA: 9] [Article Influence: 9.0] [Reference Citation Analysis (0)] |
| 25. | Pazaras N, Siargkas A, Tranidou A, Apostolopoulou A, Tsakiridis I, Bamidis PD, Stavros S, Potiris A, Chourdakis M, Dagklis T. First-Trimester Gestational Diabetes Mellitus Risk Prediction with Machine Learning Techniques: Results from the BORN2020 Cohort Study. J Clin Med. 2026;15:2461. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 1] [Reference Citation Analysis (0)] |
| 26. | Prashanthan J, Prashanthan A. Machine Learning-Based Early Prediction of Gestational Diabetes Using First-Trimester Laboratory Parameters. Cureus. 2026;18:e104782. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 1] [Reference Citation Analysis (0)] |
| 27. | Zhai H, Che L, Xu T, Li N, Li C, Xin J, Zhang X, Liu Y, Li Y, Ma Z, Li Y. Opportunistic screening data for early prediction of GDM in Northern Chinese women: a multicenter machine learning study. Sci Rep. 2026;16:12818. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 1] [Reference Citation Analysis (0)] |
| 28. | Louzoun Y, Michelson T, Bennasar M, Svirsky R, Bevilacqua E, Kugler N, Kagan K, Brown RN, Rodriguez HP, Goncé A, Borrell A, Ponce J, Geipel A, Walter A, Simonini C, Strizek B, Lennartz T, Bauer A, Meli F, Torcia E, Sharabi-Nov A, Maymon R, Nicolaides KH, Meiri H. First trimester prediction of gestational diabetes mellitus by machine learning in twin pregnancies. Arch Gynecol Obstet. 2026;313:52. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 1] [Reference Citation Analysis (0)] |
| 29. | Ni W, Chen Z, Zhu M, Li Y, Lai L, Lin B, Ouyang Z, Jiang L, Jing Y, Fan J. Gestational diabetes risk associated with early pregnancy blood pressure characteristics and trajectories: A Chinese prospective cohort study. Diabetes Obes Metab. 2025;27:5072-5084. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 1] [Cited by in RCA: 2] [Article Influence: 2.0] [Reference Citation Analysis (0)] |
| 30. | Yang M, Cao Z, Mei H, Hu L, Zhu W, Zhou J, Liu J, Zhong Y, Zhou Y, Feng X, Xiang F, Xiao H, Zhou A. Blood Pressure Levels During Pregnancy and Gestational Diabetes Mellitus: A Prospective Cohort Study and Mendelian Randomization Analysis. Diabetes Metab Syndr Obes. 2025;18:4263-4275. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 2] [Reference Citation Analysis (0)] |
| 31. | Lundberg SM, Erion G, Chen H, DeGrave A, Prutkin JM, Nair B, Katz R, Himmelfarb J, Bansal N, Lee SI. From Local Explanations to Global Understanding with Explainable AI for Trees. Nat Mach Intell. 2020;2:56-67. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 7286] [Cited by in RCA: 3366] [Article Influence: 561.0] [Reference Citation Analysis (4)] |
| 32. | Holzinger A, Langs G, Denk H, Zatloukal K, Müller H. Causability and explainability of artificial intelligence in medicine. Wiley Interdiscip Rev Data Min Knowl Discov. 2019;9:e1312. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 1211] [Cited by in RCA: 573] [Article Influence: 81.9] [Reference Citation Analysis (1)] |