BPG is committed to discovery and dissemination of knowledge
Retrospective Cohort Study Open Access
Copyright: ©Author(s) 2026. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution-NonCommercial (CC BY-NC 4.0) license. No commercial re-use. See permissions. Published by Baishideng Publishing Group Inc.
World J Diabetes. Sep 15, 2026; 17(9): 122555
Published online Sep 15, 2026. doi: 10.4239/wjd.122555
Development and validation of an interpretable machine learning model for predicting progression from prediabetes to type 2 diabetes
Zhi-Yuan Fan, Xian-Hui Ran, Na Wang, Tian-Yi Zhao, Hui Li, Jin Wu, Zhen Yang, Gang Chen, Xiao Ma, Health Checkup Center, China-Japan Friendship Hospital, Beijing 100029, China
Xiao Liu, Department of Pharmacy, China-Japan Friendship Hospital, Beijing 100029, China
Lei Yang, Health Checkup Center, Henan Honliv Hospital, Xinxiang 453499, Henan Province, China
Xiao Ma, State Key Laboratory of Respiratory Health and Multimorbidity, China-Japan Friendship Hospital, Beijing 100029, China
ORCID number: Zhi-Yuan Fan (0000-0003-3794-071X); Tian-Yi Zhao (0009-0003-6919-0746); Xiao Ma (0009-0001-5470-5095).
Co-first authors: Zhi-Yuan Fan and Xian-Hui Ran.
Co-corresponding authors: Lei Yang and Xiao Ma.
Author contributions: Fan ZY and Ran XH contributed equally to this study, including study design, data analysis, and manuscript preparation, wrote the original manuscript and revised the paper as co-first authors; Fan ZY, Ran XH, Wang N, Zhao TY, Li H, Liu X, Wu J, Yang Z, Chen G, and Yang L were responsible for data curation, methodology, and participated in formal analysis; Yang L and Ma X were the guarantor of the study and were responsible for conceptualization, project administration, supervision, methodology, writing review and editing as co-corresponding authors; all authors have read and approved the final manuscript.
AI contribution statement: Portions of this manuscript were edited using AI tools solely for language refinement. The authors carefully reviewed and verified all AI-assisted outputs and take full responsibility for the scientific content of the manuscript.
Supported by National High Level Hospital Clinical Research Funding, Elite Medical Professionals Initiative of China-Japan Friendship Hospital, No. ZRJY2025-QMPY41; Non-profit Central Research Institute Fund of Chinese Academy of Medical Sciences, No. 2022-ZHCH330-01; National High Level Hospital Clinical Research Funding, No. 2023-NHLHCRF-YXHZ-ZRMS-06; and National High Level Hospital Clinical Research Funding, No. 2025-NHLHCRF-PY-13.
Institutional review board statement: The study was approved by the Ethics Committee of China-Japan Friendship Hospital (No. 2025-KY-353).
Informed consent statement: The ethics committee agrees to waive informed consent.
Conflict-of-interest statement: All authors declare no conflict of interest in publishing the manuscript.
STROBE statement: The authors have read the STROBE Statement – checklist of items, and the manuscript was prepared and revised according to the STROBE Statement – checklist of items.
Data sharing statement: De-identified data in this study may be provided to qualified researchers upon reasonable request.
Corresponding author: Xiao Ma, MD, PhD, Dean, Health Checkup Center, China-Japan Friendship Hospital, No. 2 Yinghuayuan East Street, Chaoyang District, Beijing 100029, China. maxiaocjfh@163.com
Received: April 23, 2026
Revised: June 11, 2026
Accepted: July 15, 2026
Published online: September 15, 2026
Processing time: 135 Days and 16.6 Hours

Abstract
BACKGROUND

The rising prevalence of prediabetes is a major public health concern because of the substantial risk of progression to type 2 diabetes mellitus (T2DM). However, identifying individuals at high risk remains challenging due to the lack of reliable risk-stratification tools.

AIM

To develop a machine learning (ML)-based model to predict T2DM risk among individuals with prediabetes using routine health checkup data and to construct an interactive platform to support risk stratification.

METHODS

In this retrospective cohort study, individuals with prediabetes at baseline from two medical centers in China were included, with incident T2DM during follow-up defined as the study outcome. The development dataset comprised a training cohort (2012-2018) and a temporal testing cohort (2019-2021) from Beijing. An external validation cohort (2019-2021) from Henan Province was used to evaluate model generalizability. Four ML models were constructed to predict T2DM risk. Model performance was evaluated using the concordance index. Kaplan-Meier analysis was used to assess the risk-stratification ability of the best-performing model. Model interpretability was examined using Shapley Additive exPlanations.

RESULTS

A total of 27609 individuals with prediabetes were included. The gradient boosting survival analysis model showed the best predictive performance, with a concordance index of 0.813 (95%CI: 0.793-0.832) in the testing cohort and 0.759 (95%CI: 0.731-0.787) in the external validation cohort. The model effectively discriminated between high-risk and low-risk groups, which showed significantly different cumulative incidences of T2DM (P < 0.001). Shapley Additive exPlanations analysis identified fasting blood glucose, age, body mass index, monocyte count, and high-density lipoprotein cholesterol as the leading predictors.

CONCLUSION

An interpretable ML-based model using routine health checkup data demonstrated favorable performance for estimating T2DM risk among individuals with prediabetes and may support risk stratification in preventive care settings. Further prospective validation is warranted before broader implementation.

Key Words: Prediabetes; Type 2 diabetes mellitus; Health checkup; Machine learning; Risk prediction

Core Tip: We developed and externally validated an interpretable machine learning model to predict progression from prediabetes to type 2 diabetes mellitus using routine health checkup data. In a multicenter Chinese cohort, gradient boosting survival analysis demonstrated favorable discrimination (concordance index: 0.813 in temporal validation and 0.759 in external validation) and effectively stratified high-risk individuals. Shapley Additive exPlanations improved model transparency by identifying key predictors, including fasting blood glucose, age and body mass index. An online risk calculator was developed to support individualized risk assessment in preventive care settings.



INTRODUCTION

The development of type 2 diabetes mellitus (T2DM) is often preceded by prediabetes, an intermediate stage characterized by impaired glucose regulation[1-3]. The global prevalence of prediabetes is increasing, with projections suggesting that more than 840 million people will be affected by 2050[4,5]. Approximately 5%-10% of individuals with prediabetes progress to T2DM each year, and the lifetime risk of progression approaches 70%[6]. This growing burden represents a major public health concern because prediabetes is associated with an increased risk of both T2DM and adverse cardiovascular events[7,8].

The disease burden is particularly pronounced in China, where an estimated 35.7% of adults have prediabetes[9]. Routine health checkups, with 549 million conducted nationally in 2021, are widely used for population screening and therefore offer an important opportunity to detect this high-risk condition[10]. The clinical trajectory of prediabetes is heterogeneous, with estimated 10-year probabilities of 36.1% for reversion to normoglycemia and 12.5% for progression to T2DM[11]. Importantly, identifying individuals with prediabetes and implementing lifestyle or pharmacological interventions can delay or prevent progression to T2DM[12,13]. Accurate identification of high-risk individuals is therefore essential to enable targeted interventions and reduce the incidence and healthcare burden of T2DM. Several prediction models have been developed to estimate the risk of progression from prediabetes to T2DM. In a Danish population-based cohort, a 5-year risk model for individuals with prediabetes demonstrated the potential of long-term prognostic assessment using routinely available clinical and registry data[14]. In China, models based on health examination data have further highlighted the predictive relevance of metabolic, anthropometric, and liver-related indicators in populations with prediabetes[15,16]. Nevertheless, the broader applicability of these models remains uncertain, as most have been evaluated through internal validation, with limited evidence from external validation across geographically and socioeconomically diverse populations.

In this context, this study aimed to develop and validate a machine learning (ML)-based model to predict incident T2DM among individuals with prediabetes using longitudinal health checkup data from economically and geographically distinct regions in China. To support personalized prevention, we deployed the model as an online risk calculator that generates individualized risk estimates.

MATERIALS AND METHODS
Study design and participants

We conducted a multicenter retrospective cohort study using health checkup data from China-Japan Friendship Hospital (Beijing, China) and Henan Honliv Hospital (Henan Province, China). Eligible participants were adults (≥ 18 years) with prediabetes at baseline, defined according to the American Diabetes Association criteria as fasting blood glucose (FBG) levels of 5.6-6.9 mmol/L or hemoglobin A1c (HbA1c) levels of 5.7%-6.4%[17]. Individuals with a known history of T2DM or unavailable follow-up glycemic data were excluded.

We used temporally independent validation to evaluate model generalizability. The dataset comprised a primary analysis cohort and an external validation cohort. The primary cohort was derived from China-Japan Friendship Hospital (2012-2021) and was split into training (2012-2018) and testing (2019-2021) cohorts in temporal order. This design simulates a real-world scenario of model development followed by validation using future data. External validation was conducted in an independent, geographically distinct cohort from Henan Honliv Hospital (2019-2021).

This study was conducted in accordance with the Declaration of Helsinki and approved by the Ethics Committee of China-Japan Friendship Hospital (No. 2025-KY-353). Written informed consent was waived because of the retrospective observational design and the use of fully de-identified data.

Study outcome

Routine health checkups are offered as comprehensive health packages that enable screening for multiple diseases during a single visit. Data on demographic characteristics, medical history, and medication use were collected. The study outcome was incident T2DM during follow-up, identified according to American Diabetes Association criteria as FBG ≥ 7.0 mmol/L, or HbA1c ≥ 6.5%, a self-reported diagnosis, or use of insulin or oral hypoglycemic medication[17]. Participants were followed from baseline until the earliest occurrence of incident T2DM, loss to follow-up, or December 31, 2024.

Model variables

Candidate predictors were selected based on literature review and clinical relevance. Variables with more than 40% missing data were excluded before model development. Missing values in the retained variables were handled using multiple imputation by chained equations, generating 20 imputed datasets with 10 iterations. The proportions of missing data for the 34 retained predictors before imputation are summarized in Supplementary Table 1. The retained candidate predictors included sex, age, systolic blood pressure, diastolic blood pressure, body mass index (BMI), white blood cell count (WBC), neutrophil count, lymphocyte count, monocyte count (MONO), eosinophil count, basophil count (BASO), hemoglobin (HGB), hematocrit, mean corpuscular volume, mean corpuscular HGB, mean corpuscular HGB concentration, platelets count, plateletcrit, mean platelet volume (MPV), platelet distribution width (PDW), red blood cell distribution width-coefficient of variation, alanine aminotransferase (ALT), aspartate aminotransferase, total bilirubin, direct bilirubin, creatinine, urea, uric acid, FBG, total cholesterol, triglycerides (TG), high-density lipoprotein cholesterol (HDL-C), low-density lipoprotein cholesterol, and family history (FH) of T2DM.

Trained healthcare staff performed comprehensive assessments according to standardized protocols. Height and weight were measured using calibrated equipment, with participants wearing light clothing and no shoes. BMI was calculated as weight (kg)/height (m2). Blood pressure was measured using a certified automated electronic sphygmomanometer after participants had rested in a seated position for 5-10 minutes. Laboratory measurements were performed after a minimum fasting period of 8 hours using standardized procedures and equipment.

Feature selection for ML analysis

Feature selection was performed in the training cohort. Spearman correlation analysis was used to assess correlations among variables. When substantial collinearity was identified between two variables (Spearman’s |ρ| > 0.6), the variable deemed more clinically relevant was retained for subsequent feature selection. Univariate Cox regression analysis and least absolute shrinkage and selection operator (LASSO)-regularized Cox regression were then used to identify the most predictive variables for ML analysis while avoiding overfitting. In the univariate Cox regression analysis, variables with a P value < 0.05 were considered statistically significant. For LASSO Cox regression, 10-fold cross-validation was performed to determine the optimal regularization parameter (lambda.1se). Standardized coefficient path and cross-validation error plots were generated. Variables retained at the chosen lambda, together with their shrunken coefficients and derived hazard ratios, were summarized. Variables with nonzero coefficients and a selection frequency > 70% across the 20 imputed datasets were considered predictive features. The final features for ML-based modeling were selected by integrating results from the univariate Cox and LASSO Cox analyses, while also considering their clinical importance.

Development and evaluation of ML models

Based on the selected features, four candidate survival prediction models were developed: (1) Cox proportional hazards (CoxPH); (2) Random survival forest (RSF); (3) Extreme gradient boosting survival model (XGB); and (4) Gradient boosting survival analysis (GBSA). CoxPH was included as a conventional semi-parametric model to provide an interpretable reference for time-to-event prediction, whereas RSF was used as a non-parametric survival ensemble capable of modeling nonlinear effects and interactions. XGB and GBSA were included as boosting-based approaches to evaluate whether sequential tree-based learning could improve risk discrimination in structured health checkup data. These models were selected because they represent complementary survival modeling strategies with different assumptions and levels of flexibility, and they can be applied to censored outcomes with variable follow-up times. The models were trained independently on the 20 imputed training datasets, and performance estimates were pooled using Rubin’s rules. For RSF, GBSA, and XGB, hyperparameters were optimized using Bayesian optimization, with 5-fold cross-validation to maximize the concordance index (C-index).

The predictive performance of the ML models was assessed using the C-index and its 95%CI in the training, testing, and external validation cohorts. The model with the highest C-index in the testing cohort was selected as the final model, and its calibration in the external validation cohort was evaluated using the Brier score and calibration plot.

Incremental predictive performance was evaluated by comparing the final model with an FBG-only model and a parsimonious clinical model comprising age, BMI, WBC, TG, HDL-C, and FBG. The variables in the parsimonious clinical model were selected based on routine availability and consistency with an established T2DM risk prediction model in Chinese populations[18]. The FBG-only and parsimonious clinical models were fitted as CoxPH regression models in each of the 20 imputed training datasets, and subsequently evaluated in the testing and external validation cohorts. Incremental discrimination was quantified using the change in C-index (ΔC-index). Reclassification performance was evaluated for the 3-year predicted risk using continuous net reclassification improvement (NRI) and integrated discrimination improvement (IDI). Decision curve analysis (DCA) was performed for 3-year predicted risk to compare clinical net benefit across a range of threshold probabilities.

Model interpretation

Model interpretation was performed using Shapley Additive exPlanations (SHAP). This game-theoretic method quantifies feature importance by assigning a SHAP value to each predictor[19]. The SHAP value of a given feature represents its average contribution across all possible feature combinations. A positive SHAP value indicates that the feature was associated with a higher risk of incident T2DM, whereas a negative value suggests a lower risk. The magnitude of the SHAP value reflected the feature’s contribution to the model prediction.

Development of the web-based prediction platform

To facilitate individualized risk estimation, the final model was implemented as a web-based prediction platform using R Shiny. The deployment file incorporated the fitted final model from the imputed training datasets, the final set of predictors, categorical variable specifications, prediction horizons, pooled baseline cumulative hazards at 1 year, 3 years, and 5 years, and the empirical distribution of model-derived risk scores in the training cohort. For each user-provided profile, the platform performs basic type harmonization and aligns input variables with the model structure used during training. The fitted model generates an individual model-derived risk score, which is mapped to its percentile in the training cohort distribution. This percentile is transformed into a relative risk multiplier using a predefined percentile-based scaling function. The predicted 1-year, 3-year, and 5-year risks of incident T2DM are estimated by combining this multiplier with the pooled baseline cumulative hazards. The platform also reports the individual’s relative risk percentile and assigns a corresponding risk category. The application was deployed online using the rsconnect package.

Statistical analysis

Continuous data were reported as mean ± SD or median (interquartile range), as appropriate. For comparisons between two groups, Student’s t-test or the Wilcoxon rank-sum test was used according to variable distribution. Categorical variables were presented as n (%) and were compared using the χ2 test or Fisher’s exact test, as appropriate. Risk scores were generated by applying the best-performing model to each imputed dataset, and individual-level predictions were pooled across imputed datasets using Rubin’s rules. Individuals were stratified into high-risk and low-risk groups according to the median risk score in the training cohort. Risk stratification performance was assessed using cumulative incidence curves estimated by the Kaplan-Meier method, and between-group differences were compared using the log-rank test. To evaluate the robustness of risk stratification across populations, subgroup analyses stratified by sex, age, BMI, FH of T2DM, and hypertension status were conducted. Time-dependent receiver operating characteristic curve analysis was performed to evaluate predictive performance at 1 year, 3 years, and 5 years. The area under the curve (AUC) was calculated to assess discrimination. All statistical analyses were performed using R software (version 4.5.1). A two-tailed P value < 0.05 was considered statistically significant.

RESULTS
Baseline characteristics of study participants

A total of 27609 participants with baseline prediabetes were included in this study, comprising a training cohort (n = 15228) and a testing cohort (n = 9700) from China-Japan Friendship Hospital, and an external validation cohort (n = 2681) from Henan Honliv Hospital (Figure 1). Baseline characteristics of the three cohorts are shown in Table 1. Supplementary Table 2 summarizes the numbers and proportions of participants who met each diagnostic criterion for baseline prediabetes and incident T2DM. Baseline prediabetes was predominantly identified using FBG-based criteria, whereas HbA1c-based criteria provided additional ascertainment among participants with available data. During follow-up, incident T2DM was mainly identified based on FBG measurements, with HbA1c, hypoglycemic medication use, and self-reported clinical diagnosis providing supplementary confirmation. The availability of HbA1c measurements varied across cohorts, indicating heterogeneity in glycemic assessment within routine health checkup settings. Progression from prediabetes to T2DM occurred in 15.85% of participants in the training cohort, 9.42% in the testing cohort, and 16.93% in the external validation cohort, with corresponding median follow-up durations of 8.51 years, 3.72 years, and 4.00 years, respectively. A comparison of baseline characteristics between individuals with prediabetes who progressed to T2DM and those who did not is presented in Supplementary Table 3.

Figure 1
Figure 1  Study flowchart.
Table 1 Baseline characteristics of study population, mean ± SD/n (%)/median (interquartile range).
Variable
Training cohort (n = 15228)
Testing cohort (n = 9700)
Validation cohort (n = 2681)
Age (years)48.75 ± 13.8146.44 ± 12.4352.53 ± 12.63
Sex
Male8935 (58.67)6142 (63.32)1945 (72.55)
Female6293 (41.33)3558 (36.68)736 (27.45)
BMI (kg/m2)25.02 ± 3.4925.20 ± 3.5726.52 ± 3.33
SBP (mmHg)127.80 ± 17.69127.32 ± 16.74142.26 ± 19.46
DBP (mmHg)77.36 ± 11.6176.55 ± 11.5386.27 ± 11.91
FBG (mmol/L)5.85 (5.68, 6.11)5.66 (5.30, 5.87)5.88 (5.71, 6.17)
WBC (109/L)5.71 (4.60, 6.88)6.06 (5.18, 7.07)6.12 (5.28, 7.11)
Neutrophil count (109/L)3.48 (2.83, 4.23)3.40 (2.76, 4.12)3.50 (2.89, 4.30)
Lymphocyte count (109/L)1.99 (1.65, 2.39)2.00 (1.68, 2.40)1.90 (1.60, 2.30)
MONO (109/L)0.34 (0.28, 0.43)0.40 (0.33, 0.50)0.40 (0.40, 0.50)
EOS (109/L)0.14 ± 0.120.15 ± 0.130.14 ± 0.14
BASO (109/L)0.03 ± 0.020.03 ± 0.030.01 ± 0.03
PLT (109/L)238.68 ± 56.19243.77 ± 55.88228.17 ± 57.74
MPV (fL)8.60 (8.00, 9.50)9.70 (8.80, 10.50)8.40 (7.76, 9.15)
PDW (fL)15.31 ± 1.8713.48 ± 2.7216.86 ± 1.12
HGB (g/L)146.96 ± 14.82145.28 ± 15.16145.29 ± 14.44
TC (mmol/L)5.20 ± 0.935.61 ± 0.474.89 ± 0.94
TG (mmol/L)1.34 (0.93, 1.95)1.34 (0.93, 1.95)1.60 (1.16, 2.24)
HDL-C (mmol/L)1.28 (1.09, 1.53)1.18 (1.01, 1.41)1.20 (1.01, 1.43)
LDL-C (mmol/L)2.96 ± 0.773.09 ± 0.802.63 ± 0.67
ALT (U/L)21.00 (15.00, 30.00)23.00 (16.00, 34.00)20.00 (15.00, 29.00)
AST (U/L)20.00 (17.00, 24.00)21.00 (18.00, 25.00)20.00 (17.00, 24.00)
TBil (μmol/L)13.25 (10.28, 17.03)11.73 (9.19, 15.14)15.20 (12.00, 19.20)
FH of T2DM
Yes1579 (10.37)1136 (11.71)90 (3.36)
No11891 (78.09)6535 (67.37)2213 (82.54)
Missing1758 (11.54)2029 (20.92)378 (14.10)
Progression to T2DM
Yes2413 (15.85)914 (9.42)454 (16.93)
No12815 (84.15)8786 (90.58)2227 (83.07)
Feature selection for ML models

Variables with relatively high levels of missing data were excluded, and collinearity among the remaining variables was assessed. Absolute Spearman correlation coefficients greater than 0.6 were considered indicative of high collinearity (Supplementary Figure 1A). After seven collinear variables were excluded, 27 variables were retained for feature selection. Two complementary approaches were then applied to screen features. Univariate Cox regression identified 22 variables significantly associated with incident T2DM (Supplementary Table 4). LASSO Cox regression selected 14 candidate predictors (Supplementary Figure 1B-D), and the corresponding penalized coefficients and derived hazard ratios at the chosen lambda were summarized in Supplementary Table 5. Based on the overlap between the two methods and clinical relevance, 14 features (age, BMI, WBC, MONO, BASO, MPV, PDW, ALT, total bilirubin, FBG, total cholesterol, TG, HDL-C, and FH of T2DM) were ultimately selected for ML analysis.

Model development and evaluation

The selected features were used to train four ML models, and the optimal hyperparameters were determined. Model discrimination, as measured by the C-index, is summarized in Table 2. In the testing cohort, the GBSA model achieved the highest C-index of 0.813 (95%CI: 0.793-0.832) and was therefore selected as the final model for subsequent analyses. The four models showed comparable discrimination in the external validation cohort, with C-index values ranging from 0.756 to 0.762. The final GBSA model also showed favorable calibration, with a Brier score of 0.205 (Supplementary Figure 2).

Table 2 Predictive performance of four developed models.
C-index (95%CI)
CoxPH
XGB
RSF
GBSA
Training cohort0.789 (0.778-0.801)0.803 (0.787-0.819)0.880 (0.868-0.891)0.807 (0.795-0.819)
Testing cohort0.802 (0.783-0.821)0.806 (0.786-0.826)0.717 (0.697-0.736)0.813 (0.793-0.832)
Validation cohort0.758 (0.730-0.786)0.756 (0.727-0.785)0.762 (0.734-0.790)0.759 (0.731-0.787)

Compared with the FBG-only and parsimonious clinical models, the final GBSA model yielded higher discrimination in the testing cohort, with ΔC-index values of 0.046 and 0.013, respectively. In the external validation cohort, the corresponding ΔC-index values were 0.002 and 0.003, indicating marginal improvement in discrimination. For the 3-year predicted risk, the GBSA model showed positive NRI and IDI in both the testing and external validation cohorts, suggesting improved risk reclassification compared with the reference models. DCA indicated that the GBSA model provided higher net benefit in the testing cohort and comparable net benefit in the external validation cohort across selected threshold ranges (Supplementary Figure 3). Detailed results are summarized in Supplementary Table 6.

Risk stratification and subgroup analysis

Using the cutoff derived from the training cohort predictions, participants in the external validation cohort were stratified into high-risk and low-risk groups. The baseline characteristics of the two groups are presented in Supplementary Table 7. Kaplan-Meier analysis showed a significantly higher cumulative incidence of T2DM in the high-risk group than in the low-risk group (P < 0.001; Figure 2A and Supplementary Figure 4). Risk stratification was generally consistent across the investigated subgroups (Supplementary Figure 5). Furthermore, the time-dependent AUCs for predicting progression to T2DM at 1 year, 3 years, and 5 years were 0.875, 0.790, and 0.740, respectively, indicating favorable discriminative ability in the external validation cohort (Figure 2B). The numbers and proportions of participants at risk, cumulative outcome events, and censored observations at the 1-year, 3-year, and 5-year prediction horizons in the external validation cohort are summarized in Supplementary Table 8.

Figure 2
Figure 2 Performance of the gradient boosting survival analysis model in the external validation cohort. A: Kaplan-Meier curves comparing the cumulative incidence of type 2 diabetes mellitus in the high-risk and low-risk groups; B: Time-dependent receiver operating characteristic curves for predicting 1-year, 3-year, and 5-year progression to type 2 diabetes mellitus. AUC: Area under the curve.
Model interpretability

To interpret feature contributions to T2DM risk prediction in the GBSA model, we calculated SHAP values for each feature in the external validation cohort and ranked features according to their SHAP-based importance (Figure 3). SHAP analysis indicated that baseline FBG contributed most to the prediction of T2DM, followed by age, BMI, MONO, and HDL-C. Higher HDL-C was associated with a lower predicted risk of progression to T2DM, whereas FBG, age, BMI, and MONO were positively associated with the predicted risk.

Figure 3
Figure 3 Shapley Additive exPlanations analysis of feature importance for the gradient boosting survival analysis model in the external validation cohort. A: Beeswarm plot showing the distribution of Shapley Additive exPlanations values for each feature across individuals. Yellow indicates a positive contribution to the predicted risk, while purple indicates a negative contribution; B: Mean absolute Shapley Additive exPlanations values for each feature across 20 imputed datasets, with 95%CI. SHAP: Shapley Additive exPlanations; FBG: Fasting blood glucose; BMI: Body mass index; MONO: Monocyte count; HDL-C: High-density lipoprotein cholesterol; BASO: Basophil count; WBC: White blood cell count; TG: Triglycerides; ALT: Alanine aminotransferase; TC: Total cholesterol; TBil: Total bilirubin; FH: Family history; PDW: Platelet distribution width; MPV: Mean platelet volume.
Web-based prediction platform

To support individualized risk estimation, we developed a publicly accessible online calculator for T2DM risk assessment. As illustrated in Supplementary Figure 6, for a 51-year-old individual with prediabetes and the displayed characteristics, the tool estimated a 5-year predicted progression probability of 0.57. This estimate corresponded to the 97.4% of predicted risk, indicating a high risk of progression.

DISCUSSION

In this study, we developed an interpretable ML-based model to predict progression from prediabetes to T2DM and identify high-risk individuals who may benefit from targeted prevention. The model showed favorable performance in longitudinal cohorts from geographically and economically distinct regions in China. To our knowledge, this is among the largest multicenter studies to comprehensively apply and compare multiple advanced ML algorithms for T2DM risk prediction among individuals with prediabetes. We further deployed the model as a publicly accessible online risk calculator, facilitating its potential use in real-world preventive care.

Longitudinal evidence indicates substantial heterogeneity in prediabetes trajectories, with a higher likelihood of regression to normoglycemia than progression to T2DM[11]. This finding highlights the need for effective risk stratification to support targeted prevention in high-risk subsets. Using routinely collected health checkup variables and time-to-event modeling, our GBSA model achieved strong discrimination in the temporal testing cohort from Beijing (C-index = 0.813) and favorable performance in an external validation cohort from a geographically distinct county-level setting in Henan Province (C-index = 0.759). Previous studies predicting progression from prediabetes to T2DM have commonly relied on random data splits or geographically proximate cohorts[20,21]. In contrast, our validation strategy more rigorously evaluates transportability across distinct socioeconomic and healthcare delivery settings, thereby better approximating real-world conditions for large-scale deployment within routine checkup programs. Consistent performance across independent cohorts suggests that the model captures clinically meaningful risk signals rather than overfitting to a single institutional dataset. This is particularly relevant for risk prediction in prediabetes, for which transportability may be limited by between-setting differences in case mix, measurement procedures, and outcome ascertainment[22].

In incremental comparisons, the GBSA model showed improved discrimination over the FBG-only and parsimonious clinical models in the testing cohort, whereas the improvement in C-index was modest in the external validation cohort. This finding may be partly explained by the strong prognostic value of baseline FBG in individuals with prediabetes and its close relationship with subsequent T2DM development in routine health checkup settings. The positive 3-year NRI and IDI, together with DCA findings, further indicated that the GBSA model may provide additional value for risk reclassification, although these gains were more pronounced in the testing cohort than in the external validation cohort.

The primary goal of T2DM prevention is to tailor intervention intensity to individual risk of progression, given the resource demands of structured lifestyle programs, glycemic monitoring, and pharmacologic prevention[23,24]. This study provides a scalable approach for risk stratification in prediabetes. In the external validation cohort, the model showed clear separation in cumulative incidence between risk groups, supporting its potential utility for informing targeted preventive care pathways.

The observed decline in time-dependent AUCs over longer prediction horizons is consistent with the time-varying nature of prognostic accuracy, as baseline measurements may become progressively less informative when individuals’ exposures, behaviors, and physiology change during follow-up[25,26]. To facilitate interpretation and practical use, a publicly accessible web calculator was developed to translate model outputs into individualized risk estimates based on routine health checkup data, which may help support risk communication in preventive care settings. However, its clinical impact and implementation value require further evaluation in prospective studies.

The performance of ML algorithms depends substantially on the selection of stable and informative features[27]. In this study, we used a sequential and complementary feature selection strategy. Spearman correlation analysis was applied to reduce multicollinearity among candidate variables. Univariate Cox regression and LASSO Cox regression were then used to identify relevant features from different statistical perspectives. The overlap in features selected by both methods may enhance the stability and clinical relevance of the final feature set and reduce the risk of overfitting.

ML algorithms are often considered “black boxes” because the processes underlying their decisions are difficult to interpret[28]. SHAP, a method derived from cooperative game theory, uses Shapley values to quantify the contribution of each feature to individual predictions, providing a detailed breakdown of predicted outcomes. By combining the average prediction with the SHAP values, this approach supports both local and global model interpretability, enhances transparency, and makes complex models more understandable[29,30]. Notably, feature importance rankings derived from LASSO Cox regression and SHAP analysis showed some discrepancies. These differences may be attributed to two key factors. First, the analyses were based on different cohorts, as SHAP was applied to a geographically independent external validation cohort that may have differed from the development cohort used for LASSO-based feature selection. Second, the two methods serve different methodological purposes: LASSO Cox applies an L1 regularization penalty for feature selection during model development, whereas SHAP provides a game-theoretic explanation of predictions generated by the final GBSA model, reflecting feature contributions to model predictions in new data. Overall, the feature importance rankings generated by both methods were broadly comparable, indicating generally consistent feature contributions across analytical approaches.

The predictors identified in our study are broadly aligned with those reported in previous models for T2DM risk prediction among individuals with prediabetes. A 5-year nomogram developed in Chinese adults with prediabetes included age, TG, FBG, BMI, ALT, HDL-C, and FH of T2DM, whereas another model emphasized the combined predictive value of FBG, 2-hour postprandial plasma glucose, and HbA1c[31,32]. ML-based studies have similarly highlighted baseline glycemic measures, HDL-C, and age as relevant predictors[20,33]. Our final model retained FBG, age, BMI, TG, HDL-C, ALT, and FH of T2DM, with SHAP analysis further identifying FBG, age, BMI, and HDL-C among the leading contributors. This consistency with prior evidence supports the interpretability of the model and suggests that routinely available glycemic, anthropometric, lipidrelated, and hepatic indicators collectively capture key dimensions of progression risk in individuals with prediabetes.

However, given the retrospective cohort design, the included predictors should be interpreted primarily as predictive factors rather than evidence of causal or pathogenic determinants. The prominent contribution of FBG is consistent with existing biological evidence, as elevated FBG may reflect early glucose dysregulation, impaired insulin sensitivity, and β-cell dysfunction, which are closely related to subsequent T2DM development[34-37]. Age and BMI are well-established predictors of T2DM and may capture accumulated metabolic vulnerability and adiposity-related risk[38]. Recent evidence further suggests that body composition may provide metabolic risk information beyond BMI, as higher fat-to-muscle ratio and lower BMI-adjusted appendicular skeletal muscle mass have been associated with increased diabetes prevalence in euthyroid adults. Peripheral thyroid hormone sensitivity was also found to partially mediate the association between lower BMI-adjusted appendicular skeletal muscle mass and diabetes[39]. In our model, HDL-C contributed negatively to predicted risk, consistent with evidence linking higher HDL-C levels to more favorable glucose metabolism and lower risk of T2DM[40,41].

In addition to conventional metabolic predictors, the final model incorporated several hematologic indicators, including WBC, MONO, BASO, MPV, and PDW. Notably, MONO ranked among the top five contributors in SHAP analysis, consistent with a recent ML-based model for predicting T2DM[20]. Although routine hematologic parameters have been less commonly included in T2DM prediction models, they may provide complementary prognostic information related to systemic inflammation and hematologic status. Prior evidence has linked monocyte levels and chronic low-grade inflammation to insulin resistance and the pathophysiology of T2DM[42]. Furthermore, integrative proteomic evidence has identified subtype-specific protein associations and the enrichment of inflammatory and immune-related pathways across etiological subtypes of diabetes, supporting the relevance of inflammatory and immune processes to disease heterogeneity[43]. Collectively, these findings suggest that routine hematologic parameters may capture additional predictive information beyond conventional metabolic and clinical indicators.

Several limitations should be acknowledged. First, the retrospective observational design may have introduced selection bias, and the lack of lifestyle-related variables may lead to residual confounding. In addition, excluding individuals without follow-up glycemic measurements may have reduced the representativeness of the study population relative to the general health checkup population, thereby limiting the model’s applicability to broader and more heterogeneous populations. Second, the limited and uneven availability of HbA1c measurements across cohorts hindered its reliable inclusion as a predictor. Additionally, the absence of oral glucose tolerance testing and incomplete HbA1c data meant that glycemic assessment based on routine health checkup records may not have fully captured participants’ true glycemic status, potentially leading to misclassification of baseline prediabetes or incident T2DM. Third, although time-dependent predictive performance was evaluated at 1 year, 3 years, and 5 years, the median follow-up durations of the testing and external validation cohort were approximately 4 years. Consequently, the 5-year predictive performance was estimated with fewer participants remaining at risk and a relatively high censoring rate, and should therefore be interpreted with caution. Fourth, although a geographically distinct external validation cohort was used, both participating centers were hospital-based, which may have introduced bias toward health-conscious individuals and limit generalizability to community-based populations or other ethnic groups.

CONCLUSION

We developed and externally validated an interpretable ML-based model to estimate the risk of progression from prediabetes to T2DM using routinely available health checkup data. The model showed favorable performance in both temporal testing and external validation cohorts, supporting its potential value for risk stratification in preventive care settings. Further validation in prospective, community-based cohorts and more diverse ethnic populations is warranted to confirm the robustness and transportability of the model.

References
1.  Bailey CJ. Prediabetes: never too early to intervene. Lancet Diabetes Endocrinol. 2023;11:529-530.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 9]  [Reference Citation Analysis (0)]
2.  Tabák AG, Herder C, Rathmann W, Brunner EJ, Kivimäki M. Prediabetes: a high-risk state for diabetes development. Lancet. 2012;379:2279-2290.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 2403]  [Cited by in RCA: 2138]  [Article Influence: 152.7]  [Reference Citation Analysis (3)]
3.  Echouffo-Tcheugui JB, Selvin E. Prediabetes and What It Means: The Epidemiological Evidence. Annu Rev Public Health. 2021;42:59-77.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 321]  [Cited by in RCA: 286]  [Article Influence: 57.2]  [Reference Citation Analysis (1)]
4.  Rooney MR, Fang M, Ogurtsova K, Ozkan B, Echouffo-Tcheugui JB, Boyko EJ, Magliano DJ, Selvin E. Global Prevalence of Prediabetes. Diabetes Care. 2023;46:1388-1394.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 434]  [Cited by in RCA: 372]  [Article Influence: 124.0]  [Reference Citation Analysis (0)]
5.  Rooney MR, He JH, Salpea P, Genitsaridi I, Magliano DJ, Boyko EJ, Wallace AS, Fang M, Selvin E. Global and Regional Prediabetes Prevalence: Updates for 2024 and Projections for 2050. Diabetes Care. 2025;48:e142-e144.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 30]  [Cited by in RCA: 24]  [Article Influence: 24.0]  [Reference Citation Analysis (0)]
6.  Zhou W, Sailani MR, Contrepois K, Zhou Y, Ahadi S, Leopold SR, Zhang MJ, Rao V, Avina M, Mishra T, Johnson J, Lee-McMullen B, Chen S, Metwally AA, Tran TDB, Nguyen H, Zhou X, Albright B, Hong BY, Petersen L, Bautista E, Hanson B, Chen L, Spakowicz D, Bahmani A, Salins D, Leopold B, Ashland M, Dagan-Rosenfeld O, Rego S, Limcaoco P, Colbert E, Allister C, Perelman D, Craig C, Wei E, Chaib H, Hornburg D, Dunn J, Liang L, Rose SMS, Kukurba K, Piening B, Rost H, Tse D, McLaughlin T, Sodergren E, Weinstock GM, Snyder M. Longitudinal multi-omics of host-microbe dynamics in prediabetes. Nature. 2019;569:663-671.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 536]  [Cited by in RCA: 423]  [Article Influence: 60.4]  [Reference Citation Analysis (0)]
7.  Wallace AS, Rooney MR, Fang M, Echouffo-Tcheugui JB, Grams M, Selvin E. Natural History of Prediabetes and Long-term Risk of Clinical Outcomes in Middle-aged Adults: The Atherosclerosis Risk in Communities (ARIC) Study. Diabetes Care. 2023;46:e67-e68.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 1]  [Cited by in RCA: 9]  [Article Influence: 3.0]  [Reference Citation Analysis (0)]
8.  Schlesinger S, Neuenschwander M, Barbaresko J, Lang A, Maalmi H, Rathmann W, Roden M, Herder C. Prediabetes and risk of mortality, diabetes-related complications and comorbidities: umbrella review of meta-analyses of prospective studies. Diabetologia. 2022;65:275-285.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 263]  [Cited by in RCA: 235]  [Article Influence: 58.8]  [Reference Citation Analysis (0)]
9.  Wang L, Gao P, Zhang M, Huang Z, Zhang D, Deng Q, Li Y, Zhao Z, Qin X, Jin D, Zhou M, Tang X, Hu Y, Wang L. Prevalence and Ethnic Pattern of Diabetes and Prediabetes in China in 2013. JAMA. 2017;317:2515-2523.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 1518]  [Cited by in RCA: 1436]  [Article Influence: 159.6]  [Reference Citation Analysis (7)]
10.  Gong W, Cheng KK. Challenges in screening and general health checks in China. Lancet Public Health. 2022;7:e989-e990.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 12]  [Reference Citation Analysis (0)]
11.  Davoodian N, Lotfaliany M, Huxley RR, Lee CMY, Pasco JA, Adams RJ, Azizi F, Bertoni AG, Björkelund C, Colagiuri S, Fahimfar N, Gabriel R, Giedraitis V, Gill TK, González C, Gregg EW, Hadaegh F, Hange D, Harland JW, Hodge AM, Holloway-Kew KL, Hosseini SR, Jacobs DR Jr, Khalili D, Magliano DJ, Mongraw-Chaffin M, Mishra GD, Lissner L, Mehlig K, Najafipour H, Nieto-Martinez R, Ostovar A, Shadkam M, Shaw JE, Sundh V, Schreiner PJ, Sakurai M, Wittert GA, Yatsuya H, Zethelius B, Mohebbi M. Prediabetes transitions to normoglycaemia or type 2 diabetes and associated risk factors in the Obesity, Diabetes and Cardiovascular Disease Collaboration: an individual-level pooled analysis of 19 prospective cohort studies. Lancet Glob Health. 2025;13:e1533-e1542.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 13]  [Cited by in RCA: 15]  [Article Influence: 15.0]  [Reference Citation Analysis (0)]
12.  Gong Q, Zhang P, Wang J, Ma J, An Y, Chen Y, Zhang B, Feng X, Li H, Chen X, Cheng YJ, Gregg EW, Hu Y, Bennett PH, Li G; Da Qing Diabetes Prevention Study Group. Morbidity and mortality after lifestyle intervention for people with impaired glucose tolerance: 30-year results of the Da Qing Diabetes Prevention Outcome Study. Lancet Diabetes Endocrinol. 2019;7:452-461.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 501]  [Cited by in RCA: 454]  [Article Influence: 64.9]  [Reference Citation Analysis (4)]
13.  Zhang L, Zhang Y, Shen S, Wang X, Dong L, Li Q, Ren W, Li Y, Bai J, Gong Q, Kuang H, Qi L, Lu Q, Cheng W, Liu Y, Yan S, Wu D, Fang H, Hou F, Wang Y, Yang Z, Lian X, Du J, Sun N, Ji L, Li G; China Diabetes Prevention Program Study Group. Safety and effectiveness of metformin plus lifestyle intervention compared with lifestyle intervention alone in preventing progression to diabetes in a Chinese population with impaired glucose regulation: a multicentre, open-label, randomised controlled trial. Lancet Diabetes Endocrinol. 2023;11:567-577.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 35]  [Cited by in RCA: 56]  [Article Influence: 18.7]  [Reference Citation Analysis (0)]
14.  Nicolaisen SK, Thomsen RW, Lau CJ, Sørensen HT, Pedersen L. Development of a 5-year risk prediction model for type 2 diabetes in individuals with incident HbA1c-defined pre-diabetes in Denmark. BMJ Open Diabetes Res Care. 2022;10:e002946.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 1]  [Cited by in RCA: 16]  [Article Influence: 4.0]  [Reference Citation Analysis (0)]
15.  Yang J, Liu D, Du Q, Zhu J, Lu L, Wu Z, Zhang D, Ji X, Zheng X. Construction of a 3-year risk prediction model for developing diabetes in patients with pre-diabetes. Front Endocrinol (Lausanne). 2024;15:1410502.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 4]  [Reference Citation Analysis (0)]
16.  Liu Q, Zhou Q, He Y, Zou J, Guo Y, Yan Y. Predicting the 2-Year Risk of Progression from Prediabetes to Diabetes Using Machine Learning among Chinese Elderly Adults. J Pers Med. 2022;12:1055.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 15]  [Reference Citation Analysis (0)]
17.  American Diabetes Association Professional Practice Committee for Diabetes*. 2. Diagnosis and Classification of Diabetes: Standards of Care in Diabetes-2026. Diabetes Care. 2026;49:S27-S49.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 162]  [Reference Citation Analysis (5)]
18.  Chien K, Cai T, Hsu H, Su T, Chang W, Chen M, Lee Y, Hu FB. A prediction model for type 2 diabetes risk among Chinese people. Diabetologia. 2009;52:443-450.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 112]  [Cited by in RCA: 118]  [Article Influence: 6.9]  [Reference Citation Analysis (0)]
19.  Ponce-Bobadilla AV, Schmitt V, Maier CS, Mensing S, Stodtmann S. Practical guide to SHAP analysis: Explaining supervised machine learning model predictions in drug development. Clin Transl Sci. 2024;17:e70056.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 675]  [Cited by in RCA: 383]  [Article Influence: 191.5]  [Reference Citation Analysis (1)]
20.  Zhang Y, Zhang H, Wang D, Li N, Lv H, Zhang G. Development of a 5-Year Risk Prediction Model for Transition From Prediabetes to Diabetes Using Machine Learning: Retrospective Cohort Study. J Med Internet Res. 2025;27:e73190.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 6]  [Reference Citation Analysis (0)]
21.  Liu Y, Feng W, Lou J, Qiu W, Shen J, Zhu Z, Hua Y, Zhang M, Billong LF. Performance of a prediabetes risk prediction model: A systematic review. Heliyon. 2023;9:e15529.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 6]  [Reference Citation Analysis (0)]
22.  Xu S, Coleman RL, Wan Q, Gu Y, Meng G, Song K, Shi Z, Xie Q, Tuomilehto J, Holman RR, Niu K, Tong N. Risk prediction models for incident type 2 diabetes in Chinese people with intermediate hyperglycemia: a systematic literature review and external validation study. Cardiovasc Diabetol. 2022;21:182.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 2]  [Cited by in RCA: 6]  [Article Influence: 1.5]  [Reference Citation Analysis (0)]
23.  ElSayed NA, Aleppo G, Aroda VR, Bannuru RR, Brown FM, Bruemmer D, Collins BS, Hilliard ME, Isaacs D, Johnson EL, Kahan S, Khunti K, Leon J, Lyons SK, Perry ML, Prahalad P, Pratley RE, Seley JJ, Stanton RC, Gabbay RA;  on behalf of the American Diabetes Association. 3. Prevention or Delay of Type 2 Diabetes and Associated Comorbidities: Standards of Care in Diabetes-2023. Diabetes Care. 2023;46:S41-S48.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 80]  [Cited by in RCA: 134]  [Article Influence: 44.7]  [Reference Citation Analysis (0)]
24.  Sussman JB, Kent DM, Nelson JP, Hayward RA. Improving diabetes prevention with benefit based tailored treatment: risk based reanalysis of Diabetes Prevention Program. BMJ. 2015;350:h454.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 77]  [Cited by in RCA: 96]  [Article Influence: 8.7]  [Reference Citation Analysis (0)]
25.  Bansal A, Heagerty PJ. A Tutorial on Evaluating the Time-Varying Discrimination Accuracy of Survival Models Used in Dynamic Decision Making. Med Decis Making. 2018;38:904-916.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 28]  [Cited by in RCA: 55]  [Article Influence: 6.9]  [Reference Citation Analysis (0)]
26.  Kamarudin AN, Cox T, Kolamunnage-Dona R. Time-dependent ROC curve analysis in medical research: current methods and applications. BMC Med Res Methodol. 2017;17:53.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 674]  [Cited by in RCA: 615]  [Article Influence: 68.3]  [Reference Citation Analysis (0)]
27.  Pudjihartono N, Fadason T, Kempa-Liehr AW, O'Sullivan JM. A Review of Feature Selection Methods for Machine Learning-Based Disease Risk Prediction. Front Bioinform. 2022;2:927312.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 732]  [Cited by in RCA: 328]  [Article Influence: 82.0]  [Reference Citation Analysis (0)]
28.  Ali S, Akhlaq F, Imran AS, Kastrati Z, Daudpota SM, Moosa M. The enlightening role of explainable artificial intelligence in medical & healthcare domains: A systematic literature review. Comput Biol Med. 2023;166:107555.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 272]  [Cited by in RCA: 196]  [Article Influence: 65.3]  [Reference Citation Analysis (4)]
29.  Rodríguez-Pérez R, Bajorath J. Interpretation of Compound Activity Predictions from Complex Machine Learning Models Using Local Approximations and Shapley Values. J Med Chem. 2020;63:8761-8777.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 72]  [Cited by in RCA: 293]  [Article Influence: 41.9]  [Reference Citation Analysis (0)]
30.  Rodríguez-Pérez R, Bajorath J. Interpretation of machine learning models using shapley values: application to compound potency and multi-target activity predictions. J Comput Aided Mol Des. 2020;34:1013-1026.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 92]  [Cited by in RCA: 288]  [Article Influence: 48.0]  [Reference Citation Analysis (0)]
31.  Han Y, Hu H, Liu Y, Wang Z, Liu D. Nomogram model and risk score to predict 5-year risk of progression from prediabetes to diabetes in Chinese adults: Development and validation of a novel model. Diabetes Obes Metab. 2023;25:675-687.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 1]  [Cited by in RCA: 17]  [Article Influence: 5.7]  [Reference Citation Analysis (0)]
32.  Liang K, Guo X, Wang C, Yan F, Wang L, Liu J, Hou X, Li W, Chen L. Nomogram Predicting the Risk of Progression from Prediabetes to Diabetes After a 3-Year Follow-Up in Chinese Adults. Diabetes Metab Syndr Obes. 2021;14:2641-2649.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 4]  [Cited by in RCA: 9]  [Article Influence: 1.8]  [Reference Citation Analysis (0)]
33.  Aoki J, Khalid O, Kaya C, Nagymanyoki Z, Hussong J, Salama ME. Progression from Prediabetes to Diabetes in a Diverse U.S. Population: A Machine Learning Model. Diabetes Technol Ther. 2024;26:748-753.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 7]  [Cited by in RCA: 8]  [Article Influence: 4.0]  [Reference Citation Analysis (0)]
34.  Sitasuwan T, Lertwattanarak R. Prediction of type 2 diabetes mellitus using fasting plasma glucose and HbA1c levels among individuals with impaired fasting plasma glucose: a cross-sectional study in Thailand. BMJ Open. 2020;10:e041269.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 1]  [Cited by in RCA: 6]  [Article Influence: 1.0]  [Reference Citation Analysis (0)]
35.  Faerch K, Borch-Johnsen K, Holst JJ, Vaag A. Pathophysiology and aetiology of impaired fasting glycaemia and impaired glucose tolerance: does it matter for prevention and treatment of type 2 diabetes? Diabetologia. 2009;52:1714-1723.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 145]  [Cited by in RCA: 174]  [Article Influence: 10.2]  [Reference Citation Analysis (0)]
36.  Abdul-Ghani MA, Tripathy D, DeFronzo RA. Contributions of beta-cell dysfunction and insulin resistance to the pathogenesis of impaired glucose tolerance and impaired fasting glucose. Diabetes Care. 2006;29:1130-1139.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 17]  [Cited by in RCA: 269]  [Article Influence: 13.5]  [Reference Citation Analysis (0)]
37.  Li Y, Feng D, Esangbedo IC, Zhao Y, Han L, Zhu Y, Fu J, Li G, Wang D, Wang Y, Li M, Gao S, Willi SM. Insulin resistance, beta-cell function, adipokine profiles and cardiometabolic risk factors among Chinese youth with isolated impaired fasting glucose versus impaired glucose tolerance: the BCAMS study. BMJ Open Diabetes Res Care. 2020;8:e000724.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 6]  [Cited by in RCA: 15]  [Article Influence: 2.5]  [Reference Citation Analysis (1)]
38.  Hu S, Ji W, Zhang Y, Zhu W, Sun H, Sun Y. Risk factors for progression to type 2 diabetes in prediabetes: a systematic review and meta-analysis. BMC Public Health. 2025;25:1220.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 31]  [Reference Citation Analysis (0)]
39.  Yao N, Liu L, Jiang Y, Wang D, Lu Q, Zeng Y, Yu G, Ma Q, Li P, Liu L, Shen J, Wan H. Peripheral Thyroid Hormone Sensitivity Mediates the Association Between Body Composition and Diabetes in Euthyroid Adults. J Cachexia Sarcopenia Muscle. 2025;16:e70061.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 3]  [Cited by in RCA: 5]  [Article Influence: 5.0]  [Reference Citation Analysis (1)]
40.  Bodaghi AB, Ebadi E, Gholami MJ, Azizi R, Shariati A. A decreased level of high-density lipoprotein is a possible risk factor for type 2 diabetes mellitus: A review. Health Sci Rep. 2023;6:e1779.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 1]  [Cited by in RCA: 11]  [Article Influence: 3.7]  [Reference Citation Analysis (5)]
41.  Yan Z, Xu Y, Li K, Liu L. Association between high-density lipoprotein cholesterol and type 2 diabetes mellitus: dual evidence from NHANES database and Mendelian randomization analysis. Front Endocrinol (Lausanne). 2024;15:1272314.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 8]  [Cited by in RCA: 15]  [Article Influence: 7.5]  [Reference Citation Analysis (0)]
42.  Mokgalaboni K, Dludla PV, Nyambuya TM, Yakobi SH, Mxinwa V, Nkambule BB. Monocyte-mediated inflammation and cardiovascular risk factors in type 2 diabetes mellitus: A systematic review and meta-analysis of pre-clinical and clinical studies. JRSM Cardiovasc Dis. 2020;9:2048004019900748.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 11]  [Cited by in RCA: 22]  [Article Influence: 3.7]  [Reference Citation Analysis (6)]
43.  Wei J, Wu H, Wang N, Zhu J, Anjana RM, Sokolov AV, Schiöth HB, Mohan V, Tan X. Integrative proteomic analysis provides novel therapeutic insights for etiological subtypes of diabetes. Diabetes Obes Metab. 2025;27:6894-6904.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 4]  [Cited by in RCA: 2]  [Article Influence: 2.0]  [Reference Citation Analysis (0)]
Footnotes

Peer review: Externally peer reviewed.

Peer-review model: Single blind

Specialty type: Endocrinology and metabolism

Country of origin: China

Peer-review report’s classification

Scientific quality: Grade A, Grade B, Grade B, Grade B, Grade C

Novelty: Grade B, Grade B, Grade B, Grade B

Creativity or innovation: Grade B, Grade B, Grade B, Grade C

Scientific significance: Grade B, Grade B, Grade B, Grade B

P-Reviewer: Huo WQ, Associate Professor, PhD, China; Wan H, Associate Professor, China S-Editor: Luo ML L-Editor: A P-Editor: Wang WB

Write to the Help Desk