Copyright: ©Author(s) 2026.
World J Cardiol. Jun 26, 2026; 18(6): 120747
Published online Jun 26, 2026. doi: 10.4330/wjc.120747
Published online Jun 26, 2026. doi: 10.4330/wjc.120747
Table 2 Characteristics of the model used in the studies
| Ref. | Models evaluated | Internal validation strategy | External validation | Training dataset | Test/validation dataset | Performance metrics |
| Kayvanpour et al[8], 2021 | LR, kNN, LDA, NB, RF, CT, SVM, XGB, and ANN | The subjects were divided into training and test sets in the ratio of 9:1, respectively. This was repeated 10 times to enable ten-fold cross-validation | None | 90% of subjects; 121 samples per split | 10% of subjects; 13 samples per split | Accuracy, sensitivity, specificity, and ROC-AUC |
| Ren et al[13], 2024 | Regularized LR using either SCAD or LASSO | Leave-one-out cross-validation | None for the ML model; selected miRNAs were biologically evaluated in matched clinical samples using an ion-exchange membrane sensor platform | 24 subjects; 800-miRNA screening library (100%) | None; leave-one-out cross-validation was used because of small sample size | ROC curves and AUC (used to evaluate the selected miRNA combinations) |
| Samadishadlou et al[14], 2023 | A 2-layer architecture utilizing SVM (with linear, polynomial, and RBF kernels), LR, RF, kNN, GB, XGB, and DT models (layer 1 isolated healthy vs not-healthy; layer 2 separated MI vs CAD) | The data was split in a 7:3 ratio into the training and test sets, respectively. A ten-fold cross-validation followed this | None | 70% of all the samples | 30% of all the samples | AUC-ROC, accuracy, sensitivity, specificity, and confusion matrix |
| Samadishadlou et al[15], 2024 | SVM, GB, XGB and hard voting ensemble model | Done in 2 phases. In miRNA selection: The LASSO method was cross-validated using the dataset 10-fold to select the best miRNA to be used in model development. In model selection, the training dataset was split in a 7:3 ratio into training and validation datasets. The models were then cross-validated 5-fold on the datasets. The best-performing models were then tested on the independent dataset | Performed using an independent dataset (GSE29532) | GSE61741 (62 MI samples and 94 healthy samples) | GSE29532 (8 MI samples and 6 healthy samples) | Accuracy, AUC-ROC, sensitivity, and specificity |
| Reel et al[16], 2025 | J48, NB, IBk, RF, LB, LMT, SL, and SMO | The data was randomly split into training and testing sets in an 8:2 ratio for model development and validation | None | 80% of all the samples | 20% of all the samples | Balanced accuracy is the primary metric. Other metrics include sensitivity, specificity, AUC-ROC, F1 score and Kappa score |
| Sajid et al[17], 2024 | LR, SVM, nonlinear kNN, tree-based (DT, RF), GB, XGBM, CBoost, ABoost, and ensemble voting | The data subset was first split in an 8:2 ratio into a CV subset and a hold-out subset for final evaluation. The CV subset was then divided into 10 folds. Nine folds were used for training and one-fold for testing, and the process was repeated 10 times. The best models were then tested on the hold-out dataset | None | 80% of the 113 subjects (cohort: 58 CAD cases, 55 healthy controls) | 20% of the 113 subjects (hold-out subset) | Accuracy, sensitivity, specificity, AUC-ROC, performance evaluation measure, F-statistic, and P values |
| Yerukala Sathipati et al[18], 2025 | kNN, XGB, SVM, and RF | The data was split in a 8:2 ratio into training and validation datasets, respectively | Performed using an independent GEO dataset (GSE222739) | 80% of the dataset (n = 12) | 20% of the dataset (n = 3) | AUC-ROC, accuracy, specificity, and sensitivity |
| Jusic et al[19], 2023 | RF, SVM, MLP, XGB, kNN, Logit | Hyperparameter tuning used two repeated 10-fold CV. The final model was also evaluated using leave-one-out cross-validation | None | 147 subjects (89 from the validation cohort + 58 from the sequenced discovery cohort) | 23 subjects (randomly extracted as 20% of the 112-subject validation cohort) | AUC-ROC, balanced accuracy, F1 score, precision, sensitivity, and specificity |
| Errington et al[20], 2021 | RF, Rpart, LASSO, XGB, and Ensemble | The data was split into training and validation data sets. The models were then CV 10-fold in the training dataset | The models were externally validated using publicly available datasets | Two-thirds of the samples | One-third of the samples (validation set) | Sensitivity, specificity, AUC, correct classification rate (accuracy), positive predictive value, and negative predictive value |
- Citation: Popat A, Sathipati S, Sharma P. Machine learning integration in microRNA-based markers for cardiovascular diseases: A systematic review. World J Cardiol 2026; 18(6): 120747
- URL: https://www.wjgnet.com/1949-8462/full/v18/i6/120747.htm
- DOI: https://dx.doi.org/10.4330/wjc.120747