BPG is committed to discovery and dissemination of knowledge
Review
Copyright: ©Author(s) 2026.
World J Stem Cells. Aug 26, 2026; 18(8): 121077
Published online Aug 26, 2026. doi: 10.4252/wjsc.121077
Table 5 Validation checklist for leukemic stem cell/hematopoietic stem cell artificial intelligence models, adapted from TRIPOD + AI and PROBAST + AI guidelines with hematopoietic stem cell/Leukemic stem cell-specific requirements
Validation domain
Key requirement
LSC/HSC-specific considerations
Cohort representativenessThe training cohort must represent the target clinical population with respect to age, disease stage, and treatment eraLSC/HSC models should include balanced representation of ELN risk categories, stem-cell compartment measurements (CD34+CD38- frequencies), and both newly diagnosed and relapsed/refractory patients[52,106-108]
Event countsAn adequate number of outcome events (relapse, death, MRD positivity) to prevent overfittingMinimum 10-20 events per predictor variable; for LSC-specific endpoints (e.g., LSC+ vs LSC-), ensure sufficient LSC+ cases across validation sets[52,55]
Internal validationModel performance assessed on held-out data from the same source (cross-validation or hold-out split)Report performance metrics (AUROC, calibration) separately for LSC-enriched vs LSC-depleted subgroups if the model claims to encode stemness biology[52,109,110]
External validationIndependent cohort from a different institution, time period, or geographyEssential for LSC models given center-to-center variability in LSC phenotyping protocols and MRD detection thresholds[52,106,110]
Prospective evaluationForward-looking validation on newly enrolled patients before clinical deploymentRequired for LSC-targeted therapy selection models to confirm that AI predictions align with clinical outcomes under prospective conditions[55,109,111]
Dataset shift detectionAssess whether model performance degrades when applied to data with distributional differences (batch effects, assay drift)Critical for flow cytometry-based LSC models: Validate across different antibody panels, fluorophores, and cytometers; report performance stratified by batch[55,112]
Calibration assessmentPredicted probabilities should match observed event frequenciesFor LSC burden models, calibration plots should show agreement between predicted LSC frequency (or surrogate score) and directly measured LSC% by flow cytometry in the calibration subset[52,55,106]
Decision-curve analysisNet benefit of model-guided decisions compared to treat-all or treat-none strategiesFor LSC-directed therapies (venetoclax, Menin inhibitors), decision curves should quantify clinical utility across risk thresholds relevant to treatment intensification decisions[55,111]
Explainability and feature attributionUse of XAI methods (SHAP, attention weights) to identify which features drive predictionsEssential for validating LSC-AI hypothesis: Determine whether high-risk predictions are driven by known stemness genes (17-gene LSC score, HOXMEIS1 programs) or alternative pathways[52,113]
Bias and fairness evaluationAssess performance stratified by demographic subgroups and underrepresented populationsEvaluate whether LSC models perform equivalently across age groups (pediatric vs adult vs elderly AML), ancestry, and sex; report subgroup-specific metrics[107]
Missing data handlingTransparent reporting of missingness patterns and imputation strategiesLSC models often integrate multi-omics data with heterogeneous completeness (e.g., scRNA-seq available for a subset); clearly document handling of missing modalities and proteins in CITE-seq[111]
Comparator benchmarkingPerformance compared to established clinical risk systemsFor AML LSC models, benchmark against ELN 2022 risk classification, 17-gene LSC score, and LSC frequency by flow cytometry; report incremental predictive value[114]


Write to the Help Desk