BPG is committed to discovery and dissemination of knowledge
Review
Copyright: ©Author(s) 2026.
World J Stem Cells. Aug 26, 2026; 18(8): 121077
Published online Aug 26, 2026. doi: 10.4252/wjsc.121077
Table 3 Summary of key machine learning algorithms, applications, and trade-offs in hematopoietic stem cell research
Machine learning algorithm
Primary HSC/hematology applications
Key advantages (Pros)
Key limitations (Cons)
Ref.
CNNMorphological classification of HSCs vs MPPs; 3D chromatin age prediction (ChromAgeNet); automated quality control imaging in biomanufacturingUnparalleled performance on spatial and image data; extracts features autonomously without requiring manual gating or human-defined parametersHighly opaque “black box” nature requiring XAI for interpretability; demands massive, accurately annotated image datasets to train[6,11,34]
Random forest/ensemble treesFlow cytometric LSC phenotyping; early relapse detection; predicting cord blood CD34+ cell yieldHighly robust to overfitting; handles tabular clinical and multi-omics data effectively; naturally provides feature importance rankingsLess effective than deep learning for highly unstructured data (like raw images or free text); can struggle with extrapolating data outside the training range[22,35]
Gradient Boosting (e.g., XGBoost)Automated MRD detection in flow cytometry (MAGIC-DR); predicting synergistic drug combinations; venetoclax response predictionExceptional predictive accuracy on structured clinical and omics data; handles missing data well; highly scalable for large patient cohortsProne to overfitting on very small sample sizes; hyperparameter tuning is complex and computationally expensive[23,56]
Deep learning/ANNMulti-omics integration (e.g., totalVI); modeling age-dependent HSC self-renewal; multi-center AML survival predictionCan capture extremely complex, non-linear biological relationships across massive, high-dimensional datasets (e.g., integrating RNA and protein expression)Computationally intensive; high risk of learning artifactual batch effects rather than true biology if data is not strictly harmonized[10,15,17]
NLPExtracting HLA genotypes from unstructured electronic health records; predicting and phenotyping acute and chronic GVHD from clinical notesUnlocks vast amounts of unstructured, historical clinical data that is otherwise inaccessible to standard statistical modelsPerformance is heavily dependent on the quality, consistency, and language of physician documentation; it struggles with implicit clinical context[40-42]
RLDynamic optimization of ex vivo HSC expansion; adaptive control of bioreactor parameters (cytokines, perfusion, metabolic flux)Enables continuous, autonomous, real-time process optimization without requiring a pre-defined static protocolRequires highly accurate “digital twins” or simulation environments to train the agent safely; initial validation in GMP environments is regulatory complex[37,77,78]


Write to the Help Desk