Abd El Ghaffar HA, Arafat AMA, Khattab EHA, Khattab MA, Khallaf AM, Mahgoub SMA. Artificial intelligence in hematopoietic stem cell research and associated malignancies: From disease modeling to cell manufacturing. World J Stem Cells 2026; 18(8): 121077 [DOI: 10.4252/wjsc.121077]
Corresponding Author of This Article
Aya Mohamed Adel Arafat, MD, Lecturer, Department of Clinical and Chemical Pathology, Faculty of Medicine, Cairo University, Al-Saray Street, Cairo 11956, Egypt. aya.arafat@kasralainy.edu.eg
Research Domain of This Article
Hematology
Article-Type of This Article
review-article
Open-Access Policy of This Article
This article is an open-access article which was selected by an in-house editor and fully peer-reviewed by external reviewers. It is distributed in accordance with the Creative Commons Attribution Non Commercial (CC BY-NC 4.0) license, which permits others to distribute, remix, adapt, build upon this work non-commercially, and license their derivative works on different terms, provided the original work is properly cited and the use is non-commercial. See: http://creativecommons.org/licenses/by-nc/4.0/
Baishideng Publishing Group Inc, 7041 Koll Center Parkway, Suite 160, Pleasanton, CA 94566, USA
Share the Article
Abd El Ghaffar HA, Arafat AMA, Khattab EHA, Khattab MA, Khallaf AM, Mahgoub SMA. Artificial intelligence in hematopoietic stem cell research and associated malignancies: From disease modeling to cell manufacturing. World J Stem Cells 2026; 18(8): 121077 [DOI: 10.4252/wjsc.121077]
Heba Adel Abd El Ghaffar, Aya Mohamed Adel Arafat, Shirihan Mahmoud Anwar Mahgoub, Department of Clinical and Chemical Pathology, Faculty of Medicine, Cairo University, Cairo 11956, Egypt
Eman Hazem AbdelTawab Khattab, Department of Clinical Pharmacy, Faculty of Pharmacy, Sinai University, Ismailia 41522, Egypt
Mustafa Ashraf Khattab, Department of Computer Science, Faculty of Media Engineering and Technology, German University in Cairo, Cairo 11835, Egypt
Ahmed M Khallaf, Department of Internal Medicine and Clinical Hematology, Faculty of Medicine, Beni-Suef University, Beni-Suef 62511, Egypt
Author contributions: Abd El Ghaffar HA, Arafat AMA, and Mahgoub SMA contributed equally to this work; Abd El Ghaffar HA, Khattab EHA, and Mahgoub SMA designed the research study; Khattab MA contributed to analytic tools; Arafat AMA, Khattab MA, and Khallaf AM wrote the manuscript; and all authors have read and approved the final manuscript.
AI contribution statement: Grammarly (Grammarly Inc.) was used solely for linguistic refinement and language polishing of the manuscript. No AI tool was involved in the generation of research content, data interpretation, or formulation of conclusions. All AI-assisted outputs were critically reviewed and revised by the authors. The authors take full responsibility for the accuracy, originality, and integrity of the manuscript.
Conflict-of-interest statement: All the authors report no relevant conflicts of interest for this article.
Corresponding author: Aya Mohamed Adel Arafat, MD, Lecturer, Department of Clinical and Chemical Pathology, Faculty of Medicine, Cairo University, Al-Saray Street, Cairo 11956, Egypt. aya.arafat@kasralainy.edu.eg
Received: March 16, 2026 Revised: May 12, 2026 Accepted: June 17, 2026 Published online: August 26, 2026 Processing time: 158 Days and 16.3 Hours
Abstract
Hematopoietic stem cells (HSCs) occupy the apex of the blood cell hierarchy, and artificial intelligence (AI) is fundamentally reshaping how their biology is decoded from normal self-renewal to malignant transformation and clinical transplantation. Trajectory inference algorithms applied to single-cell multi-omics have resolved continuous HSC differentiation with lineage priming detectable at the single-cell level, while convolutional neural networks trained on chromatin imaging predict HSC biological age and detect epigenetic rejuvenation signatures. In leukemic stem cell (LSC) biology, multi-omics deep learning models map treatment-resistant quiescent LSC subclones, detect minimal residual disease with an area under the curve of 0.97 and predict venetoclax sensitivity in LSC-enriched niches. Deep learning also maps myeloma stem cell niche interactions and spatial heterogeneity in bone marrow biopsies. Virtual screening powered by AI speeds up LSC-targeted drug discovery, and reinforcement learning and digital twin models improve ex vivo HSC manufacturing and industrial-scale production of chimeric antigen receptor-T cells. In transplantation, natural language processing extraction and hybrid models classify risk for graft-vs-host disease into clinically relevant subgroups. Overall, top-performing AI models serve as computational surrogates for stemness biology but challenges in dataset diversity, interpretability and regulatory compliance need to be addressed before being used in a clinical setting.
Core Tip: Artificial intelligence is redefining the study of hematopoietic stem cells (HSCs) by serving as a computational guide to stemness biology. Machine learning and deep learning delineate HSC fate trajectories, unravel age-dependent HSC self-renewal, and identify enriched leukemic stem cell sub-compartments that sustain treatment-resistant leukemia and relapse in acute myeloid leukemia and multiple myeloma. The same approaches expedite leukemic stem cell-targeted drug discovery, model HSC biomanufacturing for digital twins and predict transplant risk using natural language processing. High-performing artificial intelligence models are therefore quantitative surrogates for stem cell state - and span single-cell biology and decision-making across the HSC continuum.
Citation: Abd El Ghaffar HA, Arafat AMA, Khattab EHA, Khattab MA, Khallaf AM, Mahgoub SMA. Artificial intelligence in hematopoietic stem cell research and associated malignancies: From disease modeling to cell manufacturing. World J Stem Cells 2026; 18(8): 121077
Hematopoietic stem cells (HSCs) are a central element of regenerative medicine and hematological therapies, with the unique ability to self-renew and differentiate into all blood cell lineages. Conventional HSC studies have largely relied on experimental biology, histopathology, and low-throughput molecular profiling[1]. But the rapid advancement of high-throughput technologies, such as next-generation sequencing (NGS), mass spectrometry proteomics, and single-cell multi-omics, has produced a deluge of multidimensional biological information that is beyond human processing capacity[2]. Artificial intelligence (AI), including machine learning (ML), deep learning (DL) and natural language processing (NLP), has emerged as a disruptive paradigm in biomedical research, allowing the extraction of valuable insights from complex data. AI has found numerous applications in hematology across the diagnostic, prognostic and therapeutic continuum, reshaping the future of precision medicine. Although there has been an explosion of high-dimensional high-throughput data, a significant gap has emerged in the ability to translate these data into clinically relevant information for hematopoietic cancers. Current risk stratification and experimental approaches often do not capture and quantify the resistant sub-populations of the leukemic stem cell (LSC) compartment that drive disease relapse. In turn, there is a pressing need for innovative computational approaches to unravel this complex biology and move beyond using bulk blast evaluation to directly address stemness and enhance targeted therapies[3]. In-depth reviews have recently outlined the different uses of AI in hematology, including for diagnosis, decision making and management of HSC transplantation (HSCT)[2,4]. More importantly, the application of AI to normal HSC biology has provided fundamental insights required for its understanding. ML modeling of HSC differentiation, aging and epigenetic control establishes the normal baseline upon which disease-associated changes can be mapped. Inference of cell trajectories from single-cell data has shown the dynamic process of hematopoiesis to involve lineage priming[5] and analysis of chromatin structure by DL has allowed prediction of HSC age and state from morphological data[6]. The application of AI in HSC research is promoting a transformative shift from “hypothesis-driven” to “data-driven” research to discover new disease mechanisms, therapeutic targets, and predictive biomarkers. Specifically, the use of AI in hematological cancers, such as acute myeloid leukemia (AML) and multiple myeloma (MM), has shown tremendous promise in the prediction of risk and treatment response by measuring LSC/MM stem cells (MMSCs)-enriched sub-compartments responsible for relapse. To avoid confusion of the scope of the present review, the term “associated malignancies” refers specifically to myeloid and plasma cell lineages, mainly AML and MM, where the biological concepts of cancer stem cells (LSCs and MMSCs) are most firmly established and targetable. Although the application of AI in lymphoid and mature lymphomas is rapidly evolving, they have different architectural and microenvironmental drivers and therefore are not the primary focus of this discussion, which is tightly focused on direct malignant products of the hematopoietic stem and progenitor compartments. Moreover, AI-led strategies are transforming drug discovery processes, especially for LSC-selective targets and stem-cell-sparing combination therapies, via virtual screening, molecular repurposing, and target identification, substantially accelerating and de-risking the drug development process[7]. AI technologies are also revolutionizing the biomanufacturing of cell-based therapies, such as chimeric antigen receptor-T (CAR-T) cells and HSC products, beyond disease modeling and drug discovery[8,9]. The integration of computer vision, digital twin modeling and Industry 4.0 principles allows real-time quality monitoring, process optimization and large-scale production of these complex biologics[8,9]. This review provides a comprehensive overview of the application of AI to HSC research, with a specific focus on recent breakthroughs (2020-2026). The purpose of this review is to conceptualize the gap between computational modeling and stem cell biology - to show how high-performing AI models work not as statistical risk factors, but as measurable computational representations of therapy-resistant stem-cell populations. Additionally, we assess the challenges in translating these technologies to the clinic and describe the required validation strategies for incorporating these technologies in hematological practice.
SCOPE AND METHODS
Search strategy and selection criteria
This narrative review was invited and does not adhere to a systematic review protocol such as the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA); we undertook a structured, but non-systematic, search of the literature to identify recent advances. We focused on recent (2020-2026) studies that exemplify the role of computational biology in HSC research. We preferred studies that focused on stem or progenitor populations (HSCs, LSCs, MMSCs) or HSCT processes, rather than the bulk disease. We searched several databases (PubMed/MEDLINE, Google Scholar, ScienceDirect, EMBASE, IEEE Xplore) for articles published between January 2020 and January 2026. Search terms combined Medical Subject Headings and keywords such as: “artificial intelligence”, “machine learning”, “deep learning”, “natural language processing”, “hematopoietic stem cells”, “acute myeloid leukemia”, “multiple myeloma”, “drug discovery”, “biomanufacturing”, “CAR-T cells”, “digital twin”, and “Industry 4.0”.
Data extraction and quality assessment
We extracted information on study design, AI techniques used, clinical use cases, performance measures, and outcomes. We assessed the quality of the studies using criteria for computational biology studies, focusing on the rigor of the algorithms, validation approaches, and reproducibility. Table 1 illustrates the correspondence between each section of the review and stem-cell biological objects, data types used, AI tasks, computational results, clinical or biological outcomes and the current validation status in normal HSC biology, LSC modeling, measurable residual disease (MRD) detection and HSCT processes, with an emphasis on dataset size and performance measures[4-42].
Table 1 Stem-cell relevance map linking each application area to stem-cell biological object, data modality, AI methodology, output, clinical/biological endpoint, and validation status.
Application
Stem-cell object
Data modality
AI method/model
Dataset/scale
Key performance metrics
Clinical/biological endpoint
Validation status
Ref.
A: Normal HSC biology
HSC differentiation modeling
Normal HSCs, MPPs
scRNA-seq, scATAC-seq
VIA (Voyager) - lazy-teleporting MCMC trajectory inference
Human CD34+ hematopoiesis (multi-site)
F1 > 0.9 for rare lineage populations; robust across scRNA-seq and scATAC-seq
AI applications to normal HSC biology are a critical step in understanding the molecular processes underlying stem cell identity, differentiation, aging and epigenetic regulation. These applications offer a critical context for the interpretation of AI findings in hematological cancers, as disease-related changes need to be interpreted against the backdrop of normal HSC biology.
AI-driven HSC differentiation modeling and lineage trajectory inference
The integration of single-cell RNA sequencing (scRNA-seq) with ML has transformed our understanding of normal HSC differentiation by allowing topologies of continuous differentiation trajectories to be mapped at a single-cell resolution. Previous models of hematopoiesis portrayed binary, discrete differentiation steps through progenitor stages. Yet, ML-based trajectory inference algorithms have shown hematopoiesis to be a dynamic process where lineage priming occurs in phenotypically defined HSC populations.
The graph-based trajectory inference method, Voyager in scRNA-seq analysis (VIA), which uses lazy-teleporting random walks combined with Markov chain Monte Carlo sampling, has outperformed other methods in capturing complex differentiation topologies of HSCs, including multifurcating trees and cycles[5]. Applied to human CD34+ hematopoiesis, VIA was repeatedly able to identify hierarchical bifurcations leading to monocytic, lymphoid, erythroid, classical, and plasmacytoid dendritic cell lineages, and megakaryocytes, with sustained sensitivity to rare cell types (F1-scores > 0.9) across varying computational parameters. It is important to note that VIA has been shown to be robust across a variety of single-cell omics modalities, including scRNA-seq and single-cell assay of transposase-accessible chromatin sequencing datasets. ML analysis of paired daughter cell expression profiles of single HSC divisions has shown that age-dependent changes in self-renewal dynamics and niche sensitivity[10].
An artificial neural network trained on single-cell gene expression patterns from hematopoietic stem and progenitor cells achieved 89% overall accuracy in simultaneously predicting cell type identity and donor age (young, adult, or aged mice), with accuracy increasing to 96% for regenerative status prediction[10,43]. This study demonstrated that HSC self-renewal ability declines with age, with HSC divisions being deterministic and intrinsically regulated in young and old age but variable and niche-sensitive in mid-life[10]. The balance between intrinsic and extrinsic regulation of stem cell activity thus alters substantially with age[10,44].
DL-based morphological classification has made it possible to identify subpopulations of HSCs and multipotent progenitors (MPPs) based on morphology[11]. A convolutional neural network (CNN) trained on bright-field microscopy images accurately classified HSCs and MPPs under steady-state conditions, demonstrating that subtle morphological features invisible to human observers encode functional stem cell states. This offers a rapid, non-invasive HSC functional analysis that does not require molecular profiling[11].
Computational modeling of HSC aging
AI applications have illuminated the complex cellular and molecular changes underlying HSC aging, a process characterized by functional decline, myeloid-biased differentiation, and accumulation of damaged cells. ML integrated with high-resolution chromatin imaging has enabled the prediction of HSC biological age from nuclear architecture. ChromAgeNet, a CNN trained on 3D 4’,6-diamidino-2-phenylindole-stained confocal microscopy images of HSC nuclei, achieved an area under the receiver operating characteristic of 0.77 ± 0.03 in distinguishing young from aged murine HSCs, outperforming classical ML models trained on handcrafted chromatin features (area under the receiver operating characteristic: 0.73 ± 0.04)[6]. Explainable AI (XAI) techniques revealed that the model learned age-associated chromatin patterns, including: (1) Chromatin entropy (increased in aged HSCs, consistent with loss of epigenomic integrity); (2) Peripheral heterochromatin organization (continuous thin line along nuclear envelope in young HSCs vs thicker, irregular signal in aged HSCs); and (3) Nucleolar size and chromatin condensates[6,45]. When applied to drug-treated aged HSCs, ChromAgeNet detected rejuvenation signatures, with histone H3K9 methylation inhibitors (UNC0646, IOX1) restoring youthful chromatin architecture scores (0.56 ± 0.13 and 0.55 ± 0.16) comparable to untreated young HSCs (0.55 ± 0.19)[6], suggesting potential therapeutic applications.
Single-cell transcriptomic studies combined with ML have identified specific aged HSC subpopulations that expand with age and contribute to hematopoietic dysfunction. scRNA-seq analysis revealed: (1) An expansion of interferon-primed HSCs prepared to respond to inflammatory stimuli; (2) Accumulation of platelet/megakaryocyte-primed HSCs contributing to myeloid bias; and (3) Increased transforming growth factor-beta (TGF-β) signature HSCs potentially linked to differentiation blockage[46]. ML inference of gene regulatory networks from these data has enabled the construction of Boolean models predicting how aging-associated transcriptional changes drive myeloid-biased differentiation.
Epigenetic regulation and chromatin architecture modeling
Epigenetic alterations represent a hallmark of HSC aging, with AI approaches enabling quantitative modeling of chromatin organization changes over time. Integration of scRNA-seq with single-cell assay of transposase-accessible chromatin sequencing through ML has revealed coordinated transcriptional and epigenetic dynamics during HSC differentiation. Network-based ML approaches, particularly single-cell regulatory network inference and clustering, have reconstructed transcription factor (TF) regulatory networks governing HSC self-renewal and differentiation by identifying TF binding motifs enriched in accessible chromatin regions of co-expressed gene sets. In HSC differentiation, single-cell regulatory network inference and clustering identified lineage-specific TF activities such as: (1) The activation of GATA1 and the commitment of HSCs to the erythroid lineage; (2) The activation of CEBPD and the specification of HSCs to the myeloid lineage; and (3) The activation of IRF8 and the differentiation of HSCs to the dendritic cell lineage, all of which were validated by chromatin accessibility data. Combining chromatin conformation data (high-throughput chromosome conformation capture) with ML has started to uncover the role of 3D genome architecture in HSC fate determination. While in its infancy, these methods will yield insights into the roles of topologically associating domains and chromatin loops in regulating gene expression during HSC differentiation and aging.
HSC quiescence and activation dynamics
HSC quiescence and activation are key to maintaining stem cell numbers and providing a reserve for tissue repair. ML approaches to single-cell data have revealed signatures of quiescent vs activated HSCs and predicted the conditions that lead to HSC activation. Boolean modeling combined with scRNA-seq data has been particularly useful in studying HSC quiescence. A Boolean network model of HSC quiescence/activation that included niche signals [thrombopoietin (TPO), stem cell factor (SCF), angiopoietin-1], cell cycle regulators (p21, p53), and metabolic factors produced stable long-term HSC, short-term HSC, and proliferating HSC states in response to different combinations of niche signals. The model identified a new regulatory role of p53 in homeostasis through reactive oxygen species- and RAS-activated TF regulators, which were experimentally confirmed[12]. Inference of Boolean networks from scRNA-seq pseudotrajectories has allowed the development of predictive models of HSC dynamics. Treating pseudotrajectories as observations of system trajectories, constraint programming and model checking have been used to infer logical rules of HSC state transitions. This approach has successfully identified Boolean networks of embryonic hematopoiesis and HSC differentiation into MPP and megakaryocyte-erythroid progenitor states[46].
Multi-omic integration for systems-level HSC understanding
Integration of multiple single-cell omics (transcriptomics, epigenomics, proteomics) with AI-based techniques is facilitating systems-level analyses of HSC regulation. ML-based methods for multi-omic data integration, such as weighted nearest neighbor and multi-omic factor analysis, have uncovered coordinated molecular programs that regulate HSC identity and function that are not apparent from individual data types. Lineage tracing, scRNA-seq and ML have shown that HSCs with similar transcriptomes can have different fate biases, with epigenetic features being more predictive of cell fate than transcriptomes[47,48]. This highlights the need for multi-modal profiling and AI-driven integration for HSC fate prediction and to distinguish normal HSCs from pre-LSCs.
Implications for understanding disease
AI applications to normal HSC biology establish essential baselines for interpreting disease-associated alterations and for defining LSC-specific deviations from normal stemness programs. The morphological, transcriptional, and epigenetic signatures of normal HSCs defined by ML provide reference points for identifying LSC characteristics. DL has demonstrated that normal HSCs in leukemic bone marrow (BM) environments undergo AI-recognizable morphological changes (> 98% classification accuracy), distinct from both LSCs and normal HSCs in healthy BM[18]. This finding suggests that monitoring normal HSC populations in leukemia patients through AI-based image analysis could provide prognostic information and assess therapeutic responses. The trajectory inference and lineage commitment models developed for normal hematopoiesis provide frameworks for understanding aberrant differentiation in myeloid malignancies. For example, the early lineage priming observed in normal HSCs through AI analysis helps explain the myeloid bias and differentiation blockages characteristic of aged and leukemic HSCs.
AI applications in hematological disease modeling
AML is a heterogeneous hematologic malignancy that is increasingly understood as a stem/progenitor-driven disease organized as a cellular hierarchy. Seminal xenotransplantation work demonstrated that human AML contains a rare leukemia-initiating population with self-renewal capacity (often termed LSCs), supporting the concept that relapse and therapy resistance can be rooted in “stemness” biology rather than bulk blast counts alone[49]. In this framework, AI models are most informative when they resolve the HSC → pre-leukemic HSC → LSC trajectory, rather than treating blasts as a homogeneous endpoint (Figure 1)[50].
Consistent with a stem-cell origin model, pre-leukemic clones can reside in HSCs and persist through therapy, providing a biological bridge between molecular genetics and stemness-linked relapse risk[51]. Conventional clinical risk stratification systems [e.g., European Leukemia Net (ELN)] are based on cytogenetic and molecular characteristics to inform prognosis and intensity of treatment. However, these clinical-genetic categories do not directly quantify therapy-resistant “stem-like” disease compartments. Since AML originates from the malignant transformation of hematopoietic progenitors, AI-driven risk stratification effectively functions as a surrogate for identifying LSC burden. This interpretation is strengthened by the observation that stemness gene-expression programs (derived from functionally validated LSC fractions) strongly predict induction failure and survival, implying that many high-performing computational models are, at least in part, learning LSC-enriched transcriptional/epigenetic states rather than “clinical risk” alone[52]. In parallel, ELN risk stratification remains a cornerstone for AML management, with the ELN 2022 update reflecting major advances in AML genomics, response criteria, and treatment guidance. These advances also increase the dimensionality of inputs available to ML systems that integrate clinical, cytogenetic, and molecular data[3].
Recent AI-driven approaches have demonstrated improved prognostic modeling in AML beyond guideline-only approaches. A recent narrative review summarized that ML models integrating clinical, cytogenetic, and molecular variables can outperform conventional ELN-based approaches, and highlighted DL performance for morphology and genetic-variant prediction[53]. In a large, concrete example of AML ML risk modeling, supervised ML models were trained to predict complete remission and 2-year overall survival in a multicenter cohort of 1383 intensively treated AML patients using clinical/Laboratory/cytogenetic/molecular features, with external validation in an independent cohort. From an LSC perspective, many of the predictive variables in such models (genetic risk lesions and downstream transcriptomic consequences) plausibly track the presence of resistant stem-like subclones that drive relapse risk[17]. Single-cell sequencing combined with AI further connects “risk prediction” to stemness biology by resolving AML heterogeneity into therapy-resistant compartments. In particular, scRNA-seq has been highlighted as enabling identification of quiescent stem-like cells/Leukemia stem cells associated with resistance and relapse, exactly the biology that clinical risk scores approximate indirectly and that AI models may capture more sensitively when trained on high-dimensional data[54].
AI-enhanced MRD monitoring and LSC-informed residual disease assessment
MRD monitoring has become a cornerstone of treatment response evaluation and post-remission risk stratification in AML. Multiparameter flow cytometry (MFC) and molecular polymerase chain reaction/NGS techniques can detect residual disease far below morphological limits, with standard MFC having a sensitivity of about 0.1%. Nevertheless, a critical and clinically relevant limitation remains: A large percentage of patients who attain MRD negativity by conventional flow cytometry still relapse, because of the persistence of therapy-resistant LSCs that fall below current assay detection limits or express immunophenotypes not detected by standard MRD panels[55]. This biological gap, between the inability to detect any bulk leukemia and the persistence of the relapse-initiating LSC reservoir, has been the driving force behind the development of LSC-specific flow cytometry assays, as well as the application of AI to enhance the sensitivity, specificity, and clinical interpretability of residual disease monitoring.
The machine-learning guided approach to AML MRD detection in residual disease framework is an example of how AI is applied to flow cytometric MRD analysis. Machine-learning guided approach to AML MRD detection in residual disease demonstrated an area under the curve (AUC) of 97% in differentiating MRD-positive and MRD-negative samples using Extreme Gradient Boosting (XGBoost) classification with uniform manifold approximation and projection dimensionality reduction. Importantly, the model revealed immature monocytic subpopulations, which could be potential computational proxies of LSC-like residual disease, which were systematically ignored during traditional manual gating, and which AI-automated MRD analysis may be able to recover biologically meaningful residual disease signals that are not visible to standard operator-driven methods[56]. This finding directly connects AI-driven MRD assessment to the biology of LSC persistence: The same quiescent, niche-protected cell populations that drive relapse may be detectable by AI-powered pattern recognition before they manifest clinically.
In addition to specific MRD systems, a ML system combining heterogeneous 53-marker flow cytometry data had an accuracy of 94.92% (random forest; AUC: 94.83) in the differentiation of active AML, complete remission, and normal BM in 194 patients[22]. The feature importance analysis using SHapley Additive exPlanations (SHAP) revealed that CD117, CD34, and human leukocyte antigen-DR isotype (HLA-DR) are the most discriminative markers in disease states, as they have been shown to play a role in defining the LSC-enriched CD34+CD38-compartment. Interestingly, the model showed that AML-complete remission has a unique immunophenotypic signature compared to both active disease and normal marrow - a result that has immediate clinical implications in early relapse detection, with immunophenotypic recovery to the AML signature preceding overt morphological relapse potentially being detected by these AI-enhanced systems.
A combination of LSC-specific phenotyping and conventional MRD assessment adds prognostic value to either method. It has been shown in clinical data that patients who attain MRD negativity, but still harbor detectable LSC-enriched populations (LSC+), are at significantly increased risk of relapse compared to truly double-negative patients (MRD-/LSC-), with 3-year overall survival rates of about 80% vs 45% in double-positive (MRD+/LSC+) patients[50]. In line with this, the ELN MRD Working Party has approved LSC-directed flow cytometry as an adjunctive prognostic instrument, especially in instances where bulk MRD analysis is not available alone is vague or on the edges of perception[50]. More recent spectral flow cytometry systems, which now permit simultaneous MRD and LSC detection in a single high-event-count tube (targeting 1-4 million events to achieve high sensitivity LSC enumeration), are further lowering inter-laboratory variability and technical limits to routine LSC-integrated MRD monitoring[55].
In the future, combining AI-based MRD analysis with LSC phenotyping, genomic MRD (NGS-based), and clinical variables into comprehensive multimodal predictive models will form the next step in precision AML monitoring. These platforms may facilitate individualized post-remission decision-making, which may be determined by which patients should undergo treatment escalation, consolidative allogeneic HSCT or LSC-targeted maintenance founded upon an overall computational evaluation of residual disease burden in both bulk leukemic and stem-cell compartments.
The MMSC niche
Accumulating data suggest that the BM microenvironment is an extra and separate tier of biological complexity in MM, where angiogenic, stromal, and immune responses determine how the disease responds to therapy, resistance to treatment, and the clinical outcome. In MM, this “niche” should be conceptualized not only as a growth-permissive ecosystem for malignant plasma cells, but also as a spatially organized set of microanatomical habitats that can differentially protect tumor subpopulations while simultaneously suppressing normal hematopoiesis. AI-resolved spatial heterogeneity as a map of “sanctuaries” for quiescent, stem-like MM cells. Recent advances are beginning to unlock the spatial architecture of MM BM biopsies at a resolution that is difficult to capture with aspirate-based profiling alone. Hagos et al[30] developed DL pipelines, MoSaicNet (habitat/tissue segmentation) and AwareNet (rare-cell-aware detection/classification), to enable spatial mapping of multiplex-immunohistochemistry BM trephine biopsies in monoclonal gammopathies of undetermined significance and newly diagnosed MM. Importantly, they report that the most significant monoclonal gammopathies of undetermined significance, newly diagnosed MM distinction was not cell density, but spatial heterogeneity, including differences in spatial proximity between BLIMP1+ tumor cells and CD8+ cells. These AI outputs should be interpreted biologically as candidate “sanctuary” regions, i.e., spatially delimited microenvironments where tumor-immune distances and local tissue habitats indicate immune access vs immune exclusion, and therefore where quiescent/drug-persistent tumor fractions are most likely to be protected in situ. In this framework, “spatial heterogeneity” and “niche interactions” detected by MoSaicNet/AwareNet are not merely descriptive patterns; they can be used to localize and quantify protective microanatomical niches that function as sanctuaries for quiescent, stem-like myeloma populations[30]. This interpretation is strengthened by independent in vivo evidence that quiescent, stem-like MM cells show niche preference. By monitoring quiescence using PKH dye retention, Chen et al[57] have shown that quiescent PKH+ MM cells selectively inhabit endosteal/osteoblastic niches (osteoblastic niche) as opposed to vascular niches, and that these quiescent cells exhibit increased stem-like properties and tumorigenicity. Thus, when AI segmentation/detection frameworks (MoSaicNet/AwareNet) delineate distinct BM habitats and tumor-immune spatial relationships in trephines, the biologically explicit read-through is that AI is helping identify the microanatomical sanctuaries that are most consistent with quiescent, stem-like MM cell persistence, including osteoblastic/endosteal protective zones already implicated experimentally[57].
One of the clinical implications is that the MM niche is not merely tumor-supportive; it can be functionally hostile to normal hematopoietic stem/progenitor cells (HSPCs), which offers a mechanistic linkage between niche remodeling and MM-associated cytopenias (including anemia, and in some cases, leukopenia/neutropenia). Bruns et al[58] conducted quantitative, molecular, and functional studies on BM-derived CD34+ HSPCs in de novo MM and demonstrated that hematopoietic impairment cannot be attributed to crowding out alone. Rather, they state that proliferation, colony formation, and long-term self-renewal are inhibited due to activated TGF-β signaling, and that the decreased stem-cell-supporting capacity of MM-derived mesenchymal stromal cells can be restored by TGF-β blockade-supporting a model where microenvironmental cues mediate reversible HSPC suppression. That is, MM reconfigures the BM niche to a condition that is capable of both protecting tumor persistence and inhibiting normal HSC/HSPC activity, which offers a direct mechanistic pathway to anemia/neutropenia[58].
Combined, spatial AI platforms on trephine biopsies can be explicitly placed as dual-purpose niche readouts: (1) To detect and quantify sanctuary habitats that may protect quiescent, stem-like MM cells (and thus inform treatment resistance and relapse risk); and (2) To map niche states that correlate with suppressed normal hematopoiesis, providing a spatially based explanation of MM-associated hematopoietic failure[30].
Disease modeling and computational biology
The wider context of AI in hematology studies has been well-documented. Kilic Gunes[59] performed a comprehensive review of AI applications in hematology, revealing the main trends since 2020, such as ML, DL, NLP, and clinical decision support systems, which enhance diagnostic and prognostic functions. Likewise, Nazha et al[60] surveyed AI applications in the field of hematology, specifically NLP of text, image processing, and predictive clinical decision-making algorithms, showing how AI/ML can help resolve the problem of growing amounts of patient data. Silva-Sousa et al[61] gave an extensive overview of AI and systems biology in stem cell studies and therapeutic development, noting their use in adult stem and progenitor cells, such as HSCs/HSPCs. This article highlighted the revolutionary possibilities of AI in comprehending stem cell biology and creating new therapeutic approaches. Raghav et al[62] have reviewed the application of computational methods in genomic and transcriptomic analysis in detail and outlined how HSC potential can be unlocked by integrative computational methods, such as scRNA-seq, inference algorithms, and ML to analyze gene regulatory networks at single-cell resolution. These methods allow unprecedented understanding of the mechanisms of HSC differentiation, self-renewal, and lineage commitment.
AI-driven prediction of therapy resistance: Leukemic stem-cell persistence as the biological substrate
The fundamental barrier to curative therapy in AML, and the unifying mechanism of resistance across most hematologic malignancies, is not the failure to achieve initial cytoreduction but the inability to eradicate LSCs. LSCs are infrequent, dormant cells that inhabit protective BM niches, overexpress anti-apoptotic proteins (BCL-2, MCL-1), use adenosine triphosphate-binding cassette transporter-mediated drug efflux, and are predominantly in a G0/quiescent cell-cycle state, making them largely resistant to traditional cytotoxic agents targeting rapidly dividing blast populations[50,63]. BCR::ABL1-independent LSC survival mechanisms mediate tyrosine kinase inhibitor resistance in chronic myeloid leukemia; MMSC compartments also support disease relapse following deep remissions in MM[50]. In all these malignancies, resistance to therapy is thus inseparable from the biology of stem-cell persistence. The combination of genomic, transcriptomic, proteomic, and clinical data at scale has made AI a key instrument in identifying the molecular signature of LSC-enriched disease states and predicting clinical resistance with unprecedented accuracy.
AI decoding of LSC-enriched resistance states via single-cell transcriptomics
The analysis of scRNA-seq data using AI has proven especially effective in mapping the transcriptional architecture of therapy resistance, as it can resolve cell-population-level heterogeneity that cannot be seen using bulk multi-omics methods. scRNA-seq and AI can be used to identify quiescent stem-like cells and LSC populations that cause therapeutic resistance and relapse, as well as the definition of functionally distinct LSC subtypes, such as CD36-high cells with predominant self-renewal capacity and CD69-high cells with predominant proliferative capacity, the proportions of which predict clinical treatment outcomes[54]. These AI-solved transcriptional cell states are direct biological representations of the biological substrate of resistance: LSCs that endure induction chemotherapy by enforced stemness programs (HOX/MEIS1, BCL-2 dependence, metabolic oxidative phosphorylation dependence) and niche-mediated quiescence, not by stochastic mutational drug resistance. The analysis of the scRNA-seq has also shed more light on how chemotherapy itself reprograms proliferating stem/progenitor-like cells into quiescent stem-like cells - a process whereby doses of chemotherapy that are not sufficient to kill cells transform them into more resilient dormant cells that later lead to relapse[54]. These AI-defined dormancy programs can now be therapeutically targeted: Quiescent LSC subpopulations with surface markers CD52, LGALS1, CD99, and CD49d have been discovered and confirmed as candidate therapeutic targets in chemoresistant AML[54].
Multi-omics DL: Capturing the LSC stemness signature in patient cohorts
Multi-omics DL approaches have demonstrated strong predictive performance, achieving high accuracy (R2 ≈ 0.90) in cell-line settings for drug response prediction, establishing a foundation for modeling therapy resistance. The clinical significance of these models extends beyond binary resistance classification: By simultaneously processing genomic, transcriptomic, and proteomic inputs, these algorithms implicitly learn LSC-enriched gene expression programs, including the 17-gene LSC stemness score, that characterize relapse-prone disease even when LSC burden is not explicitly measured. A ML analysis of 1383 patients supervised by multiple centers showed that AI models with multi-omics capabilities, such as LSC burden quantification, significantly outperform traditional ELN risk stratification in predicting complete remission and 2-year overall survival. CNN morphology models have also achieved > 98% accuracy in classifying LSC-enriched vs normal progenitor populations in BM imagery, which could be incorporated into routine diagnostic workflows. ML-based on networks that integrate genomic and proteomic data has also identified core resistance-driving signaling modules, such as BCL-2 family regulation and metabolic reprogramming pathways, which act specifically within the LSC compartment and are fundamentally different from the resistance mechanisms of bulk blast populations.
From resistance prediction to LSC-targeted therapy selection
The latest AI applications in this field have translated the predicted resistance states into actionable, personalized therapeutic plans. A ML model that combines scRNA-seq profiles with ex vivo single-agent drug sensitivity data (XGBoost) was able to identify synergistic drug combinations targeting therapy-resistant cell populations, enriched with LSC markers and blast-like transcriptional states, in paired diagnosis/relapse samples of patients with AML[23]. The predicted combinations exhibited relapse-specific synergy and generated minimal co-inhibition of normal lymphoid populations, indicating that AI can decode LSC-specific vulnerabilities, which are not evident based on standard pharmacological screens. In preliminary prospective validation using VenEx clinical trial samples, this scRNA-seq/ML platform accurately predicted clinical responses to venetoclax-azacitidine combination therapy, establishing a direct and clinically actionable link between AI-detected LSC biology and patient outcome[23].
In chronic myeloid leukemia, AI models integrating deregulated miRNA expression patterns and BCR::ABL1-independent metabolic signatures have identified quiescent LSC-like subpopulations as the dominant substrate of tyrosine kinase inhibitor resistance, consistent with clinical observations that deep molecular responses rarely translate to LSC eradication. In MM, an AI-derived five-gene transcriptomic signature retains clinical validity as a biomarker of bortezomib resistance and MMSC persistence in the BM niche. Spatial multi-omics AI tools, including MoSaicNet and AwareNet, enable mapping of BM niche interactions, crosstalk between LSCs, mesenchymal stromal cells, and immune populations, that actively sustain the therapy-resistant LSC state, opening new avenues for niche-disruption strategies that complement direct LSC-targeted agents.
Limitations of current LSC AI models
Despite encouraging performance, current AI models that aim to infer LSC biology or LSC-associated resistance signals still face important translational limitations. Many models are trained on retrospective and/or single-centre cohorts with limited numbers of outcome events, and frequently lack rigorous external evaluation; for leukemia prediction models overall, a recent systematic review found that 52% did not report internal validation and 57.4% had no external validation, with most studies judged at high risk of bias, issues that directly constrain generalisability to new centres, assays, and treatment eras[64]. In addition, a large proportion of LSC-focused modelling still relies on archived (“bulk”) datasets (e.g., banked omics and historical clinical records), which can encode selection effects, missingness patterns, and population under-representation; such data limitations are a known driver of biased or non-portable medical AI performance[65]. Consistent with contemporary guidance for AI-based prediction modelling, these constraints underscore the need for transparent reporting of study limitations (including sample representativeness and sample size) and, most importantly, prospective multicentre evaluation before LSC AI tools are used to guide clinical decisions[66,67].
AI-driven drug discovery and development
AI has profoundly accelerated drug discovery in hematologic malignancies, particularly by shortening early discovery phases through virtual screening, AI-enhanced docking, and in silico target validation[68,69]. AI-driven virtual screening and design can compress hit identification and lead optimization from months to weeks and substantially reduce the number of compounds requiring experimental testing[70,71]. Target identification platforms that integrate transcriptomic and other omics datasets with ML models on protein-protein interaction networks have expanded the druggable landscape beyond classical kinase and epigenetic targets, highlighting structural proteins, metabolic enzymes, and regulatory non-coding RNAs as candidate vulnerabilities[68,69]. However, the main limitation of the initial AI drug-discovery programs in AML has been their emphasis on bulk blast cytoreduction instead of LSC eradication: Models that are trained mostly on blast-level viability data identify compounds that are effective at reducing overall disease burden without depleting the relapse-initiating stem-cell reservoir. It is now being fundamentally re-oriented, with AI platforms explicitly designed to target agents that disrupt LSC-specific survival programs but spare normal HSCs. Three LSC-directed drug axes have emerged as the most clinically mature products of this AI-assisted precision oncology approach.
The BCL-2 axis - AI-predicted LSC apoptotic dependencies: LSCs are highly dependent on anti-apoptotic BCL-2 family proteins, particularly BCL-2 itself, as a metabolic and survival strategy linked to their mitochondrial oxidative phosphorylation reliance. ML models integrating scRNA-seq and ex vivo drug sensitivity profiles from primary patient samples can predict individual patient responses to venetoclax with Pearson correlation coefficients of 0.71-0.84 (after conformal prediction filtering), identifying LSC-enriched blast subpopulations with BCL-2-high transcriptional profiles that confer venetoclax sensitivity[23]. Critically, these models simultaneously identify patients whose dominant resistant cell type, monocytic AML blasts with reduced BCL-2 dependence and MCL-1-high profiles, predict primary venetoclax resistance, enabling prospective therapy stratification. In the VenEx clinical trial, an XGBoost platform using scRNA-seq and 19-compound ex vivo response profiles accurately classified clinical responders vs non-responders to venetoclax-azacitidine, demonstrating the translational readiness of AI-based LSC-vulnerability prediction for prospective clinical application[23].
The HOX/MEIS1-Menin axis - AI-guided disruption of LSC self-renewal programs: Menin is a scaffold nuclear protein that, in AML, sustains leukemogenic transcriptional programs through its interaction with KMT2A-associated complexes, maintaining HOX/MEIS1-high expression patterns that are tightly linked to LSC self-renewal and differentiation arrest. Because disrupting the menin-KMT2A interaction induces differentiation and reduces LSC-enriched compartments, Menin inhibition represents a prototypical LSC-directed therapy rather than a bulk cytoreduction strategy[50,72]. AI-driven single-cell multi-omics has been instrumental in decoding the clonal architecture of KMT2A-rearranged and NPM1-mutated AML, the two major Menin-inhibitor-sensitive genotypes, by mapping transcriptional heterogeneity, clonal evolution under therapeutic pressure, and the relationship between HOX/MEIS1 program activity and clinical response[50]. This AI-enabled mechanistic understanding has informed rational combination strategies to overcome primary and secondary Menin inhibitor resistance, a critical challenge as the first approved agents (revumenib and ziftomenib) enter broader clinical practice[50].
LSC surface target identification - AI-ranked immunotherapeutic candidates: The immunophenotypic heterogeneity of AML LSCs, expressing variable combinations of CD123 (IL-3Rα), T-cell immunoglobulin and mucin-domain containing-3, C-type lectin-like molecule-1 (CLL-1), CD47, CD70, CD96, and interleukin-1 receptor accessory protein against a background of normal HSC antigens, creates a complex optimization problem for immunotherapeutic target selection. Systematic ranking of LSC surface targets by their specificity (differentiated expression on LSCs vs normal HSCs), expression stability across disease stages, and therapeutic window has been systematically ranked using ML analyses of large MFC datasets[55,72]. Computational comparison of marker expression across AML subtypes and at diagnosis, remission, and relapse has identified CD123 and CLL-1 (CLEC12A) as consistently high-priority targets due to their broad expression across LSC compartments and relative absence on HSCs; these analyses have directly informed ongoing clinical trials of CD123-directed and CLL-1-directed cellular therapies[50].
AI-powered personalized drug combination design for LSC eradication: Perhaps the most transformative application of AI in this domain is personalized, LSC-focused drug combination design for relapsed/refractory AML, a setting where both disease heterogeneity and resistance evolution render standard combinations inadequate. An XGBoost-based computational-experimental strategy that integrates paired scRNA-seq (from diagnosis and relapse) with ex vivo single-agent drug sensitivity profiles successfully predicted cancer cell-selective and synergistic drug combinations targeting treatment-resistant LSC-enriched subpopulations in R/R AML[23]. The approach identified that cell population compositions evolve dynamically and uniquely between the diagnostic and relapsed stages in each patient, requiring individualized rather than histology-based combination strategies to target LSC-like resistant populations. Experimentally validated combinations demonstrated relapse-specific synergy in leukemic cell populations with minimal co-inhibition of normal T and natural killer cells, providing a preclinical framework for clinical translation. The entire pipeline, from sample collection through scRNA-seq, computational prediction, and flow cytometry validation, was completed within approximately two weeks, a clinically actionable timeframe consistent with AML disease management requirements[23].
Metabolic and drug-repurposing strategies targeting LSC vulnerabilities: AI-driven drug repurposing programs have further identified approved agents that exploit metabolic and epigenetic vulnerabilities unique to LSCs. scRNA-seq with AI trajectory analysis has shown that quiescent AML LSCs rely heavily on oxidative phosphorylation over glycolysis, and exploit fatty acid and amino acid catabolism as alternative energy sources, a metabolic dependence not seen in normal HSCs and can be targeted by approved mitochondrial translation inhibitors (e.g., tigecycline). Similarly, scRNA-seq/AI analysis has identified galectin-1 (LGALS1) inhibitors as selective against quiescent stem-like cells that cause chemotherapy residual disease, which represents a repurposing opportunity with potential to eliminate a population orthogonal to the targets of both venetoclax and Menin inhibitors[54,73]. These opportunities of metabolic repurposing, computationally prioritized amongst existing drug libraries, provide a resource-efficient route to agents capable of eliminating LSC subpopulations that are resistant to current-generation targeted therapies.
Biomanufacturing and Industry 4.0 applications
The HSC biology-to-clinical-products translation requires manufacturing systems that can produce consistent, potent, and regulatory-compliant cell grafts at scale. The overlap of AI with Industry 4.0 principles, including real-time sensing, digital connectivity, automation and data-driven decision-making, is fundamentally transforming the biomanufacturing landscape of HSC and other advanced therapy medicinal products. In contrast to the previous manufacturing paradigms, which were based on manual sampling and retrospective analysis of batches, AI-enabled platforms can continuously track culture dynamics, predict product quality in real time, and autonomously adjust process parameters in order to maximize yield and minimize variability. This section discusses the application of AI in four key dimensions of HSC biomanufacturing, including product quality control (QC), ex vivo expansion optimization, bioreactor modeling using digital twins, and automation that complies with Good Manufacturing Practice (GMP).
AI-driven QC of HSC products
The HSC-based cellular products QC poses special analytical difficulties. The critical quality attributes of a hematopoietic graft, including CD34+ cell dose, viability, engraftment potency, and absence of aberrant subpopulations, are traditionally measured by a combination of flow cytometry, colony-forming unit assays, and sterility testing. These techniques are time-consuming, operator-dependent, and in many cases, the results are only known after the manufacturing process has been completed and thus, there is little time to take corrective action. AI is changing this paradigm by allowing real-time, predictive, and label-free quality evaluation across the manufacturing workflow.
A comprehensive review of AI-driven quality monitoring in stem cell cultures, specifically demonstrating that reinforcement learning (RL) systems can use real-time cytokine feedback to dynamically optimize HSC expansion and, at the same time, act as an online quality sentinel[37]. Their architecture demonstrated that DL models implemented on high-dimensional culture monitoring data could detect deviations in cell phenotype earlier than conventional endpoint assays, and provided a proactive rather than reactive quality management strategy. Likewise, Gramatiuk et al[74] reported AI-based QC platforms that combine computer vision and DL to automatically characterize cell lines based on microscopic imaging data, with classification results validated against gold-standard flow cytometry that demonstrated comparable performance to experienced laboratory analysts and removed inter-operator variability.
One of the most effective applications is related to the prediction of the dose of CD34+ cells in cord blood units that are to be cryopreserved. Leung et al[35] created a set of parametric and non-parametric ML models, such as multivariate linear regression, random forest, and back-propagation neural networks, and trained them on 802 cord blood units of the HealthBaby cord blood bank. The neural network of back-propagation gave the highest forecast accuracy (56.99%), with the most predictive variables being the post-processing total nucleated cell count, pre-processing leukocyte counts, and collection-to-processing time interval. This model provides cord blood banks with an intelligent decision-support tool to estimate graft potency before cryopreservation, thus allowing selection of the highest-quality units to be used in clinical practice without necessarily depending on costly direct CD34+ enumeration at each step.
In addition to this predictive method, Wang et al[11] were the first to apply DL to morphology-based classification of HSC subpopulations. They used a ResNet-50 CNN trained on differential interference contrast microscopy images of murine HSCs and MPPs to develop the so-called LSM model, which achieved an overall F1 score of 0.74 and receiver operating characteristic-AUC values of more than 0.85 across all subclasses. Importantly, the model identified long- and short-term HSCs, two populations that share the same surface markers but are in different functional states, showing that morphological AI analysis can capture biological information that is invisible to traditional flow cytometric gating. This label-free, antibody-independent method may significantly decrease the reagent cost and technical load of routine HSC characterization during manufacturing. Desa et al[75] also determined that label-free optical imaging modalities, when integrated with ML algorithms, can detect viable cells and predict optimal manufacturing conditions without disrupting the culture environment, positioning AI-optical QC as a scalable, non-invasive strategy to continuous in-process monitoring.
ML is also enhancing the yield of peripheral blood stem cell apheresis at the clinical collection stage. Qi et al[36] constructed ML models to predict the outcome of CD34- cell harvesting by autologous and allogeneic donors, identifying pre-apheresis CD34- blood counts, the type of mobilization regimen, and donor characteristics as key determinants of yield. These predictive tools allow proactive scheduling of further mobilization cycles or alternative donor selection, directly reducing the rate of failed harvests and subtherapeutic grafts, a major source of clinical and economic inefficiency in existing HSCT programs.
AI-optimized ex vivo expansion of HSCs
The ability to generate clinically adequate numbers of HSCs using limited-dose sources, especially umbilical cord blood, has long been one of the most critical limitations in HSCT. The 2023 regulatory approval of Omisirge® (omidubicel), the first approved ex vivo expanded HSC product, demonstrates that AI-assisted expansion platforms can deliver grafts that accelerate neutrophil engraftment and reduce infection-related complications compared to unmanipulated cord blood transplantation, validating the clinical utility of expansion technology. Nevertheless, it is necessary to optimize the complex interaction between cytokine concentrations, oxygen tension, pH, nutrient gradients, and culture vessel geometry in expansion systems based on bioreactors, which is a daunting challenge that AI is uniquely poised to solve. The significance of multi-parameter process optimization to obtain reproducible primitive cell expansion was highlighted by a comprehensive review of the current clinical-grade expansion platforms of HSCs and hematopoietic progenitor cells[76]. The composition of the cytokine cocktail, which usually includes SCF, FMS-like tyrosine kinase 3 ligand (FLT3 L), TPO, and small-molecule agonists such as SR1 or UM729, should be carefully adjusted to maintain HSC self-renewal instead of inducing premature differentiation. Traditional one-factor-at-a-time optimization is severely insufficient to model the non-linear interactions between these variables, and requires AI-based experimental design and predictive modelling strategies. RL has become one of the most promising approaches to dynamic process control in HSC bioreactors. Rather than using fixed protocols, predictive AI and RL models continuously optimize accurate bioprocessing parameters in real time. In particular, these algorithms dynamically optimize the concentrations of cytokines (e.g., titration of SCF, TPO, and FLT3 L depending on the real-time detection of cell cycle states), adjust the bioreactor perfusion rates to clear the inhibitory byproducts, and regulate the metabolic flux by balancing the glucose delivery with the continuous production of lactate readouts[37,77]. As shown by Kim and Kim[77], RL-based control structures were capable of controlling shear stress and culture stimulation conditions in real time within stem cell culture systems, which was superior to classical proportional-integral-derivative controllers in both response time and stability. Similar results were demonstrated by Sharma et al[78], who demonstrated that RL may be an effective way to control the parameters of microfluidic bioreactors to downstream biopharmaceutical processing, such as regulation of stem cell-niche interactions, with the agent learning near-optimal feeding strategies through iterative environmental interaction. Simultaneously, Williams et al[79] showed that a new process analytical technology (PAT) could be implemented with the help of ML combined with metabolic modelling structure in cell and gene therapy manufacturing, making real-time pH and metabolic monitoring actionable process control inputs, a direct template to HSC bioreactor management. Their ML-PAT methodology allowed feed-forward control adjustments that kept cells in optimal metabolic states throughout the expansion cycle, which was not possible with time-point sampling methods. The idea of intelligent bioprocessing of HSC cultures has been discussed since the seminal work of Lim et al[80], who explained the critical importance of real-time monitoring and intelligent design of experiments to ensure optimal conditions of HSC expansion. Modern applications have far surpassed this initial conception: AI systems can now combine multi-parameter sensor streams, dissolved oxygen, glucose, lactate, cell density, and cytokine concentrations, to build predictive models of culture trajectory and intervene proactively when deviations are detected. Lin et al[81] surveyed developments in ML-assisted multi-parameter bioreactor process control and noted the convergence of mechanistic modeling and data-driven methods as the most appropriate approach to managing the complexity of living cell manufacturing systems.
Digital twin bioreactor modeling
One of the most radical changes in the biopharmaceutical manufacturing process is the concept of the digital twin. It creates a constantly updated virtual replica of a physical manufacturing process that is synchronized with real-world sensor data. In the case of HSC bioreactors, digital twins can provide the capacity to simulate culture dynamics in silico, test process modifications without risking physical batches, predict batch outcomes hours before harvest, and support scale-up decisions with computational evidence instead of extensive empirical experimentation.
Kanwar et al[82] designed a digital twin framework with an autonomous linear quadratic regulator control of bioreactor systems used in immune cell expansion, which proved that the virtual model could maintain cell expansion trajectories within specified specification windows under different perturbation conditions. The framework explicitly tackled the problem of limited empirical data in new cell therapy bioreactor systems by integrating mechanistic kinetic modeling with data-driven parameter estimation, a hybrid approach that is especially well-suited to the HSC culture, where reference datasets are sparse compared to established industrial bioprocesses.
Hengelbrock et al[83] took this paradigm a step further by creating a digital twin of human mesenchymal stem cell-derived extracellular vesicle production in a 3D bioreactor, a directly translatable architecture of HSC manufacturing, incorporating PAT and the virtual model to achieve autonomous operation with chromatographic purification. Their digital twin attained precise estimation of process conditions and allowed closed-loop feedback control, which was a demonstration of a fully autonomous cell therapy manufacturing process. The combination of PAT sensors with digital twins is becoming the technical architecture of choice in next-generation HSC manufacturing: Real-time spectroscopic and electrochemical sensors provide the data streams that drive the virtual model coordinated with the physical process, and the predictive outputs of the digital twin are used to make automated control decisions.
Raudenbush et al[8] have reviewed the development of pharmaceutical digital twins at various manufacturing scales, stating that the final value of the technology is that it will enable the development of what-if simulations during the process development of manufacturing processes, and provide risk-free training environments to manufacturing personnel. Isoko et al[9] put these advances in the broader Bioprocessing 4.0 framework, describing how convergence of Internet of Things sensors, AI-driven analytics, and digital twins is enabling adaptive process control strategies that would have been computationally infeasible with traditional mathematical methods.
GMP automation and closed-loop manufacturing
The GMP framework has strict requirements on the reproducibility, traceability, control of contamination, and documentation of all clinical-grade cell products. Traditionally, these demands have been fulfilled by intensive manual workflows with a high level of human control, which is not a scaling process, introduces variability, and creates access barriers because of high labor costs.
An automated AI-driven CAR-T cell manufacturing platform that integrates robotics, closed-loop bioreactors, automated QC modules and an AI framework into a concept of a smart manufacturing hospital has been described as developed within the European Union AIDPATH project[38]. While the immediate application was CAR-T cell manufacturing, architecture is explicitly cell-type agnostic and directly applicable to the manufacture of HSC products. The AI components of the platform have several functions: A digital twin tracks the cell product through the entire manufacturing process; RL controls the adaptive bioreactor scheduling; and a comprehensive data management layer, based on the Observational Medical Outcomes Partnership Common Data Model, integrates patient data, process parameters and sensor data streams to support continuous model training and GMP-compliant batch records. The authors showed that large-scale economies of scale could be achieved by parallelized production by use of interchangeable bioreactor cartridges and that the one product, one patient integrity of autologous therapies could be achieved by use of interchangeable bioreactor cartridges. Reviewing the regulatory environment to deploy AI/ML in GMP biopharmaceutical manufacturing, Panjwani et al[39] found that although formal regulatory frameworks are beginning to address the topic of AI validation, they are rapidly lagging behind the field. Their analysis found areas to be critical that AI-specific GMP considerations are required, including: Model drift management, explainable requirements of automated release decisions, cybersecurity in connected manufacturing systems, and data governance of continuously learning systems. Importantly, they demonstrated that AI applications in manufacturing can enhance GMP compliance by eliminating transcription errors in batch records, flagging out-of-real-time specification events, and providing detailed audit trails, functions that are categorically better than manual systems.
Jayaraman et al[84] demonstrated that the closure of the manufacturing process by automation and integration of digital workflow documentation is not only an easy way to comply with GMP but also a faster way to translate the manufacturing process of cell therapy products between research and clinical scale, a model that is directly applicable to HSC manufacturing programs. Their analysis quantified the reduction in the time of contact between personnel and products, the risk of contamination, and the rate of batch failures, which can be achieved through process closure, which forms the safety and economic case to invest in automation.
Niazi[85], also, on a regulatory and technical basis, asserted that properly qualified AI/ML systems implemented in GMP processes, such as automated visual inspection, predictive maintenance, real-time environmental and process monitoring, and AI-assisted deviation/corrective and preventive action management, can enhance process robustness, reduce deviations, and strengthen regulatory compliance in pharmaceutical manufacturing. Collectively, these advances characterize an imminent paradigm of HSC manufacturing in which AI-enabled QC systems continuously define the evolving graft product, digital twins maintain bioreactor conditions at computationally-optimal set-points, RL agents adapt feeding and cytokine delivery strategies dynamically, and closed GMP-compliant systems deliver complete electronic records of batch without manual transcription. The clinical implications are enormous: Improved quality of grafts manufactured with increased reproducibility, shorter vein-to-vein times, and lower manufacturing costs that increase patient access to potentially curative HSC therapies. In the HSC research, Figure 2 depicts the end-to-end AI workflow, starting with the acquisition and preprocessing of raw data, through model training, validation, and clinical deployment. Incorporation of AI into cell manufacturing that is more patient-specific (including autologous CAR-T and HSC grafts) also increases privacy and ethical concerns that should be strictly monitored.
Figure 2 Schematic representation of the artificial intelligence-driven biomanufacturing pipeline for cell therapies incorporating Industry 4.0 principles.
The manufacturing process integrates three key artificial intelligence technologies: (1) Digital twin technology for virtual process simulation, real-time optimization, and predictive maintenance of bioreactors; (2) Computer vision systems powered by deep learning for automated quality assessment, real-time morphological analysis, and contamination detection with precision exceeding manual inspection; (3) Artificial intelligence-driven analytics for predictive maintenance and adaptive process control. These technologies enable closed-loop manufacturing systems that minimize manual intervention, enhance reproducibility and productivity of chimeric antigen receptor-T and hematopoietic stem cell products, and facilitate scalable production suitable for commercial manufacturing. CAR-T: Chimeric antigen receptor-T; HSC: Hematopoietic stem cell; AI: Artificial intelligence.
The use of AI in the production of the individualised cell (e.g., autologous CAR-T and HSC grafts) also intensifies the privacy and ethical issues that require strong supervision. The continuous processing of highly sensitive, donor-derived biological data, ranging in sensitivity and information amount from genomic profiles to real-time cellular metabolic phenotypes, has an inherent risk of data breach or deanonymization along the digital supply chain. The most important thing is to ensure secure and end-to-end encrypted transmission of data between clinical apheresis centers and smart biomanufacturing facilities. Moreover, ethical considerations concerning algorithmic bias would need to be addressed; predictive manufacturing models that have been trained on cells that are predominantly representative of a certain demographic group may not optimally expand the population of cells that are underrepresented in the dataset that the model was trained on[39,85-87].
NLP in transplant medicine and donor selection
NLP is a specific area of AI that aims to extract structured information out of unstructured clinical text, such as electronic health records, pathology reports, and transplantation documentation. NLP applications have become essential in HSCT to extract complex genetic information and predict transplant-related complications that would otherwise remain undisclosed in free-text clinical narratives.
HLA typing extraction from clinical documentation: The automated derivation of HLA genotype data from unstructured clinical reports is one of the most clinically relevant uses of NLP in HSCT. Lee et al[40] developed a method to extract HLA typing information in electronic medical records and they processed more than 70000 HLA reports stored in free-text form. Python regular expression was used in the extraction pipeline functions to discover patient data, clinical features and exact HLA genotypes in five HLA genes (A, B, C, DR, and DQ). The system has achieved precision of 0.892-0.999 and recall of 0.795-0.998 on the extraction of HLA genotype as well as the cleaning of the data to enable the conversion of the variable-resolution data into the standard nomenclature. This solution is a serious clinical issue. Despite HLA typing, they tend to be stored in unstructured forms that limit their second use in clinical decision support and research.
Predicting graft-vs-host disease from clinical notes: NLP methods have been effectively used to forecast acute graft-vs-host disease (aGVHD), a life-threatening complication following allogeneic HSCT. Jo et al[41] created a CNN-based prediction model with NLP to process HLA information of the Japanese Transplant Registry (18763 patients). The model used word2vec, an NLP application, to treat HLA antigens and alleles as natural language, converting complex HLA typing data into computer-friendly vector representations. This CNN-NLP hybrid model was found to have better risk stratification than more traditional Cox proportional hazard models, and was able to predict grade II-IV and III-IV aGVHD with the capacity to utilize detailed, raw HLA data without the arbitrariness of more traditional binary matching (matched vs mismatched) classifications. The model showed that aGVHD risk is not only determined by the HLA disparity but also by detailed HLA information combined with various clinical factors, with prediction scores stratifying cumulative incidence between 31.8% to 54.8% between high-risk and low-risk groups.
Besides prediction, recent developments have been aimed at deriving complex clinical phenotypes out of unstructured documentation. Large language models have been used to identify various manifestations of chronic graft-vs-host disease (cGVHD) in clinical notes, and it is pleomorphic in nature, affecting multiple organ systems. ML combined with clinical feature extraction has revealed seven distinct cGVHD phenotypes with contrasting clinical risks, demonstrating that computational analysis of organ involvement patterns can stratify survival more effectively than traditional composite severity scores[42]. These phenotypes, derived from analyzing mouth, eye, liver, gastrointestinal, joint, fascia, and skin involvement documented in clinical narratives, identified high-risk patient groups with 2.24-fold higher mortality compared to low-risk groups, independent of National Institutes of Health consensus severity criteria.
Clinical decision support in transplantation: NLP-powered systems in HSCT extend beyond data extraction to active clinical decision support. Turki et al[4] reviewed AI applications in allogeneic HSCT care, emphasizing that NLP enables automated parsing of transplant-related complication documentation, including infection management, donor selection criteria, and non-relapse mortality prediction. The integration of NLP with ML algorithms enables real-time analysis of longitudinal clinical narratives, thereby facilitating early warning systems for transplant complications.
Colas et al[88] demonstrated a proof-of-concept application of NLP on electronic medical records for cGVHD assessment, showing that automated extraction of clinical features from free-text documentation can support standardized outcome measurement and quality-of-life assessments. These applications exemplify how NLP transforms transplantation practice by unlocking clinically actionable information embedded in narrative documentation that would be impractical to extract through manual chart review. Stem-cell-wise, the accurate prediction of graft-vs-host disease directly supports the improved management of the allogeneic graft of the HSC stem cells, to ensure the viability and long-term engraftment of the transplanted stem cells. By doing so, AI-based graft-vs-host disease risk models can be used both as clinical outcome predictors and as a means to maintain graft stem-cell functionality and ensure long-term hematopoietic reconstitution.
Computational frameworks and data resources
The availability of high-quality, well-annotated datasets and powerful computational frameworks is essential to the successful application of AI in hematology. Recent work has been on creating specialized databases, analytical pipelines, and standardized methodologies. Yi et al[2] surveyed biological data sources and ML models to study hematology, describing the Atlas of Blood Cells project and specialized single-cell databases devoted to blood and immune cells, such as the Atlas of Human Development of Hematopoietic Stem Cell. These resources allow researchers access to curated datasets that are appropriate for training and validating AI models and allow quick development of clinically relevant applications.
Integrative multi-omics and AI as a novel approach to systems biology, based on integrative multi-omics data, were thoroughly reviewed by Kant et al[89], who highlighted that integrative multi-omics and AI are a new paradigm of systems biology, as they provide deep insights into the complexity of life and allow the innovation of medicine and biotechnology. These computational frameworks make it possible to have a comprehensive study of biological systems to unveil emergent properties and network-level phenomena that cannot be identified based on single data modalities. Table 2 summarizes publicly available data, multimodal sources, and computational systems that would be specifically useful in the AI research on HSCs[5,6,13-16,56,90-97].
Table 2 Public datasets and computational tools specifically applicable to hematopoietic stem-cell artificial intelligence research.
To synthesize the methodological landscape reviewed in these areas, Table 3 presents a comparative overview of the main AI and ML algorithms that are currently in use in the study of HSCs. It is necessary to understand the particular strengths and inherent limitations of these computational architectures to appropriately match the algorithm to the specific biological or clinical question.
Table 3 Summary of key machine learning algorithms, applications, and trade-offs in hematopoietic stem cell research.
Machine learning algorithm
Primary HSC/hematology applications
Key advantages (Pros)
Key limitations (Cons)
Ref.
CNN
Morphological classification of HSCs vs MPPs; 3D chromatin age prediction (ChromAgeNet); automated quality control imaging in biomanufacturing
Unparalleled performance on spatial and image data; extracts features autonomously without requiring manual gating or human-defined parameters
Highly opaque “black box” nature requiring XAI for interpretability; demands massive, accurately annotated image datasets to train
Highly robust to overfitting; handles tabular clinical and multi-omics data effectively; naturally provides feature importance rankings
Less effective than deep learning for highly unstructured data (like raw images or free text); can struggle with extrapolating data outside the training range
Can capture extremely complex, non-linear biological relationships across massive, high-dimensional datasets (e.g., integrating RNA and protein expression)
Computationally intensive; high risk of learning artifactual batch effects rather than true biology if data is not strictly harmonized
Dynamic optimization of ex vivo HSC expansion; adaptive control of bioreactor parameters (cytokines, perfusion, metabolic flux)
Enables continuous, autonomous, real-time process optimization without requiring a pre-defined static protocol
Requires highly accurate “digital twins” or simulation environments to train the agent safely; initial validation in GMP environments is regulatory complex
The introduction of AI into HSC research and clinical hematology has fundamentally changed the field to be a data-driven, precision medicine paradigm. The reviewed evidence shows that AI technologies, including but not limited to disease modeling, biomarker discovery and therapy, have been remarkably successful in a variety of applications: Resistance prediction, drug development, and biomanufacturing. The fact that multi-omics integration with the help of DL algorithms outperforms the traditional risk stratification systems is a paradigm shift in prognostic assessment. Multi-omics ML models have continued to outperform traditional classification schemes (including the ELN and the International Staging System frameworks), identifying novel prognostic biomarkers and robustly predicting drug responses. Such accuracy allows tailored treatment plans, maximizing the choice and timing of treatments to maximize effectiveness and minimize toxicity.
The AI-driven virtual screening and repurposing pipelines accelerate the drug discovery process by in silico prioritizing active compounds in ultra-large libraries and shortening early hit-to-lead timelines of months or years to weeks[7]. According to industry analyses, such strategies can cut early discovery timelines and costs by up to tens of percent, which can lead to significant savings and accelerate the development of promising candidates into the clinic[7]. This efficiency improvement is especially significant in hematologic malignancies like AML, where structure-based virtual screening is particularly important, and AI-assisted repurposing has been used to design multi-target agents to overcome therapeutic resistance and disease relapse[33].
The implementation of Industry 4.0 principles to cell therapy manufacturing, including computer vision, digital twins, and automation, has transformed biomanufacturing into an industrial-scale, quality-controlled process, as opposed to an artisanal process. These technological innovations are necessary to achieve the full potential of CAR-T and HSC therapies, ensure consistent product quality, scalability, and accessibility to larger patient populations. But maybe the most intellectually important conclusion to be drawn out of this review is not a performance measure or an architectural innovation - it is a biological deduction. Throughout the reviewed evidence, a convergent trend is observed: AI prognostic models in AML and related myeloid malignancies consistently outperform conventional ELN risk stratification, and do so most dramatically in the intermediate-risk category - precisely the patient group in which LSC burden is most heterogeneous and least reflected by cytogenetic and molecular biomarkers alone. The manuscript confirms that AML is structured as a cellular hierarchy that is driven by rare, self-renewing LSCs, and that stemness gene-expression programs based on functionally validated LSC-enriched fractions - most prominently the 17-gene LSC score- are stronger predictors of induction failure and survival than bulk blast counts or many standard molecular markers. The logical implication, which this review argues constitutes its most original contribution, is that high-performing AI prognostic models are not merely identifying statistical correlates of clinical risk: They are implicitly learning LSC-enriched transcriptional and epigenetic states from high-dimensional training data. In other words, AI may be functioning as a computational surrogate for LSC burden quantification.
This review reframes how we interpret AI’s superiority over ELN-based systems. ELN guidelines, despite their clinical utility, capture disease risk primarily at the level of somatic mutations and cytogenetic aberrations, the genetic lesions present in bulk leukemic populations. They do not directly interrogate the functional stemness of the malignant clone. An AI model trained on full transcriptomic, epigenomic, or multi-omics data from a large, outcome-annotated cohort is, by contrast, exposed to the full molecular landscape of the disease, including the expression footprint of the resistant LSC compartment. If the model learns that the most reliable predictors of treatment failure reside in LSC-enriched transcriptional programs, as the high performance of the 17-gene stemness score relative to ELN risk explicitly demonstrates, then the model’s learned feature representations should, in principle, be alignable with known LSC biology. This is not merely a speculative proposition: It is a testable hypothesis with immediate clinical and methodological implications that the field is only beginning to confront. As shown in Table 4[17,30-32,41,53,98], diverse ML and DL models have been implemented along the HSCT pathway, from donor selection and HLA typing to graft-vs-host disease prediction and long-term outcome stratification.
Table 4 Artificial intelligence performance vs conventional methods in acute myeloid leukemia and multiple myeloma risk stratification.
Disease
Conventional method
AI method
Key performance advantages
Ref.
AML
ELN risk stratification
Multi-omics deep learning
Outperforms ELN-based approaches. Identifies LSC burden (not quantified by ELN). > 90% accuracy in therapy resistance prediction. Integrates clinical, cytogenetic, and molecular data
Despite the remarkable progress, several critical challenges must be addressed to fully capitalize on AI in HSC research and clinical hematology.
Data heterogeneity and standardization
The heterogeneity of biological and clinical data is one of the main challenges. Data sets in various institutions, platforms, and time series are highly likely to be not standardized in the collection, processing, and annotation of data. This variability results in profound batch effects, such as changes in the median fluorescence intensity because of divergent cytometer configurations or variations in the rates of gene dropout across different scRNA-seq sequencing runs, and confounding factors that can seriously compromise the generalizability and performance of models. This means that the application of strong data harmonization pipelines (e.g., computational batch correction algorithms to single-cell integration of ML models or automated fluorescence alignment protocols) is an absolute requirement prior to ML models could ever be reliably generalized to external clinical cohorts. Standard protocols, data harmonization plans, and data-sharing programs are necessary in order to developing strong, clinically viable AI models.
Algorithm interpretability and explainability
The extreme mathematical complexity of the neural network that many DL models represent makes the underlying decision-making logic of the model opaque, poses serious challenges to the clinical implementation of such models and their translatability to real-world scenarios. To trust and act on AI-generated recommendations, clinicians not only require clear, interpretable explanations of algorithmic predictions but also lack the underpinning trust that even fans of AI-generated recommendations will be reluctant to trust and act on the recommendations provided by AI-driven algorithms.
Exceedingly precise models would not be able to infiltrate the everyday practice of hematology. XAI is a field which has sprung up to eliminate this disadvantage. The approaches such as SHAP and local interpretable model-agnostic explanations have become the instruments that are instrumental in bridging this gap. SHAP and local interpretable model-agnostic explanations can be effectively used to de-black the box. They explain the output of opaque algorithms in a way understandable to humans, and they develop a visualization of the significance of features, to produce counterfactual explanations, and to offer mechanistic insights. The ability to interpret and have confidence in clinical AI systems is a critical deployment requirement, and XAI frameworks, such as SHAP, attention-based visualization, and counterfactual explanations, emerge as key tools ensuring transparency in early diagnosis, prognosis, and treatment planning[99].
Gimeno et al[28] evaluated interpretability on various explainable methods for precision oncology and concluded that certain model types provide better interpretability and ease of implementation. Thiriveedhi et al[100] applied this principle in practice by creating ALL-Net, a CNN that integrates XAI techniques to improve both the diagnostic accuracy and interpretability of ALL classification, directly addressing the transparency and implementation issues inherent in implementing DL models in the care of leukemia. Advancing and implementing XAI systems is essential to developing confidence in clinicians and supporting regulatory approval of AI-based clinical decision support systems.
In addition to XAI visualization tools, to effectively apply AI in clinical practice, it is imperative to reduce the predictive outputs of AI (high-dimensional) into simple, user-friendly bedside tools (low-dimensional). Actionable interfaces are required by clinicians to bridge the gap between complex multi-omics neural networks and everyday decision-making. One such strong framework is the creation of clinical nomograms based on the selection of AI features. As an example, a recent study on post-surgical oncological risk stratification was successful in its attempt to develop and validate a nomogram to predict intracranial infection after high-grade glioma surgery and provide a concrete, point-of-care tool, which translates complex risk data into a format that directly guides clinical prevention[101]. To hematology, DL models that capture the LSC burden and multi-omics signatures can be similarly synthesized into digital or visual nomograms. This would enable hematologists to quickly compute individualized relapse risk and stratify patients receiving consolidative therapies without necessarily having to directly interpret the underlying high-dimensional algorithmic architecture.
Regulatory compliance and validation
The regulatory environment of AI-powered medical devices and clinical decision support systems remains dynamic. The United States Food and Drug Administration, as well as the European Medicines Agency, have described frameworks to evaluate AI/ML-based medical devices and software, including a total product lifecycle approach, good ML practice principles, and reflection papers on AI use in the medicinal product life cycle, but there are still many uncertainties about the validation requirements, post-market surveillance, and algorithm updating procedures[102-104]. The regulatory implications of advanced biotechnologies, such as AI applications, have been discussed in the context of cancer care and that the existing legal and regulatory regimes are not keeping pace with the current rapid development of AI as an abstract concept and as a practical clinical tool. The authors also observed that regulatory frameworks should strike a balance between innovation and patient safety to provide clear avenues of validation, approval, and ongoing monitoring of AI systems[105]. Strict clinical validation studies that will prove that AI-based tools enhance patient outcomes, decrease healthcare costs, or increase clinical workflows are necessary to be approved by the regulator and be adopted by clinicians. Such studies should consider possible algorithmic biases, the ability to generalize across different populations of patients, and how they can withstand changes in the data distribution over time. Table 5 is a concise validation checklist based on TRIPOD + AI and PROBAST + AI guidelines[52,55,106-114].
Table 5 Validation checklist for leukemic stem cell/hematopoietic stem cell artificial intelligence models, adapted from TRIPOD + AI and PROBAST + AI guidelines with hematopoietic stem cell/Leukemic stem cell-specific requirements.
Validation domain
Key requirement
LSC/HSC-specific considerations
Cohort representativeness
The training cohort must represent the target clinical population with respect to age, disease stage, and treatment era
LSC/HSC models should include balanced representation of ELN risk categories, stem-cell compartment measurements (CD34+CD38- frequencies), and both newly diagnosed and relapsed/refractory patients[52,106-108]
Event counts
An adequate number of outcome events (relapse, death, MRD positivity) to prevent overfitting
Minimum 10-20 events per predictor variable; for LSC-specific endpoints (e.g., LSC+ vs LSC-), ensure sufficient LSC+ cases across validation sets[52,55]
Internal validation
Model performance assessed on held-out data from the same source (cross-validation or hold-out split)
Report performance metrics (AUROC, calibration) separately for LSC-enriched vs LSC-depleted subgroups if the model claims to encode stemness biology[52,109,110]
External validation
Independent cohort from a different institution, time period, or geography
Essential for LSC models given center-to-center variability in LSC phenotyping protocols and MRD detection thresholds[52,106,110]
Prospective evaluation
Forward-looking validation on newly enrolled patients before clinical deployment
Required for LSC-targeted therapy selection models to confirm that AI predictions align with clinical outcomes under prospective conditions[55,109,111]
Dataset shift detection
Assess whether model performance degrades when applied to data with distributional differences (batch effects, assay drift)
Critical for flow cytometry-based LSC models: Validate across different antibody panels, fluorophores, and cytometers; report performance stratified by batch[55,112]
Calibration assessment
Predicted probabilities should match observed event frequencies
For LSC burden models, calibration plots should show agreement between predicted LSC frequency (or surrogate score) and directly measured LSC% by flow cytometry in the calibration subset[52,55,106]
Decision-curve analysis
Net benefit of model-guided decisions compared to treat-all or treat-none strategies
For LSC-directed therapies (venetoclax, Menin inhibitors), decision curves should quantify clinical utility across risk thresholds relevant to treatment intensification decisions[55,111]
Explainability and feature attribution
Use of XAI methods (SHAP, attention weights) to identify which features drive predictions
Essential for validating LSC-AI hypothesis: Determine whether high-risk predictions are driven by known stemness genes (17-gene LSC score, HOXMEIS1 programs) or alternative pathways[52,113]
Bias and fairness evaluation
Assess performance stratified by demographic subgroups and underrepresented populations
Evaluate whether LSC models perform equivalently across age groups (pediatric vs adult vs elderly AML), ancestry, and sex; report subgroup-specific metrics[107]
Missing data handling
Transparent reporting of missingness patterns and imputation strategies
LSC models often integrate multi-omics data with heterogeneous completeness (e.g., scRNA-seq available for a subset); clearly document handling of missing modalities and proteins in CITE-seq[111]
Comparator benchmarking
Performance compared to established clinical risk systems
For AML LSC models, benchmark against ELN 2022 risk classification, 17-gene LSC score, and LSC frequency by flow cytometry; report incremental predictive value[114]
The use of AI in healthcare provokes significant ethical issues, such as patient privacy, informed consent, algorithmic bias, and fair access to AI-based therapies. ML models that are trained on non-representative datasets can have performance differences across demographic groups, which may worsen the current healthcare inequities. To overcome the dual challenges of data siloing and patient privacy across different institutions, Federated learning has emerged as a highly promising solution. Federated learning is a decentralized computational model that enables multi-center AI models to be trained locally on institutional data; only the learned model parameters (weights) are shared and aggregated centrally, not the raw, sensitive patient data. Privacy-preserving methods, such as federated learning and differential privacy, therefore, provide promising solutions to training robust, generalizable AI models on sensitive medical data without violating patient confidentiality. As an example, Zhang et al[115] discussed fairness-conscious and privacy-conscious enhanced collaborative learning in healthcare, showing that heterogeneous federated learning systems can enable healthcare applications without centralizing sensitive data, and address fairness issues to prevent the reinforcement of disparities in AI-driven healthcare outcomes. The paper has highlighted the significance of developing AI systems that deliver fairness to a wide range of patients. Privacy-preserving solutions, such as federated learning and differential privacy, have shown potential solutions to the problem of training AI models on sensitive medical data without violating patient privacy.
Future directions and opportunities
It is didactic to consider these developments in the macro-evolutionary perspective of the greater medical AI landscape. The most recent comprehensive knowledge mapping study has shown that medical AI as a field is in a rapid, universal transition-shifting away from an initial phase of validation of basic ML models towards a more advanced stage dominated by complex, application-specific models designed to be directly put into clinical use[116]. It is along this same path that the evolution of AI in hematological malignancies is following. The future of the field is to build a highly complex biological input into a specialized and actionable tool.
Single-cell multi-omics integration
The combination of single-cell multi-omics technologies with advanced AI architectures is a frontier in understanding HSC biology and hematological disease mechanisms at an unprecedented level of detail. Combined with epigenomic, proteomic, and spatial transcriptomic profiling, scRNA-seq can be used to characterize cellular heterogeneity, lineage trajectories, and microenvironmental interactions. As pointed out by Raghav et al[62], integrative computational methods of genomic and transcriptomic analysis, such as scRNA-seq and ML, can be used to analyze gene regulatory networks at single-cell resolution, unlocking the potential of HSCs. These methods enable the discovery of rare cell groups, transitional states, and regulatory circuits that contribute to disease progression and resistance to therapy.
Foundation models and transfer learning
The emergence of large-scale, pre-trained foundation models, which are analogous to Bidirectional Encoder Representations from Transformers and Generative Pre-trained Transformer in NLP, is an exciting prospect in biomedical AI. Such models, which have been trained on large sets of multi-omics data, can be fine-tuned to perform specific tasks with limited labelled data, addressing a typical drawback of clinical AI applications. Transfer learning methods allow the use of knowledge obtained in well-characterized disease settings to less common conditions with limited data access, which may democratize access to advanced AI capabilities across a wide range of hematological disorders.
Prospective clinical trials
Although retrospective studies have shown that AI models are technically feasible and can perform analytically, prospective clinical trials are needed to establish clinical utility and inform implementation strategies. Randomized controlled trials between AI-enhanced clinical decision-making and conventional care will give conclusive results on the effects of these technologies on patient outcomes. Implementation science frameworks should be included in such trials, and the clinical efficacy is not the only aspect that should be evaluated. The effective translation of AI technologies between research prototypes and routine clinical practice critically depends on the need to address these multifaceted implementation challenges.
Automated closed-loop therapeutic systems
The integration of AI-based diagnostics, monitoring of treatment response, and selecting adaptive therapy are the vision of automated, closed-loop therapeutic systems. These systems would continuously monitor patient condition by using multi-modal data streams, predict treatment response as well as prescribe individualized treatment adjustments in real-time. Although these systems are still aspirational, gradual advancements towards this vision can be found in the development of AI-powered clinical decision support, wearable biosensors, and adaptive clinical trial designs. The development of predictive modeling, real-time data integration, and regulatory frameworks that support dynamic and personalized treatment plans will be required to achieve closed-loop therapeutic systems.
Future opportunities in AI-driven myeloma-niche research
The combination of AI and spatial multi-omics technologies offers unprecedented opportunities to gain insight into myeloma-niche interactions at a level never before seen. Future directions are: (1) Application of ML to spatial transcriptomics data to identify microenvironmental niches that facilitate therapy resistance; (2) Development of AI models that combine spatial cellular architecture with spatial transcriptomics data to identify microenvironmental niches that promote therapy resistance. single-cell multi-omics to predict which patients will respond to niche-targeted therapies; (3) The use of DL to analyze 3D BM models that recapitulate the physiologically relevant hematopoietic niche; and (4) Development of digital twin models of the myeloma microenvironment to simulate therapeutic interventions targeting niche components. With the continued feasibility and maturation of computational tools, AI-driven approaches are poised to transform MM management by enhancing patient stratification based on microenvironmental features, the identification of novel niche-directed therapeutic targets, and optimization of combination therapies that, at the same time, attack both the malignant plasma cells and their supportive microenvironment. The integration of spatial biology, state-of-the-art imaging, and AI will allow a systems-level perspective of myeloma as a disease of both tumor cells and their niche, fundamentally changing the therapeutic approach to myeloma, which has been plasma cell-centric to date.
Decoding stemness - XAI approaches to validate the LSC-AI hypothesis: The potential, systematic research of the question whether AI prognostic models in AML are implicitly capturing the biology of LSC burden and stemness is a critical priority of the field. This research program involves: (1) The use of SHAP-based explainability analyses of trained prognostic AI models in large, well-characterized AML cohorts where LSC frequency has been directly measured by MFC; (2) Correlation of AI feature attribution rankings with validated stemness indices, such as the 17-gene LSC score, broader LSC transcriptional programs in scRNA-seq studies, and epigenetic stemness scores; (3) Testing whether AI risk scores have incremental prognostic information beyond and including directly measured LSC frequency, or whether the predictive power of AI risk scores is significantly mediated by LSC-correlating features; and (4) Prospective trials where AI-derived scores of stemness attribution are utilized to inform selection of an LSC-directed therapy such as venetoclax, Menin inhibitors, or CD123-targeted therapy. Should this research agenda succeed, the result would be a new generation of AI tools that do not merely predict results but actively characterize the biological mechanism driving risk in every patient, making AI less a statistical and more an active characterization of the biological process that drives risk in each patient prognosticator, to more of a precision molecular diagnostician of the LSC.
AI as a computational proxy for leukemic stemness - clinical and methodological implications: The clinical significance of LSC quantification and its current barriers. Terwijn et al[117] demonstrated in a breakthrough study that LSC frequency, enumerated in MFC, is a powerful and independent prognostic biomarker in AML, with high LSC burden predicting significantly shorter overall survival and resistance to induction chemotherapy. This publication, followed by additional validations, suggested elevating the quantification of LSC as a research tool to a parameter that can be acted upon clinically. However, there are several serious pitfalls in practice: Special antibody panels to determine LSC surface markers (CD34, CD38, CD123, CD117, etc.), trained personnel to distinguish between LSC phenotypes and normal hematopoietic stem and progenitor cells and residual normal elements, and standardized procedures that are not always accessible outside academic centers. As a result, LSC burden is practically never measured in standard clinical hematology, and its prognostic value is effectively unavailable at most centers that provide AML treatment. This information can be democratized by AI models trained on standard diagnostic data, such as RNA sequencing, targeted molecular panels, or even morphologic data.
The 17-gene stemness score as a validation scaffold. Bill et al[106] demonstrated that the 17-gene LSC score derived from the expression profiles of functionally validated LSC-enriched (LSC+) vs LSC-depleted (LSC-) cell fractions substantially refines AML prognosis within and across ELN risk categories, independently predicting induction failure, relapse, and survival in multiple cohorts. Kim et al[109] extended this validation to the allogeneic HSCT context, showing that the LSC17 score retained independent prognostic significance for relapse after transplantation, a setting in which immune-mediated disease control is operating, suggesting that stemness captures biological features of the leukemic clone that persist even under allogeneic immune pressure. Crucially, this gene signature was derived not from bulk disease features but from the transcriptome of the rare LSC-enriched fraction itself: It is, by construction, a molecular fingerprint of the therapy-resistant stem-like compartment. AI models that outperform ELN risk classification in predicting the same endpoints (induction failure, relapse and survival) are plausibly learning a higher-dimensional, implicit representation of the same biological information. The alignment between AI model performance and LSC17-predicted risk zones constitutes a directly testable research question.
The methodological mandate - applying XAI to decode stemness in prognostic models: The hypothesis that AI prognostic models are encoding stemness biology is not merely an intellectual observation. It generates a specific methodological research program. Explainability methods, including SHAP and attention-based visualization techniques, allow researchers to decompose a model’s prediction for individual patients into ranked feature contributions. Gimeno et al[27] demonstrated that XAI approaches applied to AML drug sensitivity data yielded clinically interpretable genotype-drug stratification guidelines that could be directly translated into precision treatment algorithms. Janizek et al[29] went further, applying ensemble-based XAI with SHAP attributions to drug response modeling in AML and showing that their explainability framework could directly assess AML “stemness scores” as features contributing to synergistic drug sensitivity predictions, providing direct proof-of-concept that XAI methods can interrogate stemness biology within AI models. Applying comparable XAI analyses to trained AI prognostic models, specifically asking: Do the features with the highest SHAP values in high-risk predictions overlap with the 17-gene LSC signature or broader stemness transcriptional programs? Would provide the first systematic evidence either confirming or refuting the LSC-AI thesis. Silva et al[118] demonstrated that ML models trained on transcriptome k-mer data achieved prognostic performance matching or exceeding ELN classification in both younger and older AML patients, with distinct age-group transcriptomic complexity not captured by genomic risk criteria alone, precisely the scenario in which LSC heterogeneity is predicted to diverge from standard molecular markers. Li and Wang[119] further demonstrated that integrative ML analysis of epigenetic subtypes in AML identified malignant HSC-associated epigenetic score elevations as central drivers of prognosis in their multi-center model, directly linking the epigenetic biology of the HSC-LSC axis to AI-derived risk stratification.
From implicit learning to explicit LSC-guided therapy: The clinical implication of the validation of the LSC-AI thesis is not limited to prognostication but also to treatment selection. Therapeutic strategies directed by LSCs now include multiple approved and investigational agents: Venetoclax, which selectively disrupts oxidative phosphorylation in LSCs; Menin inhibitors (revumenib, ziftamenib) targeting KMT2A-rearranged and NPM1-mutant LSC programs; CD123-targeted antibody-drug conjugates; and CD47 “don’t eat me” signal blockade, which takes advantage of LSC immune evasion. To select patients rationally to take these agents, it is necessary to identify patients with high LSC burden. This very parameter cannot be delivered at scale by the flow cytometry assay. An AI model whose predictions demonstrably encode LSC-enriched transcriptional states could serve as a computational triage layer: Patients whose AI risk attribution is dominated by stemness features would be prioritized for LSC-directed intensification, while patients whose AI risk reflects alternative mechanisms, immune microenvironment failure, clonal evolution, treatment pharmacodynamics, might be better served by different interventional strategies. This therapeutic directionality is absent from current AI clinical decision support tools, which output risk categories without a biological mechanism, and represents a concrete opportunity to translate the LSC-AI thesis into bedside impact.
The unifying insight
The models of the trajectory inferences that define normal HSC lineage commitment, the DL classifiers that predict LSC-associated morphological states in leukemic BM with > 98% accuracy, the multi-omics AI systems that predict therapy resistance with > 90% accuracy, and the manufacturing platforms that optimize HSC expansion to clinical use are not independent technological achievements. They are aspects of one conceptual revolution: AI is systematically encoding the biology of the HSC, and its malignant counterpart, the LSC, at a depth and dimensionality that cannot be approached by human-curated risk systems. When this latent encoding is decoded with XAI decoding and tested against functionally defined stemness indices, AI in hematologic oncology will evolve from a prediction engine into a biological discovery tool, one capable of revealing what LSCs are doing in each patient and how to stop them.
CONCLUSION
AI shows tremendous potential to become an indispensable component of modern HSC research and clinical hematology, transitioning from an experimental adjunct toward an integral element of the research and translational ecosystem. Across normal hematopoiesis and malignancy, ML, DL, and NLP are enabling disease modeling, biomarker discovery, therapy-resistance prediction, drug development, and increasingly standardized cell manufacturing workflows. In particular, multi-omics DL approaches have been reported to outperform conventional risk stratification systems in AML and MM, while AI-driven discovery pipelines and Industry 4.0-style automation (computer vision, digital twins, and closed-loop process control) are accelerating the path from data to actionable decisions in both precision hematology and cell therapy production. Nevertheless, this hope should be mitigated by the fact that there are methodological and translational obstacles which are presently exist. Critical obstacles still exist to routine clinical implementation: Heterogeneity and incompleteness of real-world data, interpretability of models, regulatory compliance, and ethical issues, including algorithmic bias and patient privacy. Although AI prognostic and resistance models with high performance have a great potential in that they implicitly learn LSC-enriched transcriptional and epigenetic states by learning high-dimensional data, and they are mostly computational surrogates to LSC burden quantification at this stage. To overcome these issues, transparent reporting, fit-for-purpose will be needed, validation plans (such as external and prospective evaluation), reproducible pipelines, and governance frameworks that align model development with clinical workflows and decision points. Above all, these AI models are yet to be integrated into the routine clinical practice before large-scale, prospective, multi-center randomized clinical trials can be conducted. In this regard, new reporting and appraisal systems of AI prediction modelling (e.g., TRIPOD + AI and PROBAST + AI)[66,67] offer a convenient framework on how to enhance reliability, interpretability, and generalisability as the field transitions away from retrospective performance to real-world impact. High-performing AI prognostic and resistance models are not just learning statistical correlates of clinical risk, but implicitly learning LSC-enriched transcriptional and epigenetic states with high-dimensional data, functioning as a computational surrogate to LSC burden quantification.
Döhner H, Wei AH, Appelbaum FR, Craddock C, DiNardo CD, Dombret H, Ebert BL, Fenaux P, Godley LA, Hasserjian RP, Larson RA, Levine RL, Miyazaki Y, Niederwieser D, Ossenkoppele G, Röllig C, Sierra J, Stein EM, Tallman MS, Tien HF, Wang J, Wierzbowska A, Löwenberg B. Diagnosis and management of AML in adults: 2022 recommendations from an international expert panel on behalf of the ELN.Blood. 2022;140:1345-1377.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 115][Cited by in RCA: 2420][Article Influence: 605.0][Reference Citation Analysis (0)]
Raudenbush K, Malinov N, Reddy JV, Ding C, Tian H, Ierapetritou M. Towards the Development of Digital Twin for Pharmaceutical Manufacturing.Syst Control Trans. 2024;3:67-74.
[PubMed] [DOI] [Full Text]
Trac QT, Pawitan Y, Mou T, Erkers T, Östling P, Bohlin A, Österroos A, Vesterlund M, Jafari R, Siavelis I, Bäckvall H, Kiviluoto S, Orre LM, Rantalainen M, Lehtiö J, Lehmann S, Kallioniemi O, Vu TN. Prediction model for drug response of acute myeloid leukemia patients.NPJ Precis Oncol. 2023;7:32.
[RCA] [PubMed] [DOI] [Full Text][Cited by in RCA: 16][Reference Citation Analysis (0)]
Jum'ah HA, Otteson GE, Timm MM, Weybright MJ, Shi M, Horna P, Jevremovic D, Reichard KK, Olteanu H. Measurable Residual Disease Analysis by Flow Cytometry: Assay Validation and Characterization of 385 Consecutive Cases of Acute Myeloid Leukemia.Cancers (Basel). 2025;17:1155.
[RCA] [PubMed] [DOI] [Full Text][Cited by in RCA: 1][Reference Citation Analysis (0)]
Chen Y, He L, Ianevski A, Nader K, Ruokoranta T, Linnavirta N, Miettinen JJ, Vähä-Koskela M, Vänttinen I, Kuusanmäki H, Kontro M, Porkka K, Wennerberg K, Heckman CA, Giri AK, Aittokallio T. A Machine Learning-Based Strategy Predicts Selective and Synergistic Drug Combinations for Relapsed Acute Myeloid Leukemia.Cancer Res. 2025;85:2753-2768.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 13][Cited by in RCA: 9][Article Influence: 9.0][Reference Citation Analysis (0)]
Ferchen K, Zhang X, Thakkar K, Li G, Bernardicius D, Sen S, Rawat P, Olsson A, Bennett SN, Potter C, Finkelman FD, Croteau J, Morris S, Singh H, Salomonis N, Grimes HL. A unified multimodal single-cell framework reveals a discrete state model of hematopoiesis in mice.Nat Immunol. 2025;26:2086-2099.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in RCA: 4][Reference Citation Analysis (0)]
Hagos YB, Lecat CSY, Patel D, Mikolajczak A, Castillo SP, Lyon EJ, Foster K, Tran TA, Lee LSH, Rodriguez-Justo M, Yong KL, Yuan Y. Deep Learning Enables Spatial Mapping of the Mosaic Microenvironment of Myeloma Bone Marrow Trephine Biopsies.Cancer Res. 2024;84:493-508.
[RCA] [PubMed] [DOI] [Full Text][Cited by in RCA: 8][Reference Citation Analysis (0)]
Ledergor G, Weiner A, Zada M, Wang SY, Cohen YC, Gatt ME, Snir N, Magen H, Koren-Michowitz M, Herzog-Tzarfati K, Keren-Shaul H, Bornstein C, Rotkopf R, Yofe I, David E, Yellapantula V, Kay S, Salai M, Ben Yehuda D, Nagler A, Shvidel L, Orr-Urtreger A, Halpern KB, Itzkovitz S, Landgren O, San-Miguel J, Paiva B, Keats JJ, Papaemmanuil E, Avivi I, Barbash GI, Tanay A, Amit I. Single cell dissection of plasma cell heterogeneity in symptomatic and asymptomatic myeloma.Nat Med. 2018;24:1867-1876.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 126][Cited by in RCA: 205][Article Influence: 25.6][Reference Citation Analysis (3)]
Cohen YC, Zada M, Wang SY, Bornstein C, David E, Moshe A, Li B, Shlomi-Loubaton S, Gatt ME, Gur C, Lavi N, Ganzel C, Luttwak E, Chubar E, Rouvio O, Vaxman I, Pasvolsky O, Ballan M, Tadmor T, Nemets A, Jarchowcky-Dolberg O, Shvetz O, Laiba M, Shpilberg O, Dally N, Avivi I, Weiner A, Amit I. Identification of resistance pathways and therapeutic targets in relapsed multiple myeloma patients through single-cell sequencing.Nat Med. 2021;27:491-503.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 208][Cited by in RCA: 192][Article Influence: 38.4][Reference Citation Analysis (4)]
Leung CK, Zhu P, Loke I, Tang KF, Leung HC, Yeung CF. Development of a quantitative prediction algorithm for human cord blood-derived CD34(+) hematopoietic stem-progenitor cells using parametric and non-parametric machine learning models.Sci Rep. 2024;14:25085.
[RCA] [PubMed] [DOI] [Full Text][Cited by in RCA: 1][Reference Citation Analysis (0)]
Qi J, Chen Y, Jin X, Wang R, Wang N, Yan J, Huang C, Huang J, Wei Y, Xie F, Yu Z, Huang D. Predicting apheresis yield and factors affecting peripheral blood stem cell harvesting using a machine learning model.J Int Med Res. 2024;52:3000605241305360.
[RCA] [PubMed] [DOI] [Full Text][Cited by in RCA: 3][Reference Citation Analysis (0)]
Jo T, Arai Y, Kanda J, Kondo T, Ikegame K, Uchida N, Doki N, Fukuda T, Ozawa Y, Tanaka M, Ara T, Kuriyama T, Katayama Y, Kawakita T, Kanda Y, Onizuka M, Ichinohe T, Atsuta Y, Terakura S. A convolutional neural network-based model that predicts acute graft-versus-host disease after allogeneic hematopoietic stem cell transplantation.Commun Med (Lond). 2023;3:67.
[RCA] [PubMed] [DOI] [Full Text][Cited by in RCA: 8][Reference Citation Analysis (0)]
Fidanza A, Stumpf PS, Ramachandran P, Tamagno S, Babtie A, Lopez-Yrigoyen M, Taylor AH, Easterbrook J, Henderson BEP, Axton R, Henderson NC, Medvinsky A, Ottersbach K, Romanò N, Forrester LM. Single-cell analyses and machine learning define hematopoietic progenitor and HSC-like cells derived from human PSCs.Blood. 2020;136:2893-2904.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 49][Cited by in RCA: 47][Article Influence: 7.8][Reference Citation Analysis (0)]
Cui X, Dong Y, Zhan Q, Huang Y, Zhu Q, Zhang Z, Yang G, Wang L, Shen S, Zhao J, Lin Z, Sun J, Su Z, Xiao Y, Zhang C, Liang Y, Shen L, Ji L, Zhang X, Yin J, Wang H, Chen Z, Ju Z, Jiang C, Le R, Gao S. Altered 3D genome reorganization mediates precocious myeloid differentiation of aged hematopoietic stem cells in inflammation.Sci China Life Sci. 2025;68:1209-1225.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 3][Cited by in RCA: 5][Article Influence: 5.0][Reference Citation Analysis (0)]
Bruns I, Cadeddu RP, Brueckmann I, Fröbel J, Geyh S, Büst S, Fischer JC, Roels F, Wilk CM, Schildberg FA, Hünerlitürkoglu AN, Zilkens C, Jäger M, Steidl U, Zohren F, Fenk R, Kobbe G, Brors B, Czibere A, Schroeder T, Trumpp A, Haas R. Multiple myeloma-related deregulation of bone marrow-derived CD34(+) hematopoietic stem and progenitor cells.Blood. 2012;120:2620-2630.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 58][Cited by in RCA: 84][Article Influence: 6.0][Reference Citation Analysis (2)]
Silva-Sousa T, Nakanishi Usuda J, Al-Arawe N, Hinterseher I, Catar R, Luecht C, Vallecillo Garcia P, Riesner K, Hackel A, Schimke LF, Dutra Dias H, Salerno Filgueiras I, Nakaya HI, Camara NOS, Fischer S, Riemekasten G, Ringdén O, Penack O, Winkler T, Duda G, Fonseca DLM, Cabral-Marques O, Moll G. Artificial intelligence and systems biology analysis in stem cell research and therapeutics development.Stem Cells Transl Med. 2025;14:szaf037.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 6][Cited by in RCA: 7][Article Influence: 7.0][Reference Citation Analysis (7)]
Serrano DR, Luciano FC, Anaya BJ, Ongoren B, Kara A, Molina G, Ramirez BI, Sánchez-Guirales SA, Simon JA, Tomietto G, Rapti C, Ruiz HK, Rawat S, Kumar D, Lalatsa A. Artificial Intelligence (AI) Applications in Drug Discovery and Drug Delivery: Revolutionizing Personalized Medicine.Pharmaceutics. 2024;16:1328.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in RCA: 196][Reference Citation Analysis (0)]
Li K, Du Y, Cai Y, Liu W, Lv Y, Huang B, Zhang L, Wang Z, Liu P, Sun Q, Li N, Zhu M, Bosco B, Li L, Wu W, Wu L, Li J, Wang Q, Hong M, Qian S. Single-cell analysis reveals the chemotherapy-induced cellular reprogramming and novel therapeutic targets in relapsed/refractory acute myeloid leukemia.Leukemia. 2023;37:308-325.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 1][Cited by in RCA: 83][Article Influence: 27.7][Reference Citation Analysis (0)]
Lin X, Li K, Wu C, Zhang C, Zhang G, Huo X. Advances in modeling analysis for multi-parameter bioreactor process control.Biotechnol Bioproc E. 2025;30:235-261.
[PubMed] [DOI] [Full Text]
Kanwar B, Wang B, Roy K, Mazumdar A, Balakirsky S.
Digital Twin Design for hMSC Expansion in Hollow-fiber Bioreactors. Proceedings of 2023 American Control Conference (ACC); 2023 May 31-June 2; San Diego, CA, United States. United States: IEEE, 2003.
[PubMed] [DOI] [Full Text]
Colas S, Gilles V, Chantepie S, Daguindau E, Huynh A, Loschi M, Robin M, Raus N, Requillard C, Chuttoo L, Poplu A, Allali N, Kiprijanovski D, Leouay F, Jeanbat V, Kirion J, Cottin J, Buchbinder N, François S, Villate A, Bouee S. P47 Evaluation of Natural Language Processing (NLP) on Electronic Medical Records: A Proof of Concept on Chronic Graft Versus Host Disease (cGVHD) in France.Value Health. 2024;27:S11-S12.
[PubMed] [DOI] [Full Text]
Hanemaaijer ES, Müskens KF, Kal IJ, Chen LT, Te Pas BM, Fryzik P, Epskamp N, van der Meulen M, Balwierz AK, Saikumar Jayalatha AK, Scheijde-Vermeulen M, Heidenreich O, Candelli T, de Jonge WJ, Margaritis T, Belderbos ME. Single-cell multiomic atlas of healthy pediatric bone marrow reveals age-dependent differences in lineage differentiation driven by stromal signaling.Nat Immunol. 2026;27:613-623.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 3][Cited by in RCA: 3][Article Influence: 3.0][Reference Citation Analysis (0)]
Bill M, Nicolet D, Kohlschmidt J, Walker CJ, Mrózek K, Eisfeld AK, Papaioannou D, Rong-Mullins X, Brannan Z, Kolitz JE, Powell BL, Archer KJ, Dorrance AM, Carroll AJ, Stone RM, Byrd JC, Garzon R, Bloomfield CD. Mutations associated with a 17-gene leukemia stem cell score and the score's prognostic relevance in the context of the European LeukemiaNet classification of acute myeloid leukemia.Haematologica. 2020;105:721-729.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 12][Cited by in RCA: 25][Article Influence: 3.6][Reference Citation Analysis (0)]
Lachowiez CA, Long N, Saultz J, Gandhi A, Newell LF, Hayes-Lattin B, Maziarz RT, Leonard J, Bottomly D, McWeeney S, Dunlap J, Press R, Meyers G, Swords R, Cook RJ, Tyner JW, Druker BJ, Traer E. Comparison and validation of the 2022 European LeukemiaNet guidelines in acute myeloid leukemia.Blood Adv. 2023;7:1899-1909.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 84][Cited by in RCA: 77][Article Influence: 25.7][Reference Citation Analysis (0)]
Kim DDH, Novitzky Basso I, Kim TS, Yi SY, Kim KH, Murphy T, Chan S, Minden M, Pasic I, Lam W, Law A, Michelis FV, Gerbitz A, Viswabandya A, Lipton J, Kumar R, Ng SWK, Stockley T, Zhang T, King I, Mattsson J, Wang JCY. The 17-gene stemness score associates with relapse risk and long-term outcomes following allogeneic haematopoietic cell transplantation in acute myeloid leukaemia.EJHaem. 2022;3:873-884.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 2][Cited by in RCA: 9][Article Influence: 2.3][Reference Citation Analysis (0)]
Ng SWK, Murphy T, King I, Zhang T, Mah M, Lu Z, Stickle N, Ibrahimova N, Arruda A, Mitchell A, Mai M, He R, Madala BS, Viswanatha DS, Dick JE, Chan S, Virtanen C, Minden MD, Mercer T, Stockley T, Wang JCY. A clinical laboratory-developed LSC17 stemness score assay for rapid risk assessment of patients with acute myeloid leukemia.Blood Adv. 2022;6:1064-1073.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 5][Cited by in RCA: 21][Article Influence: 4.2][Reference Citation Analysis (0)]