Copyright: ©Author(s) 2026.
Artif Intell Cancer. Sep 8, 2026; 7(1): 114273
Published online Sep 8, 2026. doi: 10.35713/aic.v7.i1.114273
Published online Sep 8, 2026. doi: 10.35713/aic.v7.i1.114273
Table 1 Diagnostic performance of artificial intelligence systems for gastric cancer and precancerous lesions
| No. | Research focus | Ref. | Modality | Study design | Sample size/dataset | AI model/system | Gold standard | Key performance metrics | Results | Comparison with endoscopists |
| 1 | EGC diagnosis | Chen et al[20] | WLE, NBI | Systematic review (12 studies) | 11685 cases | Various | Pathology | Pooled sensitivity, specificity, AUC | Sensitivity: 0.86 (95%CI: 0.75-0.92); specificity: 0.90 (95%CI: 0.84-0.93); AUC: 0.94 | Not compared |
| 2 | Upper GI tumors | Arribas et al[21] | NBI | Meta-analysis (19 studies) | Not specified | Various | Pathology | Overall sensitivity, specificity, AUC | Sensitivity: 90%; specificity: 89%; AUC: 0.95 | Not compared |
| 3 | Gastric neoplasia | Lui et al[22] | WLE, NBI | Systematic review and meta-analysis (23 studies) | 969318 images | Various | Pathology | AUC | AUC: 0.96 | Superior (AUC 0.98 vs 0.87, P < 0.001) |
| 4 | GPLs diagnosis | Dilaghi et al[23] | Not Specified | Systematic review and meta-analysis (4 studies) | Not specified | Various | Pathology | Accuracy | Accuracy: 90.3% | Not compared |
| 5 | CAG diagnosis | Shi et al[24] | Not specified | Systematic review and meta-analysis (8 studies) | 25216 patients, > 90000 images | Various | Pathology | Pooled sensitivity, specificity, AUC | Sensitivity: 94%; specificity: 96%; AUC: 0.98 | Significantly higher accuracy |
| 6 | CAG diagnosis | Zhang et al[25] | Not specified | Diagnostic study | 5470 antral images | CNN | Pathology | Accuracy, sensitivity, specificity | Accuracy: 0.942; sensitivity: 0.945; specificity: 0.940 | Exceeded three experts |
| 7 | CAG diagnosis | Shi et al[26] | Not specified | Diagnostic study | Not specified | GAM-efficient net | Pathology | Accuracy | External image: 93.5%; video: 92.37% | Outperformed endoscopists |
| 8 | GIM diagnosis | Yan et al[27] | Not specified | Diagnostic study | Not specified | Intelligent diagnostic system | Pathology | AUC, sensitivity, specificity, accuracy | AUC: 0.928; sensitivity: 91.9%; specificity: 86.0%; Accuracy: 88.8% | Not compared |
| 9 | Mucosal lesion DDx | Nam et al[28] | Not specified | Diagnostic study | Not specified | AI-DDx | Pathology | AUROC | AUROC: 0.86 | Comparable to experts (0.89, P = 0.12); Superior to novices and intermediates |
| 10 | CAG and IM diagnosis | Lin et al[29] | WLE | Multicenter diagnostic study | 7037 images (14 hospitals) | CNN | Pathology | AUC, Accuracy | CAG: AUC 0.98, Acc 96.4%; IM: AUC 0.99, Acc 97.6% | Not compared |
| 11 | Atrophy and IM detection | Yang et al[31] | WLE, LCI | Diagnostic study | 21420 images | Novel DL method | Pathology | Accuracy | Atrophy: 97.12%; IM: 99.18% | Not compared |
| 12 | GA and IM diagnosis | Xu et al[32] | Image-enhanced endoscopy | Multicenter diagnostic study | 6250 images, 98 videos (5 hospitals) | ENDOANGEL (DCNN) | Pathology | Accuracy | GA: 86.4%; IM: 85.9% | Comparable to experts; Superior to non-experts |
| 13 | Precursor detection | Xu et al[33] | Not specified | Prospective single-center clinical trial | Not specified | Not specified | Pathology | Detection Rate | IM: 14.23% vs 9.15%; atrophy: 22.76% vs 17.28% | Effect more pronounced in junior physicians |
| 14 | Invasion depth | Nam et al[28] | EUS | Diagnostic study | Not specified | AI-ID | Post-operative histology | AUROC | AUROC: 0.73 | Superior to EUS experts (0.56, P < 0.001) |
| 15 | Cancer vs ulcer | Namikawa et al[36] | Not specified | Diagnostic study | Not specified | A-CNN | Pathology | Sensitivity, specificity, accuracy | Sensitivity: 99.0%; specificity: 93.3%; accuracy: 95.9% | Not compared |
| 16 | GISTs vs leiomyomas | Dong et al[37] | EUS | Multicenter diagnostic study | Not specified | Real-time AI-assisted EUS system | Pathology/histology | AUC, accuracy | AUC: 0.948; accuracy: 91.7% | Significantly outperformed |
| 17 | Pathology analysis | Yang et al[40] | Macroscopic specimen | Diagnostic study | Gastric cancer surgery specimens | AI algorithm | Histology | mAP, Accuracy | Lesion localization mAP: 95.90%; LN metastasis prediction Acc: 75.00% | Not compared |
| 18 | IM scoring | Iwaya et al[41] | Pathology slides | Diagnostic study | Not specified | AI system | Expert pathologist | Scoring difference | Difference with pathologists: 7.6% | Identified missed foci by pathologists |
Table 2 Role of artificial intelligence in reducing inter-observer variability and improving diagnostic consistency
| No. | Research focus | Ref. | Modality | Study design | Sample size/dataset | AI model/system | Key performance metrics | Results without AI | Results with AI assistance | Key finding |
| 1 | EGC diagnosis with ME | Li et al[42] | Magnifying image-enhanced endoscopy | Diagnostic study | Not specified | ENDOANGEL-LA | Accuracy | Novice: 71.63% | Novice: 87.45% | AI assistance bridged the gap between novices and experts |
| 2 | Real-time AI assistance | Dong et al[43] | WLE | Diagnostic study | Not specified | ENDOANGEL-ED | Accuracy | Endoscopists: 70.61% | Endoscopists: 79.63% (P < 0.001) | AI significantly improved endoscopist diagnostic accuracy |
| 3 | CAG diagnosis | Zhao et al[45] | Not specified | Prospective nested case-control | 1306 patients | Not specified | Kappa, accuracy, sensitivity, specificity | Endoscopists' kappa: 0.291; Acc: 68.89%; Sens: 67.56%; Spec: 70.23% | AI kappa: 0.816; Acc: 89.89%; Sens: 89.31%; Spec: 90.46% | AI agreement with pathology was substantially higher |
| 4 | Pathology Dx (atrophy/IM) | Fang et al[46] | Pathology slides | Observer study (10 pathologists) | Not specified | GasMIL (SDL algorithm) | AUC, weighted kappa | Pathologists’ performance (baseline) | Pathologists’ performance significantly improved | AI assistance improved pathologists’ diagnostic metrics |
Table 3 Major limitations and challenges in the clinical integration of artificial intelligence for diagnosing gastric precancerous lesions
| Category | Specific challenge/issue | Supporting evidence/explanation | Ref. |
| Generalizability and robustness | Performance degradation across centers | Most studies are single-center, leading to models that may not perform well on data from different hospitals, equipment, or populations | [47] |
| A 2023 external validation study in Singapore found endoscopists’ average accuracy was superior to the AI system, highlighting generalizability issues | [48] | ||
| Clinical validation | Lack of high-level evidence | The majority of studies are retrospective with inherent limitations (selection bias, overfitting). Large-scale, prospective multicenter RCTs are needed | [9] |
| Scarcity of RCTs in gastric cancer | A 2025 systematic review found only 11% of GI oncology AI RCTs focused on gastric cancer | [49] | |
| Unclear impact on patient outcomes | The feasibility, effectiveness, safety, and long-term impact on patient outcomes require validation | [34,35] | |
| Inconsistent conclusions on utility | Studies comparing AI against endoscopists of varying experience levels have yielded inconsistent results | [48] | |
| Data and algorithms | High heterogeneity | High heterogeneity in algorithms, imaging techniques, and study designs makes comparing models difficult | [50] |
| Lack of standardized, public datasets | Absence of public, standardized large-scale databases (e.g., “EndoNet”) introduces selection bias, limits reproducibility, and hinders fair comparison | [1,9,13] | |
| Suboptimal training data ratios | The positive-to-negative sample ratio in training sets influences performance, with a suggested optimal ratio between 1:1 and 1:2 | [51] | |
| Potential for missed diagnosis | Models trained primarily on typical lesions may miss rare, atypical, or minute early lesions | [41] | |
| Interpretability and trust | “Black box” problem | The decision-making process of many AI models lacks transparency, hindering clinician trust, acceptance, and error troubleshooting | [1] |
| Although techniques like heatmaps can help, the underlying logic remains insufficiently transparent, posing an obstacle to adoption | [47,52] | ||
| Regulatory and workflow | Underdeveloped regulatory framework | Lack of unified data standards, reliable evaluation systems, and comprehensive regulatory/ethical frameworks | [13] |
| Undefined human-AI collaboration | The specific role of AI (e.g., primary screening, second opinion) and how to optimize human-computer interaction require further exploration | [48] |
Table 4 Proposed future research directions for artificial intelligence in diagnosing gastric precancerous lesions
| Research direction | Specific goals/actions | Expected outcomes/rationale | Ref. |
| High-quality clinical trials | Conduct large-scale, multicenter, prospective RCTs | Evaluate real-world effectiveness, safety, generalizability, cost-effectiveness, and impact on patient outcomes to provide high-level evidence for clinical adoption | [34,49] |
| Validate clinical utility in risk stratification and prediction | Bridge the gap between basic research and clinical application, strengthening translational research | [53] | |
| Data standardization and infrastructure | Create public, standardized, large-scale image databases | Contain diverse images from different regions, equipment, and pathologies. Enhance research transparency, reproducibility, and facilitate fair algorithm comparison and improvement | [1] |
| Develop international, multicenter, annotated databases (e.g., “EndoNet”) | Make high-quality data accessible to the research community to overcome a major current limitation | ||
| Establish guidelines for data acquisition and model validation | Improve the reproducibility and reliability of AI research through academia-industry collaboration | [53,56] | |
| Targeted algorithm development | Enhance detection of atypical/minute lesions | Employ augmented datasets enriched with rare cases for training to reduce miss rates and improve robustness | |
| Establish rigorous evaluation frameworks and ethical guidelines | Ensure model safety, reproducibility, and validate clinical translation potential | [54,55] | |
| Comparative and optimization studies | Head-to-head comparison of AI algorithms and imaging techniques | Identify the optimal AI technical solution for specific clinical scenarios (e.g., screening vs depth assessment) | [20] |
| Explore optimal human-AI collaboration models | Compare AI performance against endoscopists of varying experience; explore modes like real-time assistance, second reader, or quality control monitor to clarify AI’s best clinical role | [48] | |
| Optimize human-computer interaction interface and workflow | Maximize diagnostic efficiency and accuracy in clinical practice | ||
| Explainable AI | Develop interpretable AI models using heatmaps, attention mechanisms, etc. | Visualize the diagnostic basis of models, make the decision-making process transparent to clinicians, and enhance trust and efficiency in human-AI collaboration | [43,47,57] |
| Multimodal data fusion | Integrate endoscopic images, pathology, genomics, and clinical data | Build more comprehensive, robust, and accurate fused AI models for diagnosis, risk stratification, and prognosis prediction, enabling personalized medicine | [14,47] |
Table 5 Major limitations and challenges of organoid models in gastric precancerous lesion research
| Category | Specific challenge/issue | Supporting evidence/explanation | Ref. |
| Research focus | Relative scarcity of precancerous lesion models | Current studies predominantly focus on advanced gastric cancers, with limited research on constructing organoid models for atrophic gastritis and intestinal metaplasia | [17] |
| Model fidelity and complexity | Lack of tumor microenvironment (TME) components | Existing models primarily consist of epithelial cells and lack critical TME components (immune cells, stromal cells, intratumoral microbiota), unable to fully replicate essential interactions (e.g., with H. pylori) | [16-18,54,64,65] |
| Inability to recapitulate systemic physiology | Constraints in fully recapitulating vascular systems, innervation, and interactions with systemic physiological processes | [17] | |
| Standardization and reproducibility | Lack of standardized protocols | Significant challenges exist in organoid culture methodologies, analytical techniques, and data interpretation, compromising reproducibility and reliability | [17,67] |
| Model qualification gap | Absent standardized qualification processes undermine confidence in models' physiological relevance | [55] | |
| Scalability and practicality | Challenges in scalability and cost | Limitations in scalability, reproducibility, cost-effectiveness, and time efficiency. Relatively long culture cycle, variable success rates, batch-to-batch variations, and high costs limit large-scale application | [54,55,68-70] |
| Model validation and comparison | Unclear representativeness | Whether organoids fully represent all characteristics of the original lesional tissue requires further validation. Inconsistent culture success rates and extended cycles are current drawbacks | [15] |
| Lack of comparative studies | Notable absence of head-to-head studies comparing organoids against more established models (e.g., animal models, ALI models) to clarify their unique advantages and optimal applications | [6,67] | |
| Clinical translation | Limited direct clinical evidence | Organoids are primarily used in basic research. Direct evidence for application in clinical diagnostics (e.g., predicting progression risk) remains scarce, and the technology remains distant from direct clinical implementation | [17,61,62,67] |
Table 6 Proposed future research directions for organoid models in gastric precancerous lesion research
| Research direction | Specific goals/actions | Expected outcomes/rationale | Ref. |
| Model development | Establish precancerous lesion organoid biobanks | Develop organoid models from patient tissues (e.g., with IM) and validate their ability to simulate malignant transformation in vitro, enabling study of key molecular events and driver genes | [17] |
| Enhanced complexity (co-culture) | Develop complex co-culture systems | Create co-culture organoid or “organoid-on-a-chip” systems incorporating vascular networks, immune cells (T cells, macrophages), stromal cells (fibroblasts), and H. pylori | [54,64,65] |
| Utilize microfluidic technology for “tumor-on-a-chip” models to accurately mimic the complex TME for studying immune escape and prevention | |||
| Employ 3D bioprinting to construct precise TMEs that better simulate in vivo responses to drugs, particularly immunotherapies | |||
| Standardization and biobanking | Establish SOPs and quality control systems | Promote the development of biobanks with detailed clinical/pathological information. Formulate standardized SOPs for culture, qualification, functional analysis, and data handling to enhance consistency and comparability | [54,55,68] |
| High-throughput screening (HTS) | Develop automated HTS platforms | Enhance the efficiency of organoid culture and drug testing through automation, microfluidics, acoustic manipulation, and high-content imaging | [70-72] |
| Reduce costs | Explore low-cost alternative materials (e.g., synthetic hydrogels) to replace Matrigel, reducing the economic burden for large-scale screening and promoting translation | [70,72] | |
| Comparative studies | Conduct multi-model comparison studies | Directly compare organoids vs GEMMs, chemical animal models, and ALI models in simulating the “Correa cascade” to determine the best model for specific research objectives and clarify organoid applications | [6,67] |
| Predictive diagnostics | Explore application in risk stratification | Use organoids from patients with different risk grades (e.g., LGIN vs HGIN) integrated with scRNA-seq to identify molecular features predictive of progression, enabling personalized risk assessment |
Table 7 Challenges and future directions for the integration of artificial intelligence and organoids in gastric precancerous lesion research
| Category | Specific challenge/opportunity | Explanation/proposed Action | Expected outcome/goal | Ref. |
| Current challenges | ||||
| Research gap | Lack of integrated studies | Almost no studies combine AI and organoid technologies. A complete lack of empirical data on their integrated application specifically to GPLs | N/A | [1] |
| Data dependency and validation | Limited training data | AI model performance depends on high-quality, large-scale data, yet public organoid pharmacogenomic data linked to clinical outcomes remain limited | N/A | [69] |
| Data standardization | Lack of unified standards | Effective integration requires standardized processes for data acquisition, processing, and annotation (imaging and organoid data). Currently, such standards are lacking | N/A | [1,13,73] |
| Data integration | Handling data heterogeneity | Integrating multi-omics and multi-dimensional data from different platforms/batches presents technical difficulties due to heterogeneity, sparsity, and high dimensionality | N/A | [53,76,77] |
| Clinical translation | Translational gap | Seamlessly integrating laboratory findings from AI and organoids into clinical workflows to improve patient outcomes remains a key challenge | N/A | [53,56,75] |
| Future directions | ||||
| Direct integration | AI for organoid dynamic monitoring | Initiate research applying AI for real-time, label-free quantitative analysis of organoid morphological changes, proliferation, and differentiation | To high-throughput screen for factors or drugs influencing precancerous lesion progression | [38] |
| Prospective studies | Integrated clinical research projects | Design prospective, multicenter studies to collect imaging data and tissue samples. Build PDOs for screening/simulation, then use AI to integrate imaging and organoid data | To validate the value of synergy in risk prediction and intervention decision-making | |
| Database construction | Multicenter standardized database | Establish a large-scale database containing clinical info, endo/pathology images, multi-omics, organoid culture/analysis, drug response, and outcome data | To provide a high-quality data foundation for developing and validating generalizable AI models | [1,69,91] |
| Targeted translational studies | Address specific unmet clinical needs | Use CRISPR-Cas9 in organoids to simulate mutations (e.g., abnormal folate metabolism, high RAMP1). Use AI-high-throughput platforms to screen/optimize targeted drugs | To validate efficacy for personalized treatment strategies | |
| Algorithm development | Advanced data integration algorithms | Develop AI algorithms capable of handling data heterogeneity, imputing sparsity, and integrating multimodal data (imaging, spatial omics, scRNA-seq) | To build more comprehensive disease models and achieve better integration with organoid co-culture systems | [75,76] |
| Platform establishment | AI-driven HTS organoid analysis platform | Develop a dedicated AI image analysis platform for automated, quantitative dynamic monitoring of organoid growth, differentiation, and death | To enable high-throughput interpretation of drug screening results and accelerate translation from bench to bedside |
- Citation: Wu CH, Qiu JX, Jia YB, Quan Y, Liu C, Ling JH. Synergistic applications of artificial intelligence and organoid technology in gastric precancerous lesion research: Mechanisms, translation, and challenges. Artif Intell Cancer 2026; 7(1): 114273
- URL: https://www.wjgnet.com/2644-3228/full/v7/i1/114273.htm
- DOI: https://dx.doi.org/10.35713/aic.v7.i1.114273