BPG is committed to discovery and dissemination of knowledge
Review
Copyright: ©Author(s) 2026.
Artif Intell Cancer. Sep 8, 2026; 7(1): 114273
Published online Sep 8, 2026. doi: 10.35713/aic.v7.i1.114273
Table 3 Major limitations and challenges in the clinical integration of artificial intelligence for diagnosing gastric precancerous lesions
Category
Specific challenge/issue
Supporting evidence/explanation
Ref.
Generalizability and robustnessPerformance degradation across centersMost studies are single-center, leading to models that may not perform well on data from different hospitals, equipment, or populations[47]
A 2023 external validation study in Singapore found endoscopists’ average accuracy was superior to the AI system, highlighting generalizability issues[48]
Clinical validationLack of high-level evidenceThe majority of studies are retrospective with inherent limitations (selection bias, overfitting). Large-scale, prospective multicenter RCTs are needed[9]
Scarcity of RCTs in gastric cancerA 2025 systematic review found only 11% of GI oncology AI RCTs focused on gastric cancer[49]
Unclear impact on patient outcomesThe feasibility, effectiveness, safety, and long-term impact on patient outcomes require validation[34,35]
Inconsistent conclusions on utilityStudies comparing AI against endoscopists of varying experience levels have yielded inconsistent results[48]
Data and algorithmsHigh heterogeneityHigh heterogeneity in algorithms, imaging techniques, and study designs makes comparing models difficult[50]
Lack of standardized, public datasetsAbsence of public, standardized large-scale databases (e.g., “EndoNet”) introduces selection bias, limits reproducibility, and hinders fair comparison[1,9,13]
Suboptimal training data ratiosThe positive-to-negative sample ratio in training sets influences performance, with a suggested optimal ratio between 1:1 and 1:2[51]
Potential for missed diagnosisModels trained primarily on typical lesions may miss rare, atypical, or minute early lesions[41]
Interpretability and trust“Black box” problemThe decision-making process of many AI models lacks transparency, hindering clinician trust, acceptance, and error troubleshooting[1]
Although techniques like heatmaps can help, the underlying logic remains insufficiently transparent, posing an obstacle to adoption[47,52]
Regulatory and workflowUnderdeveloped regulatory frameworkLack of unified data standards, reliable evaluation systems, and comprehensive regulatory/ethical frameworks[13]
Undefined human-AI collaborationThe specific role of AI (e.g., primary screening, second opinion) and how to optimize human-computer interaction require further exploration[48]


Write to the Help Desk