Copyright: ©Author(s) 2026.
Artif Intell Cancer. Sep 8, 2026; 7(1): 114273
Published online Sep 8, 2026. doi: 10.35713/aic.v7.i1.114273
Published online Sep 8, 2026. doi: 10.35713/aic.v7.i1.114273
Table 3 Major limitations and challenges in the clinical integration of artificial intelligence for diagnosing gastric precancerous lesions
| Category | Specific challenge/issue | Supporting evidence/explanation | Ref. |
| Generalizability and robustness | Performance degradation across centers | Most studies are single-center, leading to models that may not perform well on data from different hospitals, equipment, or populations | [47] |
| A 2023 external validation study in Singapore found endoscopists’ average accuracy was superior to the AI system, highlighting generalizability issues | [48] | ||
| Clinical validation | Lack of high-level evidence | The majority of studies are retrospective with inherent limitations (selection bias, overfitting). Large-scale, prospective multicenter RCTs are needed | [9] |
| Scarcity of RCTs in gastric cancer | A 2025 systematic review found only 11% of GI oncology AI RCTs focused on gastric cancer | [49] | |
| Unclear impact on patient outcomes | The feasibility, effectiveness, safety, and long-term impact on patient outcomes require validation | [34,35] | |
| Inconsistent conclusions on utility | Studies comparing AI against endoscopists of varying experience levels have yielded inconsistent results | [48] | |
| Data and algorithms | High heterogeneity | High heterogeneity in algorithms, imaging techniques, and study designs makes comparing models difficult | [50] |
| Lack of standardized, public datasets | Absence of public, standardized large-scale databases (e.g., “EndoNet”) introduces selection bias, limits reproducibility, and hinders fair comparison | [1,9,13] | |
| Suboptimal training data ratios | The positive-to-negative sample ratio in training sets influences performance, with a suggested optimal ratio between 1:1 and 1:2 | [51] | |
| Potential for missed diagnosis | Models trained primarily on typical lesions may miss rare, atypical, or minute early lesions | [41] | |
| Interpretability and trust | “Black box” problem | The decision-making process of many AI models lacks transparency, hindering clinician trust, acceptance, and error troubleshooting | [1] |
| Although techniques like heatmaps can help, the underlying logic remains insufficiently transparent, posing an obstacle to adoption | [47,52] | ||
| Regulatory and workflow | Underdeveloped regulatory framework | Lack of unified data standards, reliable evaluation systems, and comprehensive regulatory/ethical frameworks | [13] |
| Undefined human-AI collaboration | The specific role of AI (e.g., primary screening, second opinion) and how to optimize human-computer interaction require further exploration | [48] |
- Citation: Wu CH, Qiu JX, Jia YB, Quan Y, Liu C, Ling JH. Synergistic applications of artificial intelligence and organoid technology in gastric precancerous lesion research: Mechanisms, translation, and challenges. Artif Intell Cancer 2026; 7(1): 114273
- URL: https://www.wjgnet.com/2644-3228/full/v7/i1/114273.htm
- DOI: https://dx.doi.org/10.35713/aic.v7.i1.114273