Yu HH, Chan IN, Wang JH, Qin YY, Chan IW, Wong PK. Artificial intelligence for endoscopic correlates of Correa’s cascade in gastric precancerous lesions and early neoplasia. World J Gastrointest Oncol 2026; 18(10): 123447 [DOI: 10.4251/wjgo.123447]
Corresponding Author of This Article
Pak Kin Wong, PhD, Professor, Department of Biomedical Engineering, University of Macau, Avenida da Universidade, Taipa, Macau 999078, China. fstpkw@um.edu.mo
Research Domain of This Article
Gastroenterology & Hepatology
Article-Type of This Article
review-article
Open-Access Policy of This Article
This article is an open-access article which was selected by an in-house editor and fully peer-reviewed by external reviewers. It is distributed in accordance with the Creative Commons Attribution Non Commercial (CC BY-NC 4.0) license, which permits others to distribute, remix, adapt, build upon this work non-commercially, and license their derivative works on different terms, provided the original work is properly cited and the use is non-commercial. See: http://creativecommons.org/licenses/by-nc/4.0/
Baishideng Publishing Group Inc, 7041 Koll Center Parkway, Suite 160, Pleasanton, CA 94566, USA
Share the Article
Yu HH, Chan IN, Wang JH, Qin YY, Chan IW, Wong PK. Artificial intelligence for endoscopic correlates of Correa’s cascade in gastric precancerous lesions and early neoplasia. World J Gastrointest Oncol 2026; 18(10): 123447 [DOI: 10.4251/wjgo.123447]
Author contributions: Yu HH and Chan IN contributed to collecting and critically reviewing the relevant literature, developing the scope and structure of the manuscript, drafting the initial manuscript, designing the figures and tables, synthesizing the key findings, and revising the manuscript, they contributed equally to this article, they are the co-first authors of this manuscript; Wang JH, Qin YY, and Chan IW conducted the investigation and contributed to writing, review, and editing; Wong PK provided supervision, project administration, and funding acquisition; and all authors read and approved the final version of the manuscript.
AI contribution statement: ChatGPT 5.2 and ChatGPT 5.5 developed by OpenAI, were used solely for language polishing, grammar refinement, formatting assistance, and improvement of manuscript clarity.
Supported by the Science and Technology Development Fund of Macau, No. 0026/2022/A.
Conflict-of-interest statement: All the authors report no relevant conflicts of interest for this article.
Corresponding author: Pak Kin Wong, PhD, Professor, Department of Biomedical Engineering, University of Macau, Avenida da Universidade, Taipa, Macau 999078, China. fstpkw@um.edu.mo
Received: May 19, 2026 Revised: July 21, 2026 Accepted: August 14, 2026 Published online: October 15, 2026 Processing time: 143 Days and 19.7 Hours
Abstract
Early detection is a major determinant of survival in patients with gastric cancer. However, subtle mucosal abnormalities may be missed under white-light imaging, with detection performance influenced by endoscopist experience and examination conditions. Although deep-learning methods have driven rapid advances in artificial intelligence (AI) research for upper endoscopy, much of the available evidence remains image-based and is not consistently anchored to clinical decision points along Correa’s cascade, within which adjacent stages may coexist and reference standards vary. This review synthesizes AI applications for the detection, characterization, and extent assessment of visible endoscopic correlates of non-atrophic gastritis, multifocal atrophic gastritis, intestinal metaplasia, dysplasia, and early gastric cancer using a cascade-aware framework. Although reported accuracy is often highest in curated datasets, translation into clinical practice requires decision-aligned outputs that are robust to label noise, heterogeneous reference standards, and continuous-video artifacts. Future progress depends on prospective, multicenter, video-based validation with patient-level and procedure-level analyses. It also requires standardized reporting of operating points, workflow impact, false-positive burden per procedure, and real-time feasibility, together with studies designed to determine whether AI -assisted endoscopy improves patient outcomes beyond surrogate performance metrics.
Core Tip: This article applies Correa’s cascade as a clinical framework for evaluating artificial intelligence (AI) in upper endoscopy. Rather than directly identifying biological progression, current AI systems recognize visible endoscopic phenotypes associated with gastritis, atrophy, intestinal metaplasia, dysplasia, and early gastric cancer. By integrating histopathology, endoscopic appearance, AI task design, reference standards, validation level, and clinical decision impact, this review demonstrates why high image-level accuracy may overestimate clinical utility. Future research should emphasize prospective, multicenter, video-based validation together with patient-level and procedure-level outcome assessment.
Citation: Yu HH, Chan IN, Wang JH, Qin YY, Chan IW, Wong PK. Artificial intelligence for endoscopic correlates of Correa’s cascade in gastric precancerous lesions and early neoplasia. World J Gastrointest Oncol 2026; 18(10): 123447
Gastric cancer ranks fifth in both incidence and mortality worldwide, with particularly high prevalence in Eastern Asia, Eastern Europe, and South-Central Asia[1-3]. Patient survival is determined largely by the disease stage at detection. Early gastric cancer (EGC) demonstrates an excellent prognosis, with 5-year survival rates exceeding 90% after curative resection[3-5], whereas advanced disease remains associated with poor survival outcomes. This marked survival disparity underscores the importance of timely screening and early detection as key strategies for reducing mortality. Gastric adenocarcinoma accounts for the vast majority of gastric cancers, and the intestinal subtype is classically associated with the multistep progression described by Correa’s cascade[6-8]. As shown in Figure 1A, Correa’s cascade progresses from normal mucosa through non-atrophic gastritis, multifocal atrophic gastritis, intestinal metaplasia (IM), dysplasia, and adenocarcinoma. Given this protracted progression over many years, endoscopic surveillance provides an opportunity to detect precursor lesions at potentially reversible or curable stages, making it a cost-effective strategy for reducing mortality[6,9-11].
Figure 1 Overview of gastric carcinogenesis and artificial intelligence-assisted endoscopic imaging.
A: Schematic illustration of Correa’s cascade; B: Modular endoscopic artificial intelligence tasks, including computer-aided detection, computer-aided diagnosis, and segmentation (extent mapping), which may be used independently or in combination; C: Representative anonymized endoscopic images from three different patients, with each row showing paired white-light imaging and narrow-band imaging images from the same patient. The left column shows white-light imaging, and the right column shows narrow-band imaging. Arrows indicate corresponding subtle mucosal or vascular abnormalities in the paired images, which are generally more conspicuous with image-enhanced endoscopy. IM: Intestinal metaplasia; WLI: White-light imaging; NBI: Narrow-band imaging; CADe: Computer-aided detection.
The success of endoscopy fundamentally depends on accurate lesion detection, yet conventional white-light imaging (WLI) frequently overlooks subtle mucosal abnormalities in EGC, with documented miss rates ranging from 8.3% to 25.8% and increasing to 32.4% among less experienced endoscopists[12,13]. Although advanced imaging techniques, including chromoendoscopy, narrow-band imaging (NBI), and magnifying endoscopy, improve lesion visualization[14], they cannot fully overcome operator variability and interpretive bias. Therefore, integrating artificial intelligence (AI) into endoscopic practice has become increasingly important for bridging differences in operator experience, standardizing detection performance, and reducing fatigue-related diagnostic errors, thereby improving diagnostic consistency.
Recent reviews have examined AI applications in gastric cancer, EGC, gastric precancerous lesions, and gastric cancer risk stratification, although their scopes and organizational approaches differ from those of the present review. Alsallal et al[15] systematically reviewed machine learning and deep learning applications across gastric cancer management using endoscopy, computed tomography, pathology, and multimodal data. In that review, endoscopy was considered one component of a broader multimodal framework rather than the focus of a dedicated, cascade-oriented synthesis. Lei et al[16] focused on AI applications for EGC, including detection, differentiation-type prediction, invasion-depth estimation, and boundary recognition, with the discussion organized primarily by imaging modality and EGC-specific tasks. Hiramatsu et al[17] reviewed endoscopic diagnosis for gastric cancer risk stratification and included selected AI applications, although AI was presented as one component of a broader risk-stratification framework. Yan et al[18] reviewed deep learning applications for upper gastrointestinal precancerous lesions, including gastric cascade-related entities such as Helicobacter pylori (H. pylori) infection, atrophic gastritis, gastric IM, and dysplasia. Since the publication of that review, the field has expanded substantially, particularly with respect to external validation, video-based evaluation, real-time assistance, and deployment-oriented study designs. Moreover, these disease entities were primarily treated as individual diagnostic targets rather than as biologically interconnected stages within Correa’s cascade. Collectively, these reviews provide valuable summaries organized by disease category, imaging modality, model type, application, or risk stratification. However, they do not specifically map endoscopic AI studies across the endoscopic correlates of Correa’s cascade while simultaneously considering biological progression, visible endoscopic phenotype, AI task formulation, reference standard, validation design, clinical maturity, and clinically relevant decision points.
Such cascade-aware mapping is necessary because gastric precancerous and early neoplastic lesions are biologically interconnected rather than isolated findings. Although Correa’s cascade describes stepwise progression, its stages are not always encountered as discrete categories during endoscopy. Multiple cascade stages may coexist within the same patient, anatomical region, or mucosal field, resulting in overlapping visual phenotypes and heterogeneous reference labels. Atrophy alters the background against which IM is recognized. The distribution of IM modifies cancer risk, surveillance strategies, and biopsy targeting. Dysplasia may arise within metaplastic fields and requires distinction from inflammatory changes or early malignancy. Characterization of EGC requires integration of lesion detection, margin delineation, differentiation status, and invasion-depth estimation. Consequently, similar performance metrics may represent fundamentally different clinical tasks. Image-level IM classification differs substantially from patient-level extent mapping. Likewise, binary neoplasia detection differs from dysplasia grading, EGC delineation, and invasion-depth prediction.
Therefore, rather than using Correa’s cascade merely as a chronological framework, this review uses it as an organizing framework for interpreting AI tasks in the context of coexisting disease stages, overlapping endoscopic phenotypes, heterogeneous reference standards, and clinically distinct decision points. Importantly, endoscopic AI systems should be viewed as detecting visible mucosal, vascular, and lesion-level phenotypes associated with stages of Correa’s cascade rather than as directly identifying the underlying histological stage. Because the cascade itself is defined by histopathological and pathogenetic progression, endoscopic phenotypes serve as proxies for cascade stages but cannot provide histological confirmation. Consequently, the clinical value of AI-assisted detection or characterization depends on the use of reference standards that are reliable, clinically meaningful, and aligned with the intended clinical decision point, particularly for atrophy, IM, dysplasia, and EGC. Within this framework, this minireview integrates histopathological progression, visible endoscopic phenotype, AI task formulation, reference standard, validation design, clinical maturity, and clinical action. Rather than ranking models according to architecture or isolated accuracy metrics, this review evaluates where AI systems are clinically actionable and where reported performance may be inflated or difficult to interpret. Performance claims become unreliable when studies rely on coarse labels, imperfect histological verification, image-level validation without corresponding patient-level or procedure-level assessment, limited external testing, or inadequate real-time evaluation. By contrast, rigorous validation across these dimensions strengthens confidence in clinical utility. To apply this framework in practice, this review maps AI studies across key transition points from H. pylori infection and atrophy to IM, dysplasia, EGC delineation, and invasion-depth estimation. In doing so, it provides a clinically oriented synthesis of how endoscopic AI may support surveillance, biopsy targeting, lesion characterization, and treatment planning across the gastric carcinogenesis continuum.
LITERATURE SEARCH AND SELECTION STRATEGY
This article presents a narrative, clinically oriented synthesis rather than a formal systematic review or meta-analysis. A targeted literature search was conducted primarily through Google Scholar and PubMed/MEDLINE to identify studies that evaluated AI-assisted upper gastrointestinal endoscopy for evaluation of gastric carcinogenesis and related endoscopic tasks. To ensure comprehensive coverage, the reference lists of relevant original articles and reviews were manually screened to identify additional clinically important and foundational studies. With respect to scope, the review encompassed studies that developed, validated, or clinically evaluated AI models using endoscopic images or videos, including white-light endoscopy, image-enhanced endoscopy (IEE), magnifying endoscopy, and real-time video streams intended for clinical application.
The search strategies incorporated a broad range of keyword combinations across three conceptual domains. AI-related terms included “artificial intelligence”, “AI”, “deep learning”, “machine learning”, “convolutional neural network”, “computer-aided detection”, “computer-aided diagnosis”, “CADe”, and “CADx”. Endoscopy-related search terms included “endoscopy”, “gastroscopy”, “upper gastrointestinal endoscopy”, “esophagogastroduodenoscopy”, “endoscopic image”, “endoscopic video”, and “real-time endoscopy”. Disease-related terms comprised “Helicobacter pylori”, “H. pylori”, “gastritis”, “atrophic gastritis”, “gastric atrophy”, “intestinal metaplasia”, “gastric intestinal metaplasia”, “dysplasia”, “precancerous lesion”, “early gastric cancer”, “gastric cancer”, “gastric neoplasm”, and “Correa cascade”. The searches also incorporated IEE-specific terminology, including “narrow-band imaging”, “linked-color imaging”, “blue laser imaging”, “white-light imaging”, and “image-enhanced endoscopy”.
The retrieved records were screened using a dual-stage process that began with a title and abstract review, followed by a full-text assessment of potentially relevant articles. Eligible studies were required to evaluate AI models using endoscopic images or videos for detection, characterization, segmentation, extent assessment, risk stratification, or treatment-relevant prediction of gastric precancerous conditions or early neoplasia. Because the literature in this field has rapidly expanded, studies were selected for detailed discussion based on their clinical relevance, methodological rigor, recency, representation of key AI tasks along the gastric carcinogenesis cascade, and potential impact on endoscopic decision-making. Priority was given to studies that incorporated real-time evaluation, video-based testing, external validation, a multicenter or prospective design, randomized assessment, or clinically meaningful patient-level or procedure-level endpoints. To emphasize evidence that reflects current deployment-oriented practice, particular attention was given to studies published from 2022 onward, although earlier studies were retained when they represented landmark models, randomized trials, foundational datasets, or clinically influential work. When multiple iterative publications originated from the same research group, priority was given to the most recent, comprehensively validated, or clinically relevant study. Studies focusing exclusively on advanced gastric cancer without stage-stratified analysis, non-endoscopic imaging modalities, non-gastric lesions, pathology-only or radiology-only models, or purely technical model development without clinically interpretable endpoints were excluded from detailed discussion. The literature search was finalized in March 2026.
Within this framework, AI tasks are grouped into section-indexed tiers, including Tier 4A to 4D, Tier 5A to 5C, and Tier 6A to 6D, to accommodate heterogeneous endpoints across cascade-associated clinical conditions. Rather than representing discrete stages of histopathological progression, these tiers are intended to organize endoscopic AI tasks according to their clinical relationship to cascade-associated conditions. Accordingly, they should not be interpreted as evidence that endoscopic AI directly identifies or resolves histopathological progression without appropriate pathological or clinical reference standards. Table 1 provides a cascade-aware framework that distinguishes cascade-related targets and histopathological entities from visible endoscopic phenotypes, AI task formulations, reference standards, clinical decision impacts, and expected sources of bias.
Table 1 Cascade-aware framework linking cascade-related targets with visible endoscopic phenotypes, artificial intelligence task formulation, reference standards, clinical decision impacts, and comparability limitations in gastric endoscopy.
Cascade-related target/histopathological entity
Visible endoscopic phenotype
Typical AI task
Reference standard
Clinical decision impact
Main comparability limitation
H. pylori-associated chronic gastritis
Diffuse or spotty redness; enlarged folds; altered RAC pattern
CADx for infection-status classification
UBT, RUT, histopathology, serology, stool antigen testing, or composite testing1
Supports infection diagnosis and guides eradication therapy
Figure 1B illustrates three primary modular endoscopic AI tasks: Computer-aided detection, computer-aided diagnosis, and segmentation: (1) Computer-aided detection functions as a real-time “second observer” by flagging suspicious regions that might otherwise be overlooked during systematic mucosal inspection. Its outputs typically consist of visual prompts (e.g., boxes or heatmaps) designed to reduce the miss rate rather than provide a definitive diagnosis; (2) Computer-aided diagnosis assists with lesion characterization once a target has been identified by predicting a class label with an associated confidence score. In gastric practice, this primarily involves differentiating non-neoplastic mucosa from early neoplasia but may also include estimation of actionable attributes when available; and (3) Segmentation provides pixel-level classification that differentiates abnormal tissue from normal tissue, producing a spatial map or mask of suspected disease. For the purposes of this review, delineation refers more specifically to boundary or margin definition, such as estimating the lateral border of an EGC before endoscopic resection. Although delineation is derived from segmentation output, its primary clinical purpose is margin definition rather than whole-region classification.
Conventionally, WLI is used primarily for initial inspection and detection, whereas IEE modalities, such as NBI, enhance visualization of the mucosal surface and microvascular patterns, thereby supporting lesion characterization and delineation. Accordingly, reported AI performance should be interpreted in the context of both the imaging modality used and the intended clinical step. Figure 1C provides a representative comparison of WLI and NBI.
In addition to task type, studies are categorized by model architecture, which influences how AI systems process endoscopic images or videos. The major model families reviewed in this article include convolutional neural networks, Transformers, Mamba and other state-space models, traditional machine learning, broad learning systems, automated deep learning, and hybrid pipelines, as summarized in Table 2. Regardless of architecture, model performance should be interpreted in conjunction with task formulation, label quality, imaging modality, validation design, and deployment setting rather than on the bases of the architecture name alone.
Table 2 Model families commonly used in endoscopic artificial intelligence.
Model family
Core idea
Practical implication
CNN
Learns local image features, including texture, edges, color, and mucosal patterns
Common backbone for classification, detection, and real-time applications
Transformer
Uses attention mechanisms to capture broader spatial or temporal context
Useful for context-aware analysis but more data-intensive and computationally intensive
Mamba
Uses state-space modeling for efficient long-sequence processing
Promising for video analysis and long-context tasks, although clinical validation remains limited
Traditional ML
Uses handcrafted features with classifiers such as support vector machines or random forests
Useful in selected applications but less flexible than end-to-end deep learning
Broad learning system
Uses a wide, expandable network structure
Enables rapid model updating in research settings, although clinical evidence remains limited
Automated deep learning
Automates model selection, hyperparameter tuning, or architecture search
Reduces development burden but may limit model transparency
Hybrid model
Combines multiple model families, such as CNN-transformer or CNN-ML architectures
May improve flexibility but remains dependent on the specific task, training data, labels, and validation strategy
Correa’s cascade begins with H. pylori-associated chronic active non-atrophic gastritis[8]. To standardize endoscopic diagnosis, the Kyoto classification of gastritis associates infection status with characteristic mucosal findings[17,19], including the regular arrangement of collecting venules[20], which is characteristic of uninfected fundic mucosa, as well as infection-associated changes such as diffuse or spotty redness and enlarged folds[21]. In clinical practice, these subtle color and textural features observed on WLI are operator-dependent[14,22], providing a clear rationale for AI-assisted assessment.
Diagnostic formulation in H. pylori AI research reflects the clinical workflow, with most studies using procedure-level computer-aided diagnosis to determine infection status across the entire procedure (Table 3). Binary classification (infected vs uninfected; Tier 4A) remains the predominant approach[23-27], whereas a smaller body of work has extended prediction to three classes (uninfected, currently infected, and post-eradication; Tier 4B) to better reflect clinically meaningful phenotypes[28-30]. Methodologically, newer pipelines have focused on improving workflow realism through video-derived supervision and case-level aggregation, including weakly supervised multi-instance learning with Top-k instance selection and aggregation[24,27], along with quality-aware frame selection to reduce artifacts such as blur and reflections[24]. The quality of the evidence has been further strengthened by external validation and randomized evaluation of AI-assisted endoscopy[25]. In multiclass settings, enhanced imaging modalities such as linked-color imaging (LCI), together with eradication history, can improve the discrimination between post-eradication and current infection phenotypes[29,30]. Shichijo et al[28] also investigated the impact of a markedly negative-skewed dataset during evaluation. When diagnosis was defined by the highest predicted probability, predictions were dominated by the “negative” class. To address this issue, they applied a post hoc decision rule that reduced the negative-class score by subtracting an offset (X) (floored at zero), after which the final diagnosis was assigned using patient-level mean scores across the three classes (positive, negative, and eradicated). This adjustment was intended to provide a more clinically meaningful assessment under conditions of class imbalance, with X = 0.9 selected using the kappa statistic.
Table 3 Summary of artificial intelligence studies targeting endoscopic assessment of Helicobacter pylori infection.
Interpreting performance across H. pylori studies requires careful consideration of reference-standard heterogeneity, infection phenotype, image selection, validation strategy, and pipeline design rather than model architecture alone. Ground-truth labels vary substantially across studies, encompassing urea breath testing, serology, stool antigen testing, rapid urease testing, histology, culture, and composite definitions. Because these methods do not capture identical biological or clinical states, their results are not directly interchangeable. Moreover, post-eradication mucosa presents a particular diagnostic challenge because residual atrophic, metaplastic, or inflammatory changes may resemble current or previous infection despite microbiological clearance. Although this problem has been addressed using diverse architectures and pipelines such as convolutional neural networks, Transformer-based models, and other methods, direct comparison among studies is limited by heterogeneous evaluation contexts. Instead, more clinically relevant distinctions concern methodological and contextual factors. Frame selection and multi-instance learning enable case-level aggregation from video-derived images, whereas stomach-site recognition reduces regional sampling imbalance. Similarly, integration of LCI models and eradication-history information improves discrimination between active infection and post-eradication mucosal changes. Furthermore, external validation, video-based assessment, prospective study design, and randomized evaluation provide more realistic (albeit more demanding) tests than internal still-image validation. Consequently, lower performance in multicenter, video-based, or three-class settings may reflect the complexity of a more clinically authentic task rather than inferior model design. Apparent performance differences across studies therefore warrant contextual interpretation rather than direct ranking.
Prolonged H. pylori-associated inflammation may progress to multifocal atrophic gastritis, a key precancerous step in Correa’s cascade characterized by gland loss[8]. This step is clinically important because atrophic changes regress slowly and frequently persist despite eradication therapy, thereby guiding surveillance intensity in populations at higher risk[31,32]. Because atrophy is commonly multifocal across the antrum and gastric body, severity grading is generally site-aware and relies on multi-view assessment[6,8]. In routine practice, severity is described using two complementary approaches. The first approach is endoscopic extent, most commonly the Kimura-Takemoto system, which grades closed-type and open-type atrophy by localizing the atrophic border and therefore generally requires multiple views across the stomach[17,33,34]. The second approach is histology-based grading and staging, typically using the updated Sydney system[35] and the Operative Link on Gastritis Assessment system[36], which are intrinsically patient-level because they rely on multi-site biopsies[32]. In addition, the Kyoto classification provides a complementary feature-based framework for endoscopic gastritis assessment (including atrophy), with findings that may be assessed on individual images or aggregated across multiple views for patient-level risk stratification.
While histology-based systems provide risk-oriented reference standards, routine clinical practice relies on the availability of endoscopic staging at the time of upper endoscopy. Consequently, AI models have been developed to support tasks ranging from binary classification of chronic atrophic gastritis to procedure-level severity or extent staging and patient-level risk stratification (Table 4). Multiple WLI-based studies in Tier 4C have formulated atrophy assessment as binary classification of chronic atrophic gastritis vs chronic non-atrophic gastritis using histopathology-referenced labels derived from biopsy findings and the updated Sydney system[37-39]. Beyond binary classification, studies in Tier 4D increasingly support graded severity assessment at two operational levels. Single-image analyses predict the presence or severity of atrophy within a specific view or site[40,41], whereas multi-frame aggregation studies capture the distributed, site-dependent nature of atrophy, thereby enabling alignment with endoscopic staging systems (e.g., Kimura-Takemoto, Kyoto) or histology-based risk frameworks (e.g., Operative Link on Gastritis Assessment) to predict patient-level staging or risk[42-44]. Some models additionally provide segmentation or intermediate severity outputs to support standardized documentation and staging[43,44]. Accordingly, endoscopic atrophy staging represents a practical procedure-level endpoint for routine endoscopy, whereas assessment of IM shifts the objective from global staging to mapping the distribution and extent of precancerous change.
Table 4 Summary of artificial intelligence studies targeting endoscopic assessment of gastric atrophy.
Direct comparison of atrophy models is constrained by heterogeneity in anatomical sampling, label source, outcome level, and validation design. Binary chronic atrophic gastritis-vs-chronic non-atrophic gastritis classifiers using selected WLI images and histology-referenced labels address a fundamentally different clinical task from systems estimating Kimura-Takemoto extent, Kyoto features, Operative Link on Gastritis Assessment-related risk, or patient-level severity across multiple gastric sites. Moreover, because atrophy is multifocal and topographically heterogeneous, single-image classification may underestimate patient-level disease burden. By comparison, multi-site aggregation and video-based assessment better reflect clinical staging but introduce greater measurement variability. Reported image-level performance should therefore be interpreted with consideration of the biopsy protocol, anatomical sampling scheme, and whether labels were assigned at the image, site, or patient level. Performance differences across studies therefore reflect differences in task formulation, outcome level, and validation design rather than the superiority of a particular model architecture. Accordingly, models that predict image-level appearance, site-level disease severity, mucosal extent, or patient-level risk stratification address fundamentally different clinical endpoints and should not be directly compared.
MAPPING THE RISK FIELD FOR GASTRIC IM
IM is an advanced lesion in the atrophy-metaplasia sequence, in which chronic mucosal injury leads to gland loss and subsequent replacement of native gastric glands by intestinal-type glands during chronic mucosal injury[8]. It represents a transition toward an intestinal epithelial phenotype and is commonly classified as complete (small-intestinal) IM or incomplete (colonic) IM, with the incomplete subtype carrying greater malignant potential and with risk further increasing as IM extent increases[6,8]. In clinical practice, the reference standard for IM diagnosis is histopathology[45]; this must be obtained through endoscopic biopsy because IM often appears as flat mucosa with subtle or nonspecific changes under conventional WLI[45,46]. Several endoscopic indicators, such as whitish discoloration, villous appearance, and patchy redness, have been associated with IM[47]. However, these features overlap with gastritis and atrophy, contributing to interobserver variability. This limitation is reflected in reports of low sensitivity for purely endoscopic recognition (approximately 24% in both the antrum and body) despite high specificity[48]. Consequently, IEE, including NBI, magnifying NBI and color-contrast techniques such as LCI and blue-light imaging, has been adopted to improve visualization. Beyond improving per-frame visibility, these modalities enable structured assessment of disease extent, an important clinical consideration because cancer risk depends on the distribution of IM throughout the stomach. To support real-time assessment of IM extent, the IEE-based Endoscopic Grading of Gastric IM system assigns site-level grades at multiple predefined gastric landmarks and combines them into a patient-level score, which inherently depends on multi-site image sampling during the examination[49,50]. This endoscopic extent score aligns closely with the histology-based Operative Link Gastric IM Assessment (OLGIM) system, which similarly uses topographic distribution to determine a patient-level stage[51]. Nevertheless, because histology-based frameworks (the updated Sydney System and OLGIM) remain the reference standard for staging, there is strong motivation to develop objective, real-time methods for endoscopic detection and extent quantification.
As summarized in Table 5, AI research on IM mirrors clinical workflows, progressing from diagnosis to severity grading and extent mapping for risk stratification. Diagnosis represents the natural entry point because it can be formulated as a binary or coarse multiclass task, with histopathologic findings from biopsy specimens providing a relatively accessible reference standard (Tier 5A)[52-55]. Although several studies have reported image-level accuracies exceeding 90%, direct comparison remains limited by differences in imaging modality, labeling strategy, dataset composition, thresholding methods, validation design, and model architecture[52-54]. For example, the study by Ligato et al[55], which achieved image-level accuracy below 80% using a validation-tuned patch-to-image thresholding scheme, illustrates how methodological choices can substantially influence sensitivity and specificity across heterogeneous datasets.
Table 5 Summary of artificial intelligence studies targeting endoscopic assessment of intestinal metaplasia.
IM severity grading is more clinically aligned but methodologically heterogeneous (Tier 5B). Some approaches generate image-level grades from individual frames through patch aggregation or weak supervision, whereas others perform site-level grading using the Endoscopic Grading of Gastric IM system or aggregate images to predict patient-level OLGIM-oriented risk[56-59]. Because these endpoints are not interchangeable, model performance must be interpreted within the appropriate clinical context. IM is patchy, and biopsy-based labels may be adversely affected by sampling error or spatial mismatch with endoscopic images. Moreover, cancer risk depends on topographic extent rather than isolated image appearance. Although diverse architectures and pipelines have been applied to these tasks, architectural designation alone provides limited insight when models are trained and validated for different clinical endpoints. For example, the performance decline observed by Niu et al[58] from internal to external validation demonstrates that clinical transfer can be more demanding than development-set testing. Therefore, apparent performance differences should be interpreted according to task formulation, aggregation strategy, validation rigor, and clinical alignment rather than architectural superiority. Likewise, high image-level accuracy alone does not indicate lesion location, distribution, or extent, underscoring the shift in this field from presence detection toward spatial quantification.
Extent mapping via localization and segmentation has consequently been explored using WLI[60], combined WLI and NBI in multimode settings[61], and LCI (Tier 5C)[62]. Such studies are clinically important because segmentation can support point-of-care quantification of IM distribution; however, reported performance depends strongly on dataset composition, pixel-level annotation quality, boundary definition, imaging modality, input resolution, deployment setting, model architecture, and metric selection. Pornvoraphat et al[61] emphasized practical multimode deployment for WLI and NBI using automatic mode routing and mode-specific adaptation, whereas Zhang et al[62] extended segmentation to LCI by releasing an expert-annotated dataset with pixel-level labels, thereby addressing an important data bottleneck for color-enhanced extent mapping. Accordingly, offline segmentation on curated datasets, real-time multimode deployment, and modality-specific LCI extent mapping represent distinct clinical objectives with different performance implications. The reported Dice coefficient, Intersection over Union (IoU), mean IoU, and inference speed should therefore be interpreted within the specific task context rather than compared directly across heterogeneous studies. Taken together, these studies illustrate how advances in segmentation support clinically actionable quantification and point-of-care assessment.
The heterogeneity of imaging modalities and task formulations across these studies highlights an important principle: IEE and WLI represent distinct technical and clinical contexts rather than interchangeable imaging inputs. High-quality IEE appears particularly well suited for training AI systems to detect subtle IM patterns or map lesion extent, for which enhanced visualization is critical. WLI-based systems, by contrast, offer greater practical utility for broad screening and real-time inspection in routine clinical workflows. Currently available evidence does not support the assumption that robust models trained on multimodal datasets will reliably generalize across imaging modalities without explicit validation. Consequently, future AI development should incorporate imaging modality as a structured variable, validate performance separately for WLI and IEE, and prospectively evaluate whether models trained on one modality can be successfully transferred to another. Collectively, these findings underscore the technical reality that the imaging modality fundamentally shapes both the training signal and the clinical deployment context.
TARGETING EARLY NEOPLASIA: FROM DYSPLASIA TO EGC
Dysplasia, also known as glandular intraepithelial neoplasia, is a noninvasive neoplastic lesion confined to the epithelium; by definition, it does not extend into the lamina propria[6,8]. Histologically, dysplasia is classified as low-grade dysplasia (LGD) or high-grade dysplasia (HGD) according to the degree of atypia, consistent with the revised Vienna classification[63,64]. Compared with LGD, HGD carries a substantially greater risk of synchronous or subsequent carcinoma. EGC is defined as adenocarcinoma limited to the mucosa or submucosa (T1 in tumor, node, metastasis staging system[65]), regardless of lymph node involvement. As the endpoint of Correa’s cascade, EGC is often detected before deeper invasion and therefore represents the principal target for curative endoscopic resection. Endoscopically, both dysplasia and EGC typically present as superficial (type 0) lesions[66] and frequently exhibit non-polypoid morphology, appearing flat or slightly depressed (Paris 0-II)[67]. Under WLI, these lesions may demonstrate only subtle color or surface-pattern changes with variable demarcation, often resembling inflammatory or metaplastic mucosa. This broad spectrum of endoscopic appearances, together with the need to infer staging-relevant features (e.g., depth of invasion from surface and fold morphology), contributes to substantial operator dependence[68] and underscores the need for AI-assisted detection and characterization.
As summarized in Table 6, the development of diagnostic AI for dysplasia and early gastric neoplasia has been largely constrained by label availability in routine clinical practice. Consequently, dysplasia is rarely treated as a standalone endpoint. Instead, it is either incorporated into broader multiclass staging frameworks spanning benign, precancerous, and neoplastic conditions (Tier 6A)[69-71] or combined with EGC or carcinoma to form a composite neoplasia label contrasted against non-neoplastic conditions (Tier 6B)[72-75]. In addition, some Tier 6A systems incorporate lesion localization or auxiliary segmentation of background risk lesions, such as IM or atrophy[70,71]. Rather than distinguishing LGD, HGD, and carcinoma individually, Tier 6B studies typically collapse these lesions into a binary neoplasia-vs-non-neoplasia task. This approach reduces label sparsity and diagnostic ambiguity while enabling evaluation across larger datasets, including both image-based and video-based modalities, as well as prospective or randomized clinical assistance trials[72-75].
Table 6 Summary of artificial intelligence studies targeting endoscopic diagnosis of early gastric neoplasia.
Although this endpoint simplification improves feasibility, it reduces diagnostic granularity and limits assessment of clinically important transitions between LGD, HGD, intramucosal carcinoma, and submucosal invasive cancer. Composite neoplasia endpoints are most useful as safety-oriented detection tasks because they can reduce the likelihood of missing clinically significant lesions. Their value for cascade-specific decision-making, however, is more limited because LGD, HGD, intramucosal carcinoma, and submucosal invasive cancer do not carry identical management implications. For example, LGD may warrant intensified surveillance, repeat assessment, targeted biopsy, or local resection depending on lesion characteristics and local practice, whereas HGD and EGC generally require more urgent therapeutic evaluation. Once submucosal invasion is identified, management typically shifts from endoscopic resection to surgery because of the increased risk of lymph node metastasis. Moreover, single-label classification may obscure the coexistence of IM, atrophy, or carcinoma within the surrounding mucosa. Taken together, these limitations indicate that binary neoplasia detection is best interpreted as a safety-oriented screening tool rather than a complete cascade-aware diagnostic system. Accordingly, AI systems intended to guide clinical management should be trained and validated to distinguish clinically adjacent stages or should explicitly state when their outputs are limited to lesion detection rather than treatment selection.
Conversely, multi-class formulations usually show lower aggregate accuracy. This decline likely reflects inter-class overlap, annotation inconsistency, dataset composition, imaging modality, and greater endpoint complexity rather than inherent model inferiority. From an architectural perspective, the field encompasses image classifiers, detection-classification pipelines, auxiliary segmentation systems, feature-fusion ensembles, and real-time detectors, each designed to address different clinical tasks and deployment settings. Accordingly, apparent performance differences reflect task heterogeneity, label quality, validation rigor, and clinical study maturity rather than architectural superiority. Reported accuracy, sensitivity, and miss-rate reductions are clinically informative within their specific contexts. Therefore, these results should be interpreted according to the study design rather than compared directly across heterogeneous investigations.
Beyond diagnosis, endoscopic treatment planning requires procedure-oriented information, particularly delineation of lateral lesion margins to facilitate complete resection and estimation of invasion depth for guiding endoscopic vs surgical management. The histologic differentiation status also influences curability assessment and resection strategy in some frameworks. Accordingly, these clinical needs have driven AI development in lesion delineation and extent characterization (Tier 6C) and in depth-prediction decision support with associated pathologic endpoints (Tier 6D), as summarized in Table 7.
Table 7 Summary of artificial intelligence studies targeting endoscopic delineation and depth staging of early gastric neoplasia.
Differentiation ACC 83.3% in test set and 86.2% in man-machine comparison; margin ACC 82.7% for differentiated lesions and 88.1% for undifferentiated lesions at overlap threshold of 0.80
In Tier 6C, lesion extent is typically assessed using segmentation or heatmap-based demarcation and evaluated with overlap-based metrics such as the Dice coefficient and IoU[76-78]. Methodological approaches vary considerably and include classification-guided segmentation, sliding-window heatmap demarcation, and attention-enhanced segmentation pipelines, each reflecting different assumptions about how lesion borders should be learned and localized. Likewise, ground-truth definitions differ across studies, ranging from expert endoscopic contours to biopsy-linked labels and pathology-referenced resection margins. As a result, endoscopic visual boundaries and histopathologic margins represent related but non-interchangeable targets, making direct comparison of segmentation and detection metrics inappropriate unless the annotation unit and reference standard are explicitly specified. Reported delineation performance should therefore be interpreted in the context of dataset composition, imaging modality, reference-margin definition, annotation granularity, validation design, deployment setting, and methodological strategy rather than as evidence of the superiority of a particular architectural approach.
Tier 6D studies advance beyond detection toward clinically actionable phenotyping, primarily focusing on binary prediction of invasion depth and, in selected datasets, on prediction of histologic differentiation status[79-81]. These predictions directly align with resection-planning requirements but are predominantly evaluated using WLI, reflecting the constraints of real-world clinical deployment. At the same time, Tier 6D studies are particularly vulnerable to spectrum and verification bias because high-confidence depth and differentiation labels are derived primarily from lesions undergoing resection rather than from screening-detected populations. For this reason, external validation and patient-level evaluation are especially important for interpreting Tier 6D performance and assessing generalizability beyond resection-enriched cohorts. Meaningful comparison among Tier 6D studies is further complicated by heterogeneity in model design, ground-truth sources, external validation strategies, and expert-comparison settings, as well as by whether performance is assessed for AI alone or in conjunction with endoscopists. These methodological differences substantially influence reported performance and limit direct cross-study comparison. Therefore, Tier 6D results should be interpreted within the context of their validation strategy and potential for generalization rather than as direct evidence that one model architecture is superior to another.
LIMITATIONS AND DEPLOYMENT CONSIDERATIONS
Although AI systems have demonstrated promising performance in detecting and characterizing the endoscopic correlates of Correa’s cascade in curated research settings, their outputs remain based on visible endoscopic phenotypes rather than direct assessment of the underlying histopathological stages. Accordingly, reliable and clinically meaningful reference standards are essential for interpreting AI outputs, particularly for atrophy, IM, dysplasia, and EGC. Nevertheless, translation into routine clinical practice continues to face substantial challenges. Misaligned endpoints, imperfect labels and reference standards, selection bias, and limitations in real-time workflow all impede clinical implementation. Furthermore, reported numerical performance metrics across studies reflect heterogeneity in dataset composition, annotation strategies, imaging modalities, validation settings, and outcome levels. Without explicit methodological contextualization, direct comparison of these metrics may be misleading. These challenges become particularly important at clinically decisive transitions, such as the progression from IM to dysplasia, where subtle endoscopic patterns and uncertain ground truth can fundamentally alter surveillance and treatment decisions.
This reality highlights a fundamental gap between research metrics and clinical utility. Image-level accuracy does not necessarily translate into patient-level or procedure-level usefulness. A model may correctly classify selected still frames yet fail to detect a lesion during continuous real-time inspection, generate excessive false-positive prompts, or produce unstable predictions throughout the examination. Thus, clinically meaningful evaluation should emphasize patient-level sensitivity, lesion-level miss rate, false positives per procedure or per minute, latency, real-time video performance, and workflow impact rather than image-level accuracy alone.
Endoscopic phenotype vs histopathological stage
A central conceptual limitation is that Correa’s cascade is defined by histopathological progression, whereas endoscopic AI systems are trained on visible morphology. Rather than learning histopathological features directly, these systems learn from color, texture, surface pattern, vascular pattern, lesion contour, and video-frame characteristics that correlate with cascade-associated conditions such as H. pylori-associated gastritis, atrophy, IM, dysplasia, or EGC. Accordingly, visible features alone do not constitute histological confirmation of these conditions. Consequently, AI outputs should be interpreted as phenotype-based decision support rather than as direct confirmation of the underlying histopathological stage. This distinction is particularly important for flat or patchy lesions, coexisting cascade stages, and cases involving biopsy sampling error or clinicopathological mismatch, in which endoscopic appearance and histopathology may diverge.
Endpoint definition and label granularity
A recurring limitation of the current literature is that study endpoints frequently do not align with clinical decision-making, and the corresponding labels are often too coarse to support cascade-informed care. For example, dysplasia-specific characterization remains uncommon, with LGD and HGD frequently combined with EGC into a composite “neoplasia” label. Although this approach simplifies model development, it limits the ability to distinguish patients requiring surveillance from those requiring referral. Similarly, for precursor lesions, H. pylori status is defined inconsistently across studies, either as binary infection status or as a multi-class categorization encompassing uninfected, current infection, and post-eradication phenotypes, despite these states representing distinct mucosal phenotypes with different clinical implications. A comparable limitation applies to staging of atrophy and IM. Systems based on the Kyoto classification, the Endoscopic Grading of Gastric IM system, or the OLGIM classification are sometimes inferred from single images despite being inherently regional and dependent on assessment of disease extent.
Reference-standard variability and label noise
Heterogeneous reference standards fundamentally influence task definition and introduce label noise across cascade-associated conditions. For example, the ground truth for H. pylori may be established using urea breath testing, rapid urease testing, histology, or composite diagnostic criteria, while endoscopic appearance may be confounded by post-eradication changes. Similarly, biopsy confirmation of atrophy and IM is vulnerable to sampling error, non-standardized biopsy protocols, and incomplete spatial coverage. An additional challenge is the frequent coexistence of multiple cascade-related histological stages or lesions within the same patient, lesion, or endoscopic field. Atrophy and IM may be present in different gastric regions, whereas IM, LGD, HGD, and carcinoma may coexist within the same lesion or mucosal field. Despite this biological complexity, many AI datasets assign a single dominant label to each image, frame, lesion, or procedure. Although this strategy simplifies model training and performance reporting, but may obscure biological heterogeneity, discard information about concurrent lower-grade lesions, and introduce additional label noise. For example, an image labeled as dysplasia may also contain surrounding IM or atrophy. Similarly, a procedure labeled according to the most severe histology may not capture the spatial distribution of less severe but clinically relevant precursor lesions. Observer variability, particularly in the diagnosis of LGD, further compounds these challenges. Taken together, these factors indicate that single-label performance metrics should be interpreted cautiously, especially when adjacent cascade stages exhibit overlapping endoscopic appearances, uncertain histological boundaries, and different management implications. Variability in reference standards and label definitions further limits direct comparison among studies and complicates meta-analysis. Reported sensitivity, specificity, Dice coefficient, and IoU cannot be directly aggregated across datasets that employ different ground-truth definitions, imaging modalities, annotation units, or labeling schemes.
The imaging modality represents an additional source of phenotype and label variation. WLI, NBI, LCI, blue-light imaging, chromoendoscopy, and magnifying endoscopy each emphasize different mucosal, vascular, color, and surface-pattern features. Consequently, labels or annotations generated using one imaging modality are not necessarily interchangeable with those generated using another. This distinction is particularly important for IM, dysplasia, and EGC, for which IEE reveals subtle surface or vascular patterns that are often poorly visible with WLI. When multimodal datasets are used without explicitly accounting for imaging modality, AI models may learn modality-specific artifacts or acquisition signatures rather than disease-specific features. This represents a form of label noise that is distinct from, but closely related to, the observer and sampling variability discussed above. Conversely, AI models trained on high-quality multimodal datasets with explicit modality stratification may achieve improved performance across imaging modalities, although this possibility remains an open question requiring prospective validation. Accordingly, future studies should report modality-specific performance, clearly document how labels were generated for each imaging modality, and validate whether models generalize across WLI and IEE settings. Without such safeguards, apparent performance differences across studies may reflect modality-specific training effects rather than true differences in model robustness.
Selection, verification, and spectrum bias
Because many datasets preferentially include high-quality endoscopic images while excluding cases with poor preparation, blur, or motion artifacts, reported performance may be inflated compared with that achieved during routine video-based endoscopy. Similarly, verification and spectrum bias are particularly relevant to invasion-depth and histologic differentiation tasks, as well as EGC-focused studies, in which resection-confirmed cases are often over-represented.
Generalization and real-time deployment
Generalization remains uncertain when validation is limited to a single center or a single vendor because of variation in device generation, imaging settings, imaging modality, enhancement algorithms, and disease prevalence. Furthermore, many endoscopic AI systems described in the current literature rely on convolutional neural network-based pipelines optimized for single-frame classification or localization, reflecting both the structure of available datasets and the requirements of real-time computational efficiency. Although these approaches perform well for frame-level tasks, they may be less effective for extent-dependent staging, such as atrophy and IM, or for evaluating coexisting cascade stages, which require broader spatial context and, ideally, temporal information derived from video. Future deployment-oriented studies should therefore place greater emphasis on multimodal fusion and spatiotemporal or mapping-based approaches that can operate under real-time constraints. For successful clinical implementation, real-time AI systems must also tolerate video artifacts while minimizing the false-positive burden to reduce alarm fatigue. Consequently, studies should report false positives per minute, miss rates, latency, hardware requirements, and failure modes, preferably using continuous-video evaluation rather than retrospective still-image testing.
Studies that achieve finer-grained classification, explicit recognition of coexisting disease stages, reduced dependence on observer expertise through standardized morphologic criteria, and maintenance of real-time feasibility are likely to have the greatest clinical impact. Particularly promising directions for future research include multimodal learning across WLI, NBI, and chromoendoscopy, together with whole-stomach spatial mapping that places lesion severity within the context of regional mucosal patterns rather than isolated endoscopic fields.
Bridging the gap to clinical practice
Clinical adoption of endoscopic AI depends not only on diagnostic accuracy but also on regulatory approval, workflow integration, cost-effectiveness, medico-legal responsibility, platform interoperability, explainability, and clinician acceptance. Accordingly, the systems most likely to achieve widespread implementation are those that address well-defined, high-frequency tasks with immediate procedural relevance, including H. pylori status estimation, blind-spot monitoring, inspection-quality control, and gastric neoplasia detection. By comparison, more complex applications, such as dysplasia grading, invasion-depth prediction, and treatment-curability assessment, require substantially stronger validation because diagnostic errors may directly influence surveillance strategies, endoscopic resection, or surgical referral decisions.
When evaluating commercial or near-commercial endoscopic AI systems, clinicians should consider several key implementation factors. First, prospective and external validation must be clearly documented. Second, performance metrics should be reported at the patient or procedure level rather than at the image level. Third, false positives must be quantified using clinically interpretable measures, such as the number of false positives per procedure or per minute. Fourth, latency and hardware requirements must be compatible with real-time endoscopic practice. Fifth, the model should be validated using the same endoscope brand, processor generation, and imaging modality used in the intended clinical setting. Finally, failure modes and known limitations must be transparently reported. Collectively, these criteria provide a practical framework for determining whether a given system is suitable for local clinical implementation.
Furthermore, explainability through visual prompts such as heatmaps, bounding boxes, or segmentation masks can improve clinician confidence when these outputs correspond to recognizable mucosal, vascular, or lesion-boundary abnormalities. Nevertheless, visual explainability should not be equated with diagnostic accuracy, and all visual outputs require validation against reliable clinical and pathological reference standards. Another important challenge is platform interoperability because image color, illumination, resolution, enhancement algorithms, and frame rates vary across vendors and processor generations. Similarly, cost-effectiveness remains uncertain, particularly in low-incidence regions or for systems requiring additional hardware, software licensing, cloud-based processing, staff training, or workflow redesign. Finally, medico-legal responsibility must be clearly defined, with endoscopic AI functioning as decision-support technology rather than an autonomous diagnostic system, thereby preserving the endoscopist’s responsibility for clinical interpretation and patient management.
CONCLUSION
For the endoscopic correlates of Correa’s cascade, AI has progressed from curated image benchmarks toward procedure-level decision support at clinically important transition points. Current AI systems should be interpreted as recognizing visible mucosal, vascular, and lesion-level phenotypes associated with cascade stages rather than as directly identifying the underlying biological cascade or confirming the histopathological stage. As the field has matured, the principal barrier to clinical translation has shifted from model availability to standardization. Future progress will therefore depend on clinically aligned, cascade-aware endpoints, transparent reference standards, patient-level and procedure-level validation, and prospective video-based testing across centers, vendors, and imaging modalities. Successful clinical deployment also requires reporting of real-time feasibility, latency, hardware requirements, false-positive rates in clinically interpretable units (e.g., per procedure or per minute), workflow impact, and failure modes. Particularly important priorities for future research include whole-stomach, site-aware mapping of atrophy and IM, robust detection of early neoplasia, treatment-oriented margin delineation, and validated invasion-depth prediction. Achieving these goals will require not only continued technical innovation but also collaborative standardization across centers and vendors. Ultimately, the success of AI-assisted endoscopy should be judged not by surrogate accuracy metrics alone but by its ability to improve biopsy targeting, surveillance planning, treatment selection, interval gastric cancer rates, and, ultimately, patient outcomes.
Katai H, Ishikawa T, Akazawa K, Isobe Y, Miyashiro I, Oda I, Tsujitani S, Ono H, Tanabe S, Fukagawa T, Nunobe S, Kakeji Y, Nashimoto A; Registration Committee of the Japanese Gastric Cancer Association. Five-year survival analysis of surgically resected gastric cancer cases in Japan: a retrospective analysis of more than 100,000 patients from the nationwide registry of the Japanese Gastric Cancer Association (2001-2007).Gastric Cancer. 2018;21:144-154.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 412][Cited by in RCA: 393][Article Influence: 49.1][Reference Citation Analysis (5)]
Alsallal M, Habeeb MS, Vaghela K, Malathi H, Vashisht A, Sahu PK, Singh D, Al-Hussainy AF, Aljanaby IA, Sameer HN, Athab ZH, Adil M, Yaseen A, Farhood B. Artificial intelligence in gastric cancer: a systematic review of machine learning and deep learning applications.Abdom Radiol (NY). 2026;51:1694-1710.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 1][Cited by in RCA: 4][Article Influence: 4.0][Reference Citation Analysis (0)]
Pimentel-Nunes P, Libânio D, Marcos-Pinto R, Areia M, Leja M, Esposito G, Garrido M, Kikuste I, Megraud F, Matysiak-Budnik T, Annibale B, Dumonceau JM, Barros R, Fléjou JF, Carneiro F, van Hooft JE, Kuipers EJ, Dinis-Ribeiro M. Management of epithelial precancerous conditions and lesions in the stomach (MAPS II): European Society of Gastrointestinal Endoscopy (ESGE), European Helicobacter and Microbiota Study Group (EHMSG), European Society of Pathology (ESP), and Sociedade Portuguesa de Endoscopia Digestiva (SPED) guideline update 2019.Endoscopy. 2019;51:365-388.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 813][Cited by in RCA: 773][Article Influence: 110.4][Reference Citation Analysis (6)]
Pimentel-Nunes P, Libânio D, Lage J, Abrantes D, Coimbra M, Esposito G, Hormozdi D, Pepper M, Drasovean S, White JR, Dobru D, Buxbaum J, Ragunath K, Annibale B, Dinis-Ribeiro M. A multicenter prospective study of the real-time use of narrow-band imaging in the diagnosis of premalignant gastric conditions and lesions.Endoscopy. 2016;48:723-730.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 206][Cited by in RCA: 203][Article Influence: 20.3][Reference Citation Analysis (4)]
Lin N, Yu T, Zheng W, Hu H, Xiang L, Ye G, Zhong X, Ye B, Wang R, Deng W, Li J, Wang X, Han F, Zhuang K, Zhang D, Xu H, Ding J, Zhang X, Shen Y, Lin H, Zhang Z, Kim JJ, Liu J, Hu W, Duan H, Si J. Simultaneous Recognition of Atrophic Gastritis and Intestinal Metaplasia on White Light Endoscopic Images Based on Convolutional Neural Networks: A Multicenter Study.Clin Transl Gastroenterol. 2021;12:e00385.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 33][Cited by in RCA: 25][Article Influence: 5.0][Reference Citation Analysis (0)]
Wang C, Li Y, Yao J, Chen B, Song J, Yang X.
Localizing and Identifying Intestinal Metaplasia Based on Deep Learning in Oesophagoscope. 2019 8th International Symposium on Next Generation Electronics (ISNE); 2019 Oct 9-10; Zhengzhou, China. NJ: IEEE, 2019: 1-4.
[PubMed] [DOI]
Pornvoraphat P, Tiankanon K, Pittayanon R, Nupairoj N, Vateekul P, Rerknimitr R. Real-time gastric intestinal metaplasia segmentation using a deep neural network designed for multiple imaging modes on high-resolution images.Knowl-Based Syst. 2024;300:112213.
[RCA] [PubMed] [DOI] [Full Text][Cited by in RCA: 2][Reference Citation Analysis (0)]
Zhang M, Wang L, Yu Y, Liu L, Tao X, Gu L, Ling T. A Benchmark Dataset of Endoscopic Images and a Novel Deep Learning Method to Segment Gastric Intestinal Metaplasia Under Linked Color Imaging.Comput Intell. 2026;42:e70177.
[PubMed] [DOI] [Full Text]
Schlemper RJ, Riddell RH, Kato Y, Borchard F, Cooper HS, Dawsey SM, Dixon MF, Fenoglio-Preiser CM, Fléjou JF, Geboes K, Hattori T, Hirota T, Itabashi M, Iwafuchi M, Iwashita A, Kim YI, Kirchner T, Klimpfinger M, Koike M, Lauwers GY, Lewin KJ, Oberhuber G, Offner F, Price AB, Rubio CA, Shimizu M, Shimoda T, Sipponen P, Solcia E, Stolte M, Watanabe H, Yamabe H. The Vienna classification of gastrointestinal epithelial neoplasia.Gut. 2000;47:251-255.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 1697][Cited by in RCA: 1583][Article Influence: 60.9][Reference Citation Analysis (5)]
Wu L, Shang R, Sharma P, Zhou W, Liu J, Yao L, Dong Z, Yuan J, Zeng Z, Yu Y, He C, Xiong Q, Li Y, Deng Y, Cao Z, Huang C, Zhou R, Li H, Hu G, Chen Y, Wang Y, He X, Zhu Y, Yu H. Effect of a deep learning-based system on the miss rate of gastric neoplasms during upper gastrointestinal endoscopy: a single-centre, tandem, randomised controlled trial.Lancet Gastroenterol Hepatol. 2021;6:700-708.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 130][Cited by in RCA: 116][Article Influence: 23.2][Reference Citation Analysis (0)]
Feng J, Zhang Y, Feng Z, Ma H, Gou Y, Wang P, Feng Y, Wang X, Huang X. A prospective and comparative study on improving the diagnostic accuracy of early gastric cancer based on deep convolutional neural network real-time diagnosis system (with video).Surg Endosc. 2025;39:1874-1884.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 6][Cited by in RCA: 6][Article Influence: 6.0][Reference Citation Analysis (0)]
Ling T, Wu L, Fu Y, Xu Q, An P, Zhang J, Hu S, Chen Y, He X, Wang J, Chen X, Zhou J, Xu Y, Zou X, Yu H. A deep learning-based system for identifying differentiation status and delineating the margins of early gastric cancer in magnifying narrow-band imaging endoscopy.Endoscopy. 2021;53:469-477.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 77][Cited by in RCA: 60][Article Influence: 12.0][Reference Citation Analysis (0)]
Takemoto S, Hori K, Yoshimasa S, Nishimura M, Nakajo K, Inaba A, Sasabe M, Aoyama N, Watanabe T, Minakata N, Ikematsu H, Yokota H, Yano T. Computer-aided demarcation of early gastric cancer: a pilot comparative study with endoscopists.J Gastroenterol. 2023;58:741-750.
[RCA] [PubMed] [DOI] [Full Text][Cited by in RCA: 7][Reference Citation Analysis (0)]
Wu L, Zhou W, Wan X, Zhang J, Shen L, Hu S, Ding Q, Mu G, Yin A, Huang X, Liu J, Jiang X, Wang Z, Deng Y, Liu M, Lin R, Ling T, Li P, Wu Q, Jin P, Chen J, Yu H. A deep neural network improves endoscopic detection of early gastric cancer without blind spots.Endoscopy. 2019;51:522-531.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 229][Cited by in RCA: 182][Article Influence: 26.0][Reference Citation Analysis (4)]
Wu L, Zhang J, Zhou W, An P, Shen L, Liu J, Jiang X, Huang X, Mu G, Wan X, Lv X, Gao J, Cui N, Hu S, Chen Y, Hu X, Li J, Chen D, Gong D, He X, Ding Q, Zhu X, Li S, Wei X, Li X, Wang X, Zhou J, Zhang M, Yu HG. Randomised controlled trial of WISENSE, a real-time quality improving system for monitoring blind spots during esophagogastroduodenoscopy.Gut. 2019;68:2161-2169.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 285][Cited by in RCA: 240][Article Influence: 34.3][Reference Citation Analysis (5)]
Gong EJ, Bang CS, Lee JJ, Baik GH, Lim H, Jeong JH, Choi SW, Cho J, Kim DY, Lee KB, Shin SI, Sigmund D, Moon BI, Park SC, Lee SH, Bang KB, Son DS. Deep learning-based clinical decision support system for gastric neoplasms in real-time endoscopy: development and validation study.Endoscopy. 2023;55:701-708.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 61][Cited by in RCA: 52][Article Influence: 17.3][Reference Citation Analysis (1)]
Wu L, He X, Liu M, Xie H, An P, Zhang J, Zhang H, Ai Y, Tong Q, Guo M, Huang M, Ge C, Yang Z, Yuan J, Liu J, Zhou W, Jiang X, Huang X, Mu G, Wan X, Li Y, Wang H, Wang Y, Zhang H, Chen D, Gong D, Wang J, Huang L, Li J, Yao L, Zhu Y, Yu H. Evaluation of the effects of an artificial intelligence system on endoscopy quality and preliminary testing of its performance in detecting early gastric cancer: a randomized controlled trial.Endoscopy. 2021;53:1199-1207.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 152][Cited by in RCA: 132][Article Influence: 26.4][Reference Citation Analysis (0)]
P-Reviewer: Isakov V, AGAF, Chief, Full Professor, MD, PhD, Russia; Sipos F, Associate Professor, MD, PhD, Hungary; Wang YG, PhD, Professor, China S-Editor: Bai Y L-Editor: A P-Editor: Wang WB