Published online Sep 8, 2026. doi: 10.37126/aige.118493
Revised: January 28, 2026
Accepted: March 6, 2026
Published online: September 8, 2026
Processing time: 243 Days and 17.9 Hours
There is increasing interest in using artificial intelligence (AI) in the detection of dysplasia in patients with inflammatory bowel disease (IBD). However, the appli
Core Tip: Dataset shift impairs non-inflammatory bowel disease (IBD) trained artificial intelligence’s (AI) ability in IBD dysplasia detection as they falter in inflamed colons and therefore IBD-specific models need to be developed. However, even with the IBD-trained models, various biases (selection, annotation, device) are seen which are amplified by training-deployment mismatches and affects generalizability as evident in external validation. This review compares performances, dissects these failures, and suggests a roadmap: Multicenter datasets, consensus labelling, federated learning, and multimodal models, bolstered by regulatory oversight and clinical workflows. Such integrated steps will help to generate equitable, reliable AI to enhance surveillance and avert colorectal cancer in IBD.
- Citation: Bhandari R, Gartlan J, Oppong P, Chhabra P. Artificial intelligence for inflammatory bowel disease dysplasia detection: Current evidence and future directions. Artif Intell Gastrointest Endosc 2026; 7(2): 118493
- URL: https://www.wjgnet.com/2689-7164/full/v7/i2/118493.htm
- DOI: https://dx.doi.org/10.37126/aige.118493
Patients with inflammatory bowel disease (IBD) are at an increased risk of developing colorectal neoplasia compared to the non-IBD population, making early and regular endoscopic surveillance necessary to prevent it. Endoscopic exami
The application of artificial intelligence (AI) including machine learning and deep learning techniques, such as convolutional neural networks, has revolutionized gastrointestinal (GI) endoscopy in the average-risk population by enhancing adenoma detection rates and facilitating standardized performance among endoscopists. AI has also been shown to be useful in detecting IBD-associated dysplasia because of the subtle nature of IBD-associated dysplasia and the difficulties of identifying dysplastic lesions in real-time during endoscopy[3,4].
However, many of the AI-based systems currently being tested and developed for the detection of dysplastic lesions in IBD patients were trained using data from non-IBD patients. Therefore, these AI-based systems are susceptible to dataset shift when used in IBD patients[3]. Training data may need to be tailored in order to improve performance. This can be done either through the development of IBD-specific AI-based models or multi-centers registries; however, continued performance disparities during external validation and among various subgroups of patients indicates that there are other unaddressed issues related to bias, generalizability, and equity.
This review addresses these limitations by comparing non-IBD-trained and IBD-specific AI for IBD neoplasia; by analysing the way in which dataset shift and multiple biases act as mechanisms of failure, and by outlining the technical, regulatory, and clinical steps required for AI to become a reliable adjunct in IBD-related colorectal cancer prevention.
This review narratively synthesizes 30 key studies on AI in IBD dysplasia (primarily 2019-2025, including earlier foundational works on dataset shift and annotation variability). Studies were author-selected from PubMed/PMC and GI journals by relevance. The narrative approach acknowledges inherent selection and publication biases.
Validation framework for AI in IBD must be able to evaluate the technical accuracy and robustness of AI across the diverse centers, devices and patient populations encountered in practice. Performance data comparing non-IBD-trained vs IBD-specific models demonstrate that the study design and validation strategy employed to evaluate AI strongly influences how well the model performs relative to human experts and non-experts.
Most AI systems for IBD-associated neoplasia were initially evaluated in retrospective studies that utilized existing endoscopic images or videos with histology as the reference standard. These retrospective designs allow for efficient testing on large datasets and detailed per-frame or per-lesion analysis, but they fail to represent artifacts, inadequate examinations, and real-time decision-making pressures present in routine surveillance, thereby potentially overstating performance[3-8].
Prospective trials and registry-based evaluations of AI in live colonoscopy assess the performance of AI in real-time with endoscopists and evaluate how AI functions within the workflow and how users interact with AI. Prospective trials and registry-based evaluations are fewer and more resource-intensive than retrospective trials, but they are closer to evaluating the clinically relevant performance of AI in routine surveillance, and demonstrate decreased performance when cases are more complex or inflammation is more severe than in the training data[8,9].
AI models trained on non-IBD cohorts typically do not perform well on new hospital datasets; therefore, external and multicenter validation is essential to test whether AI models are robust to variations in patient mix and technical factors. Multi-center initiatives, including federated IBD registries that aggregate data across regions and endoscopy platforms without collecting raw images centrally, have begun to address this limitation by providing exposure to broader distributions of phenotypes, imaging protocols, and demographics during both training and testing.
Regardless of study design, the performance of AI models is commonly assessed using established metrics such as sensitivity, specificity, positive and negative predictive value, and area under the receiver-operating characteristic curve, along with subgroup assessment[10-12].
The sensitivity of IBD neoplasia requires particular emphasis on both per-lesion and per-patient sensitivities of high-inflammation segments, post-resection areas and flat lesions because it is the lack of detection of neoplastic lesions in these locations during conventional surveillance that contributes significantly to the large number of missed cancers.
Comparative pooled analyses indicate that the non-IBD-trained AI models function significantly poorly in IBD compared to sporadic colorectal neoplasia[3,8]. Commercial non-IBD trained systems EndoBRAIN-Plus (98.7% sensitivity/77.1% specificity) and GI Genius (82% sensitivity/98% specificity) excel at sporadic neoplasia detection but lack IBD-specific data (Figure 1 and Table 1). In contrast, systems specifically trained for IBD-associated neoplasia have demonstrated variable performance across retrospective image-based validation studies (Tables 1 and 2)[9,11-14]. The highest reported detection sensitivity of 95.1% (specificity 98.8%) comes from Guerrero Vinsard et al[11]-a single-center retrospective study at Mayo Clinic Rochester using 1692 curated endoscopic images from 728 IBD patients. Abdelrahim et al[12] reported 93.5% lesion detection sensitivity and 80.6% specificity in a retrospective validation using 478 images from 30 patients and subsequently validated the model prospectively during live colonoscopy in 30 consecutive patients, achieving 87.5% sensitivity and 80.6% specificity for lesion characterization. Yamamoto et al[9] in a retrospective image-based pilot across two Japanese centers, reported 72.5% sensitivity and 82.9% specificity. While heterogeneity in datasets, quality control practices, inflammation severity, validation designs, and sample sizes precludes formal meta-comparisons, these metrics illustrate a consistent pattern: IBD-specific training substantially improves detection performance over generic polyp detection systems, yet IBD-trained models still exhibit marked variability across retrospective studies and remain largely untested in large prospective multicenter settings-underscoring the need for bias mitigation and rigorous real-world validation (next section).
| AI platform | Indication | Training dataset | IBD specific training | Sensitivity (%) | Specificity (%) | Population validated |
| EndoBRAIN-Plus[13] | Colorectal CADe | Approximately 68000 endocytoscopic images | No | 98.7 | 77.1 | Japan, Europe |
| GI genius[14] | Polyp detection | MMX trial videos (150 videos, 338 polyps) | No | 82 | 98 | United States |
| Vinsard IBD-CADe[11] | IBD neoplasia | Approximately 1700 IBD dysplasia images | Yes | 95.1 | 98.8 | United States |
| EfficientNet-B3[9] | IBD neoplasia | 862 IBDN images + data augmentation | Yes | 72.5 | 82.9 | Japan |
| Ref. | Study design | Model (training data) | Key result: Sensitivity (%)/specificity (%) | Dataset shift/bias findings | External validation |
| Guerrero Vinsard et al[11], 2023 | Retrospective | Non-IBD CADe, retrained IBD | 50/65 (non-IBD trained); 95.1/98.8 (IBD-trained) | Major drop in sensitivity with non-IBD, improved after retraining | Single centre (Mayo, Rochester); United States |
| Yamamoto et al[9], 2022 | Retrospective | IBD-trained CNN (EffNet-B3) | 72.5/82.9 | Generalizability challenges despite some gains | Multicentre; Japan |
| Abdelrahim et al[12], 2024 | Retrospective + prospective | DL hybrid IBD-centric | Detection: 93.5/80.6 (Retrospective). Characterisation: 87.5/80.6 (prospective) | Prospective real-time IBD-AI validation study; small sample limits generalizability | Multicentre; United Kingdom |
Depending upon lesion type, imaging modality and validation study design, these improvements bring but do not completely close the gap between IBD-trained model performance and expert endoscopist performance, especially after subjecting models to external validation and testing on challenging real-world cases. Expert IBD endoscopists generally report a sensitivity of approximately 60.5%, specificity of about 88%, and accuracy of 77.8%, whereas non-expert endo
AI failures in IBD neoplasia stem primarily from dataset shift and bias because both create gaps between the data used for model training and the data encountered by models in the real-world during surveillance.
In this review, we distinguish dataset shift from bias. Dataset shift refers to a mismatch between the distributions of data used for training and those encountered at deployment (for example, models trained on clean, non-IBD mucosa being applied to chronically inflamed, scarred IBD colons). Bias, by contrast, reflects systematic errors that disproportionately affect subgroups (for example, under-representation of severely inflamed segments or certain demographic groups in the training data), and is discussed in the next section.
Building on this distinction, dataset shift in IBD neoplasia typically reflects the underrepresentation of chronically inflamed mucosa, pseudopolyps, post-resection scars, and complex anatomy in generic polyp datasets, as well as differences in endoscopic equipment, chromoendoscopy techniques, and patient demographics. Therefore, when models trained on clean, non-IBD colon images encounter novel visual patterns in IBD, the sensitivity for detecting IBD-associated neoplastic lesions decreases[3,14,15].
However, shift is not solely visual. The training datasets of many systems contain disproportionate numbers of patients from certain tertiary centres located in specific geographic areas, with limited representation of paediatric, elderly, and/or underrepresented ethnic groups, thus creating significant disparities in the phenotypic and demographic profiles at deployment. Thus, the combination of technical, anatomical, and population-level shift implies that the performance metrics reported in the studies where models were developed frequently will not translate to broader IBD populations when systems are implemented in those populations.
Dataset shift intersects with multiple types of biases present in AI models designed for IBD neoplasia. Selection bias exists when training sets favourably select “well-visualized” “clean” lesions and exclude examinations characterized by severe inflammation, poor bowel preparation, or previous resections[16]. These exclusions cause models to focus on relatively easier and less clinically relevant cases rather than the most challenging and clinically relevant cases. Annotation bias results from inter-rater variability among endoscopists and pathologists labelling flat or subtle dysplasia, adding random error to the ground truth, and thereby reducing the maximum achievable performance[17]. Device and vendor bias exist when models are primarily trained on images collected from a single endoscope platform, processor, or imaging mode[3]. When models trained on a single platform are subsequently evaluated using different platforms, the accuracy of the model often drops, particularly for colour-sensitive or contrast-sensitive features. Finally, reporting and publication bias contribute to the distorted evidence base by increasing the likelihood of publishing positive studies demonstrating high-performance models rather than neutral or negative studies, thus overestimating the capability of current AI systems in IBD[18]. These identified biases cause a discrepancy in training vs deployment and thus limit the ability for single center models to perform when being externally validated.
Clinical effects of both dataset shift and bias are significant. The performance of non-IBD-trained models on IBD cohorts have been found to have lower sensitivities and specificities compared to the much higher values obtained in sporadic adenomas, and the performance has been shown to be inversely related to the level of inflammatory activity within the cohort. As a result, there is an increase in the number of false negatives-missed dysplasia or early cancer in those patients who have the heaviest levels of inflammation or scarring, and false positives-those patients whose regenerative or inflammatory changes are misclassified as neoplastic and thus subjected to unnecessary biopsy, resection or intensified surveillance.
It is highly likely that these errors will not occur uniformly. Those under-represented phenotypic and demographic subgroups represented in the training data will be at greater risk of being misclassified, resulting in concerns regarding unequal outcomes if AI is to be used to determine surveillance intervals or make therapeutic decisions for patients with IBD. Clinicians repeatedly encountering unstable or biased output of an AI system will lose trust in the AI system, and either rely entirely upon recommendations generated by the AI system or completely avoid using AI systems in clinical practice; both scenarios will undermine the intended goal of AI to augment expert judgment as a calibrated tool.
A comprehensive roadmap addressing the limitations of dataset shift in the context of IBD neoplasia must couple technological innovation with regulatory framework and practical clinical workflow[19,20]. Rather than addressing validation, bias, and deployment as independent issues, the development of future AI systems must incorporate them into an integrated and continuous learning cycle beginning with data curation, followed by deployment and concluded with post-deployment monitoring of AI system performance (Figure 2).
Technically, the most effective way to address the limitations of dataset shift in IBD neoplasia is to enhance the diversity of the data[16,20]. Therefore, large-scale IBD-specific image databases which contain a variety of images of severely inflamed bowel, pseudopolyps, post-surgical fields, paediatric and geriatric cases and under-represented population groups are necessary to train AI models that can handle the diversity seen in real-world clinical settings[16,21]. International collaborations among multiple centres using federated learning methodologies allow algorithms to be trained on multiple datasets without sharing raw images and provide additional information to training distributions while maintaining patient privacy and adhering to local governance regulations[22]. However, federated learning introduces technical complexity including data heterogeneity across sites, requirements for standardized protocols, and infrastruc
Methods designed to counteract the limitations of dataset shift should be developed and implemented. For example, transfer learning (adapting pre-trained AI from non-IBD data to IBD) combined with generative adversarial networks (GANs) to create synthetic images of under-represented phenotypic subgroups may help to minimize differences in sample size between groups and improve model robustness to image appearance variations. However, GAN-generated images carry risks of mode collapse, label overfitting, and potential amplification of source dataset biases; critically, visually realistic synthetic images do not always improve downstream model performance and require task-specific validation rather than perceptual quality metrics alone[25]. Multimodal models which utilize both endoscopic images and clinical, laboratory and histological data to classify neoplasia may help to stabilize AI performance by utilizing other data types when images are poor or ambiguous. However, real-world clinical data is noisy, incomplete, and heterogeneous with inconsistent coding standards, and missing data handling and complex data integration infrastructure present barriers to deployment[26-28]. These methods must be validated using predetermined external validation protocols and calibration analyses and subgroup performance metrics to identify residual biases rather than masking them[19].
Generic guidelines for AI use must be translated into IBD-specific standards[29]. Professional organizations and regula
The approval pathway for GI Genius (Food and Drug Administration De Novo DEN200055, 2020)[14] illustrates the evidence burden for polyp detection AI in non-IBD populations and provides a template for IBD-specific systems. Food and Drug Administration clearance required: (1) Technical validation demonstrating 82% per-polyp sensitivity with object-level detection performance benchmarked against average endoscopist reaction time (GI Genius detected polyps 1270 ms earlier on average); (2) Risk analysis addressing software hazards, electromagnetic safety, and cybersecurity (moderate concern level); and (3) Explicit restriction to “prescription use” and labelling stating the device is “not intended as a stand-alone diagnostic device”. Similarly, European Union Conformity marking under the medical device regulation requires technical documentation demonstrating conformity with General Safety and Performance Requirements, supported by clinical evidence aligned with Annex XIV. For class IIa devices (including most computer-aided detection CADe systems), manufacturers must demonstrate clinical benefit-either through prospective clinical investigations or, under Article 61 (10)[30] through adequate justification that non-clinical evidence (performance evaluation, bench testing, pre-clinical evaluation) suffices when clinical data is deemed inappropriate.
Translating these requirements to IBD neoplasia, we propose-as a conservative starting framework informed by sporadic polyp CADe precedents and IBD surveillance complexity-that IBD-specific CADe systems seeking regulatory approval could reasonably be expected to meet the following minimum evidence thresholds.
Publicly documented dataset composition including number of patients, images, centres, prevalence of flat dysplasia (Paris IIb/IIc), proportion of actively inflamed segments (Mayo endoscopic subscore ≥ 2), post-resection cases, and demographic distribution.
Sensitivity ≥ 90% and specificity ≥ 85% across stratified subgroups (inflammation severity, lesion morphology, disease extent) in a multi-centre retrospective cohort (≥ 500 patients across ≥ 3 centres using ≥ 2 endoscope platforms, this threshold exceeds most single-centre IBD-AI datasets including Guerrero Vinsard et al[11] (n = 728, 1 centre) while ensuring multicentre and device heterogeneity essential for external validation).
Per-lesion sensitivity ≥ 85% in a prospective trial (≥ 100 consecutive patients undergoing IBD surveillance colonoscopy; this represents a minimum cohort size meaningful for subgroup stratification by inflammation severity, exceeding the only published prospective IBD-AI validation Abdelrahim et al[12], n = 30 by more than threefold) with transparent reporting of false-negative cases, stratified by Mayo endoscopic subscore and lesion morphology.
Explicit statement that AI functions as an adjunctive tool requiring expert verification, not as a stand-alone diagnostic or biopsy decision-maker (mirroring GI Genius Food and Drug Administration-mandated labelling).
Mandatory performance monitoring dashboards tracking sensitivity/specificity drift, false-positive rates, and subgroup disparities, with predefined thresholds triggering model retraining (e.g. sensitivity decline > 5 percentage points, chosen as approximately half the observed inter-study variability 72.5%-95.1%) or withdrawal from market.
These thresholds are deliberately conservative and intended to stimulate discussion rather than prescribe rigid standards, given that IBD surveillance involves higher-stakes decisions (surveillance interval determination, colectomy consideration) than sporadic polyp screening, and that dataset shift in inflamed, scarred colons poses greater risk of harm through missed high-grade dysplasia. Definitive thresholds will require consensus development through collaborative efforts among regulatory agencies (Food and Drug Administration, European Medicines Agency), professional societies (e.g. American College of Gastroenterology, European Crohn’s and Colitis Organisation, British Society of Gastroenterology, European Society of Gastrointestinal Endoscopy, Gastroenterological Society of Australia), and IBD-AI investigators.
Ethical frameworks for AI use must emphasize informed consent for AI-augmented surveillance, clarify responsibility for AI-supported decisions, and audit routinely for differential performance across demographic and disease subgroups[31-33].
AI for IBD neoplasia must be implemented as an adjunctive tool to be utilized within structured clinical workflows, and not as an autonomous decision-making process. Current clinically validated implementations deploy AI primarily as a real-time detection overlay that displays visual alerts, such as bounding boxes highlighting suspicious regions, or auditory signals when potential dysplasia is identified, prompting the endoscopist to pause for targeted inspection or biopsy. Educational programs for clinicians must include training on how to operate AI tools, how to interpret confidence estimates, how to recognize failure modes associated with dataset shift, and how to properly override incorrect outputs.
However, clinical implementation must anticipate and mitigate potential workflow disruptions. Alert fatigue, where excessive false-positive notifications lead endoscopists to ignore or distrust AI outputs and become desensitized to genuine alerts, has been documented with CADe systems in sporadic polyp detection and represents a significant concern for clinical deployment[33-35]. This “crying wolf effect” may be exacerbated in IBD surveillance where chroni
Institutions must also establish and maintain dashboards to track the performance of AI in their respective patient populations, including stratified metrics by inflammation severity, disease extent, and relevant demographic factors. In the long run, the path forward from dataset shift is self-reinforcing: The quality and breadth of labelled data inform the quality and awareness of bias in algorithms; the rigor of multicentre testing and regulatory oversight provide assurance of their reliability; and the real-world monitoring provides feedback to continue improving them.
This review emphasizes that AI success in detecting IBD neoplasia depends not only on novel architectures but also on how models are trained, validated, and regulated when confronted with dataset shift. Systems developed on non-IBD data exhibit substantially reduced sensitivity and specificity in inflamed and scarred colon compared to IBD-specific models, which improve performance but do not completely mitigate the risk of external validation and real-world variability. The current performance of AI systems demonstrates that AI could potentially serve as a tool to help less-experienced operators detect lesions equally as well as expert operators but not replace IBD specialists. Unlike prior studies that catalogued algorithms and pooled metrics, this review positioned dataset shift and bias within validation design and clinical implementation, synthesizing how training distributions, annotation practices, device dependency, and publication biases interact to produce failure mechanisms. It synthesized these disparate threads into a road map that spans technical, regulatory, and workflow-based domains, shifting focus from “how well does AI find lesions” to “under what circumstances, for whom, and with what safeguards can AI safely support IBD cancer prevention”.
The strengths of the current evidence include increasing multicentre datasets, emerging IBD-specific models, and demonstrations of federated learning and multimodal architectures. However, most studies utilize relatively small, highly selected collections of images that poorly represent severe inflammation, post-resection anatomy, and rare phenotypes resulting in inflated estimates of performance. Few prospective real-time assessments exist, publication bias toward positive outcomes likely creates overly optimistic views, and poor documentation of negative experiences impedes reproducibility and tool comparison. These gaps create uncertainties regarding generalizability and equity. Underrepresented patient subgroups and disease phenotypes face highest misclassification risk, yet subgroup perfor
The substantial performance variability across IBD-specific AI (72.5% to 95.1% sensitivity)[9-11]-reflects methodological differences including dataset scale and curation (Guerrero Vinsard et al[11]: 1692 quality-controlled images; Yamamoto et al[9]: 862 images; Abdelrahim et al[12]: 478 images), institutional homogeneity (single-centre with standar
Explicit deployment criteria are necessary given current evidence limitations to prevent premature implementation of AI. Unsafe applications include: (1) Autonomous surveillance or biopsy decisions without expert verification; (2) Deploy
To bridge evidence gaps, the field requires large, curated, multicentre IBD repositories with standardized protocols and consensus annotation. Future research should prioritize multimodal models, robust external validation pipelines, and mandatory stratified performance reporting aligning with the roadmap presented above. All validation studies and regulatory submissions should report sensitivity, specificity, and calibration across: (1) Demographic subgroups-age categories (paediatric/adult/elderly), gender, ethnicity, geography; (2) Disease phenotypes-IBD type, extent, inflammation severity, post-surgical anatomy, disease duration; and (3) Technical factors-endoscope platform, imaging modality. Subgroup analyses should be pre-specified and reported regardless of significance to enable transparent differential performance assessment. Finally, international organizations and regulatory agencies must converge on global reporting standards, minimum evidence thresholds, and expectations for continued monitoring and retraining, enabling transparent learning healthcare system development rather than individualized tools. Ongoing prospective multicentre trials (Table 3) may address these gaps by evaluating AI performance in diverse clinical settings[36,37]. Collectively, these efforts would position AI as a calibrated adjunct enhancing colorectal cancer prevention while acknowledging its limitations.
| Trial name | Design | Target condition | Sites | Status |
| Artificial intelligence in IBD-related dysplasia (AID) study[38] | Prospective | IBD dysplasia detection | Multicentre (Australia) | Recruiting |
| Artificial intelligence and dysplasia detection in inflammatory bowel disease (EIIDISIA study)[39] | Randomized clinical study | CADe vs Chromoendoscopy in IBD dysplasia detection | Multicentre (Spain) | Recruiting |
AI holds significant potential to enhance dysplasia detection and colorectal cancer prevention in IBD, yet current systems remain constrained by dataset shift, bias, and predominance of retrospective validations using curated datasets. The substantial performance variability across IBD-specific models and near-absence of large prospective multicentre trials necessitate explicit, conservative deployment criteria that distinguish currently unsafe applications from conditionally acceptable supervised use. To progress toward trustworthy clinical integration, the field requires large diverse multi
| 1. | De Cristofaro E, Marafini I, Franchin M, Venuto C, Savino L, Lolli E, Sena G, Neri B, Zorzi F, Troncone E, Biancone L, Del Vecchio Blanco G, Orlandi A, Calabrese E, Monteleone G. Frequency of dysplasia in endoscopically resected pseudopolyps in inflammatory bowel diseases. J Crohns Colitis. 2025;19:jjaf196. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 5] [Reference Citation Analysis (0)] |
| 2. | Velayos F. Best Practices for Dysplasia Detection, Surveillance and Management in IBD. Pract Gastroenterol. 2024;48:38-42, 48. |
| 3. | Urquhart SA, Christof M, Coelho-Prabhu N. The impact of artificial intelligence on the endoscopic assessment of inflammatory bowel disease-related neoplasia. Therap Adv Gastroenterol. 2025;18:17562848251348574. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 11] [Cited by in RCA: 9] [Article Influence: 9.0] [Reference Citation Analysis (0)] |
| 4. | Tariq R, Afzali A. Artificial intelligence in inflammatory bowel disease: innovations in diagnosis, monitoring, and personalized care. Therap Adv Gastroenterol. 2025;18:17562848251357407. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 7] [Reference Citation Analysis (1)] |
| 5. | Stidham RW, Liu W, Bishu S, Rice MD, Higgins PDR, Zhu J, Nallamothu BK, Waljee AK. Performance of a Deep Learning Model vs Human Reviewers in Grading Endoscopic Disease Severity of Patients With Ulcerative Colitis. JAMA Netw Open. 2019;2:e193963. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 256] [Cited by in RCA: 209] [Article Influence: 29.9] [Reference Citation Analysis (6)] |
| 6. | Fan Y, Mu R, Xu H, Xie C, Zhang Y, Liu L, Wang L, Shi H, Hu Y, Ren J, Qin J, Wang L, Cai S. Novel deep learning-based computer-aided diagnosis system for predicting inflammatory activity in ulcerative colitis. Gastrointest Endosc. 2023;97:335-346. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 50] [Cited by in RCA: 41] [Article Influence: 13.7] [Reference Citation Analysis (0)] |
| 7. | Luo X, Zhang J, Li Z, Yang R. Diagnosis of ulcerative colitis from endoscopic images based on deep learning. Biomed Signal Proces Contr. 2022;73:103443. [RCA] [DOI] [Full Text] [Cited by in Crossref: 6] [Cited by in RCA: 24] [Article Influence: 6.0] [Reference Citation Analysis (0)] |
| 8. | Liu X, Reigle J, Prasath VBS, Dhaliwal J. Artificial intelligence image-based prediction models in IBD exhibit high risk of bias: A systematic review. Comput Biol Med. 2024;171:108093. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 15] [Cited by in RCA: 13] [Article Influence: 6.5] [Reference Citation Analysis (1)] |
| 9. | Yamamoto S, Kinugasa H, Hamada K, Tomiya M, Tanimoto T, Ohto A, Toda A, Takei D, Matsubara M, Suzuki S, Inoue K, Tanaka T, Hiraoka S, Okada H, Kawahara Y. The diagnostic ability to classify neoplasias occurring in inflammatory bowel disease by artificial intelligence and endoscopists: A pilot study. J Gastroenterol Hepatol. 2022;37:1610-1616. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 32] [Cited by in RCA: 34] [Article Influence: 8.5] [Reference Citation Analysis (0)] |
| 10. | Dos Santos CEO, Malaman D, Sanmartin IDA, Leão ABS, Leão GS, Pereira-Lima JC. Performance of artificial intelligence in the characterization of colorectal lesions. Saudi J Gastroenterol. 2023;29:219-224. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 2] [Cited by in RCA: 17] [Article Influence: 5.7] [Reference Citation Analysis (0)] |
| 11. | Guerrero Vinsard D, Fetzer JR, Agrawal U, Singh J, Damani DN, Sivasubramaniam P, Poigai Arunachalam S, Leggett CL, Raffals LE, Coelho-Prabhu N. Development of an artificial intelligence tool for detecting colorectal lesions in inflammatory bowel disease. iGIE. 2023;2:91-101.e6. [RCA] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 31] [Cited by in RCA: 30] [Article Influence: 10.0] [Reference Citation Analysis (0)] |
| 12. | Abdelrahim M, Siggens K, Iwadate Y, Maeda N, Htet H, Bhandari P. New AI model for neoplasia detection and characterisation in inflammatory bowel disease. Gut. 2024;73:725-728. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 26] [Cited by in RCA: 26] [Article Influence: 13.0] [Reference Citation Analysis (0)] |
| 13. | Mori Y, Kudo SE, Misawa M, Hotta K, Kazuo O, Saito S, Ikematsu H, Saito Y, Matsuda T, Kenichi T, Kudo T, Nemoto T, Itoh H, Mori K. Artificial intelligence-assisted colonic endocytoscopy for cancer recognition: a multicenter study. Endosc Int Open. 2021;9:E1004-E1011. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 23] [Cited by in RCA: 20] [Article Influence: 4.0] [Reference Citation Analysis (0)] |
| 14. | De novo classification request for GI genius. [cited 27 February 2026]. Available from: https://www.accessdata.fda.gov/cdrh_docs/reviews/DEN200055.pdf. |
| 15. | Moreno-Torres JG, Raeder T, Alaiz-Rodríguez R, Chawla NV, Herrera F. A unifying view on dataset shift in classification. Pattern Recogn. 2012;45:521-530. [DOI] [Full Text] |
| 16. | Lodola I, D'Amico F, Danese S, Parigi TL. Artificial intelligence in inflammatory bowel disease endoscopy - a review of current evidence and a critical perspective on future challenges. Therap Adv Gastroenterol. 2025;18:17562848251350896. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 3] [Cited by in RCA: 5] [Article Influence: 5.0] [Reference Citation Analysis (0)] |
| 17. | Lee YC, Cook MB, Bhatia S, Chow WH, El-Omar EM, Goto H, Lin JT, Li YQ, Rhee PL, Sharma P, Sung JJ, Wong JY, Wu JC, Ho KY; Asian Barrett's Consortium. Interobserver reliability in the endoscopic diagnosis and grading of Barrett's esophagus: an Asian multinational study. Endoscopy. 2010;42:699-704. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 53] [Cited by in RCA: 53] [Article Influence: 3.3] [Reference Citation Analysis (1)] |
| 18. | Shi Y, Wei N, Wang K, Tao T, Yu F, Lv B. Diagnostic value of artificial intelligence-assisted endoscopy for chronic atrophic gastritis: a systematic review and meta-analysis. Front Med (Lausanne). 2023;10:1134980. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 23] [Cited by in RCA: 20] [Article Influence: 6.7] [Reference Citation Analysis (1)] |
| 19. | Ahmad HA, East JE, Panaccione R, Travis S, Canavan JB, Usiskin K, Byrne MF. Artificial Intelligence in Inflammatory Bowel Disease Endoscopy: Implications for Clinical Trials. J Crohns Colitis. 2023;17:1342-1353. [RCA] [PubMed] [DOI] [Full Text] [Cited by in RCA: 36] [Reference Citation Analysis (1)] |
| 20. | Ahmed M, Stone ML, Stidham RW. Artificial Intelligence and IBD: Where are We Now and Where Will We Be in the Future? Curr Gastroenterol Rep. 2024;26:137-144. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 13] [Cited by in RCA: 14] [Article Influence: 7.0] [Reference Citation Analysis (1)] |
| 21. | Testoni SGG, Albertini Petroni G, Annunziata ML, Dell'Anna G, Puricelli M, Delogu C, Annese V. Artificial Intelligence in Inflammatory Bowel Disease Endoscopy. Diagnostics (Basel). 2025;15:905. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 8] [Reference Citation Analysis (1)] |
| 22. | Puca P, Lopetuso LR, Laterza L, Papa A, Danese S, Cesario A, Damiani A, Gasbarrini A, Arcuri G, Scaldaferri F. Federated learning in inflammatory bowel disease: The future of privacy-preserving Artificial Intelligence. Best Pract Res Clin Gastroenterol. 2025;78:102050. [RCA] [PubMed] [DOI] [Full Text] [Cited by in RCA: 1] [Reference Citation Analysis (0)] |
| 23. | Rieke N, Hancox J, Li W, Milletarì F, Roth HR, Albarqouni S, Bakas S, Galtier MN, Landman BA, Maier-Hein K, Ourselin S, Sheller M, Summers RM, Trask A, Xu D, Baust M, Cardoso MJ. The future of digital health with federated learning. NPJ Digit Med. 2020;3:119. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 2283] [Cited by in RCA: 1059] [Article Influence: 176.5] [Reference Citation Analysis (5)] |
| 24. | Lobanovs S, Aleksejeva J, Rūtiņa AK, Krustiņš E, Čižovs J, Bļizņuks D. Machine learning in gastrointestinal endoscopy: challenges and opportunities. BMJ Open Gastroenterol. 2025;12:e001923. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 2] [Reference Citation Analysis (0)] |
| 25. | Skandarani Y, Jodoin PM, Lalande A. GANs for Medical Image Synthesis: An Empirical Study. J Imaging. 2023;9:69. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 73] [Cited by in RCA: 66] [Article Influence: 22.0] [Reference Citation Analysis (0)] |
| 26. | Wu YM, Tang FY, Qi ZX. Multimodal artificial intelligence technology in the precision diagnosis and treatment of gastroenterology and hepatology: Innovative applications and challenges. World J Gastroenterol. 2025;31:109802. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in CrossRef: 17] [Cited by in RCA: 13] [Article Influence: 13.0] [Reference Citation Analysis (0)] |
| 27. | Zhang L, Shen Y, Gu W, Wu P. Multimodal learning in gastrointestinal diseases. Gastroentero Endosc. 2025;3:251-258. [DOI] [Full Text] |
| 28. | Wasi AT, Ridoy SZ. Position: Real-World Clinical AI Requires Multimodal, Longitudinal, and Privacy-Preserving Corpora. 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: EurIPS Wprkshop on Multimodal Representation Learning for Healthcare. [cited 27 February 2026]. Available from: https://multimodal-rep-learning-for-health.github.io/papers/5_Position_Real_World_Clinical.pdf. |
| 29. | Noor NM, Daperno M, de Laffolie J, Mookhoek A, Sinonquel P, Kopylov U, Verstockt B, El-Hussuna A, Sahnan K, Allocca M, Bossuyt P, Carter D, Ensari A, Iacucci M, Marigorta UM, Noviello D, Pellino G, Soriano A, Cleynen I, Raine T, Sebastian S, Baumgart DC. Results of the 9th Scientific Workshop of the European Crohn's and Colitis Organisation (ECCO): Artificial Intelligence in IBD: Regulatory and Methodological Considerations. J Crohns Colitis. 2025;jjaf136. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 3] [Cited by in RCA: 5] [Article Influence: 5.0] [Reference Citation Analysis (0)] |
| 30. | Kearney B, Partridge S. Article 61.10. Clinical Evaluation based on non-clinical data. [cited 27 February 2026]. Available from: https://www.bsigroup.com/siteassets/pdf/en/insights-and-media/insights/white-papers/bsi-md-article-61.10-clinical-evaluation-whitepaper-en-gb.pdf. |
| 31. | Ahmad OF, Mori Y, Bretthauer M, Dourado DA, Hassan C, Bisschops R, Bhandari P, Byrne MF, Dekker E, Mahadevan U, May FP, Messmann H, Misawa M, Ogata H, Saito Y, Silverman AL, Wang P, Yano T, Aabakken L, Berzin TM. The Legal and Ethical Framework for Artificial Intelligence in Gastrointestinal Endoscopy: A World Endoscopy Organization International Consensus Statement. Ann Intern Med. 2026;179:270-275. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 2] [Cited by in RCA: 7] [Article Influence: 7.0] [Reference Citation Analysis (0)] |
| 32. | Ramoni D, Scuricini A, Carbone F, Liberale L, Montecucco F. Artificial intelligence in gastroenterology: Ethical and diagnostic challenges in clinical practice. World J Gastroenterol. 2025;31:102725. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in CrossRef: 17] [Cited by in RCA: 12] [Article Influence: 12.0] [Reference Citation Analysis (0)] |
| 33. | Hsieh YH, Tang CP, Tseng CW, Lin TL, Leung FW. Computer-Aided Detection False Positives in Colonoscopy. Diagnostics (Basel). 2021;11:1113. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 12] [Cited by in RCA: 18] [Article Influence: 3.6] [Reference Citation Analysis (1)] |
| 34. | Chung GE, Lee J, Lim SH, Kang HY, Kim J, Song JH, Yang SY, Choi JM, Seo JY, Bae JH. A prospective comparison of two computer aided detection systems with different false positive rates in colonoscopy. NPJ Digit Med. 2024;7:366. [RCA] [PubMed] [DOI] [Full Text] [Cited by in RCA: 17] [Reference Citation Analysis (0)] |
| 35. | Barua I, Bretthauer M, Mori Y. Computer-aided quality assessment (CAQ): the next step for artificial intelligence in colonoscopy? Mini-invasive Surg. 2022;6:28. [DOI] [Full Text] |
| 36. | Areia M, Mori Y, Correale L, Repici A, Bretthauer M, Sharma P, Taveira F, Spadaccini M, Antonelli G, Ebigbo A, Kudo SE, Arribas J, Barua I, Kaminski MF, Messmann H, Rex DK, Dinis-Ribeiro M, Hassan C. Cost-effectiveness of artificial intelligence for screening colonoscopy: a modelling study. Lancet Digit Health. 2022;4:e436-e444. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 201] [Cited by in RCA: 167] [Article Influence: 41.8] [Reference Citation Analysis (9)] |
| 37. | Barkun AN, von Renteln D, Sadri H. Cost-effectiveness of Artificial Intelligence-Aided Colonoscopy for Adenoma Detection in Colon Cancer Screening. J Can Assoc Gastroenterol. 2023;6:97-105. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 36] [Cited by in RCA: 36] [Article Influence: 12.0] [Reference Citation Analysis (3)] |
| 38. | Current IBD Research Projects. [cited 27 February 2026]. Available from: https://www.thegutsygroup.com.au/research/current-ibd-research-projects. |
| 39. | López-Serrano A. Artificial Intelligence and Dysplasia Detection in Inflammatory Bowel Disease (EIIDISIA Study). [cited 27 February 2026]. Clinical Research Trial Listing. Centerwatch.com. 2025. In: ClinicalTrials.gov [Internet]. Bethesda (MD): U.S. National Library of Medicine. Available from: https://clinicaltrials.gov/study/NCT06281392 ClinicalTrials.gov Identifier: NCT06281392. |