©The Author(s) 2025.
World J Gastroenterol. Oct 21, 2025; 31(39): 111353
Published online Oct 21, 2025. doi: 10.3748/wjg.v31.i39.111353
Published online Oct 21, 2025. doi: 10.3748/wjg.v31.i39.111353
Table 1 Role of artificial intelligence-based endoscopy in the evaluation of patients with inflammatory bowel diseases
| Ref. | Disease/number of patients | Type of study | Endoscopic technique | Number of training samples | Number of test samples | AI/model | Main findings |
| Stidham et al[22] | UC/3082 patients | Retrospective, single center | WLE | 14862 images | 1652 images | DL-CNN | Discriminating ER (MES ≤ 1) from moderate-severe disease (MES ≥ 2) (AUC = 0.966, sensitivity = 83.0%, specificity = 96.0%). AI and pathologist agreement (κ = 0.840 vs κ = 0.860) |
| Maeda et al[38] | UC/187 patients | Retrospective, single center | Endocytoscopy | 12900 still images | 525 segments | CAD | Prediction of HR (GS ≥ 3.1) (sensitivity = 74.0%, specificity = 97.0%, precision = 91.0%, κ = 1.000) |
| Ozawa et al[40] | UC/955 patients | Retrospective, single center | WLE | 26304 still images | 3981 still images | CAD-CNN | AI performance for mucosal healing (MES ≤ 1, AUC = 0.980) |
| Takenaka et al[25] | UC/875 patients | Prospective, single center | WLE | 40789 still images | 4187 still images | DNUC | Evaluation of ER (UCEIS ≤ 2) (accuracy = 90.1%, ICC = 0.917). Evaluation of HR (GS < 3.1) (accuracy = 92.9%, κ = 0.859) |
| Bossuyt et al[41] | UC/35 patients | Prospective, multicenter | Prototype endoscope | NR | NR | CAD | RD for endoscopic/histological inflammation: Correlation with MES (r = 0.76), UCEIS (r = 0.74), RHI (r = 0.74). RD score (≤ 60) predicts HR (AUC = 0.950, sensitivity = 96.0%, specificity = 80.0%) |
| Yao et al[27] | UC/157 patients | Prospective, multicenter | WLE | NR | 264 videos of high resolution | DL-CNN | The still image informative classifier had excellent performance (sensitivity = 0.902, specificity = 0.870). Correct prediction of MES: 78% of videos (κ = 0.840) |
| Gottlieb et al[30] | UC/249 patients | Prospective, multicenter | WLE | 629 videos | 157 videos | RNN | Endoscopic healing evaluation according to UCEIS (accuracy = 97.0%) and MES (accuracy = 95.5%). Agreement of the model with human experts for MES (QWK = 0.844) and UCEIS (QWK = 0.855) |
| Bossuyt et al[42] | UC/58 patients | Prospective, single center | SWE | NR | 113 still images | CAD | AI algorithm yielded better HR accuracy (86.0%) than MES (74.0%) or UCEIS (79.0%) |
| Huang et al[43] | UC/54 patients | Retrospective, single center | Endoscopy HD | 600 still images | 256 still images | DNN, SVM, k-NN | Performance of the combined model for differentiation between MES ≤ 1 and MES 2 (AUC = 0.927, accuracy = 94.5%, sensitivity = 89.2%, specificity = 96.3%) |
| Takenaka et al[44] | UC/770 patients | Prospective, multicenter | WLE | NR | NR | DNUC | Prediction of HR (sensitivity = 97.9%, specificity = 94.6%). Agreement between the DNUC and experts for endoscopic assessment (ICC = 0.927) |
| Patel et al[45] | UC/73 patients | Prospective, single center | Endoscopy HD | 55 video images | 18 video images | MLA | Differentiation between: Remission (UCEIS: 0-1) and active inflammation (UCEIS ≥ 2) (accuracy = 90.0%, κ = 0.900); Mild (UCEIS: 2-3); And moderate-to-severe inflammation (UCEIS ≥ 4) (accuracy = 98.0%, κ = 0.960) |
| Kim et al[19] | UC/492 patients | Retrospective, single center | WLE | 904 still images | 80 still images | DL-CNN | Difference between MES 0 vs MES 1: Internal test. IBD experts (F1 score = 0.92, AUC = 0.970); External test. Hyper Kvasir dataset (F1 score = 0.89, AUC = 0.860) |
| Polat et al[20] | UC/564 patients | Retrospective, single center | WLE | 11276 still images | 1658 still images | DL-CNN | Excellent concordance between the five CNN networks and endoscopists for: MES evaluation (QWK: 0.847-0.854); And classification of remission cases (QWK: 0.834-0.852) |
| Wang et al[21] | UC/308 patients | Retrospective, single center | WLE | 37515 still images | 3191 still images | CNN | Diagnosis of ER (MES ≤ 1) (AUC = 0.980, accuracy = 95.1%, sensitivity = 92.9%, specificity = 95.4%, κ = 0.884) |
| Iacucci et al[46] | UC/283 patients | Prospective, multicenter | WLE, VCE | 239 video images; 245 video images | 242 video images; 244 video images | CNN | Detection of ER using VCE (PICaSSO ≤ 3) (AUC = 0.940, sensitivity = 79.0%, specificity = 95.0%, κ = 0.730) achieved better performance than WLE (UCEIS ≤ 1) (AUC = 0.850, sensitivity = 72.0%, specificity = 87.0%, κ = 0.510) |
| Byrne et al[47] | UC/NR | Prospective, single center | HD endoscopy | 134 video images | NR | DL-CNN | Performance for disease severity discrimination: MES ≤ 1 vs MES ≥ 2 (AUC = 0.941, accuracy = 94.0%, sensitivity = 96.7%, specificity = 91.3%, QWK = 0.880); UCEIS ≤ 3 vs UCEIS > 3 (AUC = 0.936, accuracy = 94.0%, sensitivity = 93.9%, specificity = 93.4%, QWK = 0.870) |
| Stidham et al[31] | UC/748 patients | Prospective, multicenter | WLE | NR | NR | ML | CDS had better performance for detecting endoscopic changes than MES (Hedges’ g: 0.743 vs 0.460, P < 0.001) |
| Takabayashi et al[48] | UC/812 patients | Retrospective, multicenter | WLE | 14208 still images | 13826 still images | CNN | Disease severity grading-correlation between: UCEGS and MES (ρ = 0.890, P < 0.001); UCEGS and IBD experts (ρ: 0.960-0.987, P < 0.001) |
| Ogata et al[49] | UC/110 patients | Prospective, single center | WLE | 74713 still images | 11452 still images | CNN | Performance for evaluation ER based MES (sensitivity = 96.9%, specificity = 78.4%, accuracy = 93.4%). Interobserver/intraobservator agreement with AI/without AI (ICC: 0.84-0.86/0.89 vs 0.64-0.76/0.76) |
| Sinonquel et al[50] | UC/36 patients | Prospective, single center | SWE | NR | NR | CAD | Histological assessment using SWE-CAD (sensitivity = 96.1%, specificity = 85.5%, accuracy = 96.4%). The accuracy of classification into mild, moderate, and severe disease was 97.7%, 62.8% and 95.0%, respectively |
| Aoki et al[51] | CD/131 patients | Retrospective, single center | CE | 5360 images | 10440 images | CNN | Ulcer recognition in small bowel video frames (AUC = 0.958, sensitivity = 88.2%, specificity = 90.9%, accuracy = 90.8%) |
| Klang et al[52] | CD/49 patients | Retrospective, single center | CE | 14112 images | 3528 images | DL-CNN | Increased performance in ulcer detection (AUC = 0.990, accuracy: 95.4%-96.7%) |
| Klang et al[32] | CD/145 patients | Retrospective, single center | CE | 27892 images | 1449 images | DNN | Performance for: Stricture detection (AUC = 0.971, accuracy = 93.5%); Differential diagnosis between strictures and normal mucosa (AUC = 0.989); Discrimination between strictures and ulcers (AUC = 0.942) |
| Barash et al[53] | CD/49 patients | Retrospective, single center | CE | 1242 images | 248 images | CNN | Ability of ulcerative lesion classification: Grade 1 vs 3 (AUC = 0.958, accuracy = 91.0%, κ = 0.910); Grade 2 vs 3 (AUC = 0.939, accuracy = 79.0%, κ = 0.790); Grade 1 vs 2 (AUC = 0.565, accuracy = 62.4%, κ = 0.670) |
| Majtner et al[54] | CD/38 patients | Retrospective, single center | CE | 5421 images | 1549 images | DL | Performance in ulcer detection (sensitivity = 95.7%, specificity = 99.8%, accuracy = 98.4%). Agreement between the model and manual reading of ulcerations (κ = 0.720) |
| Udristoiu et al[55] | CD/54 patients | Retrospective, single center | pCLE | 5081 images | 1124 images | CNN | Differentiation between inflammation and intact colonic mucosa (AUC = 0.980, accuracy = 95.3%, specificity = 92.8%, sensitivity = 94.6%) |
| de Maissin et al[56] | CD/63 patients | Retrospective, multicenter | CE | 2449 images | 700 images | RNN | Performance for discriminating pathological vs non-pathological images (accuracy = 93.7%, sensitivity = 93.0%, specificity = 95.0%, κ = 0.790) |
| Ribeiro et al[57] | CD/124 patients | Retrospective, multicenter | CE | 37319 images | 124 images | CNN | Identification of colonic ulcerations and erosions (AUC = 1.000, accuracy = 99.6%, sensitivity = 96.9%, specificity = 99.9%) |
| Ferreira et al[58] | CD/NR | Retrospective, multicenter | CE | 19740 images | 4935 images | DL-CNN | Performance of the model for lesion detection (sensitivity = 90.0%, specificity = 96.0%, precision = 97.1%, accuracy = 92.4%) |
| Afonso et al[59] | CD/NR | Retrospective, single center | CE | 4904 images | 1226 images | CNN | Detection of ulcers and erosions in the small intestine mucosa (accuracy = 95.6%, sensitivity = 90.8%, specificity = 97.1%) |
| Martins et al[60] | CD/250 patients | Retrospective, single center | DAE | 250 DAE images | 6772 images | CNN | Identification of colonic ulcerations and erosions (AUC = 1.000, accuracy = 98.7%, sensitivity = 88.5%, specificity = 99.7%) |
| Brodersen et al[34] | CD/131 patients | Prospective, multicenter | CE | NR | NR | DL | The identification capacity for CD (sensitivity: 92.0%-96.0% and specificity: 90.0%-93.0%) and IBD (sensitivity: 97.0% and specificity: 90.0%-91.0%) |
| Xie et al[61] | CD/628 patients | Retrospective, single center | DBE | NR | 28155 images | DL | The accuracy for detection of ulcers (96.3%), inflammatory stenosis (95.7%), and non-inflammatory stenosis (96.7%). The grading of ulcers based on surface area, size, and depth (precision between 85.2% and 87.8%) |
- Citation: Minea H, Singeap AM, Minea M, Chiriac S, Stanciu C, Trifan A. Artificial intelligence in inflammatory bowel disease: Current applications and future directions. World J Gastroenterol 2025; 31(39): 111353
- URL: https://www.wjgnet.com/1007-9327/full/v31/i39/111353.htm
- DOI: https://dx.doi.org/10.3748/wjg.v31.i39.111353