Copyright: ©Author(s) 2026.
World J Gastroenterol. Nov 7, 2026; 32(41): 122556
Published online Nov 7, 2026. doi: 10.3748/wjg.122556
Published online Nov 7, 2026. doi: 10.3748/wjg.122556
Table 1 Characteristics of 58 included studies
| Number | Ref. | Domain | Study design | Sample size | AI method | TML | Key finding |
| 1 | Gong et al[1], 2020 | Withdrawal speed monitoring | RCT | 704 patients | CNN (ENDOANGEL) | A | ENDOANGEL ADR 16.3% vs 7.7% control (intention-to-treat) |
| 2 | Su et al[2], 2020 | Withdrawal speed monitoring | RCT | 659 patients | CNN (real-time quality control) | A | ADR 28.9% vs 16.5%; polyp detection rate 383% vs 25.4% |
| 3 | Yao et al[9], 2022 | Withdrawal speed monitoring | RCT | 1076 patients | CNN (ENDOANGEL CAQ) | A | COMBO ADR 30.6% vs CADe-only 21.3% vs CAQ-only 24.5% vs control 148% |
| 4 | Liu et al[27], 2025 | Withdrawal speed monitoring | RCT | 1254 patients | CNN (ENDOANGEL) | A | 6-center ADR improved 22.6% to 32.7% in moderate/Low detectors |
| 5 | Barua et al[28], 2023 | Withdrawal speed monitoring | Prospective | 332 patients | CNN (speedometer) | B | No benefit at high-ADR centers (45.8% vs 45.2%); ceiling effect |
| 6 | Lu et al[29], 2023 | Withdrawal speed monitoring | Retrospective | 1780 patients | CNN (ENDOANGEL) | C | AI-assisted arm (CADe + CAQ + combined) eliminated time-of-day quality decline (13.7%-5.7% unassisted; stable 22%-23% with AI) |
| 7 | Liu et al[55], 2022 | Withdrawal speed monitoring | Prospective | 103 colonoscopies | CNN | B | AI-based fold examination quality correlated with ADR (r = 0.852) |
| 8 | Lux et al[56], 2023 | Withdrawal speed monitoring | Multicenter | 100 colonoscopy videos | DL (documentation) | C | AI prototype for automated withdrawal time measurement and photo-documentation (5 centers) |
| 9 | Lui et al[57], 2024 | Withdrawal speed monitoring | Retrospective | 350 videos | DL (real-time) | C | AI real-time monitoring of effective withdrawal time |
| 10 | Li et al[58], 2024 | Withdrawal speed monitoring | Retrospective | 472 videos | YOLOv5 | C | Novel withdrawal time indicator based on YOLOv5 |
| 11 | Wu et al[30], 2019 | Coverage/blind spot mapping | RCT | 324 patients | CNN (WISENSE) | A | EGD blind spots 5.9% vs 22.5% control |
| 12 | Wu et al[31], 2021 | Coverage/blind spot mapping | RCT | 1050 patients | CNN (ENDOANGEL) | A | 5-hospital blind spots 5.38 vs 9.82 |
| 13 | Chen et al[60], 2020 | Coverage/blind spot mapping | RCT | 437 patients | DNN (ENDOANGEL) | A | AI reduced blind spots; sedated 3.4% vs 22.4% in EGD |
| 14 | Freedman et al[32], 2020 | Coverage/blind spot mapping | Algorithm dev | Synthetic + real videos | CNN (depth-based) | E | Coverage quantification; algorithm 0.075 vs expert 0.177 MAE; 93% agreement |
| 15 | Wu et al[59], 2019 | Coverage/blind spot mapping | Algorithm dev | 3170 gastric cancer + 5981 benign images | DNN (ENDOANGEL) | E | EGC detection 92.5% accuracy, 94.0% sensitivity |
| 16 | Li et al[61], 2021 | Coverage/blind spot mapping | Algorithm dev | 170297 images + 5779 videos | DL (IDEA) | D | Real-time 31-site gastric anatomical recognition; 95.3% video accuracy in EGD |
| 17 | Cao et al[33], 2023 | Workflow recognition | Retro + animal | 201026 labeled frames | CNN (AI-Endo) | D | 83.5% real-time ESD phase recognition across centers |
| 18 | Furube et al[34], 2024 | Workflow recognition | Retrospective | 94 videos | CNN | C | Esophageal ESD phase recognition 90% accuracy |
| 19 | Liu et al[35], 2025 | Workflow recognition | Multicenter | 195 videos | CNN | C | International 7-center esophageal ESD workflow recognition |
| 20 | Chen et al[36], 2025 | Workflow recognition | Dataset | 66656 frames | Transformer | F | Renji ESD benchmark dataset |
| 21 | Biffi et al[37], 2025 | Workflow recognition | Dataset | 2.7M frames | TCN | F | REAL-colon temporal segmentation benchmark |
| 22 | Ward et al[62], 2021 | Workflow recognition | Retrospective | 50 videos | CNN (LSTM) | D | Automated POEM phase identification |
| 23 | Zhang et al[78], 2026 | Workflow recognition | Algorithm dev | 385 videos | Mamba (SPRMamba) | E | State-space model for ESD phase recognition |
| 24 | Nerup et al[38], 2015 | Skill assessment | Prospective | 10 experienced + 11 trainees in colonoscopy | ML (kinematic) | D | MEI-based kinematic scoring discriminated expert vs trainee |
| 25 | Vilmann et al[39], 2020 | Skill assessment | Prospective | 24 endoscopists | ML (simulation) | D | Computerized assessment validated in simulation |
| 26 | Yao et al[40], 2024 | Skill assessment | RCT | 685 patients | CNN (ENDOANGEL) | A | AI novice miss rate 188% vs control novice 437% vs expert 27.0% |
| 27 | Wittbrodt et al[63], 2024 | Skill assessment | Retrospective | 50 colonoscopies | ML (random forest) | D | ML colonoscopy skill; withdrawal time r = 0.99 |
| 281 | Cold et al[79], 2024 | Skill assessment | Prior systematic review (cross-reference) | 13 studies | Various | - | Systematic review of computer-aided colonoscopy competence assessment |
| 29 | Martin et al[41], 2020 | Autonomous navigation | Animal study | 2 pigs; 10 novices | CNN | E | Completion rates 58% direct vs 96% intelligent teleop vs 100% semi-autonomous |
| 30 | Hwang et al[42], 2026 | Autonomous navigation | Simulation | 50 experiments | Supervised DL | F | Autonomous robotic colonoscopy; 90% success rate |
| 31 | Corsi et al[43], 2023 | Autonomous navigation | Simulation | Simulation | Constrained RL | F | Safe colonoscopy navigation with formal verification |
| 32 | Prendergast et al[64], 2018 | Autonomous navigation | Phantom | Phantom | ML (localization) | F | Autonomous haustral fold detection for robotic endoscopy |
| 33 | Huang et al[65], 2021 | Autonomous navigation | Simulation | Simulation | ML (magnetic) | F | Autonomous navigation of magnetic colonoscope |
| 34 | Lazo et al[66], 2022 | Autonomous navigation | Simulation | Simulation | DL visual servoing | F | Autonomous soft-robot intraluminal navigation |
| 35 | Pore et al[67], 2022 | Autonomous navigation | Simulation | Simulation | End-to-end RL | F | Deep visuomotor control for colonoscopy |
| 36 | Tan et al[68], 2025 | Autonomous navigation | Simulation | Simulation | Human-intervention RL | F | Safe navigation via human-intervention-based RL |
| 37 | Ma et al[44], 2021 | SLAM/3D reconstruction | Algorithm dev | Real + phantom sequences | RNN-SLAM | E | 38%-46% drift reduction in 3D colon reconstruction |
| 38 | Ozyoruk et al[45], 2021 | SLAM/3D reconstruction | Dataset + algo | 42700 frames | Self-supervised | E | EndoSLAM dataset; unsupervised depth estimation |
| 39 | Bonilla et al[46], 2024 | SLAM/3D reconstruction | Algorithm dev | Endoscopy videos | 3DGS | E | Gaussian Pancakes for endoscopic reconstruction |
| 40 | Wang et al[47], 2024 | SLAM/3D reconstruction | Algorithm dev | Endoscopy videos | 3DGS + SLAM | E | EndoGSLAM: Real-time dense reconstruction (> 100 fps) |
| 41 | Recasens et al[69], 2021 | SLAM/3D reconstruction | Algorithm dev | Endoscopy videos | DL (depth + motion) | E | Endo-Depth-and-Motion tracking |
| 42 | Shao et al[70], 2022 | SLAM/3D reconstruction | Algorithm dev | Endoscopy videos | Self-supervised | E | Self-supervised monocular depth in endoscopy |
| 43 | Hayoz et al[71], 2023 | SLAM/3D reconstruction | Algorithm dev | Endoscopy videos | DL (pose est) | E | Robust camera pose estimation in endoscopy |
| 44 | Azagra et al[72], 2023 | SLAM/3D reconstruction | Dataset | 96 procedures | SLAM baseline | F | Endomapper dataset of calibrated procedures |
| 45 | Bobrow et al[73], 2023 | SLAM/3D reconstruction | Dataset | 10015 frames | NeRF (2D-3D reg) | E | Colonoscopy 3D dataset with paired depth |
| 46 | Shi et al[74], 2023 | SLAM/3D reconstruction | Algorithm dev | Colonoscopy videos | NeRF (ColonNeRF) | E | High-fidelity long-sequence colonoscopy reconstruction |
| 47 | Elvira et al[75], 2024 | SLAM/3D reconstruction | Algorithm dev | Full procedures | CudaSIFT-SLAM | E | Multiple-map SLAM for full procedure mapping |
| 48 | Guo et al[80], 2025 | SLAM/3D reconstruction | Algorithm dev | Endoscopy videos | NeRF (UC-NeRF) | E | Uncertainty-aware NeRF from sparse endoscopic views |
| 49 | Kaleta et al[81], 2025 | SLAM/3D reconstruction | Algorithm dev | Endoscopy videos | 3DGS (PR-ENDO) | E | Physically based relightable Gaussian Splatting |
| 50 | Trovato et al[48], 2010 | Robotic control/instrument tracking | Simulation | Simulation | Q-learning | F | First RL-based colon-endoscope-robot locomotion |
| 51 | Turan et al[49], 2019 | Robotic control/instrument tracking | Simulation | Simulation | Deep RL | F | Learning to navigate endoscopic capsule robots |
| 52 | İncetan et al[50], 2021 | Robotic control/instrument tracking | Simulation | Simulation | Deep RL (VR-caps) | F | Virtual environment for capsule endoscopy |
| 53 | Jha et al[51], 2020 | Robotic control/instrument tracking | Dataset | 590 frames | DL (segmentation) | F | Kvasir-Instrument: GIE tool-segmentation dataset |
| 54 | Ali et al[52], 2021 | Robotic control/instrument tracking | Challenge | 2531 detection + 643 segmentation frames | Transformer/CNN | F | DL for artefact and disease detection and segmentation |
| 55 | Brand et al[53], 2022 | Robotic control/instrument tracking | Multicenter retro | 580 videos | DL | D | DL model to improve polyp-detection usability |
| 56 | Jha et al[54], 2025 | Robotic control/instrument tracking | Challenge | Multiple datasets | Various DL | F | Polyp and instrument segmentation benchmark |
| 57 | Zhang et al[76], 2022 | Robotic control/instrument tracking | Simulation | Simulation | Deep RL | F | RL-based stomach coverage scanning of WCE |
| 58 | Ng et al[77], 2024 | Robotic control/instrument tracking | Simulation | Simulation | Deep RL | F | RL navigation of tendon-driven flexible robotic endoscope |
Table 2 Multi-dimensional evidence imbalance between diagnostic artificial intelligence (computer-aided detection/ computer-aided diagnosis) and kinematic artificial intelligence (computer-aided quality) in gastrointestinal endoscopy
| Evidence indicator | Diagnostic AI (CADe/CADx) | Kinematic AI (CAQ) | Ratio or qualitative gap |
| RCTs published | > 40 | 8 | Approximately 5:1 |
| Patients enrolled in RCTs | > 27000 | Approximately 6200 | Approximately 4:1 |
| Systematic reviews/meta-analyses | > 10 | 0 (this is the first) | > 10:1 |
| Regulatory approvals (FDA/CE/MFDS) | ≥ 6 devices | 0 standalone | Qualitative |
| GRADE certainty for primary outcome | High (ADR improvement) | Moderate (2 of 8 domains) | 2 levels lower |
| Domains with ≥ 1 clinical trial | 3/3 (detection, classification, characterization) | 2/8 (withdrawal speed, coverage) | Qualitative |
| Domains with zero patient-level data | 0/3 | 6/8 (75%) | Qualitative |
| Annual research output (2020-2024 PubMed) | Approximately 150 studies/year | Approximately 12 studies/year | Approximately 12:1 |
| Industry investment in device development | Multiple companies (Medtronic, Fujifilm, NEC, Olympus, etc.) | Single academic platform dominates (ENDOANGEL) | Qualitative |
Table 3 Risk of bias assessment (Cochrane Risk of Bias 2 for randomized controlled trials)
| Ref. | D1: Randomization | D2: Deviations from intended interventions | D3: Missing outcome data | D4: Measurement of outcome | D5: Selection of reported result | Overall |
| Gong et al[1], 2020 | Low (computer-generated, sealed envelopes) | High (real-time AI display; inherently unblindable) | Low (ITT; 2.1% attrition) | Low (histopathology reference; blinded pathologist) | Low (pre-registered; all outcomes reported) | Some concerns (unblinding inevitable) |
| Su et al[2], 2020 | Low (computer-generated randomization) | High (real-time AI overlay visible to endoscopist) | Low (ITT; < 3% missing) | Low (histopathology reference; blinded assessment) | Low (pre-registered; primary/secondary endpoints reported) | Some concerns (unblinding inevitable) |
| Yao et al[9], 2022 | Low (computer-generated; 4-arm parallel design) | High (CAQ display visible; endoscopist aware of AI arm) | Low (ITT; complete follow-up) | Low (histopathology; blinded pathologist) | Low (pre-registered 4-arm design; all outcomes reported) | Some concerns (unblinding inevitable) |
| Wu et al[30], 2019 | Low (computer-generated; concealed allocation) | High (WISENSE display visible during EGD) | Low (complete data; no attrition) | Low (blind spot count by independent reviewer) | Low (pre-registered; all endpoints reported) | Some concerns (unblinding inevitable) |
| Chen et al[60], 2020 | Low (computer-generated randomization) | High (AI display visible during EGD) | Low (< 2% missing; ITT) | Low (blind spot assessment by blinded reviewer) | Low (all pre-specified outcomes reported) | Some concerns (unblinding inevitable) |
| Wu et al[31], 2021 | Low (centralized randomization; 5-hospital) | High (real-time AI feedback visible to endoscopist) | Low (ITT; minimal attrition) | Low (blind spot mapping by independent reviewer) | Low (pre-registered multicenter; all outcomes reported) | Some concerns (unblinding inevitable) |
| Yao et al[40], 2024 | Low (computer-generated; tandem design) | High (AI assistance visible during colonoscopy) | Low (complete follow-up; tandem design ensures paired data) | Some concerns (miss rate depends on tandem sequence; learning effect possible) | Low (pre-registered tandem RCT; all outcomes reported) | Some concerns (unblinding + potential tandem sequence effect) |
| Liu et al[27], 2025 | Low (centralized randomization; 6-center) | High (real-time AI quality control visible) | Low (ITT; < 3% attrition across 6 centers) | Low (histopathology; blinded pathologist) | Low (pre-registered multicenter; all outcomes reported) | Some concerns (unblinding inevitable) |
Table 4 Risk of bias assessment (Quality Assessment of Diagnostic Accuracy Studies-2 for diagnostic accuracy studies)
| Ref. | Patient selection | Index test | Reference standard | Flow and timing | Overall |
| Barua et al[28], 2023 | Low (consecutive patients in implementation trial) | Low (AI speedometer threshold pre-defined) | Low (withdrawal time by independent timer) | Low (all patients received same assessment) | Low (all domains low risk) |
| Lu et al[29], 2023 | Unclear (single-center convenience sample; time-of-day subgroups) | Low (ENDOANGEL AI output pre-specified) | Low (histopathology for ADR; blinded pathologist) | Low (all patients received colonoscopy and pathology) | Unclear (patient selection concern) |
| Liu et al[55], 2022 | Unclear (single-center; convenience sampling) | Low (AI withdrawal assessment pre-defined) | Low (expert endoscopist consensus) | Low (all patients assessed by both AI and experts) | Unclear (patient selection concern) |
| Lui et al[57], 2024 | Unclear (retrospective video selection; single center) | Low (AI withdrawal monitoring pre-specified) | Low (manual review by experienced endoscopist) | Low (all videos analyzed by both methods) | Unclear (patient selection concern) |
| Li et al[58], 2024 | Unclear (retrospective convenience sample) | Low (YOLOv5 threshold pre-specified) | Low (manual withdrawal time measurement) | Low (all videos assessed) | Unclear (patient selection concern) |
| Li et al[61], 2021 | Unclear (single-center convenience sample) | Low (IDEA system output pre-defined) | Low (expert annotation of anatomical landmarks) | Low (all patients received same EGD protocol) | Unclear (patient selection concern) |
| Cao et al[33], 2023 | Unclear (retrospective video collection; single center + animal) | Low (AI-endo phase output pre-specified) | Low (expert surgeon frame-level annotation) | Low (all videos fully annotated) | Unclear (patient selection concern) |
| Furube et al[34], 2024 | Unclear (retrospective single-center video selection) | Low (AI phase recognition threshold pre-defined) | Low (expert endoscopist annotation) | Low (all videos received complete annotation) | Unclear (patient selection concern) |
| Liu et al[35], 2025 | Low (multicenter; prospective + retrospective cohorts) | Low (AI workflow recognition pre-specified) | Low (expert panel annotation consensus) | Low (all cases assessed by AI and experts) | Low (all domains low risk) |
| Ward et al[62], 2021 | Unclear (retrospective single-center video selection) | Low (AI POEM phase output pre-defined) | Low (expert surgeon phase annotation) | Low (all videos fully annotated) | Unclear (patient selection concern) |
| Nerup et al[38], 2015 | High (small convenience sample; 12 endoscopists only) | Unclear (MEI kinematic thresholds derived from same cohort) | Low (expert/trainee classification pre-defined) | Low (all endoscopists assessed) | High (small sample + index test concern) |
| Vilmann et al[39], 2020 | High (small convenience sample; 27 endoscopists in simulation) | Unclear (computerized metrics derived from training data overlap) | Low (GRS expert assessment as reference) | Low (all participants completed assessment) | High (small sample + index test concern) |
Table 5 GRADE certainty of evidence assessment across 8 kinematic artificial intelligence domains
| Domain | n | Risk of bias | Inconsistency | Indirectness | Imprecision | Pub. bias | GRADE |
| Withdrawal speed | 10 | Not serious1 | Not serious | Not serious | Not serious | Unlikely | Moderate |
| Coverage mapping | 6 | Not serious1 | Not serious | Not serious | Serious2 | Unlikely | Moderate |
| Workflow recognition | 7 | Serious3 | Not serious | Serious4 | Serious | Undetected | Low |
| Skill assessment | 5 | Serious | Serious5 | Very serious6 | Very serious | Undetected | Very low |
| Navigation (pre-clinical) | 8 | Serious | Not serious | Very serious7 | Very serious | Undetected | Very low |
| SLAM/3D (pre-clinical) | 13 | Serious8 | Not serious | Very serious7 | Very serious | Suspected9 | Very low |
| Robotic control (pre-clinical) | 5 | Serious | Serious | Very serious7 | Very serious | Undetected | Very low |
| Instrument tracking (pre-clinical) | 4 | Serious | Not serious | Very serious7 | Very serious | Undetected | Very low |
- Citation: Gong EJ, Bang CS, Lee JJ. Artificial intelligence for kinematic (procedural motion) analysis in gastrointestinal endoscopy: A systematic review. World J Gastroenterol 2026; 32(41): 122556
- URL: https://www.wjgnet.com/1007-9327/full/v32/i41/122556.htm
- DOI: https://dx.doi.org/10.3748/wjg.122556