BPG is committed to discovery and dissemination of knowledge
Systematic Reviews
Copyright: ©Author(s) 2026.
World J Gastroenterol. Nov 7, 2026; 32(41): 122556
Published online Nov 7, 2026. doi: 10.3748/wjg.122556
Table 1 Characteristics of 58 included studies
Number
Ref.
Domain
Study design
Sample size
AI method
TML
Key finding
1Gong et al[1], 2020Withdrawal speed monitoringRCT704 patientsCNN (ENDOANGEL)AENDOANGEL ADR 16.3% vs 7.7% control (intention-to-treat)
2Su et al[2], 2020Withdrawal speed monitoringRCT659 patientsCNN (real-time quality control)AADR 28.9% vs 16.5%; polyp detection rate 383% vs 25.4%
3Yao et al[9], 2022Withdrawal speed monitoringRCT1076 patientsCNN (ENDOANGEL CAQ)ACOMBO ADR 30.6% vs CADe-only 21.3% vs CAQ-only 24.5% vs control 148%
4Liu et al[27], 2025Withdrawal speed monitoringRCT1254 patientsCNN (ENDOANGEL)A6-center ADR improved 22.6% to 32.7% in moderate/Low detectors
5Barua et al[28], 2023Withdrawal speed monitoringProspective332 patientsCNN (speedometer)BNo benefit at high-ADR centers (45.8% vs 45.2%); ceiling effect
6Lu et al[29], 2023Withdrawal speed monitoringRetrospective1780 patientsCNN (ENDOANGEL)CAI-assisted arm (CADe + CAQ + combined) eliminated time-of-day quality decline (13.7%-5.7% unassisted; stable 22%-23% with AI)
7Liu et al[55], 2022Withdrawal speed monitoringProspective103 colonoscopiesCNNBAI-based fold examination quality correlated with ADR (r = 0.852)
8Lux et al[56], 2023Withdrawal speed monitoringMulticenter100 colonoscopy videosDL (documentation)CAI prototype for automated withdrawal time measurement and photo-documentation (5 centers)
9Lui et al[57], 2024Withdrawal speed monitoringRetrospective350 videosDL (real-time)CAI real-time monitoring of effective withdrawal time
10Li et al[58], 2024Withdrawal speed monitoringRetrospective472 videosYOLOv5CNovel withdrawal time indicator based on YOLOv5
11Wu et al[30], 2019Coverage/blind spot mappingRCT324 patientsCNN (WISENSE)AEGD blind spots 5.9% vs 22.5% control
12Wu et al[31], 2021Coverage/blind spot mappingRCT1050 patientsCNN (ENDOANGEL)A5-hospital blind spots 5.38 vs 9.82
13Chen et al[60], 2020Coverage/blind spot mappingRCT437 patientsDNN (ENDOANGEL)AAI reduced blind spots; sedated 3.4% vs 22.4% in EGD
14Freedman et al[32], 2020Coverage/blind spot mappingAlgorithm devSynthetic + real videosCNN (depth-based)ECoverage quantification; algorithm 0.075 vs expert 0.177 MAE; 93% agreement
15Wu et al[59], 2019Coverage/blind spot mappingAlgorithm dev3170 gastric cancer + 5981 benign imagesDNN (ENDOANGEL)EEGC detection 92.5% accuracy, 94.0% sensitivity
16Li et al[61], 2021Coverage/blind spot mappingAlgorithm dev170297 images + 5779 videosDL (IDEA)DReal-time 31-site gastric anatomical recognition; 95.3% video accuracy in EGD
17Cao et al[33], 2023Workflow recognitionRetro + animal201026 labeled framesCNN (AI-Endo)D83.5% real-time ESD phase recognition across centers
18Furube et al[34], 2024Workflow recognitionRetrospective94 videosCNNCEsophageal ESD phase recognition 90% accuracy
19Liu et al[35], 2025Workflow recognitionMulticenter195 videosCNNCInternational 7-center esophageal ESD workflow recognition
20Chen et al[36], 2025Workflow recognitionDataset66656 framesTransformerFRenji ESD benchmark dataset
21Biffi et al[37], 2025Workflow recognitionDataset2.7M framesTCNFREAL-colon temporal segmentation benchmark
22Ward et al[62], 2021Workflow recognitionRetrospective50 videosCNN (LSTM)DAutomated POEM phase identification
23Zhang et al[78], 2026Workflow recognitionAlgorithm dev385 videosMamba (SPRMamba)EState-space model for ESD phase recognition
24Nerup et al[38], 2015Skill assessmentProspective10 experienced + 11 trainees in colonoscopyML (kinematic)DMEI-based kinematic scoring discriminated expert vs trainee
25Vilmann et al[39], 2020Skill assessmentProspective24 endoscopistsML (simulation)DComputerized assessment validated in simulation
26Yao et al[40], 2024Skill assessmentRCT685 patientsCNN (ENDOANGEL)AAI novice miss rate 188% vs control novice 437% vs expert 27.0%
27Wittbrodt et al[63], 2024Skill assessmentRetrospective50 colonoscopiesML (random forest)DML colonoscopy skill; withdrawal time r = 0.99
281Cold et al[79], 2024Skill assessmentPrior systematic review (cross-reference)13 studiesVarious-Systematic review of computer-aided colonoscopy competence assessment
29Martin et al[41], 2020Autonomous navigationAnimal study2 pigs; 10 novicesCNNECompletion rates 58% direct vs 96% intelligent teleop vs 100% semi-autonomous
30Hwang et al[42], 2026Autonomous navigationSimulation50 experimentsSupervised DLFAutonomous robotic colonoscopy; 90% success rate
31Corsi et al[43], 2023Autonomous navigationSimulationSimulationConstrained RLFSafe colonoscopy navigation with formal verification
32Prendergast et al[64], 2018Autonomous navigationPhantomPhantomML (localization)FAutonomous haustral fold detection for robotic endoscopy
33Huang et al[65], 2021Autonomous navigationSimulationSimulationML (magnetic)FAutonomous navigation of magnetic colonoscope
34Lazo et al[66], 2022Autonomous navigationSimulationSimulationDL visual servoingFAutonomous soft-robot intraluminal navigation
35Pore et al[67], 2022Autonomous navigationSimulationSimulationEnd-to-end RLFDeep visuomotor control for colonoscopy
36Tan et al[68], 2025Autonomous navigationSimulationSimulationHuman-intervention RLFSafe navigation via human-intervention-based RL
37Ma et al[44], 2021SLAM/3D reconstructionAlgorithm devReal + phantom sequencesRNN-SLAME38%-46% drift reduction in 3D colon reconstruction
38Ozyoruk et al[45], 2021SLAM/3D reconstructionDataset + algo42700 framesSelf-supervisedEEndoSLAM dataset; unsupervised depth estimation
39Bonilla et al[46], 2024SLAM/3D reconstructionAlgorithm devEndoscopy videos3DGSEGaussian Pancakes for endoscopic reconstruction
40Wang et al[47], 2024SLAM/3D reconstructionAlgorithm devEndoscopy videos3DGS + SLAMEEndoGSLAM: Real-time dense reconstruction (> 100 fps)
41Recasens et al[69], 2021SLAM/3D reconstructionAlgorithm devEndoscopy videosDL (depth + motion)EEndo-Depth-and-Motion tracking
42Shao et al[70], 2022SLAM/3D reconstructionAlgorithm devEndoscopy videosSelf-supervisedESelf-supervised monocular depth in endoscopy
43Hayoz et al[71], 2023SLAM/3D reconstructionAlgorithm devEndoscopy videosDL (pose est)ERobust camera pose estimation in endoscopy
44Azagra et al[72], 2023SLAM/3D reconstructionDataset96 proceduresSLAM baselineFEndomapper dataset of calibrated procedures
45Bobrow et al[73], 2023SLAM/3D reconstructionDataset10015 framesNeRF (2D-3D reg)EColonoscopy 3D dataset with paired depth
46Shi et al[74], 2023SLAM/3D reconstructionAlgorithm devColonoscopy videosNeRF (ColonNeRF)EHigh-fidelity long-sequence colonoscopy reconstruction
47Elvira et al[75], 2024SLAM/3D reconstructionAlgorithm devFull proceduresCudaSIFT-SLAMEMultiple-map SLAM for full procedure mapping
48Guo et al[80], 2025SLAM/3D reconstructionAlgorithm devEndoscopy videosNeRF (UC-NeRF)EUncertainty-aware NeRF from sparse endoscopic views
49Kaleta et al[81], 2025SLAM/3D reconstructionAlgorithm devEndoscopy videos3DGS (PR-ENDO)EPhysically based relightable Gaussian Splatting
50Trovato et al[48], 2010Robotic control/instrument trackingSimulationSimulationQ-learningFFirst RL-based colon-endoscope-robot locomotion
51Turan et al[49], 2019Robotic control/instrument trackingSimulationSimulationDeep RLFLearning to navigate endoscopic capsule robots
52İncetan et al[50], 2021Robotic control/instrument trackingSimulationSimulationDeep RL (VR-caps)FVirtual environment for capsule endoscopy
53Jha et al[51], 2020Robotic control/instrument trackingDataset590 framesDL (segmentation)FKvasir-Instrument: GIE tool-segmentation dataset
54Ali et al[52], 2021Robotic control/instrument trackingChallenge2531 detection + 643 segmentation framesTransformer/CNNFDL for artefact and disease detection and segmentation
55Brand et al[53], 2022Robotic control/instrument trackingMulticenter retro580 videosDLDDL model to improve polyp-detection usability
56Jha et al[54], 2025Robotic control/instrument trackingChallengeMultiple datasetsVarious DLFPolyp and instrument segmentation benchmark
57Zhang et al[76], 2022Robotic control/instrument trackingSimulationSimulationDeep RLFRL-based stomach coverage scanning of WCE
58Ng et al[77], 2024Robotic control/instrument trackingSimulationSimulationDeep RLFRL navigation of tendon-driven flexible robotic endoscope
Table 2 Multi-dimensional evidence imbalance between diagnostic artificial intelligence (computer-aided detection/ computer-aided diagnosis) and kinematic artificial intelligence (computer-aided quality) in gastrointestinal endoscopy
Evidence indicator
Diagnostic AI (CADe/CADx)
Kinematic AI (CAQ)
Ratio or qualitative gap
RCTs published> 408Approximately 5:1
Patients enrolled in RCTs> 27000Approximately 6200Approximately 4:1
Systematic reviews/meta-analyses> 100 (this is the first)> 10:1
Regulatory approvals (FDA/CE/MFDS)≥ 6 devices0 standaloneQualitative
GRADE certainty for primary outcomeHigh (ADR improvement)Moderate (2 of 8 domains)2 levels lower
Domains with ≥ 1 clinical trial3/3 (detection, classification, characterization)2/8 (withdrawal speed, coverage)Qualitative
Domains with zero patient-level data0/36/8 (75%)Qualitative
Annual research output (2020-2024 PubMed)Approximately 150 studies/yearApproximately 12 studies/yearApproximately 12:1
Industry investment in device developmentMultiple companies (Medtronic, Fujifilm, NEC, Olympus, etc.)Single academic platform dominates (ENDOANGEL)Qualitative
Table 3 Risk of bias assessment (Cochrane Risk of Bias 2 for randomized controlled trials)
Ref.
D1: Randomization
D2: Deviations from intended interventions
D3: Missing outcome data
D4: Measurement of outcome
D5: Selection of reported result
Overall
Gong et al[1], 2020Low (computer-generated, sealed envelopes)High (real-time AI display; inherently unblindable)Low (ITT; 2.1% attrition)Low (histopathology reference; blinded pathologist)Low (pre-registered; all outcomes reported)Some concerns (unblinding inevitable)
Su et al[2], 2020Low (computer-generated randomization)High (real-time AI overlay visible to endoscopist)Low (ITT; < 3% missing)Low (histopathology reference; blinded assessment)Low (pre-registered; primary/secondary endpoints reported)Some concerns (unblinding inevitable)
Yao et al[9], 2022Low (computer-generated; 4-arm parallel design)High (CAQ display visible; endoscopist aware of AI arm)Low (ITT; complete follow-up)Low (histopathology; blinded pathologist)Low (pre-registered 4-arm design; all outcomes reported)Some concerns (unblinding inevitable)
Wu et al[30], 2019Low (computer-generated; concealed allocation)High (WISENSE display visible during EGD)Low (complete data; no attrition)Low (blind spot count by independent reviewer)Low (pre-registered; all endpoints reported)Some concerns (unblinding inevitable)
Chen et al[60], 2020Low (computer-generated randomization)High (AI display visible during EGD)Low (< 2% missing; ITT)Low (blind spot assessment by blinded reviewer)Low (all pre-specified outcomes reported)Some concerns (unblinding inevitable)
Wu et al[31], 2021Low (centralized randomization; 5-hospital)High (real-time AI feedback visible to endoscopist)Low (ITT; minimal attrition)Low (blind spot mapping by independent reviewer)Low (pre-registered multicenter; all outcomes reported)Some concerns (unblinding inevitable)
Yao et al[40], 2024Low (computer-generated; tandem design)High (AI assistance visible during colonoscopy)Low (complete follow-up; tandem design ensures paired data)Some concerns (miss rate depends on tandem sequence; learning effect possible)Low (pre-registered tandem RCT; all outcomes reported)Some concerns (unblinding + potential tandem sequence effect)
Liu et al[27], 2025Low (centralized randomization; 6-center)High (real-time AI quality control visible)Low (ITT; < 3% attrition across 6 centers)Low (histopathology; blinded pathologist)Low (pre-registered multicenter; all outcomes reported)Some concerns (unblinding inevitable)
Table 4 Risk of bias assessment (Quality Assessment of Diagnostic Accuracy Studies-2 for diagnostic accuracy studies)
Ref.
Patient selection
Index test
Reference standard
Flow and timing
Overall
Barua et al[28], 2023Low (consecutive patients in implementation trial)Low (AI speedometer threshold pre-defined)Low (withdrawal time by independent timer)Low (all patients received same assessment)Low (all domains low risk)
Lu et al[29], 2023Unclear (single-center convenience sample; time-of-day subgroups)Low (ENDOANGEL AI output pre-specified)Low (histopathology for ADR; blinded pathologist)Low (all patients received colonoscopy and pathology)Unclear (patient selection concern)
Liu et al[55], 2022Unclear (single-center; convenience sampling)Low (AI withdrawal assessment pre-defined)Low (expert endoscopist consensus)Low (all patients assessed by both AI and experts)Unclear (patient selection concern)
Lui et al[57], 2024Unclear (retrospective video selection; single center)Low (AI withdrawal monitoring pre-specified)Low (manual review by experienced endoscopist)Low (all videos analyzed by both methods)Unclear (patient selection concern)
Li et al[58], 2024Unclear (retrospective convenience sample)Low (YOLOv5 threshold pre-specified)Low (manual withdrawal time measurement)Low (all videos assessed)Unclear (patient selection concern)
Li et al[61], 2021Unclear (single-center convenience sample)Low (IDEA system output pre-defined)Low (expert annotation of anatomical landmarks)Low (all patients received same EGD protocol)Unclear (patient selection concern)
Cao et al[33], 2023Unclear (retrospective video collection; single center + animal)Low (AI-endo phase output pre-specified)Low (expert surgeon frame-level annotation)Low (all videos fully annotated)Unclear (patient selection concern)
Furube et al[34], 2024Unclear (retrospective single-center video selection)Low (AI phase recognition threshold pre-defined)Low (expert endoscopist annotation)Low (all videos received complete annotation)Unclear (patient selection concern)
Liu et al[35], 2025Low (multicenter; prospective + retrospective cohorts)Low (AI workflow recognition pre-specified)Low (expert panel annotation consensus)Low (all cases assessed by AI and experts)Low (all domains low risk)
Ward et al[62], 2021Unclear (retrospective single-center video selection)Low (AI POEM phase output pre-defined)Low (expert surgeon phase annotation)Low (all videos fully annotated)Unclear (patient selection concern)
Nerup et al[38], 2015High (small convenience sample; 12 endoscopists only)Unclear (MEI kinematic thresholds derived from same cohort)Low (expert/trainee classification pre-defined)Low (all endoscopists assessed)High (small sample + index test concern)
Vilmann et al[39], 2020High (small convenience sample; 27 endoscopists in simulation)Unclear (computerized metrics derived from training data overlap)Low (GRS expert assessment as reference)Low (all participants completed assessment)High (small sample + index test concern)
Table 5 GRADE certainty of evidence assessment across 8 kinematic artificial intelligence domains
Domain
n
Risk of bias
Inconsistency
Indirectness
Imprecision
Pub. bias
GRADE
Withdrawal speed10Not serious1Not seriousNot seriousNot seriousUnlikelyModerate
Coverage mapping6Not serious1Not seriousNot seriousSerious2UnlikelyModerate
Workflow recognition7Serious3Not seriousSerious4SeriousUndetectedLow
Skill assessment5SeriousSerious5Very serious6Very seriousUndetectedVery low
Navigation (pre-clinical)8SeriousNot seriousVery serious7Very seriousUndetectedVery low
SLAM/3D (pre-clinical)13Serious8Not seriousVery serious7Very seriousSuspected9Very low
Robotic control (pre-clinical)5SeriousSeriousVery serious7Very seriousUndetectedVery low
Instrument tracking (pre-clinical)4SeriousNot seriousVery serious7Very seriousUndetectedVery low


Write to the Help Desk