Gong EJ, Bang CS, Lee JJ. Artificial intelligence for kinematic (procedural motion) analysis in gastrointestinal endoscopy: A systematic review. World J Gastroenterol 2026; 32(41): 122556 [DOI: 10.3748/wjg.122556]
Corresponding Author of This Article
Chang Seok Bang, MD, PhD, Department of Internal Medicine, Hallym University College of Medicine, Sakju-ro 77, Chuncheon 24253, Gangwon-do, South Korea. cloudslove@naver.com
Research Domain of This Article
Gastroenterology & Hepatology
Article-Type of This Article
research-article
Open-Access Policy of This Article
This article is an open-access article which was selected by an in-house editor and fully peer-reviewed by external reviewers. It is distributed in accordance with the Creative Commons Attribution Non Commercial (CC BY-NC 4.0) license, which permits others to distribute, remix, adapt, build upon this work non-commercially, and license their derivative works on different terms, provided the original work is properly cited and the use is non-commercial. See: http://creativecommons.org/licenses/by-nc/4.0/
Baishideng Publishing Group Inc, 7041 Koll Center Parkway, Suite 160, Pleasanton, CA 94566, USA
Share the Article
Gong EJ, Bang CS, Lee JJ. Artificial intelligence for kinematic (procedural motion) analysis in gastrointestinal endoscopy: A systematic review. World J Gastroenterol 2026; 32(41): 122556 [DOI: 10.3748/wjg.122556]
Co-corresponding authors: Chang Seok Bang and Jae Jun Lee.
Author contributions: Gong EJ and Bang CS were responsible for writing-original draft; Bang CS was responsible for conceptualization, formal analysis, methodology, project administration, resources; Bang CS and Lee JJ were responsible for writing-review and editing as co-corresponding authors; Lee JJ was responsible for funding acquisition; Gong EJ, Bang CS, and Lee JJ were responsible for data curation, investigation; all of the authors read and approved the final version of the manuscript to be published.
AI contribution statement: We did not use AI in the preparation of this manuscript.
Supported by the Bio and Medical Technology Development Program of the National Research Foundation (NRF) funded by the Korean government (MSIT), No. RS-2023-00223501.
Conflict-of-interest statement: All authors declare no conflict of interest in publishing the manuscript.
PRISMA 2009 Checklist statement: The authors have read the PRISMA 2009 Checklist, and the manuscript was prepared and revised according to the PRISMA 2009 Checklist.
Corresponding author: Chang Seok Bang, MD, PhD, Department of Internal Medicine, Hallym University College of Medicine, Sakju-ro 77, Chuncheon 24253, Gangwon-do, South Korea. cloudslove@naver.com
Received: April 22, 2026 Revised: May 19, 2026 Accepted: June 24, 2026 Published online: November 7, 2026 Processing time: 149 Days and 14.3 Hours
Abstract
BACKGROUND
Artificial intelligence (AI) in gastrointestinal endoscopy has focused on computer-aided detection (CADe) for recognition errors; AI addressing endoscope motion, coverage, workflow, and skill – collectively kinematic analysis – has not been systematically synthesized.
AIM
To map kinematic AI evidence across eight predefined domains and assessed its quality relative to CADe and surgical AI.
METHODS
MEDLINE/PubMed, EMBASE-OVID, Cochrane, and IEEE Xplore were searched following Preferred Reporting Items for Systematic reviews and Meta-Analyses checklist (CRD420261320279). Studies applying AI to kinematic data in gastrointestinal endoscopy were eligible. Risk of bias used Cochrane Risk of Bias 2, Quality Assessment of Diagnostic Accuracy Studies-2, and PROBAST + AI; certainty of evidence used GRADE.
RESULTS
Fifty-eight studies were included. Evidence generation was asymmetric: (1) Randomized controlled trial ratio approximately 5:1 (> 40 vs 8); (2) Patient ratio approximately 4:1 (> 27000 vs approximately 6200); (3) Annual publication ratio approximately 12:1; and (4) ≥ 6 vs 0 regulatory approvals favoring diagnostic AI. Withdrawal speed monitoring and coverage/blind-spot mapping reached moderate GRADE certainty, but all 8 randomized controlled trials were from China and 6 of 8 used the ENDOANGEL platform. The remaining six domains have not reached moderate certainty: (1) Workflow recognition and skill assessment are approaching clinical readiness; and (2) Autonomous navigation, robotic control, simultaneous localization and mapping/three-dimensional reconstruction, and instrument tracking remain at the engineering stage. Comparison with surgical AI indicated that the gap reflects ecosystem-level factors – absence of open datasets, benchmarking challenges, and regulatory pathways for motion-acting AI – rather than technical immaturity.
CONCLUSION
Kinematic AI addresses exposure errors that CADe cannot correct and has shown additive benefit alongside CADe. Two domains are ready for geographically diverse validation; six remain at the engineering stage. Priorities are open datasets, benchmarking infrastructure, and dedicated regulatory pathways for motion-acting AI.
Core Tip: Artificial intelligence (AI) in gastrointestinal (GI) endoscopy has been dominated by computer-aided detection (CADe) of lesions – targeting recognition errors – with over 40 randomized controlled trials (RCTs) and multiple regulatory approvals, whereas AI addressing endoscope motion, coverage, and procedural workflow has developed along a separate, largely engineering-focused trajectory. No prior systematic review has mapped the distribution, translational maturity, and certainty of evidence for kinematic AI across GI endoscopy domains. Compared with diagnostic AI, kinematic AI has produced approximately one-fifth the number of RCTs, one-quarter the number of enrolled patients, about one-twelfth the annual publication output, and no standalone regulatory approvals. Only two of eight domains reached moderate GRADE certainty; all eight RCTs were conducted in Chinese centers and six used the ENDOANGEL platform, indicating pronounced geographic and platform concentration. The remaining six domains are at the pre-clinical/engineering stage; the gap with surgical AI is best explained by ecosystem-level factors rather than by technical immaturity alone. The four-arm RCT demonstrated that computer-aided quality and CADe address independent failure modes with additive benefit on adenoma detection rate, supporting integration of kinematic monitoring into existing CADe platforms as the most direct translational pathway. Priority investments include building open kinematic datasets, establishing GI-specific benchmarking challenges analogous to the EndoVis series, conducting colonoscopy three-dimensional coverage RCTs outside China, and defining regulatory pathways for AI that acts on motion rather than on images.
Citation: Gong EJ, Bang CS, Lee JJ. Artificial intelligence for kinematic (procedural motion) analysis in gastrointestinal endoscopy: A systematic review. World J Gastroenterol 2026; 32(41): 122556
When a gastric cancer is missed during gastroscopy [esophagogastroduodenoscopy (EGD)] or a colorectal adenoma escapes detection during colonoscopy, two distinct failure mechanisms may be responsible: (1) The endoscopist may have seen the lesion but failed to recognize it (a recognition error); and (2) The endoscopist may never have adequately examined the relevant mucosal area (an exposure error)[1,2]. This distinction has important implications for how artificial intelligence (AI) can improve endoscopy.
The past decade of AI development has focused predominantly on recognition errors. Computer-aided detection (CADe) and computer-aided diagnosis (CADx) systems analyze endoscopic images to highlight or classify lesions, with more than 40 randomized controlled trials (RCTs) enrolling over 27000 patients, multiple regulatory approvals, and real-time clinical decision support systems covering esophageal, gastric, and colorectal neoplasia[3-7]. However, these systems cannot detect exposure errors: No polyp detector, however sensitive, can find lesions behind folds that are never visualized or in gastric areas examined too briefly[8].
A separate category of AI applications addresses this gap by analyzing endoscope kinematics. In this systematic review, we define kinematic AI as any AI system whose primary function is to analyze or act upon the motion, spatial trajectory, coverage, procedural phase, or motor performance of the endoscope or its instruments. This definition is deliberately broader than pure motion-sensor analysis; it encompasses: (1) Direct motion/speed measurement (e.g., withdrawal speed quantification, kinematic skill scoring); (2) Spatial coverage estimation derived from visual scene reconstruction [e.g., simultaneous localization and mapping (SLAM)-based three-dimensional (3D) colon coverage, depth-based coverage mapping]; (3) Proxy kinematic measures inferred from anatomical landmark recognition (e.g., WISENSE-style blind-spot monitoring, which uses learned landmark categories to infer whether the endoscope has traversed each mapped region); and (4) Higher-level procedural intelligence derived from motion-aware video understanding [e.g., endoscopic submucosal dissection (ESD) workflow-phase recognition]. We acknowledge that [proxy kinematic measures inferred from anatomical landmark recognition (e.g., WISENSE-style blind-spot monitoring, which uses learned landmark categories to infer whether the endoscope has traversed each mapped region)] and [higher-level procedural intelligence derived from motion-aware video understanding (e.g., ESD workflow-phase recognition)] rely in part on image-classification models; they are included as kinematic AI because the variable they ultimately quantify – whether each anatomical region was adequately traversed, or which procedural motion phase is underway – is kinematic rather than diagnostic in nature. To operationalize this distinction at the algorithmic level for borderline cases, we treat an image-based model as kinematic when its per-frame output is a label of anatomical position, procedural phase, or coverage state rather than a lesion identity, and when those frame-level labels are aggregated over time to yield a procedure-level metric of trajectory, coverage, or workflow. A model whose per-frame output is a lesion characterization is classified as diagnostic even when it shares the same convolutional backbone; the decisive criterion is what the algorithm ultimately quantifies and whether that quantity is integrated across the procedure, not how it processes each individual frame. The implications of alternative, narrower classification boundaries are examined in the Discussion.
Yao et al[9] provided direct proof of concept that this category of AI is clinically non-redundant with CADe: In a four-arm RCT, adding computer-aided quality (CAQ) monitoring to CADe improved adenoma detection rate (ADR) by 26% beyond what CADe alone achieved (30.6% vs 21.3%), demonstrating that kinematic and diagnostic AI address independent error mechanisms with additive benefit.
Despite this importance, kinematic AI has developed across disconnected disciplines – gastroenterology, computer science, and robotics – with no unified synthesis. In the parallel field of minimally invasive surgery, AI-driven kinematic analysis has matured far more rapidly[10-14]: Standardized skill-assessment datasets (JIGSAWS, 2014), annual benchmarking competitions (the EndoVis series, 9 editions from 2015 to 2024), workflow-recognition benchmarks exceeding 91% accuracy (Cholec80), and autonomous robotic procedures in living animals have all been established, leading to Food and Drug Administration (FDA)-cleared surgical AI systems[15-19]. Whether gastrointestinal endoscopy (GIE) can replicate this ecosystem – and what structural barriers have prevented equivalent progress – has not been systematically examined. This systematic review provides the first systematic mapping and certainty assessment of kinematic AI evidence in GIE, quantifies the evidence imbalance relative to diagnostic AI and to surgical AI, and identifies the translational bottlenecks that must be addressed for clinical adoption.
MATERIALS AND METHODS
Protocol and registration
This systematic review followed the Preferred Reporting Items for Systematic reviews and Meta-Analyses 2020 statement (Supplementary material) and was prospectively registered in PROSPERO (CRD420261320279)[20].
Search strategy and information sources
MEDLINE/PubMed, EMBASE-OVID, the Cochrane Central Register of Controlled Trials, and IEEE Xplore were searched from January 2010 to January 2026. Controlled vocabulary and free-text terms covering AI, deep learning (DL), machine learning, reinforcement learning (RL), kinematics, motion analysis, coverage, trajectory, workflow, skill assessment, SLAM, and GIE were combined; the full search strategy is provided in Supplementary material. Reference lists of included studies and relevant reviews were hand-searched to identify additional records.
Eligibility criteria
We included studies that applied AI or machine learning to endoscope kinematic or procedural-motion data in GIE, operationalized through the eight predefined domains above. Studies whose primary output was lesion detection, characterization, or histological prediction (pure CADe/CADx) without a motion, coverage, workflow, or skill component were excluded, as were studies using non-AI methods, studies in non-gastrointestinal (GI) settings, narrative reviews and editorials, conference abstracts without full text, and non-English publications. For hybrid systems (e.g., ENDOANGEL) that combine CADe and CAQ modules, inclusion was determined at the study level based on the primary function evaluated; studies that evaluated only the CADe component were excluded even if produced on a platform whose other modules are kinematic. Because several engineering-oriented domains – such as autonomous navigation, SLAM/3D reconstruction, robotic control, and instrument tracking – publish predominantly through preprint servers and indexed conference proceedings rather than peer-reviewed journals, full-text arXiv preprints were eligible provided they reported sufficient methodological detail for data extraction and formal risk-of-bias appraisal; conference abstracts and preprints lacking such detail were excluded. Eight included studies met this threshold as preprints. These were not exempted from quality appraisal: Each was assessed with PROBAST + AI in the same way as published technical studies, and their non-peer-reviewed status is acknowledged as a limitation.
Study selection, data extraction and outcomes
Two reviewers independently screened titles and abstracts, followed by full-text assessment, with disagreements resolved by discussion. Extracted data included study design, sample size, AI method, domain, outcomes, and performance metrics. Each study was assigned a translational maturity level (TML): (1) A (RCT with patient outcomes); (2) B (prospective non-randomized clinical study); (3) C (retrospective clinical study); (4) D (algorithm with external validation); (5) E (algorithm with internal validation only); and (6) F (simulation, phantom, animal, or dataset study). The TML scheme was not a previously validated instrument but an operational classification defined for the present review; it adapts the principle of staged technology-readiness assessment – articulated for machine-learning systems in the machine learning technology readiness levels framework[21] – to the mixed clinical and engineering literature of kinematic AI, in which conventional clinical-evidence hierarchies cannot accommodate simulation, phantom, and dataset studies. Levels A-C correspond to conventional clinical study designs and levels D-F to pre-clinical engineering maturity. Two reviewers independently assigned a TML to each of the 57 primary studies; initial agreement was reached for 52 of 57 studies (91.2%), and the 5 discrepant assignments – arising at the C/D and E/F boundaries where a study combined clinical and algorithmic elements – were resolved by discussion. The primary outcome was the mapping of AI applications for kinematic analysis in GIE across the eight predefined domains, including the number of RCTs, total patient enrollment, meta-analyses, and regulatory approvals, compared with those of CADe. Secondary outcomes were: (1) GRADE certainty per domain; (2) Risk of bias; (3) Translational maturity distribution; (4) Identification of research gaps defined as domains lacking patient-level clinical outcomes; and (5) Comparison of the kinematic AI ecosystem in GIE with that of surgical AI.
Risk of bias and certainty of evidence assessment
Risk of bias was assessed using the Cochrane Risk of Bias 2 (RoB 2) tool for RCTs[22], the Quality Assessment of Diagnostic Accuracy Studies (QUADAS)-2 tool for diagnostic accuracy studies[23], and PROBAST + AI for technical studies including algorithm-development, simulation, and dataset studies[24]. Certainty of evidence per domain was rated high, moderate, low, or very low using the GRADE approach[25], considering risk of bias, inconsistency, indirectness, imprecision, and publication bias. For engineering-stage domains without clinical outcome data, GRADE was applied to the closest clinically relevant surrogate (e.g., phantom/animal performance), and these domains were explicitly labelled as pre-clinical to avoid conflating low certainty with negative evidence. Applying GRADE to engineering-stage domains requires justification, because GRADE was developed for clinical outcomes. We retained it for two reasons: (1) GRADE explicitly permits rating certainty for surrogate outcomes, provided the surrogate is identified and the rationale for downgrading is made transparent; and (2) Applying one certainty framework across all eight domains allows the clinical-versus-engineering evidence gap to be expressed on a single comparable scale. The pre-clinical label signifies that low certainty in these domains reflects an early translational stage by design, not refuted or negative evidence.
Statistical analysis
Given the heterogeneity of study designs (RCTs, prospective and retrospective clinical studies, algorithm-development and dataset studies, and simulation or phantom experiments), AI methods, outcome measures, and clinical domains, quantitative meta-analysis was not feasible. A narrative synthesis was therefore performed following the Synthesis Without Meta-analysis reporting guideline[26]. Studies were grouped by the eight kinematic domains and synthesized within each domain. The comparative evidence gap between CADe and kinematic AI was reported as a set of individual ratio estimates (number of RCTs, enrolled patients, meta-analyses, regulatory approvals, and annual publication output) rather than as a single composite number, because individual indicators provide more interpretable comparisons to readers.
RESULTS
Study selection and characteristics
The systematic search identified 1524 records. After removal of 460 duplicates, 1064 records were screened by title and abstract, and 384 full-text articles were assessed for eligibility. Three hundred and twenty-six records were excluded (15 narrative reviews, 2 systematic reviews, and 309 studies with insufficient data or outside the eligibility criteria). Fifty-eight studies[1,2,9,27-81] met the final inclusion criteria (Figure 1 and Table 1). Publication volume increased sharply over time, with 95% (55/58) of studies published from 2019 onward and 2024 as the peak year (13 studies). The detailed year-by-year distribution was 1 study in 2010[48], 1 in 2015[38], 1 in 2018[64], 3 in 2019[30,49,59], 7 in 2020[1,2,32,39,41,51,60], 9 in 2021[31,44,45,50,52,61,62,65,69], 6 in 2022[9,53,55,66,67,70], 10 in 2023[28,29,33,43,56,71-74,76], 13 in 2024[34,40,46,47,57,58,63,68,75,77-79,81], 6 in 2025[27,35-37,54,80], and 1 in 2026[42]. Study designs were heterogeneous: (1) 8 RCTs[1,2,9,27,30,31,40,60]; (2) 4 prospective non-randomized clinical studies[28,38,39,55]; (3) 10 retrospective clinical studies[29,33-35,53,56-58,62,63]; (4) 35 algorithm-development, simulation, phantom, or dataset studies[32,36,37,41-52,54,59,61,64-78,80,81]; and (5) 1 prior systematic review[79] retained for cross-reference in the skill-assessment domain. The eight RCTs were concentrated in three domains – withdrawal speed monitoring (n = 4), coverage/blind spot mapping (n = 3), and skill assessment (n = 1) – and all eight originated from Chinese academic centers; six of eight used the ENDOANGEL platform (Wuhan University; the remaining two were conducted by the same group under the WISENSE designation and by an independent Chinese center). Sample sizes ranged from simulation experiments[42] to 170297 images[61] for technical studies, and from 100 colonoscopy videos[56] to 1780 patients[29] for clinical studies. AI methods spanned convolutional neural networks[1,9,27,32,41], transformers[36,52,74], RL[42,43,66-68], neural radiance fields[70,71,73], 3D Gaussian splatting[47,74,75], and state-space models (Mamba)[78,80]. A single AI platform, ENDOANGEL, contributed 8 studies across 3 domains[1,9,27,29-31,40,59], representing the most extensively validated kinematic AI system in the field.
Evidence imbalance between diagnostic and kinematic AI
Across the indicators pre-specified in methods (Table 2), kinematic AI has produced substantially less clinical evidence than CADe/CADx despite emerging in a comparable timeframe (2015-2019). Specific ratios were approximately 5:1 for RCTs (> 40 vs 8), 4:1 for enrolled patients (> 27000 vs approximately 6200), > 10:1 for systematic reviews/meta-analyses, and approximately 12:1 for annual PubMed publication output (approximately 150 vs approximately 12 studies per year, 2020-2024). CADe/CADx has received ≥ 6 regulatory approvals (FDA/Conformité Européenne/Ministry of Food and Drug Safety), compared with no standalone regulatory approvals for kinematic AI. The certainty gap is equally pronounced: CADe for ADR improvement carries high GRADE certainty from more than 20 concordant RCTs, whereas only 2 of 8 kinematic domains (withdrawal monitoring, coverage mapping) reached moderate certainty. Within kinematic AI itself, 3 domains accounted for all 8 RCTs (withdrawal speed monitoring, n = 4; coverage/blind spot mapping, n = 3; skill assessment, n = 1), while 5 domains had zero patient-level clinical data despite technical progress.
Table 2 Multi-dimensional evidence imbalance between diagnostic artificial intelligence (computer-aided detection/ computer-aided diagnosis) and kinematic artificial intelligence (computer-aided quality) in gastrointestinal endoscopy.
Clinically validated domains: Withdrawal monitoring and coverage mapping
Withdrawal speed monitoring (10 studies, 4 RCTs, > 5900 patients; GRADE: Moderate): ENDOANGEL-based systems provide real-time feedback on colonoscope withdrawal speed and mucosal-fold inspection. The Yao et al[9] four-arm trial established that CAQ and CADe have additive effects on ADR (COMBO 30.6% vs CAQ-only 24.5% vs CADe-only 21.3% vs control 14.8%). A six-center RCT (Liu et al[27], 2025; 1254 patients) confirmed an ADR improvement from 22.6% to 32.7% in moderate-detectors and low-detectors. Barua et al[28] (2023) found no benefit at high-baseline-ADR centers (approximately 45%), establishing a ceiling effect. In the Lu et al[29] (2023) retrospective study, AI-assisted colonoscopy was associated with elimination of the time-of-day decline in quality (unassisted ADR fell from 13.7% in the morning to 5.7% in the afternoon, whereas AI-assisted ADR remained at 22%-23%). We note explicitly that the AI-assisted arm in that study combined three subgroups (CADe, CAQ, and CADe + CAQ). The observed effect therefore cannot be attributed to kinematic AI alone but reflects the contribution of the combined AI intervention, and we interpret it accordingly in the Discussion. GRADE was downgraded from high to moderate owing to structurally unavoidable endoscopist unblinding in real-time AI trials and to single-platform dominance (ENDOANGEL).
Coverage/blind-spot mapping (6 studies, 3 RCTs, > 1800 patients; GRADE: Moderate): In this review, blind spots during EGD are defined as anatomical regions of the upper GI tract that endoscopy guidelines recommend be photographically documented (e.g., 22 or 31 predefined gastric sites in the Japanese and Chinese systematic-examination protocols) but that were not actually visualized during a given procedure. The WISENSE system (Wu et al[30], 2019) reduced EGD blind-spot rates from 22.5% to 5.9% (P < 0.001), and a similar benefit was confirmed in a five-hospital RCT (5.38 vs 9.82 mean blind spots)[31]. For colonoscopy, Freedman et al[32] (2020) developed depth-based coverage quantification. This method estimates per-pixel colonic wall depth from the monocular video feed and integrates it over the withdrawal phase to compute the fraction of mucosa that passed within the camera’s field of view. The system achieved 93% agreement with expert coverage labeling[32], although no colonoscopy-coverage RCT yet exists. GRADE was moderate owing to unblinding and to the absence of clinical trials specifically targeting colonoscopy coverage.
Emerging domains approaching clinical readiness
Workflow recognition (7 studies; GRADE: Low): ESD phase recognition is the automated identification, from endoscopic video, of the canonical procedural steps of ESD (marking, submucosal injection, mucosal incision, submucosal dissection, and hemostasis). AI-Endo (Cao et al[33], 2023) achieved 83.5% real-time ESD phase recognition with external validation on 15 independent cases. Furube et al[34] (2024) extended this to esophageal ESD with 90% phase accuracy, and Liu et al[35] (2025) performed the first international multicenter esophageal ESD workflow-recognition study with an educational RCT component. Two 2025 benchmark datasets (Renji ESD, 66656 frames; REAL-Colon, 2.7 million frames) have been released, suggesting growing methodological maturity[36,37]. GRADE was low because no prospective patient-outcome trial has been conducted and because phase-recognition accuracy, measured against expert annotation, is a surrogate rather than a clinical outcome.
Skill assessment (5 studies; GRADE: Very low): Kinematic skill scoring based on magnetic endoscopic imaging of the colonoscope tip trajectory (Nerup et al[38], 2015; Vilmann et al[39], 2020) discriminated expert from trainee colonoscopists. AI-assisted novice colonoscopy (Yao et al[40], 2024) reached adenoma miss rates comparable to experts in a tandem RCT (AI-novice 18.8% vs control novice 43.7% vs expert 27.0%). However, the combined paradigm of DL applied to concurrent kinematic-sensor data – DL-kinematic sensor integration, defined here as the joint analysis of video, force/torque, electromagnetic tracker, or robotic-encoder signals by a deep neural network to derive automated skill metrics, as popularized by JIGSAWS in robotic surgery[10] – has not been attempted for flexible GI endoscopy. GRADE was very low owing to the single available RCT, the absence of DL-kinematic sensor integration, and the lack of a standardized kinematic-measurement infrastructure for flexible endoscopes. The clinical importance of optimal technique selection and standardised skill assessment is established in adjacent surgical evidence. Comparative data after totally laparoscopic total gastrectomy show that the choice of esophagojejunal anastomotic technique directly influences anastomotic complication rates, with strictures concentrated in the more technically demanding circular-stapler arm[82] and in colorectal cancer surgery anastomotic leakage confers an approximately fivefold increase in 30-day mortality, with additional perioperative risk factors further amplifying complication and mortality risk in geriatric patients[83]. Translating this principle to flexible endoscopy provides a direct clinical rationale for AI-based kinematic skill metrics in endoscopic training, where standardised objective motion assessment is a plausible upstream intervention for the procedural failure modes that determine patient survival, leakage rates, and overall complication burden.
Engineering-stage (pre-clinical) domains
The four domains below consist predominantly of algorithm-development, simulation, phantom, or dataset studies. Their principal outputs are mathematical accuracy, phantom performance, or formally verified safety properties, rather than patient outcomes. In the IDEAL framework and the FDA Software-as-a-Medical-Device (SaMD) Total Product Lifecycle, these studies sit at stages 0-1 (idea, development, pre-clinical) by design. Accordingly, we describe them here as pre-clinical and map their translational maturity rather than evaluate them against a patient-outcome endpoint, which would constitute a category error.
Autonomous navigation (8 studies; GRADE: Very low – pre-clinical): Technical progress has been substantial, ranging from convolutional neural network-based semi-autonomous magnetic colonoscopy in pigs[41] to full-procedure autonomous robotic colonoscopy simulation[42]. RL dominates the recent literature (Proximal Policy Optimization, Soft Actor-Critic, constrained RL with formal safety verification)[43]. All studies to date used simulation, phantom, or animal models; no first-in-human clinical trial and no regulatory pathway has been established. GRADE is very low; the domain is at an engineering stage.
SLAM and 3D reconstruction (13 studies; GRADE: Very low – pre-clinical): This is the most technically active domain, with rapid methodological progression from classical SLAM to recurrent neural network (RNN)-SLAM (38%-46% drift reduction)[44], self-supervised depth estimation[45], neural radiance fields, and 3D Gaussian Splatting (> 100 fps at approximately two minutes of training)[46,47]. Real-time clinical deployment is now technically feasible, yet no study has prospectively tested reconstruction outputs in patient endoscopy. The bottleneck is not computational speed but the absence of validated clinical metrics linking coverage completeness to patient outcomes.
Robotic control and instrument tracking (9 studies; GRADE: Very low – pre-clinical): Robotic capsule control ranged from Q-learning[48] to deep RL controllers[49,50], all in simulation. For GI instrument tracking, the principal public dataset is Kvasir-Instrument (590 images)[51], compared with thousands of annotated frames in laparoscopic EndoVis challenges. No ESD-knife tracking, endoscopic retrograde cholangiopancreatography instrument segmentation, or hemostasis-clip detection study was identified[51-54]. GRADE is very low; the domain is at an engineering stage.
Risk of bias assessment
Using RoB 2, all eight RCTs received a “high risk” judgment for Domain 2 (deviations from intended interventions) because endoscopist blinding to a real-time AI display is structurally impossible. Because this limitation is inherent to all real-time AI-assisted trials rather than a methodological flaw of individual studies, the overall judgment was rated “some concerns”, consistent with prior systematic reviews of CADe. Future RCT designs could potentially mitigate this structural unblinding by employing sham AI interfaces (non-functional overlays indistinguishable from active feedback to the endoscopist) or by relying on retrospective kinematic analysis of blinded video recordings as the primary outcome measure rather than on real-time intervention; either approach would shift the bias profile from Domain 2 (deviations from intended interventions) toward Domain 4 (measurement of outcome), where blinded outcome assessment is methodologically feasible. Randomization and outcome measurement were adequate in 7 of 8 trials; the remaining trial (Yao et al[40], 2024) had some concerns for outcome measurement owing to a possible learning effect in the tandem design (Table 3)[1,2,9,27,30,31,60]. QUADAS-2 assessment of 12 diagnostic accuracy studies showed patient-selection concerns in 67% of studies (convenience sampling), while reference standard and flow/timing were generally low risk (Table 4)[28,29,33-35,38,39,55,57,58,61,62]. PROBAST + AI assessment of 38 technical studies revealed overall high risk of bias in 25 studies (66%), primarily driven by Domain 4 (Analysis; 21/38, 55%) owing to the absence of external validation. Domain 1 (Participants) was rated high in 15 studies (39%), predominantly those conducted entirely in simulation or phantom environments; Domain 2 (Predictors) was uniformly low risk, as deep-learning architectures extract features automatically (Supplementary Table 1)[32,36,37,41-54,56,59,63-81].
Table 3 Risk of bias assessment (Cochrane Risk of Bias 2 for randomized controlled trials).
Table 5 summarizes the GRADE assessment across all eight domains. Moderate certainty was achieved only for withdrawal monitoring and coverage mapping (downgraded from high owing to unblinding and single-platform dominance). Workflow recognition was rated low (no patient-outcome trial). The remaining five domains were rated very low: (1) Skill assessment (downgraded for indirectness because its only RCT used adenoma miss rate rather than a kinematic outcome, and for imprecision) and the four engineering-stage domains (autonomous navigation, SLAM/3D reconstruction, robotic control, and instrument tracking); and (2) Downgraded for absence of clinical trials, indirectness from simulation/phantom settings, and imprecision from small sample sizes. The contrast with diagnostic AI – for which GRADE certainty for CADe improving ADR is high, based on more than 20 concordant RCTs – illustrates the scale of the evidence imbalance.
Table 5 GRADE certainty of evidence assessment across 8 kinematic artificial intelligence domains.
Sensitivity analysis under a narrower kinematic definition
To address the concern that our operational definition of kinematic AI is broader than a pure motion-signal-based definition, we conducted a pre-specified sensitivity analysis using a four-tier hierarchical taxonomy (Figure 2)[9]. Tier 1 (true kinematic AI) comprised studies that analyze direct motion signals from sensors or encoders (withdrawal-speed monitoring, magnetic endoscopic imaging-based skill assessment, and encoder-based robotic control). Tier 2 (proxy kinematic inference) comprised studies that derive motion variables from monocular video (instrument tracking, SLAM/3D reconstruction, and autonomous navigation). Tier 3 (workflow intelligence) comprised ESD phase-recognition studies, which use video-frame classification to infer procedural-motion phase. Tier 4 [procedural quality-assurance (QA) systems] comprised landmark-classification studies (WISENSE blind-spot mapping and gastric anatomical landmark recognition) in which kinematic adequacy is inferred from whether predefined anatomical regions were photographed rather than from explicit motion variables.
Figure 2 Hierarchical taxonomy of artificial intelligence domains in gastrointestinal endoscopy.
Four tiers are arranged along the horizontal axis by their proximity to direct motion-signal analysis: (1) Tier 1 [true kinematic artificial intelligence (AI), direct motion signal]; (2) Tier 2 (proxy kinematic inference, vision-derived motion); (3) Tier 3 (workflow intelligence, phase recognition); and (4) Tier 4 (procedural quality-assurance systems, landmark classification). The vertical axis represents translational maturity (regulatory approval, clinical randomized controlled trial, prospective/retrospective, algorithm/external validation, engineering/simulation). The dashed-line box denotes the strict kinematic definition (Tiers 1-2) used in the pre-specified sensitivity analysis. Boxes with bold black borders mark domains with clinical-trial (randomized controlled trial) evidence. Adjacent AI ecosystems – surgical AI (left) and diagnostic AI/computer-aided detection/computer-aided diagnosis (right) – are shown for translational reference. AI: Artificial intelligence; FDA: Food and Drug Administration; RCT: Randomized controlled trial; SLAM: Simultaneous localization and mapping; 3D: Three-dimensional; ADR: Adenoma detection rate; CADe: Computer-aided detection; CADx: Computer-aided diagnosis; CAQ: Computer-aided quality; CE: Conformée Européenne; ESD: Endoscopic submucosal dissection; MFDS: Ministry of Food and Drug Safety; QA: Quality-assurance.
Under a strict kinematic definition restricted to Tiers 1-2, four of the eight original domains were retained, comprising 5 RCTs (4 withdrawal-speed and 1 skill-assessment) and approximately 4400 enrolled patients; the kinematic-to-CADe RCT ratio became approximately 8:1 (vs 5:1 under the broad definition), and the patient ratio became approximately 6:1 (vs 4:1). Moderate GRADE certainty was retained only for withdrawal-speed monitoring; coverage/blind-spot mapping was reclassified to Tier 4 and would therefore be excluded under the strict definition. Restricting still further to Tier 1 alone (3 domains) yielded an RCT ratio of approximately 13:1. Across all three definitions – broad (Tiers 1-4), strict (Tiers 1-2), and strictest (Tier 1 only) – the core conclusion of a substantial multi-fold evidence gap between kinematic and diagnostic AI was preserved, indicating that the principal finding of this review is robust to the boundary of the kinematic definition.
DISCUSSION
Principal findings
This systematic review provides four principal findings. First, kinematic AI and diagnostic AI address independent error mechanisms with additive clinical benefit but receive markedly asymmetric research investment. The multi-fold evidence imbalance (RCT ratio approximately 5:1, patient ratio approximately 4:1, meta-analysis ratio > 10:1, annual publication output approximately 12:1, regulatory approvals ≥ 6 vs 0) quantifies this asymmetry across interpretable dimensions. For device manufacturers, the direct practical implication is that integrating kinematic monitoring into existing CADe platforms – rather than further sensitivity optimization of detection modules – is a tractable translational pathway for reducing miss rates.
Second, within kinematic AI, a two-to-three-level GRADE gap separates the two validated domains from the six non-validated domains, and the dominant explanation for this gap is the absence of translational infrastructure rather than algorithmic immaturity. SLAM achieves over 100 fps real-time rendering; autonomous robotic controllers complete simulated colonoscopies; ESD workflow recognition reaches 83%-90% accuracy against expert annotation. These demonstrate technical feasibility. Their translational bottlenecks are: (1) The absence of validated clinical-outcome metrics linking kinematic measures (coverage completeness, motion smoothness, phase-transition timing) to patient outcomes beyond ADR; (2) The absence of a regulatory framework for AI that directly controls endoscope manipulation; and (3) The absence of standardized kinematic-measurement infrastructure for flexible endoscopes, which – unlike rigid laparoscopes with known geometric constraints and embedded encoders – lack equivalent sensing. Each bottleneck requires interdisciplinary collaboration between engineers, clinicians, and regulators. We justify the infrastructure framing by referencing the explicit contrast with surgical AI below, where infrastructure (datasets, challenges, regulatory precedents) demonstrably preceded clinical adoption.
Third, all eight clinical RCTs of kinematic AI were conducted in Chinese centers and six of eight used the ENDOANGEL platform. This pronounced geographic and platform concentration, which was highlighted by both handling editor and reviewers, limits external generalizability and introduces the possibility that regional health-system factors (workload, baseline ADR, documentation culture, training pathways) influenced the observed effects. External replication of CAQ concepts in Western, Japanese, and other health-system contexts, and validation of alternative (non-ENDOANGEL) platforms, should be considered prerequisites for broad clinical adoption.
Fourth, the six non-validated domains should not be interpreted as having produced “low-quality” evidence. Workflow recognition and skill assessment provide video-annotation and clinical-pilot evidence, while the four engineering-stage domains (autonomous navigation, SLAM/3D reconstruction, robotic control, and instrument tracking) produce the type of evidence their design permits – mathematical validation, phantom/animal performance, or simulation-based safety verification – and sit at early IDEAL/SaMD translational stages. The value of the current review for these domains is mapping rather than clinical evaluation.
Comparison with previous literature
The underdevelopment of kinematic AI in GIE relative to surgical AI is not primarily a natural consequence of field age but an ecosystem gap, as direct comparison with surgery demonstrates (Supplementary Table 2, Supplementary Figures 1-3)[10-18,33,51]. We make this comparison cautiously and explicitly acknowledge its limits, because the physical substrates of the two domains differ in important ways.
Laparoscopic instruments are rigid, have well-defined geometric constraints, and – in robotic platforms (e.g., da Vinci) – are equipped with built-in encoders that provide direct kinematic readouts of instrument position, velocity, and force. Flexible GIEs, by contrast, exhibit complex non-linear tip deformation, are subject to loop formation and redundant insertion configurations, and currently lack equivalent embedded kinematic sensing. Consequently, kinematic AI in flexible endoscopy must in most cases infer motion indirectly from the monocular video feed, using optical flow, depth estimation, pose reconstruction, or anatomical-landmark recognition, rather than from direct sensor readouts. The computational task of scene understanding from monocular video is comparable between the two domains, but the physical sensing substrate is not identical, and the challenge of deriving kinematic information from video alone is genuinely harder for flexible endoscopy in several respects.
Our ecosystem comparison with surgery is therefore intended as a translational – not a purely technical – parallel. Surgery established its first kinematic dataset (JIGSAWS, 2014)[10], organized annual instrument-segmentation challenges (EndoVis, 9 editions from 2015 to 2024)[11], released the Cholec80 workflow-recognition benchmark (2016)[12], and achieved autonomous in vivo intestinal anastomosis with outcomes comparable to experienced surgeons (STAR, 2016)[13]. By 2022, STAR had performed fully autonomous laparoscopic surgery on living pigs[14], and Moon Surgical received FDA 510(k) clearance for a robotic-assisted surgical system with AI-enabled camera positioning[15]. The first remote robotic surgery (Operation Lindbergh) was performed in 2001[16], more than two decades before any comparable GIE milestone. The quantitative infrastructure gap is pronounced: (1) Approximately 30000 annotated instrument-segmentation frames vs 590 in GIE (approximately 50:1); (2) 50-84 machine-learning-based skill-assessment studies vs approximately 2 (approximately 30:1)[17,18]; and (3) Six or more public kinematic datasets vs one. The key point of the analogy is not that the underlying technical problems are equally difficult, but that the combination of open datasets, annual benchmarking competitions, substantial industry investment, and dedicated regulatory pathways – which surgery has built – has not yet been replicated for GIE kinematic AI[19].
This contrast, however, conflates two barriers of fundamentally different kinds, and distinguishing them is essential before the surgical analogy can be applied. Some limitations concern the data ecosystem surrounding the field, whereas others are intrinsic sensor and physical modelling constraints of the flexible endoscope itself.
Data ecosystem limitations
The first class concerns the research infrastructure surrounding the field rather than the endoscope itself: The absence of open annotated datasets, of recurring benchmarking challenges, and of sustained industry investment and dedicated regulatory pathways. These are the gaps quantified above (one public instrument dataset vs six or more; no GI-specific challenge vs nine EndoVis editions) and they are, in principle, remediable by coordinated community action – exactly as surgery demonstrated. The surgical analogy is valid for this class of barrier, because data-ecosystem maturity is largely independent of the physical platform and can be transferred across fields.
Sensor and physical modelling limitations
The second class is intrinsic to flexible endoscopy and is not resolved by data sharing alone. Robotic laparoscopes are rigid, kinematically constrained, and instrumented with joint encoders that yield direct position, velocity, and force readouts; their forward kinematics are exactly modellable. Flexible endoscopes have no equivalent embedded sensing: The shaft deforms non-linearly, forms loops, and adopts redundant insertion configurations, so the same tip pose can arise from many shaft states and motion must be inferred indirectly from the monocular video stream by optical flow, depth estimation, or landmark recognition. This inference problem is ill-posed in ways that encoder-based surgical kinematics is not, and it would persist even with abundant data. For this class of barrier the surgical analogy is therefore one of trajectory, not of technical equivalence: Progress will require advances in endoscope-specific sensing hardware and physical modelling, not only a larger data ecosystem.
The clinical importance of preoperative anatomical mapping has been illustrated in recent laparoscopic experience. Anatomical anomalies such as adult intestinal malrotation can necessitate non-standard reconstruction techniques to avoid limb torsion, obstruction, and anastomotic leak[84], and postoperative dynamic intestinal obstruction occurs in approximately one in five patients after laparoscopic colorectal surgery, independently predicted by preoperative obstruction and prior abdominal surgery[85]. Accurate preoperative spatial mapping – the function that endoscopic SLAM/3D reconstruction is designed to support – is therefore a clinical motivation for the sensor and physical modelling advances called for above.
Implications for clinical practice, regulation, and ecosystem development
Beyond ESD, kinematic AI may also become relevant to therapeutic endoscopy in urgent settings. Endoscopic self-expanding stent placement is already an established, effective and safer alternative to emergency surgery for acute malignant colorectal obstruction, with substantially lower complication rates than emergency laparotomy in recent comparative data[86]. Workflow-recognition and phase-segmentation models analogous to those validated in ESD could, in principle, support such time-critical interventions by identifying procedural phases and flagging deviations during stent deployment.
At the clinical level, the additive benefit of kinematic monitoring is now established: The Yao et al[9] (2022) four-arm trial demonstrated that CADe alone improved ADR from 14.8% to 21.3%, but adding CAQ further increased ADR to 30.6%, confirming that exposure errors and detection errors are independent failure modes. This benefit is not uniform. Barua et al[28] (2023) showed that withdrawal-speed monitoring provided no additional benefit at high-baseline-ADR centers (baseline approximately 45%), establishing a ceiling effect with practical implications for resource allocation. Kinematic AI should be prioritized where baseline quality indicators are moderate or low, and where the marginal gain is greatest. The Lu et al[29] (2023) study observed that time-of-day decline in ADR was eliminated in the AI-assisted arm. As the reviewers correctly emphasized, the AI-assisted arm in Lu et al[29] combined CADe, CAQ, and CADe + CAQ subgroups, so the observed effect cannot be attributed to kinematic AI alone. It does, however, support the broader hypothesis that AI-assisted monitoring may attenuate the clinical impact of endoscopist fatigue; dedicated CAQ-only and CADe-only comparative trials are required to isolate the contribution of each modality. Even at high-baseline-ADR centers where the ceiling effect limits incremental ADR gains, kinematic AI may retain non-detection clinical and administrative value: Examples include sustaining inspection quality across long shifts to mitigate endoscopist fatigue, consistent with the time-of-day stabilisation reported by Lu et al[29], and providing automated time-stamped documentation of withdrawal time, photodocumentation, and inspection completeness for quality-audit and medicolegal purposes, as prototyped by Lux et al[56]. At such centers, kinematic AI is therefore better framed as a QA and procedural-audit layer than as a detection augmenter.
At the regulatory level, a structural barrier persists. Current AI-as-a-medical-device frameworks are optimized for diagnostic outputs (e.g., per-frame polyp probability), where the clinical decision pathway is well defined: (1) The AI flags; and (2) The physician decides. Kinematic AI operates on a different paradigm. Autonomous navigation requires the AI to act on physical endoscope control; SLAM-based coverage mapping must define what counts as “adequate” examination in real time; and workflow recognition must trigger alerts at procedurally critical moments. None of these functions fit neatly into existing SaMD classification tiers. The absence of any standalone regulatory approval for kinematic AI – despite 8 RCTs and technically advanced algorithms – reflects not only an evidence gap but a regulatory framework gap. Defining kinematic AI-specific regulatory pathways, analogous to what the FDA established for AI-enabled robotic surgical systems [e.g., Moon Surgical 510(k) clearance][15], is a prerequisite for clinical translation. As of late 2025, no FDA, European Medicines Agency, or European Commission instrument specifically targets motion-acting AI in flexible GIE. The FDA’s final guidance on Predetermined Change Control Plans for AI-enabled device software functions (December 2024)[87], article 6 (1) of the European AI act (Regulation 2024/1689, in force August 2024)[88], and the joint Medical Device Coordination Group/AI-Board guidance Medical Device Coordination Group 2025-6 (June 2025)[89] cover medical-device AI generically and do not name endoscopic motion as a distinct subdomain. The closest authorised motion-acting precedent – Moon Surgical’s ‘Physical AI’ ScoPilot, FDA-cleared in 2025 with a Predetermined Change Control Plan – is indicated for rigid-laparoscope robotic control, not flexible endoscopy[90], and the only regulatory authorisation globally that explicitly bundles withdrawal-speed and blind-spot quality-control functions with polyp CADe is the Chinese NMPA approval of the ENDOANGEL platform in May 2023[91]. No FDA or European Medicines Agency workshop, request for information, or task force naming kinematic endoscopic AI as a regulatory focus has been identified to date, indicating that motion-acting endoscopic AI remains a regulatory void addressed only by analogy to diagnostic AI and surgical-robotics autonomy frameworks, awaiting the first dedicated instrument.
At the ecosystem level, the comparison with surgery suggests that the rate-limiting factor is infrastructure, not algorithmic capability. Surgery’s translational trajectory was built on a self-reinforcing cycle: Open datasets (JIGSAWS, 2014)[10] enabled benchmarking competitions (EndoVis, 2015–2024)[11], which attracted industry investment[19], which funded clinical trials, which generated regulatory precedents, which in turn attracted further investment. GIE kinematic AI has not entered this cycle. The field has one public instrument dataset (590 images)[51] vs surgery’s six or more; no annual GI-specific benchmarking challenge vs nine EndoVis editions; and a single academic platform (ENDOANGEL) dominating the clinical literature. Breaking into this cycle requires coordinated, simultaneous action on multiple fronts – open dataset creation, challenge organization, and industry engagement – rather than incremental algorithm improvement.
Future perspectives
Five priority research gaps warrant focused effort: (1) Colonoscopy 3D coverage RCT: SLAM technology can now quantify mucosal coverage in real time, a trial testing whether coverage feedback reduces adenoma miss rates is feasible with current hardware and should be conducted outside China to address external validity; (2) GI instrument tracking: 590 annotated images vs thousands in surgical endoscopy, a multi-center annotation initiative for ESD knives, ERCP instruments, and hemostasis clips is needed; (3) Flexible endoscopy skill assessment: Combined DL and electromagnetic kinematic-sensor analysis should replicate the JIGSAWS paradigm; (4) Laparoscopic Endoscopic Cooperative Surgery procedure AI: Zero studies were identified despite clinical use of Laparoscopic Endoscopic Cooperative Surgery since 2008; and (5) AI-based colonoscope loop detection and correction remains unexplored. Future open datasets for flexible GIE should prioritise synchronised multimodal recordings – endoscopic video time-locked to magnetic endoscope imaging[38], force–torque traces from the insertion port, and electromagnetic-tracker pose, to enable JIGSAWS-style cross-platform skill and trajectory benchmarking comparable to that established for robotic surgery[10].
Strengths and limitations
This is the first systematic review to systematically map AI applications targeting endoscope kinematics across GIE domains. Key strengths include the coverage of four databases spanning both clinical and engineering literature (capturing the 62% of studies published in non-medical venues), an eight-domain classification enabling cross-domain translational comparison, risk-of-bias assessment across heterogeneous designs using RoB 2, QUADAS-2, and PROBAST + AI, the first GRADE certainty rating for kinematic AI, and quantitative benchmarking against surgical AI.
Limitations warrant emphasis. First, all eight RCTs of kinematic AI originated from Chinese centers and six used the ENDOANGEL platform; geographic and platform generalizability is therefore unproven, and regional health-system factors may have contributed to the observed effects. Several non-exclusive factors may account for this concentration. High colonoscopy volumes at large Chinese tertiary centres provide both the data scale needed to train motion-recognition models and the throughput needed to recruit pragmatic single-system trials rapidly. In addition, the dominant platform was developed and trialled by a single vertically integrated academic group, which compresses the conventional separation between algorithm development, clinical validation, and trial sponsorship and accelerates iteration. These same factors, however, reinforce rather than resolve the platform-dependence concern, since the observed effects remain entangled with one platform and one health-system context, and independent multi-platform Western trials are needed before the magnitude of benefit can be regarded as generalizable. Second, heterogeneous study designs precluded quantitative meta-analysis, and certainty estimates for engineering-stage domains should be interpreted as pre-clinical rather than as indicative of weak evidence per se. Third, eight included studies were arXiv preprints without peer review, reflecting the predominance of engineering literature in several domains. As detailed in the Methods, these preprints were eligible only where methodological detail was sufficient for data extraction and formal risk-of-bias appraisal with PROBAST + AI, and were not exempted from quality assessment; nonetheless, their non-peer-reviewed status remains a limitation. Fourth, our definition of kinematic AI is broader than pure motion-sensor analysis and includes image-recognition-based proxy measures (e.g., anatomical-landmark blind-spot monitoring, workflow-phase recognition); under a narrower, purely motion-based definition, coverage-mapping and workflow-recognition domains would be reclassified as primarily image-recognition tasks, which would reduce the number of domains considered “kinematic” from eight to four but would not alter the core conclusion of a substantial evidence imbalance. To make this scope explicit, we organised the included domains into a four-tier hierarchical taxonomy (Figure 2)[9] distinguishing true kinematic AI (Tier 1, direct motion signal), proxy kinematic inference (Tier 2, vision-derived motion), workflow intelligence (Tier 3, phase recognition), and procedural QA systems (Tier 4, landmark classification), and conducted a pre-specified sensitivity analysis under this Tier 1-2-only definition; the kinematic-to-CADe RCT ratio rose from approximately 5:1 to approximately 8:1, but the core conclusion of a substantial evidence imbalance was preserved, indicating robustness of the principal finding to the boundary of the kinematic definition. Fifth, the comparison with surgical AI is offered as a translational parallel; direct extrapolation of technical performance between rigid-instrument surgery and flexible endoscopy is constrained by differences in embedded kinematic sensing, as discussed above.
CONCLUSION
AI-based kinematic analysis in GIE addresses exposure errors that CADe cannot correct, yet has produced only a fraction of the clinical evidence generated for diagnostic AI. Withdrawal-speed monitoring and coverage/blind-spot mapping have reached moderate GRADE certainty and are ready for geographically diverse validation; workflow recognition and skill assessment are approaching clinical readiness; the remaining four domains are at the engineering stage and require ecosystem-level investment – open datasets, benchmarking challenges, and dedicated regulatory pathways – before meaningful clinical trials can be designed. Looking forward, the demonstrated additive benefit of co-deployed CADe and CAQ implies that integration should extend beyond software: Next-generation flexible endoscopes may need to embed dedicated kinematic sensing – built-in shaft-pose encoders, insertion-port force-torque measurement, and a unified data bus carrying synchronised video and motion streams to a single onboard detection-plus-quality AI module – so that exposure-error and recognition-error AI converge into one physically integrated procedural-quality system rather than two parallel software overlays running on the same scope.
Hassan C, Spadaccini M, Iannone A, Maselli R, Jovani M, Chandrasekar VT, Antonelli G, Yu H, Areia M, Dinis-Ribeiro M, Bhandari P, Sharma P, Rex DK, Rösch T, Wallace M, Repici A. Performance of artificial intelligence in colonoscopy for adenoma and polyp detection: a systematic review and meta-analysis.Gastrointest Endosc. 2021;93:77-85.e6.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 441][Cited by in RCA: 395][Article Influence: 79.0][Reference Citation Analysis (8)]
Spadaccini M, Iannone A, Maselli R, Badalamenti M, Desai M, Chandrasekar VT, Patel HK, Fugazza A, Pellegatta G, Galtieri PA, Lollo G, Carrara S, Anderloni A, Rex DK, Savevski V, Wallace MB, Bhandari P, Roesch T, Gralnek IM, Sharma P, Hassan C, Repici A. Computer-aided detection versus advanced imaging for detection of colorectal neoplasia: a systematic review and network meta-analysis.Lancet Gastroenterol Hepatol. 2021;6:793-802.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 111][Cited by in RCA: 103][Article Influence: 20.6][Reference Citation Analysis (4)]
Gong EJ, Bang CS, Lee JJ, Baik GH, Lim H, Jeong JH, Choi SW, Cho J, Kim DY, Lee KB, Shin SI, Sigmund D, Moon BI, Park SC, Lee SH, Bang KB, Son DS. Deep learning-based clinical decision support system for gastric neoplasms in real-time endoscopy: development and validation study.Endoscopy. 2023;55:701-708.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 61][Cited by in RCA: 52][Article Influence: 17.3][Reference Citation Analysis (1)]
Yao L, Zhang L, Liu J, Zhou W, He C, Zhang J, Wu L, Wang H, Xu Y, Gong D, Xu M, Li X, Bai Y, Gong R, Sharma P, Yu H. Effect of an artificial intelligence-based quality improvement system on efficacy of a computer-aided detection system in colonoscopy: a four-group parallel study.Endoscopy. 2022;54:757-768.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 105][Cited by in RCA: 97][Article Influence: 24.3][Reference Citation Analysis (3)]
Allan M, Shvets A, Kurmann T, Zhang ZC, Duggal R, Su YH, Rieke N, Laina I, Kalavakonda N, Bodenstedt S, Herrera L, Li WQ, Iglovikov V, Luo HL, Yang J, Stoyanov D, Maier-Hein L, Speidel S, Azizian M.
2017 Robotic Instrument Segmentation Challenge. 2019 Preprint. Available from: arXiv:1902.06426.
[PubMed] [DOI] [Full Text]
Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, Shamseer L, Tetzlaff JM, Akl EA, Brennan SE, Chou R, Glanville J, Grimshaw JM, Hróbjartsson A, Lalu MM, Li T, Loder EW, Mayo-Wilson E, McDonald S, McGuinness LA, Stewart LA, Thomas J, Tricco AC, Welch VA, Whiting P, Moher D. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews.BMJ. 2021;372:n71.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 9803][Reference Citation Analysis (0)]
Barua I, Misawa M, Glissen Brown JR, Walradt T, Kudo SE, Sheth SG, Nee J, Iturrino J, Mukherjee R, Cheney CP, Sawhney MS, Pleskow DK, Mori K, Løberg M, Kalager M, Wieszczy P, Bretthauer M, Berzin TM, Mori Y. Speedometer for withdrawal time monitoring during colonoscopy: a clinical implementation trial.Scand J Gastroenterol. 2023;58:664-670.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 18][Cited by in RCA: 15][Article Influence: 5.0][Reference Citation Analysis (11)]
Wu L, Zhang J, Zhou W, An P, Shen L, Liu J, Jiang X, Huang X, Mu G, Wan X, Lv X, Gao J, Cui N, Hu S, Chen Y, Hu X, Li J, Chen D, Gong D, He X, Ding Q, Zhu X, Li S, Wei X, Li X, Wang X, Zhou J, Zhang M, Yu HG. Randomised controlled trial of WISENSE, a real-time quality improving system for monitoring blind spots during esophagogastroduodenoscopy.Gut. 2019;68:2161-2169.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 285][Cited by in RCA: 240][Article Influence: 34.3][Reference Citation Analysis (5)]
Wu L, He X, Liu M, Xie H, An P, Zhang J, Zhang H, Ai Y, Tong Q, Guo M, Huang M, Ge C, Yang Z, Yuan J, Liu J, Zhou W, Jiang X, Huang X, Mu G, Wan X, Li Y, Wang H, Wang Y, Zhang H, Chen D, Gong D, Wang J, Huang L, Li J, Yao L, Zhu Y, Yu H. Evaluation of the effects of an artificial intelligence system on endoscopy quality and preliminary testing of its performance in detecting early gastric cancer: a randomized controlled trial.Endoscopy. 2021;53:1199-1207.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 152][Cited by in RCA: 132][Article Influence: 26.4][Reference Citation Analysis (0)]
Liu R, Yuan X, Huang K, Peng T, Pavlov PV, Zhang W, Wu C, Feoktistova KV, Bi X, Zhang Y, Chen X, George J, Liu S, Liu W, Zhang Y, Yang J, Pang M, Hu B, Yi Z, Ye L. Artificial intelligence-based automated surgical workflow recognition in esophageal endoscopic submucosal dissection: an international multicenter study (with video).Surg Endosc. 2025;39:2836-2846.
[RCA] [PubMed] [DOI] [Full Text][Cited by in RCA: 7][Reference Citation Analysis (0)]
Hwang B, Sohn DK, Joe S, Kim B. Autonomous Robotic Colonoscopy: A Supervised Learning Approach for Enhanced Navigation and Collision Detection.Adv Intell Syst. 2026;8:e202500841.
[PubMed] [DOI] [Full Text]
Corsi D, Marzari L, Pore A, Farinelli A, Casals A, Fiorini P, Dall'alba D.
Constrained Reinforcement Learning and Formal Verification for Safe Colonoscopy Navigation. 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); 2023 October 1-5, Detroit, United States. Bengaluru: IEEE, 2023.
[PubMed] [DOI] [Full Text]
Wang K, Yang C, Wang Y, Li S, Wang Y, Dou Q, Yang X, Shen W.
EndoGSLAM: Real-Time Dense Reconstruction and Tracking in Endoscopic Surgeries Using Gaussian Splatting. In: Linguraru MG, Dou Q, Feragen A, Giannarou S, Glocker B, Lekadir K, Schnabel JA, editors. Medical Image Computing and Computer Assisted Intervention – MICCAI 2024. Cham: Springer, 2024.
[PubMed] [DOI] [Full Text]
Turan M, Almalioglu Y, Gilbert HB, Mahmood F, Durr NJ, Araujo H, Sari AE, Ajay A, Sitti M. Learning to Navigate Endoscopic Capsule Robots.IEEE Robot Autom Lett. 2019;4:3075-3082.
[PubMed] [DOI] [Full Text]
Jha D, Ali S, Emanuelsen K, Hicks S, Thambawita VL, Garcia-ceja E, Riegler M, de Lange T, Schmidt PT, Johansen HD, Johansen D, Halvorsen P.
Kvasir-Instrument: Diagnostic and therapeutic tool segmentation dataset in gastrointestinal endoscopy. 2020 Preprint. Available from: Open Science Framework.
[PubMed] [DOI] [Full Text]
Ali S, Dmitrieva M, Ghatwary N, Bano S, Polat G, Temizel A, Krenzer A, Hekalo A, Guo YB, Matuszewski B, Gridach M, Voiculescu I, Yoganand V, Chavan A, Raj A, Nguyen NT, Tran DQ, Huynh LD, Boutry N, Rezvy S, Chen H, Choi YH, Subramanian A, Balasubramanian V, Gao XW, Hu H, Liao Y, Stoyanov D, Daul C, Realdon S, Cannizzaro R, Lamarque D, Tran-Nguyen T, Bailey A, Braden B, East JE, Rittscher J. Deep learning for detection and segmentation of artefact and disease instances in gastrointestinal endoscopy.Med Image Anal. 2021;70:102002.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 140][Cited by in RCA: 79][Article Influence: 15.8][Reference Citation Analysis (0)]
Jha D, Sharma V, Banik D, Bhattacharya D, Roy K, Hicks SA, Tomar NK, Thambawita V, Krenzer A, Ji GP, Poudel S, Batchkala G, Alam S, Ahmed AMA, Trinh QH, Khan Z, Nguyen TP, Shrestha S, Nathan S, Gwak J, Jha RK, Zhang Z, Schlaefer A, Bhattacharjee D, Bhuyan MK, Das PK, Fan DP, Parasa S, Ali S, Riegler MA, Halvorsen P, de Lange T, Bagci U. Validating polyp and instrument segmentation methods in colonoscopy through Medico 2020 and MedAI 2021 Challenges.Med Image Anal. 2025;99:103307.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 8][Cited by in RCA: 7][Article Influence: 7.0][Reference Citation Analysis (0)]
Wu L, Zhou W, Wan X, Zhang J, Shen L, Hu S, Ding Q, Mu G, Yin A, Huang X, Liu J, Jiang X, Wang Z, Deng Y, Liu M, Lin R, Ling T, Li P, Wu Q, Jin P, Chen J, Yu H. A deep neural network improves endoscopic detection of early gastric cancer without blind spots.Endoscopy. 2019;51:522-531.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 229][Cited by in RCA: 182][Article Influence: 26.0][Reference Citation Analysis (4)]
Chen D, Wu L, Li Y, Zhang J, Liu J, Huang L, Jiang X, Huang X, Mu G, Hu S, Hu X, Gong D, He X, Yu H. Comparing blind spots of unsedated ultrafine, sedated, and unsedated conventional gastroscopy with and without artificial intelligence: a prospective, single-blind, 3-parallel-group, randomized, single-center trial.Gastrointest Endosc. 2020;91:332-339.e3.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 87][Cited by in RCA: 80][Article Influence: 13.3][Reference Citation Analysis (3)]
Lazo JF, Lai C, Moccia S, Rosa B, Catellani M, de Mathelin M, Ferrigno G, Breedveld P, Dankelman J, De Momi E.
Autonomous Intraluminal Navigation of a Soft Robot using Deep-Learning-based Visual Servoing. 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); 2022 October 23-27, Kyoto, Japan. Bengaluru: IEEE, 2022.
[PubMed] [DOI] [Full Text]
Pore A, Finocchiaro M, Dall'alba D, Hernansanz A, Ciuti G, Arezzo A, Menciassi A, Casals A, Fiorini P.
Colonoscopy Navigation using End-to-End Deep Visuomotor Control: A User Study. 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); 2022 October 23-27, Kyoto, Japan. Bengaluru: IEEE, 2022.
[PubMed] [DOI] [Full Text]
Tan M, Tao Y, Zheng B, Xie G, Feng L, Xia Z, Xiong J. Safe navigation for robotic digestive endoscopy via human intervention-based reinforcement learning.Expert Syst Appl. 2025;294:128841.
[PubMed] [DOI] [Full Text]
Shi Y, Lu B, Liu JW, Li M, Shou MZ.
ColonNeRF: High-Fidelity Neural Reconstruction of Long Colonoscopy. 2023 Preprint. Available from: arXiv:2312.02015.
[PubMed] [DOI] [Full Text]
Elvira R, Tardós JD, Montiel JMM.
CudaSIFT-SLAM: multiple-map visual SLAM for full procedure mapping in real human endoscopy. 2024 Preprint. Available from: arXiv:2405.16932.
[PubMed] [DOI] [Full Text]
Zhang Y, Bai L, Liu L, Ren H, Meng MQ.
Deep Reinforcement Learning-Based Control for Stomach Coverage Scanning of Wireless Capsule Endoscopy. 2022 IEEE International Conference on Robotics and Biomimetics (ROBIO); 2022 December 5-9, Jinghong, China. Bengaluru: IEEE, 2022.
[PubMed] [DOI] [Full Text]
Ng C, Gao H, Ren T, Lai J, Ren H.
Navigation of Tendon-driven Flexible Robotic Endoscope through Deep Reinforcement Learning. 2024 IEEE International Conference on Advanced Robotics and Its Social Impacts (ARSO); 2024 May 20-22, Hong Kong, China. Bengaluru: IEEE, 2024.
[PubMed] [DOI] [Full Text]
Zhang X, Zhang Q, Chen J, Zhou C, Wang Y, Zhang Z, Li X, Qian D. SPRMamba: Surgical Phase Recognition for Endoscopic Submucosal Dissection With Mamba.IEEE Sensors J. 2026;26:8887-8898.
[PubMed] [DOI] [Full Text]
Kaleta J, Smolak-dyżewska W, Malarz D, Dall’alba D, Korzeniowski P, Spurek P.
PR-ENDO: Physically Based Relightable Gaussian Splatting for Endoscopy. In: Gee JC, Alexander DC, Hong J, Iglesias JE, Sudre CH, Venkataraman A, Golland P, Kim JH, Park J, editors. Medical Image Computing and Computer Assisted Intervention – MICCAI 2025. Cham: Springer, 2025.
[PubMed] [DOI] [Full Text]
United States Food and Drug Administration.
Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions. Final Guidance for Industry and Food and Drug Administration Staff. 2024. Available from: https://www.fda.gov/media/166704/download.
[PubMed] [DOI]
European Parliament and Council of the European Union.
Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence. 2024. Available from: https://pacmap.dev/regulation/eu-ai-act-2024.
[PubMed] [DOI]
United States Food and Drug Administration.
510(k) Premarket Notification K242323: Maestro System (Moon Surgical SAS), cleared 14 March 2025; and K250984: Maestro System with ScoPilot Predetermined Change Control Plan, cleared 27 June 2025. 2025. Available from: https://www.accessdata.fda.gov/scripts/cdrh/cfdocs/cfPMN/pmn.cfm?ID=K242323.
[PubMed] [DOI]
National Medical Products Administration.
Computer Aided Detection Software for Intestinal Polyps in Lower Gastrointestinal Endoscopy Approved for Marketing. 2023. Available from: https://english.nmpa.gov.cn/2023-05/12/c_923513.htm.
[PubMed] [DOI]
Footnotes
Peer review: Externally peer reviewed.
Peer-review model: Single blind
Specialty type: Gastroenterology and hepatology
Country of origin: South Korea
Peer-review report’s classification
Scientific quality: Grade B, Grade B
Novelty: Grade A, Grade C
Creativity or innovation: Grade A, Grade B
Scientific significance: Grade B, Grade B
P-Reviewer: Liu S, MD, China; Vaithiyam V, Assistant Professor, DM, MD, India S-Editor: Luo ML L-Editor: A P-Editor: Wang CH