BPG is committed to discovery and dissemination of knowledge
Minireviews Open Access
Copyright: ©Author(s) 2026. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution-NonCommercial (CC BY-NC 4.0) license. No commercial re-use. See permissions. Published by Baishideng Publishing Group Inc.
World J Transl Med. Sep 28, 2026; 12(3): 125559
Published online Sep 28, 2026. doi: 10.5528/wjtm.125559
Artificial intelligence for fluoroscopic cholangiogram interpretation during endoscopic retrograde cholangiopancreatography
Ahmed Salman, Internal Medicine, Kasr Alainy School of Medicine, Cairo 11562, Egypt
Mohamed AbdAlla Salman, General Surgery, Kasr Alainy School of Medicine, Cairo 11562, Egypt
ORCID number: Ahmed Salman (0000-0003-0026-0841); Mohamed AbdAlla Salman (0000-0001-5445-6415).
Author contributions: Salman A conceived the review, conducted the literature search, interpreted the evidence, drafted the manuscript, and critically revised it; Salman MA appraised the literature, contributed surgical and procedural interpretation, revised the manuscript, and approved the final version; both authors accept accountability for the work.
AI contribution statement: OpenAI Codex (GPT-5) was used to assist with language editing, structural organization, and drafting revisions in response to peer-review comments. The authors critically reviewed and revised all artificial-intelligence-assisted text, verified the cited sources and numerical statements, and retain full responsibility for the manuscript’s interpretation, accuracy, originality, integrity, citations, and conclusions. No artificial intelligence tool was used to generate research data or perform clinical decision-making.
Conflict-of-interest statement: All the authors report no relevant conflicts of interest for this article.
Corresponding author: Ahmed Salman, FRACP, FRCP, Internal Medicine, Kasr Alainy School of Medicine, 1 Al-Saray Street, Al-Manial, Cairo 11562, Egypt. awea844@gmail.com
Received: July 10, 2026
Revised: July 23, 2026
Accepted: July 29, 2026
Published online: September 28, 2026
Processing time: 55 Days and 11.7 Hours

Abstract

Endoscopic retrograde cholangiopancreatography (ERCP) is a dominant therapeutic approach for biliary and pancreatic ductal diseases, but it remains technically demanding, operator-dependent, and primarily fluoroscopy-guided. The cholangiogram produced during each ERCP is the central image on which intraprocedural decisions rest, but it is interpreted subjectively, with appreciable interobserver variability. Unlike luminal mucosa or cross-sectional imaging, fluoroscopic cholangiography has remained a relatively underexploited area for artificial intelligence (AI). This narrative review outlines the potential applications of AI, particularly deep learning, in interpreting fluoroscopic cholangiograms during ERCP. Reported applications include differentiating malignant from benign biliary strictures, detecting and quantifying common bile duct stones, automated scoring of stone-extraction difficulty, stent-length selection, real-time cannulation and papilla guidance, predicting post-ERCP pancreatitis, and reducing radiation dose. Although early evidence is promising across several of these domains, it is largely based on small, single-center, retrospective datasets, with limited external validation and no regulatory-approved systems. Because erroneous output may immediately influence an invasive decision, the most credible near-term role is assistive: Confidence-aware, physician-in-the-loop support evaluated prospectively for safety, workflow, generalizability, governance, equity, privacy, and patient-important outcomes.

Key Words: Artificial intelligence; Biliary tract; Cholangiography; Endoscopic retrograde cholangiopancreatography; Fluoroscopy; Patient safety

Core Tip: Artificial intelligence may help transform fluoroscopic cholangiogram interpretation during endoscopic retrograde cholangiopancreatography (ERCP) from subjective visual assessment to objective, reproducible decision support. Current applications include biliary stricture characterization, bile duct and stone segmentation, stone-extraction difficulty scoring, stent-length selection, papilla and cannula localization, post-ERCP pancreatitis prediction, and potential radiation-dose reduction. However, most evidence remains retrospective and early-stage. Future progress requires prospective validation, external testing, explainability, workflow integration, data governance, and multimodal physician-in-the-loop systems that combine fluoroscopy with endoscopic, endosonographic, and clinical data.



INTRODUCTION

Endoscopic retrograde cholangiopancreatography (ERCP) is a cornerstone therapeutic procedure for biliary and pancreatic ductal disease, but it remains one of the most technically challenging and operator-dependent procedures in gastrointestinal endoscopy[1]. In contrast to luminal endoscopy, which allows direct examination of the mucosa, ERCP relies largely on fluoroscopy, and the cholangiogram serves as the principal map by which the endoscopist localizes pathology, plans cannulation, and selects therapy[1,2]. The interpretive burden is high. Interpretive error may contribute to incomplete stone clearance, missed malignant obstruction, inappropriate device selection, or procedure-related harm, although its contribution has not been quantified[1,2]. As case complexity increases and access to advanced expertise remains uneven, interest is increasing in tools that could support more objective and reproducible interpretation of cholangiographic images.

Despite its central role, the cholangiogram remains subjective, experience-dependent, and subject to interobserver variability. Indeterminate biliary strictures remain difficult to characterize confidently on imaging alone, and assessment of stone size, number, and ductal configuration, which are key factors influencing extraction difficulty, remains largely subjective and experience-dependent[2,3]. These limitations are not failures of the modality itself, but rather reflect the difficulty of extracting quantitative information from a low-contrast, two-dimensional projection in real time.

Artificial intelligence (AI), particularly deep learning, has been increasingly adopted in luminal endoscopy, with computer-aided detection systems for colorectal neoplasia having entered clinical practice in some settings[3]. Nevertheless, in pancreaticobiliary endoscopy, AI remains at a nascent stage, and the literature to date has focused mainly on endoscopic ultrasound, digital cholangioscopy, and white-light images of the papilla rather than on the fluoroscopic image itself[1,3]. Thus, the cholangiogram, which is generated during ERCP and may be stored as still images or video, remains largely underexploited.

This is beginning to change. Initial studies have shown that deep learning models can differentiate malignant from benign biliary strictures directly from ERCP fluoroscopy images[4] and that automated systems trained on multicenter cholangiogram datasets can segment the duct, stones, and the duodenoscope to assess the technical difficulty of stone extraction[5]. Overall, this suggests that fluoroscopic cholangiograms are amenable to computational analysis, as in other domains of endoscopy, although questions remain regarding dataset size, external validation, and clinical readiness[6].

This narrative review focuses on the emerging use of AI in interpreting fluoroscopic cholangiograms during ERCP. We outline the theoretical foundations of the field, briefly summarize current diagnostic, procedural, and predictive applications, and critically evaluate the limitations and validation gaps that impede the translation of promising pilot data into routine clinical practice.

METHODS

This article was planned as a narrative review rather than a systematic review or meta-analysis, given the nature of the available evidence. Studies applying AI to fluoroscopic cholangiography are relatively new and methodologically heterogeneous, spanning diverse clinical settings, neural network architectures, and outcome measures that are not suitable for statistical pooling. To preserve transparency and methodological rigor despite the narrative design, the review was structured according to the Scale for the Assessment of Narrative Review Articles, with particular attention to the rationale for the study question, description of the literature search, and appraisal of the quality of the available evidence[7].

A structured literature search was performed in PubMed/MEDLINE, Scopus, and Web of Science, with support from IEEE Xplore, to capture engineering and computer science articles that may not be fully indexed in biomedical databases. Search strings combined technology terms, including AI, deep learning, machine learning, convolutional neural network (CNN), and radiomics, with clinical context terms, including ERCP, cholangiography, cholangiogram, and fluoroscopy, and relevant target terms, including bile duct, stone, stricture, stent, papilla, cannulation, radiation, ethics, governance, liability, bias, and physician in the loop. The search was limited to full-text English-language publications from January 2018 through July 16, 2026, corresponding to the period during which deep-learning techniques became increasingly used in endoscopic imaging.

Eligible studies applied an AI-based approach to fluoroscopic, cholangiographic, or peri-procedural ERCP imaging for diagnostic, procedural, or predictive purposes. Studies restricted to digital single-operator cholangioscopy images or magnetic resonance cholangiopancreatography (MRCP) were excluded from the primary synthesis and cited only when they provided important comparative or contextual framing, as their source data differ substantially from the projectional fluoroscopic images that are the focus of this review. Reference lists of included studies and relevant review articles were manually searched to identify additional sources.

Instead of pooling results, the evidence was organized qualitatively and by clinical task. For each study, particular attention was given to dataset size, single-versus multicenter provenance, the presence or absence of external validation, real-time feasibility, and overall clinical readiness. This task-focused, appraisal-based approach was selected to communicate not only what AI can accomplish with the cholangiogram but also how close each application is to bedside use.

TECHNICAL FOUNDATIONS: TEACHING MACHINES TO READ FLUOROSCOPY

The fluoroscopic cholangiogram is not a stand-alone component, but one element of an increasingly integrated biliopancreatic imaging workup, in which ERCP and endoscopic ultrasound are deliberately combined to maximize diagnostic and therapeutic yield[8]. Therefore, understanding how AI may interpret the cholangiogram begins with recognizing the computational principles shared across this imaging landscape. Conventional machine learning models rely on features that investigators hand-engineer and select before training, whereas deep learning models ingest images directly and generate their own hierarchical representations, an approach better suited to the variability of medical imaging[6]. CNNs are a widely used deep-learning paradigm in this context; they learn spatial features across successive layers and can be trained to classify an entire image, localize structures within it, or label each pixel. This flexibility makes the projectional cholangiogram, despite its noise and operator-dependent acquisition, a tractable computational target.

Four architectural families are particularly relevant to fluoroscopic interpretation. Image classification networks assign a global label to a cholangiogram, such as malignant vs benign strictures[4]. Semantic segmentation networks, such as U-Net and D-LinkNet, segment structures, including the bile duct, stones, and duodenoscope, pixel-by-pixel and provide the geometric information on which downstream tasks, such as difficulty scoring, depend[5]. Object-detection systems, including those based on transformer-assisted object-detection models, can localize small targets such as the duodenal papilla and cannula in real time during the procedure[9]. Finally, radiomics generates numerous quantitative texture and shape descriptors from a focal region of interest, which can then be fed into predictive models, as applied to the papilla region to predict the risk of complications[10].

Fluoroscopic images have challenges that distinguish them from many other image types used in endoscopic AI, such as white-light endoscopy and cross-sectional imaging. Fluoroscopic images are two-dimensional projections that collapse three-dimensional ductal anatomy, with low intrinsic contrast and frequent superimposition of the spine, endoscope, and dynamically filling contrast medium. Image quality is further influenced by the radiation-dose trade-off of fluoroscopy, patient movement, and substantial variability in acquisition settings across operators and units. These acquisition differences may contribute to transportability gaps. In one biliary-stricture classifier, strong internal discrimination decreased across independent cohorts[4]; this single example should not be generalized to all models or tasks.

Two additional factors determine whether such models can be trusted and reproduced. The first is interpretability. Saliency maps can show which image regions are associated with a prediction, but they do not by themselves establish causal or clinically valid reasoning[4]. The second is data: Deep-learning methods are constrained by the size, quality, and representativeness of their training datasets, while the labor-intensive expert annotation required for fluoroscopic images remains a major bottleneck. Clinical translation also requires representative data, independent external validation, prospective workflow evaluation, and analysis of failure cases[6].

DIAGNOSTIC APPLICATIONS: STONES, STRICTURES, AND DUCTAL ANATOMY

To date, the diagnostic application of cholangiographic AI has primarily focused on characterizing biliary strictures. A deep-learning classifier trained on ERCP fluoroscopy images from three German centers differentiated malignant from benign biliary strictures, achieving a cross-validation area under the receiver operating characteristic curve of 0.89. However, performance decreased to 0.72 and 0.76 in two independent external cohorts[4]. This pattern, strong internal discrimination followed by reduced performance on external validation, reflects the current maturity of the field and frames how the remaining diagnostic evidence should be interpreted.

Comparing these projectional-image results with those obtained through direct visualization is instructive. A meta-analysis of AI applied to digital single-operator cholangioscopy reported a pooled sensitivity of 95% and specificity of 88% for identifying malignancy in indeterminate biliary strictures[11]. This represents strong performance in a different input domain and should not be compared directly with fluoroscopy because patient selection, reference standards, and available imaging information differ.

Stone disease is the second major diagnostic target. The cholangiogram is particularly well suited to automated analysis in this setting because stones, ducts, and instruments are geometrically distinct. A deep-learning system trained on 1954 cholangiograms segmented the common bile duct, stones, and duodenoscope with mean intersection-over-union values of 86.4%, 68.4%, and 95.9%, respectively, providing quantitative measurements of stone size and ductal configuration on which subsequent assessment depends[5]. The lower mean intersection-over-union for stones compared with the duct or scope reflects their small size, variable contrast conspicuity, and frequent partial obscuration by contrast, which is a common challenge in projectional imaging.

Where fluoroscopy-specific evidence is limited, cross-sectional cholangiographic AI provides useful contextual information, although its input data differ substantially. In MRCP, deep-learning models have detected common bile duct stones with accuracy approaching that of radiologists, decreasing from 94% for solitary stones to 70% for multiple stones[12]. Similarly, automated recognition of primary sclerosing cholangitis-compatible ductal change has been performed on three-dimensional MRCP data, with sensitivities and specificities of 95% and 91%, respectively[13]. These studies demonstrate feasibility in different imaging modalities and should not be treated as direct performance benchmarks for fluoroscopic cholangiography.

A unifying theme across these diagnostic applications is that accurate image segmentation provides the foundation for higher-order interpretation. Whether the downstream goal is malignancy prediction, stone quantification, or anatomical mapping, the model must first distinguish the duct, lesion, and instrument from a low-contrast, superimposed background[5]. The technical feasibility of cholangiogram analysis has therefore been demonstrated, but is currently constrained by limited datasets, largely single-task models, and insufficient evidence of robustness for routine clinical application[4,11].

PROCEDURAL AND PREDICTIVE APPLICATIONS: FROM CANNULATION TO COMPLICATION RISK

In addition to diagnosis, the most important clinical utility of cholangiographic AI lies in the procedure itself, beginning with cannulation, a technically demanding step and an important determinant of adverse-event risk. From papilla localization to complication prediction, AI may assist at several procedural touchpoints, as summarised in Figure 1. An early CNN system trained on white-light endoscopic images acquired during ERCP localized the ampulla with a mean intersection-over-union score of 64.1% and classified cannulation as easy or difficult, with recall values of 71.9% and 61.1%, respectively, demonstrating feasibility for image-based localization and difficulty classification[14].

Figure 1
Figure 1 Artificial intelligence-assisted workflow for fluoroscopic cholangiogram interpretation during endoscopic retrograde cholangiopancreatography. Fluoroscopic images undergo quality control, segmentation, and feature extraction before task-specific diagnostic, procedural, or predictive output. A clinical safety gate applies confidence thresholds, out-of-distribution alerts, fail-safe abstention, audit logging, and immediate physician override. The endoscopist retains final interpretive and therapeutic authority and integrates the output with endoscopic, endosonographic, clinical, and procedural information. This original schematic was created specifically for this manuscript and contains no patient material or re-used third-party content. AI: Artificial intelligence; ERCP: Endoscopic retrograde cholangiopancreatography.

More recent transformer-assisted detectors have reported higher object-detection performance on a separate dataset, achieving a mean average precision of 93.2% while also computing the planar distance and direction of movement required for the cannula to access the orifice[9]. The result suggests potential real-time navigation support, but it does not establish clinical efficacy and is not directly comparable with the earlier model because the datasets and performance metrics differed.

After duct access is achieved, AI may guide therapeutic strategy by quantifying the difficulty of stone extraction. The previously described deep-learning difficulty-scoring system translated cholangiographic segmentations into a clinically relevant scale, in which scores of 2 or greater were associated with significantly lower complete-clearance rates (36% vs 86% for lower scores) and a greater need for endoscopic papillary large-balloon dilation[5]. By estimating the likelihood of difficult clearance before instrumentation, such scores may help guide accessory selection, referral decisions, or staged management, and illustrate how a diagnostic segmentation task can directly support procedural planning.

AI-assisted stent-length selection has moved beyond conceptual discussion. In a 2026 model development and validation study, an AI workflow identified 121 of 124 common bile duct strictures and selected a stent length within 1 cm of the reference in 104 of 121 cases (85.95%), with a mean absolute error of 0.81 cm[15]. Radiation exposure was reduced by approximately 202 mGy.cm2 per patient, with the greatest improvement among less-experienced endoscopists. These findings support the use of AI as a measurement and selection aid, but not for autonomous device deployment.

Relevant procedural labels also arise from non-artificial-intelligence clinical studies. In 596 patients with common bile duct stones, type III papillary morphology independently increased the odds of difficult cannulation (odds ratio, 2.255)[16]. This association can inform model development and risk stratification, but should not be converted into an automated treatment rule without prospective testing.

A further procedural application is radiation reduction. Evidence remains limited, but the stent-length selection study reported a mean reduction in dose-area product of approximately 202 mGy.cm per patient[15]. This is task-specific evidence from an assistive workflow and should not be interpreted as proof that AI broadly reduces radiation across ERCP.

The predictive dimension is illustrated by post-ERCP pancreatitis, a clinically important adverse event. Researchers developed a radiomics-based predictive model using quantitative features from white-light papillary images. The model reported areas under the curve of 0.825-0.857 across the training, testing, and validation cohorts[10].

García-Marmolejo et al[17] examined the timing of ERCP in acute cholangitis but did not evaluate AI[17]. Its relevance is contextual: An image-only model cannot determine urgency or procedural timing without clinical severity, laboratory, hemodynamic, and resource information. This reinforces the need for multimodal decision support.

Additionally, broader reviews in this area highlight the potential of predictive analytics that combines imaging data with clinical information. Such approaches could help identify patients at higher risk of complications and enable more personalized peri-procedural management. However, these models still require prospective testing and external validation before they can be confidently used in routine clinical practice[18].

The diagnostic, procedural, and predictive tools discussed in this and the previous section share a common developmental pathway. They have evolved from retrospective, single-center proof-of-concept studies toward the real-time, validated integration required for clinical use. Table 1 outlines the clinical tasks, computational methods, best reported performance, validation status, and readiness for implementation of these tools, providing the basis for the critical appraisal that follows.

Table 1 Artificial intelligence applications relevant to fluoroscopic cholangiogram interpretation and peri-procedural endoscopic retrograde cholangiopancreatography imaging.
Clinical task
Imaging input
AI approach
Best reported performance
Dataset and validation
Readiness
Malignant vs benign stricture characterization[4]Fluoroscopic cholangiogramClassification CNN + saliency mappingAUROC 0.89 (internal); 0.72-0.76 (external)251 patients, 3 centers; internal CV + 2 external cohortsExperimental-externally validated
CBD stone and duct segmentation[5]Fluoroscopic cholangiogramSemantic segmentation (D-LinkNet/U-Net)mIoU 86.4% (duct), 68.4% (stone)1,954 cholangiograms, 3 hospitals; internal validationExperimental-multicenter
Stone-extraction difficulty scoring[5]Fluoroscopic cholangiogramSegmentation-derived scoring scaleScore ≥ 2 → 36% vs 86% complete clearanceAs aboveExperimental-multicenter
Biliary stent-length selection[15]Fluoroscopic cholangiogramSegmentation and geometric measurement104 of 121 selections within 1 cm; mean absolute error 0.81 cm; dose-area product reduced by about 202 mGy.cm2794 images from 431 patients; model development and validationExperimental-assistive use only
Ampulla localization and cannulation-difficulty classification[14]White-light papillary imageClassification/detection CNNAmpulla mIoU 64.1%; difficulty recall 719% (easy)/61.1% (difficult)531/451 patients, single center; internalExperimental-proof-of-concept
Real-time papilla and cannula navigation[9]White-light papillary imageSwin-transformer detector (4STDH)mAP 93.2%; outputs cannula distance/direction1840 images; single dataset + video testExperimental-proof-of-concept
Post-ERCP pancreatitis prediction[10]White-light papillary imagePapilla radiomics + machine learningAUC: 0.825-0.8572372 patients, 2 centers; multicohort validationExperimental-multicohort
DOES AI INCREASE RISK DURING ERCP?

Risk is not determined by procedural complexity alone. It depends on the clinical task, model autonomy, error detectability, reversibility, and downstream consequences. Because ERCP is a high-risk interventional endoscopic procedure in which image interpretation may immediately alter guidewire or cannula positioning, sphincterotomy, dilation, stone extraction, or stent selection, an undetected algorithmic error may have more immediate and less reversible consequences than an error in low-stakes offline image review.

Near-term systems should therefore remain assistive. They should display source images and uncertainty, detect out-of-distribution inputs, abstain when image quality is inadequate, preserve immediate physician override, and record model outputs and final actions. No study identified in this review evaluated autonomous control of biliary radiofrequency ablation or another irreversible therapy.

Evaluation should progress from silent local testing to external validation and prospective human-factors studies, with prespecified stopping rules and outcomes including procedure time, cannulation attempts, radiation dose, technical success, adverse events, override frequency, and near misses.

LIMITATIONS, VALIDATION GAPS, AND FUTURE DIRECTIONS

The most important limitation of cholangiographic AI is the chasm between internal promise and external reality. At present, most studies surveyed are retrospective and single- or few-center in scope, and when independent validation has been attempted, performance has declined. For example, the stricture classifier showed a reduction in discrimination from 0.89 internally to 0.72-0.76 externally[4]. The generalizability of machine-learning models for indeterminate biliary strictures has been specifically questioned, as models calibrated to the imaging features, patient mix, and labeling conventions of selected centers may not translate well to other settings[19]. Until external robustness is demonstrated routinely, internal estimates may overstate performance in new clinical settings.

This is compounded by methodological limitations in studies making clinical claims. The available literature is dominated by diagnostic-accuracy and proof-of-concept studies rather than prospective comparative evaluations. Future clinical trial reports should follow the Consolidated Standards of Reporting Trials-Artificial Intelligence extension, and protocols should follow the Standard Protocol Items: Recommendations for Interventional Trials-Artificial Intelligence extension[20,21]. Early-stage prospective evaluation may follow the Developmental and Exploratory Clinical Investigations of Decision-support systems driven by AI guideline, which emphasizes clinical performance, safety, usability, workflow, and human-artificial-intelligence interaction before large comparative trials[22]. These frameworks improve completeness of reporting and error analysis; they are not themselves risk-of-bias instruments[23].

Several practical challenges remain between current models and bedside use. Fluoroscopic images require manual expert annotation, datasets are heterogeneous across endoscope manufacturers and acquisition environments, and most pipelines have not been designed for real-time intraprocedural use. Large self-supervised foundation models trained on millions of luminal-endoscopy images may improve downstream accuracy and reduce task-specific annotation requirements, but GastroNet-5M is a luminal-endoscopy resource and not a substitute for ERCP fluoroscopy data[24]. Similarly, semi-automated annotation of live colonoscopy video provides only methodological context and requires ERCP-specific, multicenter external validation before translation[25].

Clinical translation will depend on more than model performance alone. The reviewed literature did not identify a regulatory-approved AI system for ERCP as of July 2026, and deployment raises concerns, including false positives, alarm fatigue, automation bias, interpretability, data privacy, and algorithmic bias[6,23]. The most defensible near-term role is assistive, physician-supervised decision support, which still requires prospective safety evaluation. Multimodal integration is a promising future direction: The fluoroscopic cholangiogram may be combined with cholangioscopic, endosonographic, clinical, and laboratory data in an integrated diagnostic and predictive framework that may eventually be supported by foundation-model architectures[23,24]. The central task is therefore validation, standardization, and responsible clinical integration rather than further proof-of-concept demonstrations.

ETHICAL, LEGAL, DATA GOVERNANCE, AND EQUITY CONSIDERATIONS

The World Endoscopy Organization consensus groups the central implementation challenges into data governance, medicolegal implications, and equity and bias[26]. For ERCP, governance should define data stewardship, lawful use, retention, access, de-identification, secondary use, cybersecurity, change control, and incident response. Model, training data, and label provenance should be documented.

The World Medical Association endorses physician oversight of AI in medical care[27]. A licensed physician should review model output and retain final clinical authority, but legal responsibility may also rest with developers, institutions, and other parties according to jurisdiction and control. When AI may materially influence invasive care, patients should receive proportionate information about its role, limitations, data use, and oversight.

Aggregate performance may conceal poorer results in underrepresented groups, altered anatomy, uncommon indications, low-resource settings, or centers using different equipment. Developers should report the dataset composition and the performance of clinically relevant subgroups. Institutions should require local validation, equitable access, and clinician training that preserves the unaided skills required during system failure.

CONCLUSION

AI is an emerging area for interpreting fluoroscopic cholangiograms during ERCP. Current evidence suggests that AI may facilitate differentiation of malignant from benign biliary strictures, segmentation of the bile duct and stones, estimation of stone-extraction difficulty, stent-length selection, papilla and cannula localization, and prediction of post-ERCP pancreatitis. A task-specific assistive workflow has also been reported to reduce radiation exposure. Nevertheless, most studies remain retrospective, small or selected, single- or few-center, and incompletely externally validated; the reviewed literature did not identify a regulatory-approved ERCP system as of July 2026.

A plausible future direction is integration into a multimodal, physician-in-the-loop workflow combining fluoroscopy with cholangioscopy, endoscopic ultrasound, clinical data, and procedural outcomes. Before clinical use, future systems must demonstrate robust external validity, real-time feasibility, explainability, workflow compatibility, and benefits for clinical decision-making and patient-important outcomes. Responsible validation and standardization, rather than proof-of-concept performance alone, should define the next stage of development[23,24]. Current evidence therefore supports transparent, physician-supervised assistance rather than autonomous therapeutic decision-making.

References
1.  Low DJ, Hong Z, Lee JH. Artificial intelligence implementation in pancreaticobiliary endoscopy. Expert Rev Gastroenterol Hepatol. 2022;16:493-498.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 5]  [Reference Citation Analysis (0)]
2.  Jiang H, Ye LS, Yuan XL, Luo Q, Zhou NY, Hu B. Artificial intelligence in pancreaticobiliary endoscopy: Current applications and future directions. J Dig Dis. 2024;25:564-572.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 3]  [Reference Citation Analysis (0)]
3.  Agudo Castillo B, Mascarenhas M, Martins M, Mendes F, de la Iglesia D, Costa AMMPD, Esteban Fernández-Zarza C, González-Haba Ruiz M. Advancements in biliopancreatic endoscopy - A comprehensive review of artificial intelligence in EUS and ERCP. Rev Esp Enferm Dig. 2024;116:613-622.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 3]  [Cited by in RCA: 5]  [Article Influence: 2.5]  [Reference Citation Analysis (0)]
4.  Vu Trung K, Hollenbach M, Veldhuizen GP, Saldanha OL, Garbe J, Rosendahl J, Krug S, Michl P, Feisthammel J, Karlas T, Hampe J, Hoffmeister A, Kather JN. Deep Learning-Based Detection of Malignant Bile Duct Stenosis in Fluoroscopy Images of Endoscopic Retrograde Cholangiopancreatography. Digestion. 2025;106:287-302.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 2]  [Reference Citation Analysis (0)]
5.  Huang L, Lu X, Huang X, Zou X, Wu L, Zhou Z, Wu D, Tang D, Chen D, Wan X, Zhu Z, Deng T, Shen L, Liu J, Zhu Y, Gong D, Chen D, Zhong Y, Liu F, Yu H. Intelligent difficulty scoring and assistance system for endoscopic extraction of common bile duct stones based on deep learning: multicenter study. Endoscopy. 2021;53:491-498.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 26]  [Cited by in RCA: 25]  [Article Influence: 5.0]  [Reference Citation Analysis (2)]
6.  Bharwad AV, Ahuja R, Jain P, Wadhwa V. Artificial Intelligence in Pancreatobiliary Endoscopy: Current Advances, Opportunities, and Challenges. J Clin Med. 2025;14:7519.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 3]  [Reference Citation Analysis (0)]
7.  Baethge C, Goldbeck-Wood S, Mertens S. SANRA-a scale for the quality assessment of narrative review articles. Res Integr Peer Rev. 2019;4:5.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 1648]  [Cited by in RCA: 1680]  [Article Influence: 240.0]  [Reference Citation Analysis (2)]
8.  De Angelis CG, Dall'Amico E, Staiano MT, Gesualdo M, Bruno M, Gaia S, Sacco M, Fimiano F, Mauriello A, Dibitetto S, Canalis C, Stasio RC, Caneglias A, Mediati F, Rocca R. The Endoscopic Retrograde Cholangiopancreatography and Endoscopic Ultrasound Connection: Unity Is Strength, or the Endoscopic Ultrasonography Retrograde Cholangiopancreatography Concept. Diagnostics (Basel). 2023;13:3265.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 4]  [Cited by in RCA: 8]  [Article Influence: 2.7]  [Reference Citation Analysis (0)]
9.  Liu Y, Chen X, Zuo S. A deep learning-driven method for safe and effective ERCP cannulation. Int J Comput Assist Radiol Surg. 2025;20:913-922.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 1]  [Reference Citation Analysis (3)]
10.  Chen K, Lin H, Zhang F, Chen Z, Ying H, Cao L, Fang J, Zhu D, Liang K. Duodenal papilla radiomics-based prediction model for post-ERCP pancreatitis using machine learning: a retrospective multicohort study. Gastrointest Endosc. 2024;100:691-702.e9.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 14]  [Cited by in RCA: 17]  [Article Influence: 8.5]  [Reference Citation Analysis (0)]
11.  McCarty TR, Shah R, Allencherril RP, Moon N, Njei B. The Role of Artificial Intelligence Combined With Digital Cholangioscopy for Indeterminant and Malignant Biliary Strictures: A Systematic Review and Meta-analysis. J Clin Gastroenterol. 2025;.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 1]  [Cited by in RCA: 2]  [Article Influence: 2.0]  [Reference Citation Analysis (0)]
12.  Luo B, Li Z, Zhang K, Wu S, Chen W, Fu N, Yang Z, Hao J. Using deep learning models in magnetic resonance cholangiopancreatography images to diagnose common bile duct stones. Scand J Gastroenterol. 2024;59:118-124.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 1]  [Cited by in RCA: 3]  [Article Influence: 1.5]  [Reference Citation Analysis (0)]
13.  Ringe KI, Vo Chieu VD, Wacker F, Lenzen H, Manns MP, Hundt C, Schmidt B, Winther HB. Fully automated detection of primary sclerosing cholangitis (PSC)-compatible bile duct changes based on 3D magnetic resonance cholangiopancreatography using machine learning. Eur Radiol. 2021;31:2482-2489.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 19]  [Cited by in RCA: 16]  [Article Influence: 3.2]  [Reference Citation Analysis (0)]
14.  Kim T, Kim J, Choi HS, Kim ES, Keum B, Jeen YT, Lee HS, Chun HJ, Han SY, Kim DU, Kwon S, Choo J, Lee JM. Artificial intelligence-assisted analysis of endoscopic retrograde cholangiopancreatography image for identifying ampulla and difficulty of selective cannulation. Sci Rep. 2021;11:8381.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 24]  [Cited by in RCA: 23]  [Article Influence: 4.6]  [Reference Citation Analysis (1)]
15.  Zhang WL, Shao XJ, Dong XY, Shao HT, Li GC, Li Z, Zhong N, Ji R. Artificial intelligence-assisted biliary stent length selection for common bile duct strictures in endoscopic retrograde cholangiopancreatography: Model development and validation. Hepatobiliary Pancreat Dis Int. 2026;25:76-82.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 2]  [Reference Citation Analysis (0)]
16.  Chen C, Tao R, Hu QH, Wu ZJ. Effect of duodenal papilla morphology on biliary cannulation and complications in patients with common bile duct stones. Hepatobiliary Pancreat Dis Int. 2025;24:316-322.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 2]  [Reference Citation Analysis (0)]
17.  García-Marmolejo JP, Aguilar-Shotborgh M, Diaz-Brochero C, Leguízamo-Naranjo AM, Vargas-Rubio R. Timing of endoscopic retrograde cholangiopancreatography in acute cholangitis: A cohort study in Colombia. Hepatobiliary Pancreat Dis Int. 2026;25:225-228.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 1]  [Reference Citation Analysis (0)]
18.  Araújo CC, Frias J, Mendes F, Martins M, Mota J, Almeida MJ, Ribeiro T, Macedo G, Mascarenhas M. Unlocking the Potential of AI in EUS and ERCP: A Narrative Review for Pancreaticobiliary Disease. Cancers (Basel). 2025;17:1132.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 8]  [Reference Citation Analysis (1)]
19.  Ghandour B, Vedula SS, Akshintala VS, Khashab MA. Generalizability challenges of a machine learning model for classification of indeterminate biliary strictures. Gastrointest Endosc. 2022;95:1283-1284.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 1]  [Cited by in RCA: 3]  [Article Influence: 0.8]  [Reference Citation Analysis (0)]
20.  Liu X, Cruz Rivera S, Moher D, Calvert MJ, Denniston AK; SPIRIT-AI and CONSORT-AI Working Group. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat Med. 2020;26:1364-1374.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 340]  [Cited by in RCA: 746]  [Article Influence: 124.3]  [Reference Citation Analysis (4)]
21.  Cruz Rivera S, Liu X, Chan AW, Denniston AK, Calvert MJ; SPIRIT-AI and CONSORT-AI Working Group;  SPIRIT-AI and CONSORT-AI Steering Group;  SPIRIT-AI and CONSORT-AI Consensus Group. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. Nat Med. 2020;26:1351-1363.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 453]  [Cited by in RCA: 438]  [Article Influence: 73.0]  [Reference Citation Analysis (5)]
22.  Vasey B, Nagendran M, Campbell B, Clifton DA, Collins GS, Denaxas S, Denniston AK, Faes L, Geerts B, Ibrahim M, Liu X, Mateen BA, Mathur P, McCradden MD, Morgan L, Ordish J, Rogers C, Saria S, Ting DSW, Watkinson P, Weber W, Wheatstone P, McCulloch P; DECIDE-AI expert group. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat Med. 2022;28:924-933.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 372]  [Cited by in RCA: 402]  [Article Influence: 100.5]  [Reference Citation Analysis (3)]
23.  Alemam A, Tamanna R, Ali M, Ibraheem N, Swealem A. Artificial Intelligence in Upper Gastrointestinal Endoscopy: Current Evidence, Practice, and Future Directions. Cureus. 2025;17:e96657.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 2]  [Reference Citation Analysis (0)]
24.  Jong MR, Boers TGW, Fockens KN, Jukema JB, Kusters CHJ, Jaspers TJM, van Eijck van Heslinga RAH, Slooter FC, Struyvenberg MR, Bisschops R, van der Putten JA, de With PHN, van der Sommen F, de Groof AJ, Bergman JJGHM; BONS-AI Consortium. GastroNet-5M: A Multicenter Dataset for Developing Foundation Models in Gastrointestinal Endoscopy. Gastroenterology. 2026;170:174-187.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 4]  [Cited by in RCA: 12]  [Article Influence: 12.0]  [Reference Citation Analysis (0)]
25.  Kim Y, Keum JS, Kim JH, Chun J, Oh SI, Kim KN, Yoon YH, Park H. Real-World Colonoscopy Video Integration to Improve Artificial Intelligence Polyp Detection Performance and Reduce Manual Annotation Labor. Diagnostics (Basel). 2025;15:901.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 1]  [Reference Citation Analysis (0)]
26.  Ahmad OF, Mori Y, Bretthauer M, Dourado DA, Hassan C, Bisschops R, Bhandari P, Byrne MF, Dekker E, Mahadevan U, May FP, Messmann H, Misawa M, Ogata H, Saito Y, Silverman AL, Wang P, Yano T, Aabakken L, Berzin TM. The Legal and Ethical Framework for Artificial Intelligence in Gastrointestinal Endoscopy: A World Endoscopy Organization International Consensus Statement. Ann Intern Med. 2026;179:270-275.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 2]  [Cited by in RCA: 7]  [Article Influence: 7.0]  [Reference Citation Analysis (0)]
27.  World Medical Association  WMA Statement on Artificial and Augmented Intelligence in Medical Care. [cite 29 July 2026] Available from: https://www.wma.net/policies-post/wma-statement-on-artificial-and-augmented-intelligence-in-medical-care.  [PubMed]  [DOI]
Footnotes

Peer review: Externally peer reviewed.

Peer-review model: Single blind

Specialty type: Medicine, research and experimental

Country of origin: Egypt

Peer-review report’s classification

Scientific quality: Grade B, Grade B

Novelty: Grade A, Grade C

Creativity or innovation: Grade B, Grade B

Scientific significance: Grade B, Grade B

P-Reviewer: Huang ZT, Deputy Director, PhD, Post Doctoral Researcher, China S-Editor: Liu H L-Editor: A P-Editor: Wang CH

Write to the Help Desk