BPG is committed to discovery and dissemination of knowledge
Minireviews
Copyright: ©Author(s) 2026.
Artif Intell Gastroenterol. Aug 8, 2026; 7(2): 118476
Published online Aug 8, 2026. doi: 10.35712/aig.118476
Table 5 Limitations of artificial intelligence and potential mitigation strategies in the diagnosis of biliary strictures
Limitations
Evidence from current literature
Illustration from AI biliary stricture studies
Potential mitigation strategies
Retrospective design and selection bias[14,74] The majority of published AI studies are retrospective and single-center, limiting generalizability and inflating performance estimatesZhang et al[71] (MBSDeiT) demonstrated prospective real-time D-SOC prediction on a retrospectively trained model, but did not assess the impact on clinical outcomes or compare performance with biopsy guidance. Saraiva et al[68] similarly acknowledged that despite their large dataset for a proof-of-concept study, clinical validation requires a much larger volume of dataConduct prospective, multicenter randomized trials and pragmatic validation studies
Black-box decision making[12,68] Most high-performing CNN and ensemble models lack intrinsic interpretability, reducing clinician trust in AI-guided lesion targeting and biopsy decisionsSaraiva et al[68] demonstrated post-hoc heatmap validation of CNN attention to tumor vessels and papillary projections but lacked real-time interpretability during procedures, similarly acknowledged that despite their large dataset for a proof-of-concept study, clinical validation requires a much larger volume of dataIntegrate explainable AI techniques (e.g., Grad-CAM, SHAP) as real-time overlays on D-SOC/EUS. Develop AI systems that highlight and verbalize decision-driving features
Limited dataset size and diversity[70,74,88] Stricture-specific EUS and cholangioscopy datasets remain small (< 10000 labelled images across studies), with underrepresentation of perihilar and benign inflammatory stricturesRobles-Medranda et al[70] conducted a two-phase multicenter study (n = 164), yet disease spectrum and procedural variability remained limited, with ongoing discrepancy between operators visual impression using current classifications for indeterminate biliary stricturesCreate open-access, multi-vendor endoscopic image repositories (e.g., expanding The Cancer Imaging Archive. Employ federated learning to train models across institutions without sharing raw data
Domain shift and device dependency[70,89] Model performance may degrade when applied across different EUS, D-SOC processors, probes, contrast agents, or imaging protocolsThe multicenter validation by Robles-Medranda et al[70] may be affected by variability in D-SOC systems and protocolsImplement cross-platform training, harmonization of acquisition protocols, and external validation across vendors and geographic regions
Lack of clinical workflow integration[86,87] Many AI models are evaluated offline on static images or pre-recorded videos, without assessment of real-time feasibility, procedural impact, or endoscopist interactionOnly Marya et al[87] implemented and validated real-time D-SOC classification during live proceduresConduct human-factors studies on endoscopist-AI interaction, and conduct real-time deployment trials measuring procedural metrics
Absence of outcome-driven validation[84,85] Most studies report diagnostic accuracy but lack data on downstream clinical outcomes (time to diagnosis, avoided ERCP, surgical yield)A multicenter AI study[85] demonstrated superior diagnostic accuracy but did not report the impact on avoided ERCPs or surgical outcomesDesign endpoint-driven trials linking AI-guided diagnosis to clinical outcomes. Perform formal cost-effectiveness and patient-centered measures


Write to the Help Desk