Copyright: ©Author(s) 2026.
Artif Intell Gastroenterol. Aug 8, 2026; 7(2): 118476
Published online Aug 8, 2026. doi: 10.35712/aig.118476
Published online Aug 8, 2026. doi: 10.35712/aig.118476
Table 5 Limitations of artificial intelligence and potential mitigation strategies in the diagnosis of biliary strictures
| Limitations | Evidence from current literature | Illustration from AI biliary stricture studies | Potential mitigation strategies |
| Retrospective design and selection bias[14,74] | The majority of published AI studies are retrospective and single-center, limiting generalizability and inflating performance estimates | Zhang et al[71] (MBSDeiT) demonstrated prospective real-time D-SOC prediction on a retrospectively trained model, but did not assess the impact on clinical outcomes or compare performance with biopsy guidance. Saraiva et al[68] similarly acknowledged that despite their large dataset for a proof-of-concept study, clinical validation requires a much larger volume of data | Conduct prospective, multicenter randomized trials and pragmatic validation studies |
| Black-box decision making[12,68] | Most high-performing CNN and ensemble models lack intrinsic interpretability, reducing clinician trust in AI-guided lesion targeting and biopsy decisions | Saraiva et al[68] demonstrated post-hoc heatmap validation of CNN attention to tumor vessels and papillary projections but lacked real-time interpretability during procedures, similarly acknowledged that despite their large dataset for a proof-of-concept study, clinical validation requires a much larger volume of data | Integrate explainable AI techniques (e.g., Grad-CAM, SHAP) as real-time overlays on D-SOC/EUS. Develop AI systems that highlight and verbalize decision-driving features |
| Limited dataset size and diversity[70,74,88] | Stricture-specific EUS and cholangioscopy datasets remain small | Robles-Medranda et al[70] conducted a two-phase multicenter study (n = 164), yet disease spectrum and procedural variability remained limited, with ongoing discrepancy between operators visual impression using current classifications for indeterminate biliary strictures | Create open-access, multi-vendor endoscopic image repositories (e.g., expanding The Cancer Imaging Archive. Employ federated learning to train models across institutions without sharing raw data |
| Domain shift and device dependency[70,89] | Model performance may degrade when applied across different EUS, D-SOC processors, probes, contrast agents, or imaging protocols | The multicenter validation by Robles-Medranda et al[70] may be affected by variability in D-SOC systems and protocols | Implement cross-platform training, harmonization of acquisition protocols, and external validation across vendors and geographic regions |
| Lack of clinical workflow integration[86,87] | Many AI models are evaluated offline on static images or pre-recorded videos, without assessment of real-time feasibility, procedural impact, or endoscopist interaction | Only Marya et al[87] implemented and validated real-time D-SOC classification during live procedures | Conduct human-factors studies on endoscopist-AI interaction, and conduct real-time deployment trials measuring procedural metrics |
| Absence of outcome-driven validation[84,85] | Most studies report diagnostic accuracy but lack data on downstream clinical outcomes (time to diagnosis, avoided ERCP, surgical yield) | A multicenter AI study[85] demonstrated superior diagnostic accuracy but did not report the impact on avoided ERCPs or surgical outcomes | Design endpoint-driven trials linking AI-guided diagnosis to clinical outcomes. Perform formal cost-effectiveness and patient-centered measures |
- Citation: Majeed AA, Butt AS. Leveraging artificial intelligence to differentiate benign from malignant biliary strictures: A step toward precision diagnosis. Artif Intell Gastroenterol 2026; 7(2): 118476
- URL: https://www.wjgnet.com/2644-3236/full/v7/i2/118476.htm
- DOI: https://dx.doi.org/10.35712/aig.118476