Published online Sep 21, 2026. doi: 10.3748/wjg.119939
Revised: March 21, 2026
Accepted: May 27, 2026
Published online: September 21, 2026
Processing time: 191 Days and 21.7 Hours
Drug-induced liver injury (DILI) remains one of the most challenging diagnoses in hepatology due to its complexity, reliance on expert interpretation and exclu
Core Tip: Drug-induced liver injury remains challenging due to its vague presentation and reliance on exclusion-based techniques. This complexity is particularly evident in cases of pyrrolizidine alkaloid-induced hepatic sinusoidal obstruction syndrome, which often require invasive treatments and expert imaging interpretation. Recent advancements in deep learning applied to computed tomography are changing the diagnostic landscape by identifying subtle, diffuse parenchymal abnormalities that might be missed using standard methods. Artificial intelligence can improve clinician performance, enhance diagnostic consistency, and reduce interpretation times, as demonstrated by the validated model discussed here. These integrated technologies could improve the early identification of complex drug-induced liver injury characteristics.
- Citation: Fouad Y, Mostafa AM, Abdelhalim SM, Eslam M. Letter to the Editor: Artificial intelligence in hepatology - when deep learning meets drug-induced liver injury. World J Gastroenterol 2026; 32(35): 119939
- URL: https://www.wjgnet.com/1007-9327/full/v32/i35/119939.htm
- DOI: https://dx.doi.org/10.3748/wjg.119939
Drug-induced liver injury (DILI) remains a significant diagnostic challenge in hepatology due to the frequent absence of precise biomarkers, mimicking other liver disorders, and reliance on time-consuming, clinician-dependent exclusionary diagnostic techniques[1]. Conventional methods often lead to diagnostic uncertainty and may necessitate confirmatory invasive procedures like liver biopsy or angiography, particularly in conditions like pyrrolizidine alkaloid-induced hepatic sinusoidal obstruction syndrome (PA-HSOS), even when supported by clinical history, laboratory markers, and imaging[2]. These challenges emphasize the urgent need for tools that can enhance diagnostic consistency and accuracy, particularly as imaging and clinical scoring systems advance.
Artificial intelligence (AI), particularly deep learning, may hold the key to resolving these long-standing issues by utilizing large datasets to identify complex image patterns that are difficult for humans to discern. Recent research over the past decade has demonstrated that deep learning models, especially convolutional neural networks, can extract intricate features from medical images. This capability facilitates tasks such as lesion detection, classification, and prognosis prediction across various liver disorders. The advancements are especially relevant for liver imaging techniques like computed tomography (CT) and magnetic resonance imaging, which, despite offering excellent anatomical detail, historically suffer from reader variability[3].
The deep learning-based diagnostic model presented in the study by Wang et al[4], published in World Journal of Gastroenterology, addresses a critical unmet need in hepatology by effectively and accurately diagnosing PA-HSOS using enhanced CT images. Their approach bridges the gap between clinical reasoning and technical innovation by integrating multiscale convolutional modules with an anatomically guided region-of-interest sampling strategy to mimic the clinical decision-making process. This method aims to differentiate PA-HSOS from other diseases with similar imaging characteristics, such as hepatitis B cirrhosis and Budd-Chiari syndrome, with an accuracy comparable to that of specialists[4]. Crucially, it’s important to distinguish between what this study shows and what still needs to be proven for therapeutic application. Strong evidence of diagnostic discrimination within a case-control framework, including enhanced reader performance in aided situations and external validation, is provided by Wang et al’s work[4]. These results validate the usefulness of deep learning as a controlled diagnostic supplement. These findings should not be confused with actual efficacy, though, as process integration, case-mix variability, and disease prevalence can all have a significant impact on performance.
One of the key aspects of this study’s significance is its advanced methodology. The authors assembled a robust multicenter case-control cohort to compare patients with PA-HSOS with two critical differential diagnoses, Budd-Chiari syndrome and hepatitis B cirrhosis. Their technical strategy was meticulously crafted to tackle the dual challenges of complex feature extraction and sparse data. The two-stage framework, which employs a proprietary classification model following automated liver segmentation using a transfer-learned nnU-Net, is particularly elegant[5]. The anatomy-informed region of interest sampling strategy improves clinical interpretability by aligning with the cognitive process of consulting hepatologists or radiologists, focusing on important liver segments and areas of hepatic venous drainage[4].
With area under the curve (AUC) values close to 0.94 at the individual patient level, the model demonstrated robust discriminatory capacity in both training and validation cohorts. It significantly improved diagnostic accuracy and specificity when compared to resident physicians. Notably, when used to support clinicians, including both attending gastroenterologists and residents, the model increased diagnostic accuracy and decreased interpretation time across various skill levels. This highlights the potential translational utility of deep learning, which can serve as a supportive decision-making tool that enhances performance while still requiring human oversight. Even though the stated AUC values show good discriminative performance, clinical deployment cannot be guided solely by AUC. Discrimination shows how well cases may be ranked, but it doesn’t show how closely projected probability match actual results.
Beyond the specific diagnosis of PA-HSOS, this finding has broader implications. They provide compelling proof-of-concept for the application of deep learning to a wider array of diffuse parenchymal liver disorders, including DILI. The work demonstrates that AI can be trained to recognize complex, non-mass-like imaging abnormalities, in addition to its established role in diagnosing focused lesions, such as tumors. By offering a reliable, quantitative “second opinion,” these models can assist less experienced clinicians in high-pressure environments, reduce inter-observer variability, and accelerate the diagnostic process for patients with rare diseases. Earlier initiation of appropriate management, such as supportive care and the removal of the triggering agent, may lead to improved clinical outcomes[6,7].
This aligns with broader trends in the application of AI within hepatology. Beyond PA-HSOS, deep learning has demonstrated promise in various liver conditions, including the detection and staging of localized lesions and the quantification of hepatic steatosis and fibrosis. However, regulatory ambiguity, technological hurdles, and ethical concerns persist as ongoing hindrances to real-world implementation[8].
Using such models in clinical settings offers numerous significant advantages. First, standardizing diagnostic inter
Recent approvals and qualifications of AI tools for liver disease assessment in clinical trials highlight the increasing regulatory interest in the use of AI in liver disease diagnosis. For instance, regulatory agencies have certified AI platforms like AIM-NASH to facilitate standardized histopathologic evaluation in the drug development process for fatty liver disease. These advancements suggest a future where AI will play a significant role in both therapeutic research and cli
However, for widespread adoption of these technologies to become routine, several challenges and limitations must be addressed. One of the most commonly cited issues is the generalizability of AI models. Many AI algorithms are trained on datasets from a limited number of sites, which often lack variation in imaging techniques, disease prevalence, and types of scanners used. As a result, if these models do not undergo extensive external validation, they may not perform effectively in different clinical settings or populations[10]. Wang et al’s work[4] makes a valuable effort toward multi
The interpretability and transparency of deep learning methods present significant challenges. These models are often described as “black boxes,” particularly when they have complex architectures[11]. Because there are typically no clear explanations for how decisions are made, clinicians may be hesitant to rely on the results produced by these models. To build clinician trust and regulatory approval, it is essential to improve system explainability using methods such as attention mapping or anatomically guided sampling.
The possible spectrum bias seen in case-control methods is another drawback. Prospective, consecutive patient enrolment with expanded control groups that accurately represent the clinical differential diagnosis of PA-HSOS on CT imaging should be given top priority in future research. Hepatic venous outflow blockage (e.g., Budd-Chiari syndrome), congestive hepatopathy, acute hepatitis, other types of DILI, and hepatic graft-versus-host disease should all be included. A more accurate assessment of diagnostic performance, clinical yield, and generalizability in ordinary practice would be possible with this method.
Another ongoing issue is data heterogeneity. Factors such as variations in CT protocols, contrast phases, image quality, and preprocessing techniques can significantly affect model performance. Furthermore, while integrating AI outputs with clinical and laboratory data may enhance diagnostic accuracy, it complicates the process of model development and validation. As a result, a promising yet technically challenging advancement lies in creating multimodal frameworks that combine imaging, clinical, and laboratory data[12].
Attention must also be paid to operational and ethical issues. In many jurisdictions, issues including data privacy, AI governance, and clinician liability in the event of model error have not yet been adequately addressed. Strong frame
Deep learning’s use in hepatology is probably going to grow in the future. In addition to imaging classification, new re
The influence of domain shift, whereby differences in imaging acquisition and reconstruction techniques might result in a decline in model performance when used in the original training environment, is another factor to take into account for clinical translation. Slice thickness, phase timing, reconstruction kernel, contrast delivery techniques, and scanner vendor and model can all have a significant impact on image features and, in turn, model predictions in CT-based liver imaging. Therefore, in order to guarantee reproducibility and enable significant external evaluation, transparent re
In conclusion, Wang et al’s work[4] on a deep learning model represents a significant advancement in AI-assisted diagnosis of complex DILI variations. It exemplifies how contemporary computational techniques can enhance physician competence, standardize imaging interpretations, and potentially improve the quality of care. While there are ongoing challenges regarding validation, interpretability, and integration, this study contributes to the growing body of evidence indicating that when well-planned and thoroughly tested, AI can be a powerful ally in hepatology. Future CT-based AI models for PA-HSOS must exhibit strong calibration, resistance to domain shift, and clinically significant net benefit in prospective, consecutively enrolled cohorts; high discriminative performance is required but not sufficient. Rather than replacing clinicians, deep learning technologies are anticipated to enhance clinician efficiency, insight, and ultimately, patient outcomes. It will take carefully planned prospective research, integration into practical processes, and assessment against patient-centered outcomes to close the gap between diagnostic accuracy and clinical impact.
| 1. | Keshari AC, Thitame SN, Aher AA, Keshari UC. Drug-Induced Liver Injury: Mechanisms, Diagnosis, and Management: A Review. J Pharm Bioallied Sci. 2025;17:S55-S58. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 6] [Cited by in RCA: 11] [Article Influence: 11.0] [Reference Citation Analysis (0)] |
| 2. | Huang Z, Wu Z, Gu X, Ji L. Diagnosis, toxicological mechanism, and detoxification for hepatotoxicity induced by pyrrolizidine alkaloids from herbal medicines or other plants. Crit Rev Toxicol. 2024;54:123-133. [RCA] [PubMed] [DOI] [Full Text] [Cited by in RCA: 3] [Reference Citation Analysis (0)] |
| 3. | Niu H, Alvarez-Alvarez I, Chen M. Artificial Intelligence: An Emerging Tool for Studying Drug-Induced Liver Injury. Liver Int. 2025;45:e70038. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 12] [Cited by in RCA: 10] [Article Influence: 10.0] [Reference Citation Analysis (1)] |
| 4. | Wang SY, Yin SQ, Yang JY, Ji MY, Zeng XQ, Rao SX, Lv MZ, Bao J, Wang MN, Gao H. Development and validation of a deep-learning-based diagnostic model for drug-induced liver injury using computed tomography images. World J Gastroenterol. 2026;32:114778. [RCA] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 1] [Reference Citation Analysis (0)] |
| 5. | Isensee F, Jaeger PF, Kohl SAA, Petersen J, Maier-Hein KH. nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nat Methods. 2021;18:203-211. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 8060] [Cited by in RCA: 4218] [Article Influence: 843.6] [Reference Citation Analysis (4)] |
| 6. | Yasaka K, Akai H, Kunimatsu A, Abe O, Kiryu S. Deep learning for staging liver fibrosis on CT: a pilot study. Eur Radiol. 2018;28:4578-4585. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 113] [Cited by in RCA: 91] [Article Influence: 11.4] [Reference Citation Analysis (4)] |
| 7. | Park SH, Han K. Methodologic Guide for Evaluating Clinical Performance and Effect of Artificial Intelligence Technology for Medical Diagnosis and Prediction. Radiology. 2018;286:800-809. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 729] [Cited by in RCA: 566] [Article Influence: 70.8] [Reference Citation Analysis (3)] |
| 8. | Morel SMG, Wu S, Kendall TJ, Guha IN, Fallowfield JA. Opportunities and challenges of artificial intelligence in hepatology. npj Gut Liver. 2026;3:3. [DOI] [Full Text] |
| 9. | Pulaski H, Harrison SA, Mehta SS, Sanyal AJ, Vitali MC, Manigat LC, Hou H, Madasu Christudoss SP, Hoffman SM, Stanford-Moore A, Egger R, Glickman J, Resnick M, Patel N, Taylor CE, Myers RP, Chung C, Patterson SD, Sejling AS, Minnich A, Baxi V, Subramaniam GM, Anstee QM, Loomba R, Ratziu V, Montalto MC, Anderson NP, Beck AH, Wack KE. Clinical validation of an AI-based pathology tool for scoring of metabolic dysfunction-associated steatohepatitis. Nat Med. 2025;31:315-322. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 50] [Cited by in RCA: 36] [Article Influence: 36.0] [Reference Citation Analysis (0)] |
| 10. | Hoghooghi Esfahani H, Toyonaga S, Oyibo K. The application of explainable artificial intelligence in the prediction, diagnoses, treatment, and management of chronic diseases: A systematic review. Digit Health. 2025;11:20552076251355669. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 4] [Reference Citation Analysis (0)] |
| 11. | Şahin E, Arslan NN, Özdemir D. Unlocking the black box: An in-depth review on interpretability, explainability, and reliability in deep learning. Neural Comput Appl. 2025;37:859-965. [DOI] [Full Text] |
| 12. | Ardic N, Dinc R. Emerging trends in multi-modal artificial intelligence for clinical decision support: A narrative review. Health Informatics J. 2025;31:14604582251366141. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 11] [Cited by in RCA: 11] [Article Influence: 11.0] [Reference Citation Analysis (1)] |
| 13. | Khan S, Noor MN, Ashraf I, Masud MI, Aman M. Impact of CT Intensity and Contrast Variability on Deep-Learning-Based Lung-Nodule Detection: A Systematic Review of Preprocessing and Harmonization Strategies (2020-2025). Diagnostics (Basel). 2026;16:201. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 1] [Reference Citation Analysis (0)] |