Revised: February 16, 2026
Accepted: March 25, 2026
Published online: September 27, 2026
Processing time: 277 Days and 6.9 Hours
The study by Rech et al recently published in World Journal of Hepatology, repre
Core Tip: Machine learning is transitioning from theoretical promise to practical implementation within hepatology, particularly in high-risk conditions such as acute esophageal variceal bleeding. The study by Rech et al distinguishes itself by combining high-performing mortality prediction with prospective validation and real-world deployment as an online calculator. This editorial highlights why such efforts represent an important step toward bridging the persistent gap between algorithm development and clinical adoption. Yet we also explore the remaining challenges-model interpretability, ethical complexities surrounding race-based predictors, workflow integration, model drift, and the need for multicenter external validation. Understanding these dimensions is crucial for translating artificial intelligence into safer, more equitable, and genuinely impactful tools at the bedside.
- Citation: Othman AAA. From algorithm to bedside: Navigating the promise and perils of implementing machine learning for variceal bleeding mortality prediction. World J Hepatol 2026; 18(9): 117720
- URL: https://www.wjgnet.com/1948-5182/full/v18/i9/117720.htm
- DOI: https://dx.doi.org/10.4254/wjh.117720
This editorial refers to “Development and prospective validation of a machine learning model to predict mortality in cirrhosis with esophageal variceal bleeding” by Rech MM et al, 2026; https://dx.doi.org/10.4254/wjh.v18.i2.111099.
Acute esophageal variceal bleeding (AEVB) remains one of the most urgent and lethal complications of portal hy
Prognostication has historically relied on linear models such as the Child-Pugh and Model for end-stage liver disease (MELD) scores. While useful, these tools may not fully capture the complex, nonlinear interactions driving outcomes in AEVB[3]. This limitation has catalyzed the exploration of machine learning (ML) within hepatology. ML algorithms can model intricate relationships across high-dimensional data, from laboratory trends and hemodynamic parameters to imaging radiomics[4].
The past five years have seen accelerated integration of ML, with applications ranging from non-invasive fibrosis staging to predicting hepatic decompensation[5,6]. However, a persistent translational gap exists: Most models are retrospectively derived and lack prospective validation or real-world deployment, confining them to academic exercise.
In a recent issue of the World Journal of Hepatology, Rech et al[7] reported a retrospective cohort study. Rech et al[7] represents a concerted effort to bridge this gap. By developing a random forest model for AEVB mortality, prospectively validating it in a temporal cohort, and deploying it as an accessible online calculator, they provide a prototype for translational ML research. This editorial contextualizes their contribution, appraises its methodological and ethical dimensions, and outlines the critical next steps required for the responsible and effective integration of ML into hepatology practice.
ML is gaining traction in hepatology for its capacity to handle nonlinear interactions and high-dimensional data more effectively than conventional statistics. For variceal bleeding, recent studies have demonstrated that ML models integrating clinical, laboratory, and endoscopic variables can outperform traditional scores like MELD and Child-Pugh[8,9].
Rech et al[7] advance this field meaningfully by adhering to higher methodological standards. Their model demon
Equally important is the evaluation of clinical utility through decision-curve analysis, which quantifies net benefit across relevant risk thresholds and determines whether improved statistical performance translates into better decision-making[10]. A modest AUC increase, even if statistically significant, may not confer a meaningful clinical advantage unless it leads to changes in management or resource allocation[11,12]. Future studies should therefore prioritize these comparative analyses to demonstrate not just statistical superiority, but genuine clinical utility.
Future studies could further strengthen clinical applicability by incorporating decision-curve analysis to quantify the net benefit of the model across different risk thresholds. These results align with findings from other ML frameworks, such as gradient boosting and neural networks, which have reported similar high performance in recent analyses[13]. By making their model available via an online calculator, the authors directly address the issue of accessibility, a core tenet of implementation science that argues that utility hinges on more than statistical performance alone[14].
The prospective validation design also strengthens the work’s credibility, placing it among a minority of hepatology artificial intelligence (AI) studies that meet contemporary reporting standards like TRIPOD-AI, which are designed to improve the transparency and reproducibility of predictive model research[15].
A key strength of the study is its methodologically sound approach. The authors evaluated multiple algorithms, em
However, a critical distinction warrants emphasis: SHAP values quantify the contribution of each variable to a specific prediction, they do not establish causal relationships. The identification of race as the most influential predictor, for example, does not imply that race causally determines mortality; rather, it reflects statistical associations within the training data that may capture unmeasured social determinants, healthcare access disparities, or other confounding factors. Conflating predictive importance with causal inference risks misinterpretation and could lead to misguided clinical decisions if clinicians assume that modifying a predictive feature would alter outcomes. This distinction is particularly consequential in variceal bleeding, where variables such as bilirubin or creatinine may reflect disease severity but are not direct targets for acute intervention.
Beyond the technical distinction between prediction and causation, the practical interpretability of SHAP outputs in acute clinical settings presents additional challenges. During an active variceal bleed, clinicians require rapid, intuitive explanations that align with clinical reasoning, not a ranked list of feature contributions that must be mentally integrated under time pressure. What clinicians need are explanations that connect model predictions to recognizable clinical phenotypes or actionable factors: ‘This patient’s high predicted mortality is driven by severe encephalopathy, elevated INR, and renal dysfunction, features that suggest consideration of early transjugular intrahepatic portosystemic shunt (TIPS) or intensive care unit (ICU) escalation’. Future iterations of ML tools for acute care should therefore prioritize explanation interfaces designed for cognitive efficiency, such as grouping related features into clinically meaningful categories or highlighting potentially modifiable factors[7].
It is important to note that SHAP values explain the contribution of variables to the model's output, not necessarily their direct pathophysiological role in disease progression. Nonetheless, comparative performance analyses against traditional scores should extend beyond reporting separate AUC values. Demonstrating statistically significant improvement and clinically meaningful net benefit relative to MELD and Child-Pugh is essential before concluding incremental prognostic value. The field must move toward rigorous comparative frameworks that quantify what new models add to, rather than simply how they perform in isolation.
Despite these strengths, several limitations merit closer scrutiny. Although the authors should be commended for undertaking validation in a prospective, temporally distinct cohort, the relatively small size of this cohort (n = 24) raises the possibility of optimistic performance estimates, particularly given the model’s complexity and the number of included predictors. This concern is not unique to this study but reflects a broader challenge in clinical ML, where models trained on limited datasets may demonstrate impressive discrimination yet struggle to maintain calibration and stability when exposed to larger or more heterogeneous populations. Consequently, multicenter external validation across diverse healthcare systems remains a necessary step before widespread clinical adoption can be recommended.
Furthermore, the single-center origin of the data is a significant limitation. Clinical protocols, patient demographics, and healthcare resources vary widely across institutions and geographies. Models trained on one population often experience degraded performance when applied externally, a phenomenon noted in ML studies for other cirrhosis complications like hepatic encephalopathy and rebleeding[16].
The challenges of cross-institutional validation extend beyond simple demographic differences. Variations in how variables are defined and measured across sites, what might be termed measurement heterogeneity, can substantially degrade model performance even when patient populations appear similar. For example, the timing of laboratory collection, definitions of key clinical events like ‘active bleeding’, and thresholds for interventions such as TIPS may differ across centers, directly influencing both predictor values and outcomes[2]. These differences in clinical protocols and data capture can render a model miscalibrated when transported to a new setting, not because its predictions are wrong, but because the underlying clinical context differs[3].
Addressing these challenges requires prospective attention to data harmonization. As emphasized in the TRIPOD-AI reporting guidelines, transparent documentation of variable definitions, data sources, and handling of between-site heterogeneity is essential for meaningful external validation[15]. Without such harmonization, apparent performance degradation may reflect differences in measurement rather than true failures of model generalizability
Furthermore, the specific inclusion and exclusion criteria (e.g., for hepatorenal syndrome, concomitant infections, or prior interventions) necessarily shape the model's predictive landscape and must be carefully considered when applying it to different patient subsets.
A finding that demands particular scrutiny is the identification of race as the most influential predictor in the SHAP analysis. It is crucial to emphasize that race should not be interpreted as a direct biological cause of mortality; rather, it often serves as a proxy for unmeasured social determinants of health, such as disparities in healthcare access, socioeconomic status, and environmental exposures[17]. Incorporating race uncritically into prognostic algorithms risks perpetuating and automating existing healthcare inequities[18].
Beyond conceptual reflection, quantitative fairness evaluation is essential. This includes reporting subgroup-specific discrimination (e.g., AUC stratified by race or socioeconomic status), calibration-in-the-large and calibration slope within each subgroup, and comparison of error rates such as false-negative and false-positive proportions. Disparities in false-negative rates are particularly concerning in high-risk conditions such as acute variceal bleeding, where underestimation of mortality risk could delay escalation of care for vulnerable populations. Moreover, fairness assessment involves technical trade-offs that must be navigated transparently. Models that are well calibrated across groups may still violate equalized odds if error rates differ substantially, and enforcing strict parity constraints may reduce overall predictive performance[19]. These tensions-between calibration, equalized odds, and demographic parity-cannot be resolved mathematically but require ethical deliberation informed by clinical context and stakeholder engagement[20].
The distinction between race as a predictor and race as a proxy has direct implications for model revision. If race is serving as a proxy for socioeconomic deprivation, then collecting and incorporating granular social determinants of health (e.g., income, education, neighborhood disadvantage indices, and healthcare access metrics) could potentially attenuate or eliminate its predictive weight[21]. This approach aligns with emerging consensus in medical AI ethics that algorithms should, wherever possible, be designed to detect and mitigate inequities rather than encode them. However, such granular data are often absent from routine clinical documentation, highlighting the need for healthcare systems to systematically collect social determinants of health as part of standard care.
A further complexity is that even when race is removed as an input variable, models may still encode race through correlated proxies, a phenomenon termed the ‘fairness through unawareness’ fallacy. Variables such as zip code, insurance status, or even certain laboratory values can serve as surrogates for race, potentially recreating biased predictions despite the apparent absence of demographic identifiers[22]. Detecting such indirect encoding requires sophisticated auditing techniques, including analysis of feature correlations and counterfactual fairness assessment.
For the field to progress toward equitable AI in hepatology, prospective validation studies should predefine fairness metrics and acceptable disparity thresholds, ensuring that equity considerations are operationalized rather than merely acknowledged[23]. This includes intentional over-sampling of minority populations during data collection to enable robust subgroup analyses, pre-registration of fairness evaluation plans, and transparent reporting of model performance stratified by race, ethnicity, and socioeconomic status.
While the authors appropriately acknowledge this limitation, future iterations of the model should actively pursue race-neutral alternatives by incorporating more granular measures of social determinants of health, healthcare access, and disease severity. Achieving truly equitable, race-neutral models remains an aspirational goal that requires concerted effort to collect and integrate more precise socioeconomic and biological data. Such refinement is essential to ensure that predictive accuracy does not come at the cost of reinforcing structural inequities. As ML tools move closer to clinical deployment, proactive fairness auditing must become a standard component of hepatology-focused AI research.
An additional consideration relates to outcome selection. The use of one-year all-cause mortality provides a broad and clinically meaningful endpoint; however, it differs from the conventional six-week mortality benchmark commonly used in acute variceal bleeding research. As a result, it remains unclear to what extent the model captures bleeding-related risk vs longer-term cirrhosis progression and comorbidity burden. Clarifying this distinction is particularly important if the tool is intended to inform acute management decisions, rather than long-term prognostication alone.
Translating a validated ML model into a trusted bedside tool requires more than strong discrimination metrics. In acute esophageal variceal bleeding, clinical decisions are time-sensitive and high-stakes, encompassing triage to intensive care, escalation to early TIPS, and initiation of goals-of-care discussions. These interventions are highlighted because they are resource-intensive, high-risk, and require early, definitive decision-making. They represent precisely the scenarios where accurate prognostication could have the greatest impact on patient trajectory and resource allocation[2]. These scenarios are presented as illustrative examples of high-stakes, time-sensitive decisions where accurate prognostication is most valuable. For the model to directly guide such decisions, future iterations would need to be validated against near-term outcomes (e.g., 6-week mortality or failure to control bleeding) and integrated with explicit clinical decision pathways. As discussed, the model's current prediction of one-year all-cause mortality positions it primarily as a prognostic instrument for longer-term risk stratification. This endpoint choice inherently focuses the model’s predictive power on the composite burden of cirrhosis over a longer horizon, which differs from predicting acute, bleeding-specific decompensation. For the model by Rech et al[7] to transition from an excellent prognostic instrument to a decisive clinical support tool, its actionability would be strengthened by the explicit definition of risk thresholds linked to specific management pathways.
The transition from a validated algorithm to a trusted clinical tool requires surmounting substantial practical barriers. First is the challenge of interpretability. While SHAP values indicate feature importance at the individual prediction level, clinicians need explanations that align with pathophysiological reasoning and support rapid decision-making under pressure[24]. In the context of an acute variceal bleed, a clinician does not have time to parse a ranked list of multiple feature contributions; they need to know, in seconds, whether this prediction should change their management. This demands explanation interfaces that distill complex SHAP outputs into clinically intuitive formats, for example, highlighting the dominant drivers of risk, grouping related features into meaningful categories (e.g., ‘portal hypertension severity’, ‘hemodynamic instability’), or explicitly flagging when the model's reasoning diverges from expected clinical patterns[7]. Developing such clinically grounded explanation interfaces is, therefore, as important as optimizing the algorithm itself.
Second, model drift presents an underappreciated long-term threat. Clinical environments are dynamic; changes in treatment guidelines, laboratory assays, and patient populations can degrade a model’s accuracy over months to years[25]. This is especially relevant in hepatology, where treatment paradigms for variceal bleeding and standards of care for cirrhosis are continually evolving. However, not all drift is alike, and distinguishing between different forms is essential for appropriate remediation[26]. Concept drift occurs when the fundamental relationship between predictors and outcomes changes, for example, if newer vasoactive agents or earlier TIPS deployment alter the mortality risk associated with a given set of clinical parameters. Data drift (or covariate shift) refers to changes in the distribution of input variables themselves, such as shifts in population demographics, laboratory reference ranges, or coding practices. Outcome drift involves changes in the definition or base rate of the outcome itself, which could occur if survival improves overall due to advances in supportive care. Finally, clinical protocol drift, changes in treatment thresholds or standards, can render a model miscalibrated even when patient characteristics remain stable, because the expected outcome for a given risk profile has changed. Each drift type requires different detection strategies and remediation approaches. Concept drift may necessitate full model retraining; data drift might be addressed through recalibration; and protocol drift requires careful clinical review to determine whether the model’s assumptions remain valid[27]. Successful deployment, therefore, requires not merely a generic monitoring plan, but a nuanced framework that tracks multiple performance indicators, distinguishes between drift types, and triggers appropriate responses based on the underlying cause.
Third, workflow integration is paramount. A tool requiring manual entry of dozens of variables is impractical during an acute variceal bleed. For real-time utility, ML models must be seamlessly embedded within the electronic health record (EHR), with automated data extraction and minimal-click interfaces. Successful precedents in other acute care domains, like EHR-integrated sepsis prediction, provide a blueprint for hepatology to follow[27].
For clinicians, mature ML tools could eventually inform critical decisions regarding patient triage (e.g., direct admission to ICU), escalation of therapy (e.g., expedited TIPS consideration), and family communication about prognosis. They should function as intelligent decision-support systems, augmenting rather than replacing clinical judgment. However, for clinicians to trust and adopt these tools, evidence must extend beyond discrimination metrics to demonstrate that the model consistently adds value beyond what is already achievable with simple, familiar scores. This means that future validation studies should prioritize head-to-head comparisons using clinically meaningful endpoints, such as changes in management decisions, resource utilization, or patient outcomes, rather than relying solely on statistical measures of predictive accuracy. Decision-curve analysis and reclassification metrics should become standard components of model evaluation, providing clinicians with the information they need to assess whether adopting a new tool will meaningfully improve their decision-making.
It is important to emphasize that, in its current form, this model should not be used to withhold life-saving inter
For researchers, this study underscores the imperative of prospective validation, fairness auditing, and imple
The broader AI in medicine literature reinforces these directions, with comprehensive guides emphasizing that the strength of modern ML approaches lies in their ability to discover complex patterns that may escape traditional statistical modeling[31,32]. Furthermore, the evolving understanding of metabolic dysfunction-associated liver disease has created new opportunities for ML applications that account for complex etiological interactions in hepatology populations[33].
Emerging evidence specifically in variceal bleeding demonstrates that imaging-based ML models integrating radiomic features from cross-sectional imaging can outperform traditional clinical risk scores[34], and multimodal approaches combining endoscopic findings with structured clinical data show promise for improving point-of-care risk stratification[35].
Importantly, the contribution of Rech et al[7] should be viewed as a high-fidelity prototype rather than a finalized clinical product. Its value lies not only in its performance metrics but in its demonstration that prospective validation and real-world deployment are achievable within hepatology research. By making the model openly accessible, the authors invite the community to interrogate, validate, and refine it, a necessary step toward collective progress rather than isolated algorithm development[36].
The work by Rech et al[7] represents an important advance in translational AI for hepatology. Their combination of prospective validation, strong performance, and open deployment models a commitment to clinical relevance often absent in the field. While questions remain about the model’s incremental value relative to traditional scores, questions that can only be answered through rigorous comparative analyses using reclassification metrics and decision-curve analysis, the study provides a methodological template that moves the field closer to clinically meaningful imple
The promise of ML to refine risk stratification and improve outcomes in AEVB is clear. Realizing the promise of ML in acute esophageal variceal bleeding will depend not only on algorithmic performance but on rigorous external validation, ethical stewardship, and thoughtful integration into the realities of acute clinical care.
| 1. | Ding J, Zhao J, Huan L, Liu Y, Qiao Y, Wang Z, Chen Z, Huang S, Zhao Y, He X. Inflammation-Induced Long Intergenic Noncoding RNA (LINC00665) Increases Malignancy Through Activating the Double-Stranded RNA-Activated Protein Kinase/Nuclear Factor Kappa B Pathway in Hepatocellular Carcinoma. Hepatology. 2020;72:1666-1681. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 39] [Cited by in RCA: 57] [Article Influence: 9.5] [Reference Citation Analysis (0)] |
| 2. | Bettinger D, Sturm L, Pfaff L, Hahn F, Kloeckner R, Volkwein L, Praktiknjo M, Lv Y, Han G, Huber JP, Boettler T, Reincke M, Klinger C, Caca K, Heinzow H, Seifert LL, Weiss KH, Rupp C, Piecha F, Kluwe J, Zipprich A, Luxenburger H, Neumann-Haefelin C, Schmidt A, Jansen C, Meyer C, Uschner FE, Brol MJ, Trebicka J, Rössle M, Thimme R, Schultheiss M. Refining prediction of survival after TIPS with the novel Freiburg index of post-TIPS survival. J Hepatol. 2021;74:1362-1372. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 44] [Cited by in RCA: 140] [Article Influence: 28.0] [Reference Citation Analysis (0)] |
| 3. | Terres AZ, Balbinot RS, Muscope ALF, Eberhardt LZ, Balensiefer JIL, Cini BT, Rost GL, Longen ML, Schena B, Balbinot RA, Balbinot SS, Soldera J. Predicting mortality for cirrhotic patients with acute oesophageal variceal haemorrhage using liver‐specific scores. GastroHep. 2021;3:236-246. [DOI] [Full Text] |
| 4. | Kröner PT, Engels MM, Glicksberg BS, Johnson KW, Mzaik O, van Hooft JE, Wallace MB, El-Serag HB, Krittanawong C. Artificial intelligence in gastroenterology: A state-of-the-art review. World J Gastroenterol. 2021;27:6794-6824. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in CrossRef: 150] [Cited by in RCA: 122] [Article Influence: 24.4] [Reference Citation Analysis (0)] |
| 5. | Eskandar K. Artificial intelligence in hepatology: A comprehensive scoping review of clinical applications, challenges, and future directions. ILIVER. 2025;4:100205. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 6] [Reference Citation Analysis (0)] |
| 6. | Malik S, Tenorio BG, Moond V, Dahiya DS, Vora R, Dbouk N. Systematic review of machine learning models in predicting the risk of bleed/grade of esophageal varices in patients with liver cirrhosis: A comprehensive methodological analysis. J Gastroenterol Hepatol. 2024;39:2043-2059. [RCA] [PubMed] [DOI] [Full Text] [Cited by in RCA: 11] [Reference Citation Analysis (0)] |
| 7. | Rech MM, Corso LL, Dal Bó EF, Ferraza AD, Tomé F, Terres AZ, Balbinot RS, Balbinot RA, Balbinot SS, Soldera J. Development and prospective validation of a machine learning model to predict mortality in cirrhosis with esophageal variceal bleeding. World J Hepatol. 2026;18:111099. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 3] [Reference Citation Analysis (0)] |
| 8. | Pencina MJ, D'Agostino RB Sr, D'Agostino RB Jr, Vasan RS. Evaluating the added predictive ability of a new marker: from area under the ROC curve to reclassification and beyond. Stat Med. 2008;27:157-72; discussion 207. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 5307] [Cited by in RCA: 5213] [Article Influence: 289.6] [Reference Citation Analysis (5)] |
| 9. | Leening MJ, Vedder MM, Witteman JC, Pencina MJ, Steyerberg EW. Net reclassification improvement: computation, interpretation, and controversies: a literature review and clinician's guide. Ann Intern Med. 2014;160:122-131. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 365] [Cited by in RCA: 472] [Article Influence: 39.3] [Reference Citation Analysis (0)] |
| 10. | Vickers AJ, Van Calster B, Steyerberg EW. Net benefit approaches to the evaluation of prediction models, molecular markers, and diagnostic tests. BMJ. 2016;352:i6. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 812] [Cited by in RCA: 811] [Article Influence: 81.1] [Reference Citation Analysis (5)] |
| 11. | Pencina MJ, D'Agostino RB Sr, Steyerberg EW. Extensions of net reclassification improvement calculations to measure usefulness of new biomarkers. Stat Med. 2011;30:11-21. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 2030] [Cited by in RCA: 2140] [Article Influence: 142.7] [Reference Citation Analysis (6)] |
| 12. | Kerr KF, Brown MD, Zhu K, Janes H. Assessing the Clinical Impact of Risk Prediction Models With Decision Curves: Guidance for Correct Interpretation and Appropriate Use. J Clin Oncol. 2016;34:2534-2540. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 531] [Cited by in RCA: 536] [Article Influence: 53.6] [Reference Citation Analysis (0)] |
| 13. | Gao Y, Yu Q, Li X, Xia C, Zhou J, Xia T, Zhao B, Qiu Y, Zha JH, Wang Y, Tang T, Lv Y, Ye J, Xu C, Ju S. An imaging-based machine learning model outperforms clinical risk scores for prognosis of cirrhotic variceal bleeding. Eur Radiol. 2023;33:8965-8973. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 23] [Cited by in RCA: 25] [Article Influence: 8.3] [Reference Citation Analysis (2)] |
| 14. | Collins GS, Moons KGM, Dhiman P, Riley RD, Beam AL, Van Calster B, Ghassemi M, Liu X, Reitsma JB, van Smeden M, Boulesteix AL, Camaradou JC, Celi LA, Denaxas S, Denniston AK, Glocker B, Golub RM, Harvey H, Heinze G, Hoffman MM, Kengne AP, Lam E, Lee N, Loder EW, Maier-Hein L, Mateen BA, McCradden MD, Oakden-Rayner L, Ordish J, Parnell R, Rose S, Singh K, Wynants L, Logullo P. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 1587] [Cited by in RCA: 2241] [Article Influence: 1120.5] [Reference Citation Analysis (10)] |
| 15. | Liu X, Cruz Rivera S, Moher D, Calvert MJ, Denniston AK; SPIRIT-AI and CONSORT-AI Working Group. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Lancet Digit Health. 2020;2:e537-e548. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 334] [Cited by in RCA: 322] [Article Influence: 53.7] [Reference Citation Analysis (1)] |
| 16. | Sachan A, Kushwah S. Letter to the editor: GALAD score: Superior surveillance strategy or not? Hepatology. 2022;76:E38-E39. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 1] [Cited by in RCA: 2] [Article Influence: 0.5] [Reference Citation Analysis (0)] |
| 17. | Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366:447-453. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 1119] [Cited by in RCA: 2732] [Article Influence: 455.3] [Reference Citation Analysis (3)] |
| 18. | Vyas DA, Eisenstein LG, Jones DS. Hidden in Plain Sight - Reconsidering the Use of Race Correction in Clinical Algorithms. N Engl J Med. 2020;383:874-882. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 631] [Cited by in RCA: 1225] [Article Influence: 204.2] [Reference Citation Analysis (3)] |
| 19. | Mitchell S, Potash E, Barocas S, D’amour A, Lum K. Algorithmic Fairness: Choices, Assumptions, and Definitions. Annu Rev Stat Appl. 2021;8:141-163. [DOI] [Full Text] |
| 20. | Chen IY, Pierson E, Rose S, Joshi S, Ferryman K, Ghassemi M. Ethical Machine Learning in Healthcare. Annu Rev Biomed Data Sci. 2021;4:123-144. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 216] [Cited by in RCA: 277] [Article Influence: 55.4] [Reference Citation Analysis (4)] |
| 21. | Panagides R, Keim-Malpass J. Reconsidering the use of race, sex, and age in clinical algorithms to address bias in practice: A discussion paper. Int J Nurs Stud Adv. 2025;9:100380. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 4] [Reference Citation Analysis (0)] |
| 22. | Chen IY, Szolovits P, Ghassemi M. Can AI Help Reduce Disparities in General Medical and Mental Health Care? AMA J Ethics. 2019;21:E167-E179. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 126] [Cited by in RCA: 200] [Article Influence: 28.6] [Reference Citation Analysis (0)] |
| 23. | Ferryman K. Addressing health disparities in the Food and Drug Administration's artificial intelligence and machine learning regulatory framework. J Am Med Inform Assoc. 2020;27:2016-2019. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 12] [Cited by in RCA: 41] [Article Influence: 8.2] [Reference Citation Analysis (0)] |
| 24. | Critelli B, Hassan A, Lahooti I, Noh L, Park JS, Tong K, Lahooti A, Matzko N, Adams JN, Liss L, Quion J, Restrepo D, Nikahd M, Culp S, Lacy-Hulbert A, Speake C, Buxbaum J, Bischof J, Yazici C, Evans-Phillips A, Terp S, Weissman A, Conwell D, Hart P, Ramsey M, Krishna S, Han S, Park E, Shah R, Akshintala V, Windsor JA, Mull NK, Papachristou G, Celi LA, Lee P. A systematic review of machine learning-based prognostic models for acute pancreatitis: Towards improving methods and reporting quality. PLoS Med. 2025;22:e1004432. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 13] [Reference Citation Analysis (0)] |
| 25. | Finlayson SG, Subbaswamy A, Singh K, Bowers J, Kupke A, Zittrain J, Kohane IS, Saria S. The Clinician and Dataset Shift in Artificial Intelligence. N Engl J Med. 2021;385:283-286. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 543] [Cited by in RCA: 482] [Article Influence: 96.4] [Reference Citation Analysis (3)] |
| 26. | Vokinger KN, Feuerriegel S, Kesselheim AS. Continual learning in medical devices: FDA's action plan and beyond. Lancet Digit Health. 2021;3:e337-e338. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 19] [Cited by in RCA: 59] [Article Influence: 11.8] [Reference Citation Analysis (0)] |
| 27. | Wong A, Otles E, Donnelly JP, Krumm A, McCullough J, DeTroyer-Cooley O, Pestrue J, Phillips M, Konye J, Penoza C, Ghous M, Singh K. External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Intern Med. 2021;181:1065-1070. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 742] [Cited by in RCA: 588] [Article Influence: 117.6] [Reference Citation Analysis (0)] |
| 28. | Trebicka J, Fernandez J, Papp M, Caraceni P, Laleman W, Gambino C, Giovo I, Uschner FE, Jansen C, Jimenez C, Mookerjee R, Gustot T, Albillos A, Bañares R, Jarcuska P, Steib C, Reiberger T, Acevedo J, Gatti P, Shawcross DL, Zeuzem S, Zipprich A, Piano S, Berg T, Bruns T, Danielsen KV, Coenraad M, Merli M, Stauber R, Zoller H, Ramos JP, Solé C, Soriano G, de Gottardi A, Gronbaek H, Saliba F, Trautwein C, Kani HT, Francque S, Ryder S, Nahon P, Romero-Gomez M, Van Vlierberghe H, Francoz C, Manns M, Garcia-Lopez E, Tufoni M, Amoros A, Pavesi M, Sanchez C, Praktiknjo M, Curto A, Pitarch C, Putignano A, Moreno E, Bernal W, Aguilar F, Clària J, Ponzo P, Vitalis Z, Zaccherini G, Balogh B, Gerbes A, Vargas V, Alessandria C, Bernardi M, Ginès P, Moreau R, Angeli P, Jalan R, Arroyo V; PREDICT STUDY group of the EASL-CLIF CONSORTIUM. PREDICT identifies precipitating events associated with the clinical course of acutely decompensated cirrhosis. J Hepatol. 2021;74:1097-1108. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 237] [Cited by in RCA: 236] [Article Influence: 47.2] [Reference Citation Analysis (5)] |
| 29. | Degasperi E, Anolli MP, Lampertico P. Bulevirtide for patients with compensated chronic hepatitis delta: A review. Liver Int. 2023;43 Suppl 1:80-86. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 9] [Cited by in RCA: 17] [Article Influence: 5.7] [Reference Citation Analysis (0)] |
| 30. | Roberts M, Driggs D, Thorpe M, Gilbey J, Yeung M, Ursprung S, Aviles-Rivero AI, Etmann C, Mccague C, Beer L, Weir-Mccall JR, Teng Z, Gkrania-Klotsas E, AIX-COVNET, Ruggiero A, Korhonen A, Jefferson E, Ako E, Langs G, Gozaliasl G, Yang G, Prosch H, Preller J, Stanczuk J, Tang J, Hofmanninger J, Babar J, Sánchez LE, Thillai M, Gonzalez PM, Teare P, Zhu X, Patel M, Cafolla C, Azadbakht H, Jacob J, Lowe J, Zhang K, Bradley K, Wassin M, Holzer M, Ji K, Ortet MD, Ai T, Walton N, Lio P, Stranks S, Shadbahr T, Lin W, Zha Y, Niu Z, Rudd JHF, Sala E, Schönlieb C. Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans. Nat Mach Intell. 2021;3:199-217. [RCA] [DOI] [Full Text] [Cited by in Crossref: 282] [Cited by in RCA: 476] [Article Influence: 95.2] [Reference Citation Analysis (0)] |
| 31. | Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. 2019;25:44-56. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 6739] [Cited by in RCA: 4562] [Article Influence: 651.7] [Reference Citation Analysis (9)] |
| 32. | Esteva A, Robicquet A, Ramsundar B, Kuleshov V, DePristo M, Chou K, Cui C, Corrado G, Thrun S, Dean J. A guide to deep learning in healthcare. Nat Med. 2019;25:24-29. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 3854] [Cited by in RCA: 2052] [Article Influence: 293.1] [Reference Citation Analysis (11)] |
| 33. | Tang SY, Tan JS, Pang XZ, Lee GH. Metabolic dysfunction associated fatty liver disease: The new nomenclature and its impact. World J Gastroenterol. 2023;29:549-560. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in CrossRef: 22] [Cited by in RCA: 20] [Article Influence: 6.7] [Reference Citation Analysis (0)] |
| 34. | Chirapongsathorn S, Akkarachinores K, Chaiprasert A. Development and validation of prognostic model to predict mortality among cirrhotic patients with acute variceal bleeding: A retrospective study. JGH Open. 2021;5:658-663. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 7] [Cited by in RCA: 8] [Article Influence: 1.6] [Reference Citation Analysis (0)] |
| 35. | Wang Y, Hong Y, Wang Y, Zhou X, Gao X, Yu C, Lin J, Liu L, Gao J, Yin M, Xu G, Liu X, Zhu J. Automated Multimodal Machine Learning for Esophageal Variceal Bleeding Prediction Based on Endoscopy and Structured Data. J Digit Imaging. 2023;36:326-338. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 38] [Cited by in RCA: 34] [Article Influence: 11.3] [Reference Citation Analysis (0)] |
| 36. | Xiang Y, Yang N, Zheng TL, Huang YF, Liu TY, Ma DQ, Hu SJ, Zhang WH, Xiang HL, Zhang LY, Yuan LL, Wang X, Dang T, Zhang G, Wu B, Peng LJ, Gao M, Xia DL, Liu ZB, Li J, Song Y, Zhou XQ, Qi XS, Zeng J, Tan XY, Deng MM, Fang HM, Qi SL, He S, He YF, Ye B, Wu W, Shao JB, Wei W, Hu JP, Yong X, He CH, Bao JL, Zhang YN, Ji R, Bo Y, Yan W, Li HJ, Li SL, Geng S, Zhao L, Liu B, Qi XL. Development of a deep learning model for guiding treatment decisions of acute variceal bleeding in patients with cirrhosis. World J Gastroenterol. 2025;31:111361. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in RCA: 2] [Reference Citation Analysis (3)] |