BPG is committed to discovery and dissemination of knowledge
Retrospective Study Open Access
Copyright: ©Author(s) 2026. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution-NonCommercial (CC BY-NC 4.0) license. No commercial re-use. See permissions. Published by Baishideng Publishing Group Inc.
World J Gastroenterol. Nov 7, 2026; 32(41): 120899
Published online Nov 7, 2026. doi: 10.3748/wjg.120899
Endoscopic ultrasound-based deep learning for predicting chemotherapy response in unresectable pancreatic ductal adenocarcinoma
Ze-Hua Li, Jun Weng, Shi-Yong Lin, Shuo Li, Kun-Hao Bai, Guo-Liang Xu, Department of Endoscopy, Sun Yat-sen University Cancer Center, State Key Laboratory of Oncology in South China, Guangdong Provincial Clinical Research Center for Cancer, Guangzhou 510060, Guangdong Province, China
Yu-Hong Zeng, Department of Medical Engineering, Zhujiang Hospital of Southern Medical University, Guangzhou 510280, Guangdong Province, China
Shuo Li, State Key Laboratory of Liver Research, The University of Hong Kong, Pokfulam, Hong Kong 999077, China
ORCID number: Ze-Hua Li (0000-0001-5624-0087); Jun Weng (0000-0003-0792-1526); Yu-Hong Zeng (0000-0002-2654-5653); Shi-Yong Lin (0000-0002-3881-6422); Kun-Hao Bai (0000-0003-1184-7576); Guo-Liang Xu (0000-0002-8882-2636).
Co-first authors: Ze-Hua Li and Jun Weng.
Co-corresponding authors: Kun-Hao Bai and Guo-Liang Xu.
Author contributions: Li ZH and Weng J conceived and designed the study; they contributed equally to this work and are the co-first authors; Li ZH led the project implementation, performed the data analysis, developed the deep learning workflow, interpreted the results, and drafted the manuscript; Weng J and Zeng YH coordinated patient recruitment, acquired clinical data, completed tumor image annotation, collected follow-up information, and verified the relevant data; Lin SY assisted with model construction, computational and statistical validation, and figure preparation; Li S was responsible for data curation, image preprocessing, database management, and manuscript editing; Bai KH and Xu GL supervised the study, provided intellectual input, acquired resources and funding support, and critically revised the manuscript, they contributed equally to this article and are the co-corresponding authors; all authors read and approved the final manuscript.
AI contribution statement: The authors used artificial intelligence tools, including ChatGPT, only for language polishing, translation assistance, and writing assistance during manuscript preparation and revision. No artificial intelligence tool was used for study design, data analysis, image generation, interpretation of results, or generation of scientific conclusions. The scientific content, responses to reviewers, and final manuscript were prepared, reviewed, and approved by the authors.
Supported by the National Natural Science Foundation of China (General Program), No. 82403973, No. 82200442, and No. 82373118; Guangdong Basic and Applied Basic Research Foundation, No. 2023A1515010828; Science and Technology Program of Guangzhou, No. 2025A04J3768; Guangdong Medical Equipment Association Research Fund, No. YZXH2025KT07; and Hong Kong Scholar, Hong Kong Scholar, No. XJWQ2025016.
Institutional review board statement: This study was approved by the Medical Ethics Committee of Sun Yat-sen University Cancer Center (Approval No. SL-G2023-244-01).
Informed consent statement: The requirement for informed consent was waived by the Institutional Review Board.
Conflict-of-interest statement: All the authors report no relevant conflicts of interest for this article.
Data sharing statement: The data that support the findings of this study are available from the corresponding author upon reasonable request. Owing to institutional regulations and patient privacy considerations, the data are not publicly available.
Corresponding author: Guo-Liang Xu, MD, PhD, Professor, Chief Physician, Department of Endoscopy, Sun Yat-sen University Cancer Center, State Key Laboratory of Oncology in South China, Guangdong Provincial Clinical Research Center for Cancer, No. 651 Dongfeng East Road, Yuexiu District, Guangzhou 510060, Guangdong Province, China. xugl@sysucc.org.cn
Received: March 12, 2026
Revised: April 20, 2026
Accepted: June 3, 2026
Published online: November 7, 2026
Processing time: 191 Days and 2.7 Hours

Abstract
BACKGROUND

Pancreatic ductal adenocarcinoma (PDAC) is one of the most lethal malignancies, and chemotherapy remains the main treatment option for patients with unresectable disease. Reliable prediction of chemotherapy response is therefore critical for individualized treatment planning. Endoscopic ultrasound (EUS) provides high-resolution, tumor-specific imaging that may be well suited for deep learning-based prediction of chemotherapy response.

AIM

To develop and validate an EUS-based deep learning model for predicting first-line chemotherapy response in patients with unresectable PDAC, and to evaluate its prognostic value for overall survival (OS) stratification.

METHODS

In this retrospective, single-center study, 190 patients with histologically confirmed PDAC diagnosed between February 2016 and March 2025 were included. All patients were deemed ineligible for upfront surgical resection based on formal multidisciplinary team evaluation, including those with locally advanced or metastatic disease and those considered medically or personally unsuitable for surgery. Patients diagnosed in 2023 or earlier were assigned to the training and internal validation cohort (n = 152) using five-fold cross-validation, whereas patients diagnosed from January 2024 onward at a separate campus of the same institution constituted a temporally independent test cohort with site separation (n = 38). Tumor regions of interest were manually annotated on pre-treatment EUS images. Four convolutional neural network (CNN) architectures (VGG19, VGG19-BN, ResNet50, and ResNeXt50) were trained to predict progressive disease, as defined by RECIST version 1.1 after first-line chemotherapy. Model performance was evaluated at the patient level using the area under the receiver operating characteristic curve (AUC) and accuracy, and model-derived risk stratification was assessed for OS.

RESULTS

In the independent test cohort, ResNeXt50 demonstrated the best discriminatory performance (AUC = 0.848; sensitivity, 65.38%; specificity, 84.19%), followed by ResNet50 (AUC = 0.844), whereas VGG19-BN (AUC = 0.775) and VGG19 (AUC = 0.614) showed lower performance. Model-predicted probabilities enabled effective risk stratification, with high-risk patients exhibiting significantly shorter OS compared with low-risk patients (log-rank P < 0.01; hazard ratio = 1.77). In contrast, baseline carbohydrate antigen 19-9-based stratification did not significantly discriminate survival, and CNN-based models demonstrated higher discriminatory performance compared with the conventional serum biomarker.

CONCLUSION

CNN models based on region-of-interest-level EUS imaging can predict chemotherapy response and provide clinically meaningful prognostic stratification in PDAC. Among the evaluated architectures, ResNeXt50 demonstrated the most consistent and robust performance across independent testing and sensitivity analyses. This noninvasive imaging-based approach shows potential for integration into clinical workflows to support personalized treatment planning.

Key Words: Pancreatic ductal adenocarcinoma; Endoscopic ultrasound; Convolutional neural network; Deep learning; Chemotherapy response

Core Tip: Pancreatic ductal adenocarcinoma (PDAC) has a poor prognosis, and predicting chemotherapy response remains challenging. In this study, we developed deep learning models based on pretreatment endoscopic ultrasound images to predict chemotherapy responses in patients with PDAC. Four convolutional neural network architectures were evaluated, with ResNeXt50 demonstrating the best performance in the independent test cohort. The model also enabled effective risk stratification for overall survival, outperforming the conventional serum biomarker carbohydrate antigen 19-9. These findings suggest that endoscopic ultrasound-based convolutional neural network models may provide a noninvasive tool to support individualized treatment planning in PDAC.



INTRODUCTION

Pancreatic ductal adenocarcinoma (PDAC) remains one of the most lethal malignancies, with a 5-year survival rate below 12% despite advances in diagnostic and therapeutic strategies[1]. At diagnosis, the majority of patients present with unresectable disease, often owing to distant metastases or extensive local vascular involvement that precludes surgical resection[2]. For patients with unresectable PDAC, systemic chemotherapy constitutes the mainstay of treatment; however, therapeutic responses vary substantially, reflecting the pronounced biological and molecular heterogeneity of the disease[3]. Accurate prediction of chemotherapy response is therefore essential for optimizing treatment strategies, improving clinical outcomes, and minimizing unnecessary treatment-related toxicity.

Several approaches have been investigated to predict chemotherapy response in PDAC, including clinical parameters, serum biomarkers, and molecular profiling. Carbohydrate antigen 19-9 (CA19-9) is the most widely used biomarker in clinical practice; however, its predictive utility is limited by suboptimal sensitivity (SEN) and specificity (SPE), particularly in patients with low or undetectable baseline levels[4,5]. Conventional imaging modalities, such as computed tomography (CT) and magnetic resonance imaging (MRI), provide important anatomical and functional information but often fail to capture subtle intratumoral heterogeneity that is closely linked to chemotherapy response[6]. Consequently, there remains a need for more accurate, noninvasive, and reproducible tools to guide personalized therapeutic decision-making in PDAC.

Endoscopic ultrasound (EUS) has emerged as a highly sensitive modality for the detection and characterization of pancreatic lesions[7]. In addition to its diagnostic role, EUS offers high-resolution, tumor-specific imaging that is well suited for advanced computational analysis. Recent advances in artificial intelligence, particularly deep learning approaches such as convolutional neural networks (CNNs), have demonstrated a strong capability to extract high-dimensional imaging features from ultrasound-based modalities that are imperceptible to the human eye yet clinically relevant[8,9]. Although CNN-based models have been applied in various oncologic settings to predict treatment response and prognosis, their integration with EUS imaging for predicting chemotherapy response in pancreatic cancer remains relatively underexplored[10,11].

In this study, we developed and validated region-of-interest-based CNN models using manually annotated EUS images to predict chemotherapy response in PDAC. We further evaluated the prognostic value of model-derived probability scores by stratifying patients into distinct risk groups and compared model performance with that of the conventional serum biomarker CA19-9. Our findings highlight the potential of noninvasive, imaging-based deep learning approaches to support individualized treatment planning and inform clinical decision-making in pancreatic cancer.

MATERIALS AND METHODS
Study design and participants

This retrospective, single-center study was conducted at the Endoscopy Department of Sun Yat-sen University Cancer Center (SYSUCC). The study protocol was reviewed and approved by the Institutional Review Board of SYSUCC (Approval No. SL-G2023-244-01), and the study was conducted in accordance with the principles of the Declaration of Helsinki. Given the retrospective study design and the use of anonymized clinical and imaging data, the requirement for written informed consent was waived by the institutional ethics committee.

Between February 2016 and March 2025, a total of 4721 consecutive patients with histologically confirmed PDAC who underwent diagnostic EUS at SYSUCC were screened for eligibility. After the application of predefined inclusion and exclusion criteria, 190 patients were included in the final analysis. Eligible patients were required to have: (1) Histopathologically confirmed PDAC; (2) Available pretreatment EUS images of sufficient quality for analysis; (3) Receipt of first-line chemotherapy, with treatment response assessed according to RECIST version 1.1; and (4) Complete clinicopathological and follow-up data. Exclusion criteria included prior pancreatic surgery, radiotherapy, or local ablative therapy before chemotherapy; poor-quality or incomplete EUS images; missing key clinical or follow-up information; and the presence of concurrent primary malignancies or severe comorbidities that could confound treatment response or survival outcomes. All included patients had an Eastern Cooperative Oncology Group (ECOG) performance status of 0–2, consistent with the standard eligibility criteria for the administered chemotherapy regimens according to National Comprehensive Cancer Network (NCCN) guidelines.

All included patients were deemed ineligible for upfront surgical resection based on formal multidisciplinary team (MDT) evaluation conducted at SYSUCC, in accordance with NCCN guidelines. Surgical ineligibility was determined by MDT consensus rather than by American Joint Committee on Cancer (AJCC) clinical stage alone, as anatomical staging does not fully capture vascular relationships, functional operability, or patient-level surgical risk. Among the 18 patients with AJCC clinical stage I or II disease included in this cohort, ineligibility for upfront surgery was attributable to one or more of the following MDT-documented criteria: (1) Significant abutment or encasement of the superior mesenteric artery, superior mesenteric vein, or portal vein that was not fully reflected in the AJCC stage classification, resulting in MDT consensus for a borderline resectable or locally advanced designation; (2) Poor performance status (ECOG ≥ 2) or severe cardiopulmonary or hepatic comorbidities precluding safe surgical intervention; or (3) MDT consensus favoring primary systemic chemotherapy owing to an unacceptably high operative risk, independent of anatomical resectability. The remaining patients presented with AJCC stage III or IV disease (locally advanced or metastatic PDAC), for whom systemic chemotherapy represents the standard of care as per current guidelines.

Patients diagnosed in 2023 or earlier were assigned to the training and internal validation cohort (n = 152) using five-fold cross-validation, whereas patients diagnosed from January 2024 onward were allocated to a temporally independent test cohort (n = 38). The overall study workflow, including patient screening, application of the eligibility criteria, and cohort allocation, is summarized in Figure 1.

Figure 1
Figure 1 Flowchart of patient selection and cohort allocation. This flowchart illustrates the screening and selection process of patients with histologically confirmed pancreatic ductal adenocarcinoma who underwent endoscopic ultrasound between February 2016 and March 2025. Of the 4721 patients initially screened, predefined inclusion and exclusion criteria were applied, resulting in 190 eligible patients. These patients were subsequently allocated to a training and internal validation cohort with five-fold cross-validation (n = 152) and a temporally independent test cohort (n = 38). EUS: Endoscopic ultrasound; PDAC: Pancreatic ductal adenocarcinoma.
Chemotherapy regimens and response assessment

All patients received standard first-line chemotherapy regimens in accordance with the NCCN guidelines. The modified FOLFIRINOX (mFOLFIRINOX) regimen consisted of oxaliplatin (85 mg/m2), irinotecan (180 mg/m2), leucovorin (400 mg/m2), and 5-fluorouracil (400 mg/m2) administered as an intravenous bolus, followed by a continuous infusion of 5-fluorouracil (2400 mg/m2) over 46 h, repeated every 2 weeks. Alternatively, the gemcitabine plus nab-paclitaxel regimen comprised gemcitabine (1000 mg/m2) and nab-paclitaxel (125 mg/m2) administered intravenously on days 1, 8, and 15 of a 28-day cycle. The choice of regimen was determined by the treating oncologist based on patient-level factors, including age, ECOG performance status, organ function, and anticipated tolerability, in accordance with NCCN guideline recommendations. The distribution of chemotherapy regimens is summarized in Table 1. A complete patient-level enumeration of all individual chemotherapy regimens administered in the cohort, including the alternative regimens classified as “Other regimens” in Table 1, is provided in Supplementary Table 1.

Table 1 Baseline clinical characteristics of patients stratified by RECIST version 1.1-based chemotherapy response groups.
Characteristics
Non-PD
PD
P value
Age, years59.65 ± 8.2658.90 ± 10.090.626
Female62 (41.6)11 (26.8)0.123
T category0.500
T14 (2.7)1 (2.4)
T28 (5.4)1 (2.4)
T341 (27.5)16 (39.0)
T496 (64.4)23 (56.1)
N category0.166
N021 (14.1)3 (7.3)
N1114 (76.5)30 (73.2)
N210 (6.7)7 (17.1)
N34 (2.7)1 (2.4)
M category0.588
M060 (40.3)14 (34.1)
M189 (59.7)27 (65.9)
Clinical stage0.906
I3 (2.0)1 (2.4)
II11 (7.4)3 (7.3)
III45 (30.2)10 (24.4)
IV90 (60.4)27 (65.9)
CA19-9, U/mL1908.72 ± 4357.011884.20 ± 4388.780.975
Tumor size, mm47.59 ± 18.9342.05 ± 17.400.081
Chemotherapy0.259
mFOLFIRINOX39 (26.2)16 (39.0)
Gemcitabine + nab-PTX66 (44.3)14 (34.1)
Other regimens44 (29.5)11 (26.8)

Tumor response was assessed after two to three cycles of first-line chemotherapy using contrast-enhanced CT or MRI. Given the different administration schedules of the two regimens, response assessment was performed after 2–3 cycles (approximately 4–6 weeks) for patients receiving mFOLFIRINOX and after 2–3 cycles (approximately 8–12 weeks) for patients receiving gemcitabine plus nab-paclitaxel, consistent with the standard clinical practice at our institution. Imaging studies were independently reviewed by two board-certified radiologists, each with more than 10 years of experience in abdominal imaging, who were blinded to the clinical outcomes. Response assessment was performed before and independently of the deep learning model development; model predictions were generated retrospectively from pretreatment EUS images and had no influence on the RECIST version 1.1-based response classifications used as training labels. Treatment response was classified according to RECIST version 1.1 as progressive disease (PD) or disease control, with disease control defined as complete response, partial response, or stable disease.

EUS image acquisition and preprocessing

All EUS examinations were performed at SYSUCC by experienced endoscopists with expertise in endoscopic ultrasonography and at least 8 years of clinical experience, using a linear-array echoendoscope (GF-UCT260; Olympus Medical Systems, Tokyo, Japan) coupled with an EU-ME2 PREMIER PLUS ultrasound processor (Olympus). Patients fasted for at least 6 h before the procedure and received intravenous propofol sedation under continuous cardiorespiratory monitoring. Pancreatic lesions were systematically examined in multiple scanning planes to ensure comprehensive tumor visualization.

Representative static images and cine loops were exported in Digital Imaging and Communications in Medicine format and anonymized before subsequent analysis. Tumor regions of interest (ROIs) were manually delineated using LabelMe software (MIT CSAIL, Cambridge, MA, United States) by two endosonographers with 8 and 12 years of experience, respectively. Any discrepancies were reviewed and resolved by a senior expert with more than 15 years of experience in EUS interpretation. Interobserver agreement for ROI annotation was assessed using Cohen’s κ coefficient, with κ values > 0.80 indicating excellent reproducibility.

All annotated ROIs were cropped, resized to 224 × 224 pixels using bilinear interpolation, and normalized to have zero mean and unit variance before being input into the CNN models.

Model development and training

Four CNN architectures—VGG19, VGG19-BN, ResNet50, and ResNeXt50—were implemented using the PyTorch framework (version ≥ 2.0). To improve model generalization and mitigate overfitting, data augmentation strategies were applied during training, including random rotations (± 15°), horizontal and vertical flips, Gaussian noise addition, and random brightness and contrast adjustments.

Model training was performed on a workstation equipped with an NVIDIA CUDA-enabled graphics processing unit. The Adam optimizer was used with an initial learning rate of 0.001 and a weight decay of 1 × 10−5. Binary cross-entropy loss was adopted, and the batch size was set to 16. Training was conducted for up to 100 epochs, with early stopping applied when no improvement in validation area under the receiver operating characteristic curve (AUC) was observed for 15 consecutive epochs.

Key hyperparameters, including the learning rate, weight decay, and data augmentation intensity, were optimized using a grid-search strategy within the training and internal validation framework. Model performance during training was monitored using the AUC and accuracy (ACC) across five-fold internal validation.

Risk stratification and survival analysis

Model-predicted probabilities of PD were used to generate patient-level risk scores for survival analysis. To avoid patient-level information leakage, all dataset partitions were performed at the patient level. For the development cohort (n = 152), predictions were generated within the five-fold cross-validation framework. In each fold, the model was trained on four folds and generated predictions for the held-out fifth fold, ensuring that each patient’s prediction was produced by a model that had not been trained on that patient. For patients with multiple EUS images, the patient-level risk score was computed as the arithmetic mean of the three highest image-level predicted probabilities of PD for that patient (top-3 probability averaging). For patients with fewer than three available EUS images, the patient-level risk score was computed as the arithmetic mean of all available image-level predicted probabilities. For the independent test cohort (n = 38), patient-level risk scores were generated using the trained ResNeXt50 model and the same image-to-patient aggregation strategy.

High-risk and low-risk groups were defined according to the median patient-level risk score of the analyzed cohort. Patients with risk scores greater than or equal to the median were assigned to the high-risk group, whereas those below the median were assigned to the low-risk group. This cutoff was not optimized using survival outcomes and was not selected based on the independent test cohort, thereby reducing the risk of circularity in the prognostic analysis.

Overall survival (OS) was estimated using the Kaplan–Meier method and was defined as the time from initiation of first-line chemotherapy to death from any cause or the date of last follow-up. Differences in survival distributions between risk groups were evaluated using the log-rank test. Hazard ratios (HRs) and 95% confidence intervals (CIs) were estimated using Cox proportional hazards regression to quantify the association between risk group and OS. Survival analyses were performed on the full 190-patient cohort, including both the development cohort and the independent test cohort, using the risk group assignments described above.

Model evaluation

Model performance was evaluated in the independent test cohort exclusively at the patient level using multiple classification metrics, including the AUC, ACC, SEN, SPE, positive predictive value, and negative predictive value. All dataset splits were performed strictly at the patient level, ensuring that no patient’s images appeared in more than one cohort at any stage of the analysis. For the five-fold cross-validation within the development cohort (n = 152), each fold represented a patient-level partition, such that all images from a given patient were assigned exclusively to either the training or the validation subset within each fold, with no overlap across folds. The independent test cohort (n = 38) was entirely withheld from all stages of model training, hyperparameter tuning, and cutoff determination. This patient-level partitioning strategy ensures that no patient’s data contributed to both training and evaluation at any stage, thereby excluding the possibility of information leakage.

During model training, all available static frames and cine-loop frames from each patient were used at the image level to maximize the effective training sample size and improve model generalization. For model evaluation, the patient-level predicted probability was defined as the arithmetic mean of the three highest image-level predicted probabilities of PD for that patient (top-3 probability averaging). For patients with fewer than three available EUS images, the patient-level predicted probability was defined as the arithmetic mean of all available image-level predicted probabilities. This single aggregation rule was applied uniformly across all four CNN architectures and all cohorts (training, internal validation, and independent test), and was used for both patient-level performance evaluation and generation of patient-level risk scores in the survival analysis. The reported AUC values across all cohorts—training, internal validation, and independent test—reflect this patient-level aggregation and represent the primary performance metric throughout the manuscript.

The predictive performance of the CNN-based models was further compared with that of a CA19-9-based random forest (RF) classifier. The RF model was constructed using three baseline CA19-9-derived input features: (1) The raw baseline CA19-9 value (log-transformed to address distributional skewness); (2) A binary indicator of CA19-9 positivity (≥ 37 U/mL); and (3) A binary indicator of markedly elevated CA19-9 (≥ 1000 U/mL). The RF model was trained and evaluated using the same five-fold cross-validation framework as the CNN models, thereby ensuring that both the CNN models and the CA19-9 comparator were evaluated as data-driven classifiers under identical validation conditions, enabling a methodologically consistent comparison.

Statistical analysis

All statistical analyses were performed using Python (version 3.10) and R (version 4.2.1). Continuous variables were summarized as the mean ± SD or median with interquartile range, as appropriate. Normality of continuous variables was assessed using standard tests, and comparisons between groups were conducted using Student’s t-test for normally distributed data or the Mann–Whitney U test for non-normally distributed data. Categorical variables were compared using the χ2 test or Fisher’s exact test, as appropriate. The final analyzed cohort of 190 patients had no missing data for primary outcomes or covariates included in the analysis; patients with missing key clinical or follow-up information were excluded at the screening stage (n = 182), as specified in the exclusion criteria.

Survival outcomes were analyzed using the Kaplan–Meier method, and differences between groups were evaluated with the log-rank test. HRs and corresponding 95%CIs were estimated using Cox proportional hazards regression models. Model discrimination was assessed using the AUC, with 95%CIs computed using the DeLong method. CIs for ACC, SEN, and SPE were calculated using the Wilson score interval method. All statistical tests were two-sided, and a P value < 0.05 was considered statistically significant.

Reproducibility and reporting

This study adhered to the Standards for Reporting Diagnostic ACC Studies 2015 and the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis 2022 guidelines. To ensure reproducibility, all source code, model architectures, and trained weights have been deposited in an open-access GitHub repository. The repository includes environment configuration files, data preprocessing scripts, training, validation, and inference pipelines for all CNN models, as well as statistical analysis scripts used for survival analyses. Random seeds were fixed, and deterministic training settings were applied to ensure consistent performance across runs. Representative anonymized example datasets, together with preprocessing scripts, are provided to facilitate independent replication.

RESULTS
Patient characteristics

A total of 190 patients with histologically confirmed PDAC were included in the final analysis. The median age of the full cohort was 62 (interquartile range, 55–69) years, and 114 patients (60.0%) were male. When stratified by RECIST-based chemotherapy response group (non-PD vs PD), the mean (± SD) age was 59.65 ± 8.26 years in the non-PD group and 58.90 ± 10.09 years in the PD group. The proportion of female patients was 41.6% (n = 62) in the non-PD group and 26.8% (n = 11) in the PD group, with no significant between-group difference in either age (P = 0.626) or sex distribution (P = 0.123). Baseline demographic and clinical characteristics of patients stratified by RECIST-based chemotherapy response (non-PD vs PD) are presented in Table 1. This RECIST-based response classification is distinct from the model-derived high-risk and low-risk groups used in the survival analysis (described below). No significant differences were observed between the non-PD and PD groups with respect to age, sex, tumor-node-metastasis stage, baseline CA19-9 levels, or tumor size (all P > 0.05), indicating good baseline comparability between the groups.

Performance of CNN models

Model performance of the four CNN architectures across the training, validation, and independent test cohorts is presented in Tables 2, 3, 4, and 5. Each table corresponds to a single CNN architecture (Table 2: VGG19; Table 3: VGG19-BN; Table 4: ResNet50; Table 5: ResNeXt50) and reports the patient-level AUC, ACC, SEN, SPE, positive predictive value, negative predictive value, false negative rate, and F1-score across all three cohorts. These performance patterns are visually illustrated in Figure 2 corresponding to VGG19, VGG19-BN, ResNet50, and ResNeXt50, respectively. VGG19 (Figure 2A) showed lower performance in the independent test cohort (AUC, 0.614), whereas VGG19-BN (Figure 2B) exhibited moderate discrimination (AUC, 0.775). ResNet50 (Figure 2C) showed stronger discriminatory performance in the independent test cohort (AUC, 0.844; ACC, 81.74%; SEN, 59.62%; SPE, 84.42%), and ResNeXt50 (Figure 2D) achieved the best overall performance (AUC, 0.848; ACC, 82.16%; SEN, 65.38%; SPE, 84.19%). Overall, ResNeXt50 and ResNet50 consistently demonstrated the strongest discriminatory performance across all datasets. Figure 2B displays image-level receiver operating characteristic (ROC) curves for VGG19-BN without patient-level aggregation and is provided for reference only.

Figure 2
Figure 2 Receiver operating characteristic curves of the four convolutional neural network models for chemotherapy response prediction. Each panel displays the receiver operating characteristic curves of a single convolutional neural network architecture across the training cohort (blue solid line), internal validation cohort (orange solid line), and independent test cohort (green dashed line). All dataset splits and performance evaluations were performed at the patient level, with image-level predicted probabilities aggregated to the patient level using top-3 probability averaging, defined as the arithmetic mean of the three highest image-level predicted probabilities for each patient, or the arithmetic mean of all available image-level probabilities when fewer than three images were available. The corresponding area under the receiver operating characteristic curve values for each model and cohort are indicated within each panel. A: VGG19; B: VGG19-BN (image-level, without patient-level aggregation or test-time augmentation, shown for reference only); C: ResNet50; D: ResNeXt50. AUC: Area under the receiver operating characteristic curve; ROC: Receiver operating characteristic.
Table 2 Performance metrics of the VGG19 model for chemotherapy response prediction in the training, internal validation, and independent test cohorts.
Group
AUC (95%CI)
ACC (%) (95%CI)
SEN (%) (95%CI)
SPE (%) (95%CI)
PPV (%)
NPV (%)
FNR (%)
F1
Train0.96391.1182.8192.7369.0196.5017.190.753
Val0.75879.0058.2583.0640.1991.0541.750.476
Test0.614 (0.532–0.692)71.58 (67.84–75.31)36.54 (22.91–50.00)75.81 (71.82–79.49)15.4590.8163.460.217
Table 3 Performance metrics of the VGG19-BN model for chemotherapy response prediction in the training, validation, and independent test cohorts.
Group
AUC (95%CI)
ACC (%) (95%CI)
SEN (%) (95%CI)
SPE (%) (95%CI)
PPV (%)
NPV (%)
FNR (%)
F1
Train1.00099.77100.0099.7398.622100.0000.993
Val0.86490.3666.6794.9972.2493.5833.330.693
Test0.775 (0.705–0.841)82.78 (79.05–85.89)46.15 (32.43–59.02)87.21 (83.72–90.19)30.3893.0553.850.366
Table 4 Performance metrics of the ResNet50 model for chemotherapy response prediction in the training, internal validation, and independent test cohorts.
Group
AUC (95%CI)
ACC (%) (95%CI)
SEN (%) (95%CI)
SPE (%) (95%CI)
PPV (%)
NPV (%)
FNR (%)
F1
Train0.96292.7179.3095.3476.8795.9320.700.781
Val0.87183.6576.4985.0550.0094.8723.510.605
Test0.844 (0.798–0.884)81.74 (78.01–85.06)59.62 (45.44–72.73)84.42 (80.84–87.65)31.6394.5340.380.413
Table 5 Performance metrics of the ResNeXt50 model for chemotherapy response prediction in the training, internal validation, and independent test cohorts.
Group
AUC (95%CI)
ACC (%) (95%CI)
SEN (%) (95%CI)
SPE (%) (95%CI)
PPV (%)
NPV (%)
FNR (%)
F1
Train0.97193.3483.5195.2777.5296.7316.490.804
Val0.89688.8175.7991.3663.1695.0724.210.689
Test0.848 (0.798–0.890)82.16 (78.63–85.48)65.38 (51.06–77.78)84.19 (80.96–87.30)33.3395.2634.620.442

To provide a non-imaging clinical benchmark, an RF model based on baseline CA19-9-derived variables was constructed, and its performance metrics are summarized in Table 6. In the independent test cohort, the CA19-9-based RF model exhibited substantially lower discriminatory performance than all four CNN models, as reflected by the ROC curves (Figure 3). Notably, the CA19-9-based RF model demonstrated a marked discrepancy between training (AUC, 1.000) and validation performance (AUC, 0.517), indicating substantial overfitting likely attributable to the limited predictive signal in CA19-9-derived features. This comparator should therefore be interpreted with caution and is intended to serve as a conventional clinical benchmark rather than an optimized competing classifier. Collectively, these findings demonstrate the superior discriminative ability of EUS-based deep learning models relative to conventional CA19-9-based approaches for predicting chemotherapy response.

Figure 3
Figure 3 Comparison of receiver operating characteristic performance between the deep learning models and the carbohydrate antigen 19-9-based model. Receiver operating characteristic curves comparing the predictive performance of four endoscopic ultrasound-based convolutional neural network models and the carbohydrate antigen 19-9-based model (random forest) in the independent test cohort are shown. The corresponding area under the receiver operating characteristic curve (AUC) values for each model are indicated in the figure legend. ResNet50 (AUC = 0.844) and ResNeXt50 (AUC = 0.848) achieved higher AUC values than the carbohydrate antigen 19-9-based random forest model (AUC = 0.673), indicating superior discriminatory performance for predicting chemotherapy response. AUC: Area under the receiver operating characteristic curve; ML: Machine learning; ROC: Receiver operating characteristic.
Table 6 Performance metrics of the carbohydrate antigen 19-9-based random forest model for chemotherapy response prediction in the training, internal validation, and independent test cohorts.
Group
AUC (95%CI)
ACC (%) (95%CI)
SEN (%) (95%CI)
SPE (%) (95%CI)
PPV (%)
NPV (%)
FNR (%)
F1
Train1.000100.00100.00100.00100.00100.0001.000
Val0.51776.979.0995.8037.5079.1790.910.146
Test0.673 (0.445–0.877)78.95 (65.79–92.11)25.00 (0.00–60.00)93.33 (83.33–100.00)50.0082.3575.000.333
Risk stratification and prognostic value

Risk stratification based on ResNeXt50-derived predicted probabilities effectively separated patients into high- and low-risk groups, with 95 patients in each group based on the median risk score (cutoff = 0.3956) of the full 190-patient cohort. The equal group sizes are a direct result of using the median as the cutoff threshold and do not correspond to the RECIST-based non-PD/PD classification presented in Table 1. Kaplan–Meier survival analysis demonstrated significantly shorter OS among patients in the high-risk group compared with those in the low-risk group (log-rank P = 0.0045; Figure 4). The median OS was 13.97 months (95%CI: 10.43–19.87) in the high-risk group and 20.87 months (95%CI: 14.50–25.10) in the low-risk group. Consistently, Cox proportional hazards regression analysis showed that high-risk patients had a significantly increased risk of death (HR = 1.77; 95%CI: 1.19–2.65; P = 0.0051).

Figure 4
Figure 4 Kaplan–Meier overall survival curves stratified by ResNeXt50-derived risk groups. Kaplan–Meier curves show overall survival for patients classified into low-risk (n = 95) and high-risk (n = 95) groups based on ResNeXt50-derived predicted probabilities using the median patient-level risk score of the full cohort as the cutoff. The equal group sizes reflect the use of the cohort median as the predefined cutoff and do not correspond to the RECIST-based non-progressive disease/progressive disease classification presented in Table 1. Overall survival differed significantly between the two groups (log-rank P = 0.0045). Cox proportional hazards regression analysis demonstrated a significantly higher risk of death in the high-risk group (hazard ratio = 1.77; 95% confidence interval: 1.19-2.65; P = 0.0051). CI: Confidence interval; HR: Hazard ratio; KM: Kaplan–Meier.

To further characterize treatment response patterns according to RECIST-based chemotherapy response classification, a waterfall plot illustrating percentage changes in target lesion size according to RECIST version 1.1 criteria is presented in Figure 5. Patients in the PD group more frequently experienced PD, whereas those in the non-PD group more commonly achieved disease control, including stable disease and partial response. These response patterns are consistent with the RECIST-based response classifications used as training labels in the model development.

Figure 5
Figure 5 Waterfall plot of tumor response according to RECIST version 1.1 stratified by chemotherapy response classification. The waterfall plot illustrates the percentage change in target lesion size from baseline for individual patients, ordered from the greatest increase to the greatest decrease in tumor size. Bars are color-coded according to the RECIST-based chemotherapy response classification (blue, non-progressive disease [non-PD; disease control]; orange, progressive disease [PD]). Dashed horizontal lines indicate the RECIST version 1.1 thresholds for partial response (−30%) and PD (+20%). The distribution of tumor size changes differed between the response groups, with a higher proportion of progressive disease observed in the PD group and a higher proportion of disease control (stable disease and partial response) observed in the non-PD group. PD: Progressive disease; PR: Partial response; RECIST: Response Evaluation Criteria in Solid Tumors.

For comparison, Kaplan–Meier analysis stratified by baseline CA19-9 levels using a median cutoff did not demonstrate a statistically significant difference in OS between the high and low CA19-9 groups (log-rank P = 0.128; Figure 6). Similarly, Cox proportional hazards analysis showed no significant association between baseline CA19-9 levels and OS (HR = 1.35; 95%CI: 0.91-2.00; P = 0.130). In contrast, model-predicted prognostic stratification showed good concordance with observed survival outcomes, further supporting the clinical relevance of the ResNeXt50-derived predictions.

Figure 6
Figure 6 Kaplan–Meier overall survival curves stratified by baseline carbohydrate antigen 19-9 levels. Kaplan–Meier overall survival curves are shown for patients stratified according to the median baseline serum carbohydrate antigen 19-9 (CA19-9) level. Overall survival did not differ significantly between the high and low CA19-9 groups (log-rank P = 0.128). Cox proportional hazards regression analysis showed no significant association between baseline CA19-9 levels and overall survival (hazard ratio = 1.35; 95% confidence interval: 0.91-2.00; P = 0.130). CA19-9: Carbohydrate antigen 19-9; CI: Confidence interval; HR: Hazard ratio; OS: Overall survival.

Subgroup analyses further demonstrated that the ResNeXt50-derived risk stratification was generally preserved across clinically relevant subgroups, including baseline CA19-9 levels, clinical stage IV vs non-IV disease, T stage, and N stage (Supplementary Figure 1). Across most subgroups, patients classified as high risk exhibited poorer OS compared with those classified as low risk. Although statistical significance was not reached in all subgroup analyses, the direction of survival separation remained consistent, supporting the robustness of the model-derived risk stratification. It should be emphasized that, because this prognostic analysis was performed using the full 190-patient study population rather than on a separately reserved prognostic validation set, these survival findings should be regarded as exploratory and hypothesis-generating rather than fully validated.

SEN and robustness analyses

SEN analyses demonstrated consistent performance of the ResNeXt50 model in the independent test cohort across different training random seeds. When the model was retrained five times with independent random initializations, the resulting ROC curves showed substantial overlap, and the corresponding AUC values exhibited minimal variability, supporting the stability and reproducibility of the model (Supplementary Figure 2).

To further assess robustness to image acquisition and preprocessing variability, a series of controlled perturbation experiments were conducted, including Gaussian noise addition, Gaussian blur, contrast adjustments, small-angle rotations, mild center cropping, and downsampling-upsampling. Across most perturbation conditions, discriminative performance remained largely preserved, with ROC curves and AUC values comparable to those obtained under unperturbed conditions. A modest decline in performance was observed under pronounced contrast degradation, whereas other perturbations had minimal impact. Overall, these results indicate that the ResNeXt50 model is robust to common stochastic variations, image quality fluctuations, and minor spatial distortions encountered in routine clinical practice (Supplementary Figure 3). Representative contrast-enhanced CT images illustrating the concordance between artificial intelligence-derived risk scores and observed chemotherapy responses are provided in Supplementary Figure 4.

DISCUSSION

In this study, we demonstrated that EUS-based deep learning models can effectively predict chemotherapy response and provide prognostic stratification in patients with PDAC. Among the evaluated architectures, the ResNeXt50 model showed the most consistent and robust performance across independent testing, survival stratification, and SEN analyses. Model-derived probability scores enabled clinically meaningful risk stratification, identifying patients with significantly poorer OS, and outperformed the conventional serum biomarker CA19-9 in both predictive discrimination and prognostic assessment. Collectively, these findings support the potential clinical utility of EUS-based deep learning as a noninvasive approach to guide individualized treatment planning.

Our findings are consistent with a growing body of literature demonstrating that deep learning models can extract high-dimensional imaging features that are imperceptible to the human eye yet clinically relevant for diagnosis, risk stratification, and prognosis in oncology[12,13]. Compared with cross-sectional imaging modalities such as CT and MRI, EUS provides substantially higher spatial resolution and more detailed visualization of tumor microarchitecture, enabling the capture of subtle intratumoral heterogeneity that is closely associated with treatment response. As EUS is routinely used for diagnostic confirmation and tissue acquisition in patients with advanced PDAC, integrating artificial intelligence-based analysis into standard EUS workflows may impose minimal additional procedural burden while providing clinically actionable insights. Previous studies have demonstrated the superior SEN of EUS in detecting small pancreatic lesions and its ability to delineate internal tumor structures that are often inadequately characterized by CT or MRI[14,15]. Consistent with recent radiomics and deep learning studies, EUS-based approaches appear to yield more discriminatory and robust imaging features compared with cross-sectional modalities. Furthermore, tumor size alone did not significantly discriminate chemotherapy response in our cohort (non-PD: 47.59 ± 18.93 mm vs PD: 42.05 ± 17.40 mm; P = 0.081), further supporting the added value of EUS-based deep learning features over conventional morphological parameters.

The favorable performance of the ResNeXt50 architecture may be partly explained by its aggregated residual learning design, which facilitates efficient hierarchical feature extraction and improved generalization, particularly when trained on relatively limited datasets[16,17]. Similar advantages of ResNeXt-based models have been reported across multiple medical imaging applications, including tumor classification, treatment response assessment, and survival analysis[18,19]. The stability of model performance across internal validation, independent testing, and multiple robustness analyses further supports its reproducibility and potential for clinical translation, consistent with recent studies demonstrating the reliability of deep learning approaches in heterogeneous clinical settings[20]. Importantly, the use of probabilistic outputs enabled risk stratification associated with significant differences in OS, underscoring the combined predictive and prognostic relevance of imaging-based deep learning biomarkers[21,22].

From a methodological perspective, this work represents an early effort to apply deep learning to EUS imaging for predicting chemotherapy response in pancreatic cancer, leveraging the modality’s superior spatial resolution and tumor-specific visualization compared with conventional cross-sectional imaging modalities[23,24]. The integration of model-derived risk stratification with survival analyses provided additional prognostic context, reinforcing the potential role of imaging-based deep learning in precision oncology, as supported by prior studies linking imaging biomarkers with survival outcomes across multiple malignancies[25-28]. Moreover, the proposed workflow—from standardized image acquisition and expert-guided region-of-interest annotation to model training and validation—was designed in accordance with the principles of transparent, reproducible, and scalable artificial intelligence pipelines in medical imaging[29], providing a foundation for future prospective and multicenter validation and for potential incorporation into clinical decision-support systems[30].

Although the CA19-9-based RF comparator exhibited the most pronounced overfitting (training AUC, 1.000; validation AUC, 0.517), the CNN performance tables (Tables 2, 3, 4, and 5) also reveal architecture-dependent train-to-test performance gaps that warrant explicit acknowledgement. The two VGG-family architectures showed the largest declines, with VGG19 declining from a training AUC of 0.963 to 0.758 in internal validation and 0.614 in the independent test cohort, and VGG19-BN declining from 1.000 to 0.864 and 0.775. The corresponding deeper architectures showed substantially smaller train-to-test attenuation (ResNet50: 0.962 → 0.871 → 0.844; ResNeXt50, which was similarly preserved across cohorts), although a degree of optimistic bias on the training set was nonetheless evident across all CNN models. This pattern is consistent with the well-recognized tendency of high-capacity convolutional networks to memorize training-cohort idiosyncrasies when sample sizes are limited, and with previous reports that VGG-style architectures, in particular, are more prone to overfitting on small medical imaging datasets than residual-connection architectures. Several aspects of our analytical strategy were designed to mitigate this risk, including patient-level five-fold cross-validation, on-the-fly data augmentation, an early-stopping criterion based on validation AUC, hyperparameter tuning restricted to the development cohort, and complete withholding of the temporally and geographically separated independent test cohort from all training and optimization steps. These measures likely contributed to the relative robustness of the ResNet50 and ResNeXt50 models.

Nonetheless, the residual train-to-test attenuation observed even for the best-performing models indicates that the absolute performance estimates reported in the development cohort should be interpreted as upper bounds rather than as expected real-world performance, and that the independent test cohort metrics provide the more clinically relevant estimate. Accordingly, we advanced ResNeXt50 as the primary model based on its independent test performance rather than its training performance, and external multicenter validation will be essential to confirm the generalizability of these findings.

Several limitations of this study should be acknowledged. First, although the model-derived risk stratification was significantly associated with OS and was preserved across subgroup analyses, the prognostic analysis was conducted using the full 190-patient study population rather than a separately reserved prognostic validation cohort. The cohort-median cutoff was specified a priori and was not optimized using survival outcomes, which mitigates—but does not eliminate—the risk of optimistic bias inherent in evaluating prognostic performance in the same population from which the risk score was generated. The prognostic findings reported here should therefore be considered exploratory and hypothesis-generating rather than fully validated, and dedicated prospective evaluation in independent prognostic validation cohorts will be required before these findings can be considered definitively established.

This study predominantly included patients with unresectable PDAC who underwent chemotherapy, which may limit the generalizability of the findings to earlier-stage disease. In addition, although treatment allocation between mFOLFIRINOX and gemcitabine plus nab-paclitaxel was not randomized and the timing of response assessment was not strictly uniform across the two regimens, regimen-adjusted analyses confirmed that the prognostic value of the model-derived risk stratification was independent of chemotherapy type. Nevertheless, treatment regimens were not stratified according to platinum-containing vs non–platinum-containing chemotherapy, and such stratification may further refine model performance[31,32].

Furthermore, this study was conducted at a single center, and although the independent test cohort was temporally and geographically separated from the development cohort, this design does not fully substitute for external validation using data from independent institutions. Institution-specific factors such as EUS equipment type, operator technique, and patient demographics may influence model generalizability; future prospective multicenter studies are warranted to validate these findings across diverse clinical settings. Manual ROI annotation, although associated with excellent interobserver agreement, remains time-consuming and subject to observer variability. Automated or semiautomated segmentation methods may improve efficiency and reproducibility in future studies[33,34]. Furthermore, the current models were developed using imaging data alone, and integration of multimodal clinical, laboratory, and genomic information may enhance predictive ACC and biological interpretability[35,36]. Response assessment was based on the morphological RECIST version 1.1 criteria, which may underestimate treatment response in highly desmoplastic tumors. Functional imaging approaches such as diffusion-weighted MRI or positron emission tomography-CT may complement size-based assessment in future studies. Finally, response prediction was based solely on baseline EUS images; incorporating longitudinal imaging during treatment may enable earlier identification of nonresponders and adaptive therapeutic strategies[37,38].

Future research should focus on prospective, multicenter validation, the incorporation of explainable artificial intelligence techniques to enhance transparency and clinician trust, the development of multimodal predictive frameworks, and the integration of these models into real-world clinical workflows to support individualized treatment decision-making.

CONCLUSION

In conclusion, this study developed and validated CNN-based models using EUS imaging to predict chemotherapy response in PDAC. Among the evaluated architectures, the ResNeXt50 model demonstrated the most consistent and robust predictive performance and showed superior discriminatory ability compared with that of the conventional serum biomarker CA19-9. Risk stratification based on model-derived probability scores provided clinically relevant prognostic information for OS, supporting the potential value of noninvasive, imaging-based deep learning approaches for individualized treatment planning. Future prospective, multicenter studies and the integration of explainable artificial intelligence frameworks are warranted to further validate these findings and facilitate their clinical translation.

References
1.  Siegel RL, Giaquinto AN, Jemal A. Cancer statistics, 2024. CA Cancer J Clin. 2024;74:12-49.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 7368]  [Cited by in RCA: 7085]  [Article Influence: 3542.5]  [Reference Citation Analysis (7)]
2.  Grossberg AJ, Chu LC, Deig CR, Fishman EK, Hwang WL, Maitra A, Marks DL, Mehta A, Nabavizadeh N, Simeone DM, Weekes CD, Thomas CR Jr. Multidisciplinary standards of care and recent progress in pancreatic ductal adenocarcinoma. CA Cancer J Clin. 2020;70:375-403.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 340]  [Cited by in RCA: 368]  [Article Influence: 61.3]  [Reference Citation Analysis (10)]
3.  Mahajan UM, Oehrle B, Goni E, Strobel O, Kaiser J, Grützmann R, Werner J, Friess H, Gress TM, Seufferlein TW, Uhl W, Will U, Neoptolemos JP, Wittel UA, Vornhülz M, Sirtl S, Beyer G, Regel I, Boeck S, Heinemann V, Frost F, Steveling A, Völzke H, Petersmann A, Nauck M, Weber E, Kamlage B, Lerch MM, Mayerle J; METAPAC trial investigators. Validation of two plasma multimetabolite signatures for patients at risk of or with suspected pancreatic ductal adenocarcinoma (METAPAC): a prospective, multicentre, investigator-masked, enrichment design, phase 4 diagnostic study. Lancet Gastroenterol Hepatol. 2025;10:634-647.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 7]  [Cited by in RCA: 13]  [Article Influence: 13.0]  [Reference Citation Analysis (0)]
4.  Fahrmann JF, Schmidt CM, Mao X, Irajizad E, Loftus M, Zhang J, Patel N, Vykoukal J, Dennison JB, Long JP, Do KA, Zhang J, Chabot JA, Kluger MD, Kastrinos F, Brais L, Babic A, Jajoo K, Lee LS, Clancy TE, Ng K, Bullock A, Genkinger J, Yip-Schneider MT, Maitra A, Wolpin BM, Hanash S. Lead-Time Trajectory of CA19-9 as an Anchor Marker for Pancreatic Cancer Early Detection. Gastroenterology. 2021;160:1373-1383.e6.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 175]  [Cited by in RCA: 173]  [Article Influence: 34.6]  [Reference Citation Analysis (11)]
5.  Lambin P, Leijenaar RTH, Deist TM, Peerlings J, de Jong EEC, van Timmeren J, Sanduleanu S, Larue RTHM, Even AJG, Jochems A, van Wijk Y, Woodruff H, van Soest J, Lustberg T, Roelofs E, van Elmpt W, Dekker A, Mottaghy FM, Wildberger JE, Walsh S. Radiomics: the bridge between medical imaging and personalized medicine. Nat Rev Clin Oncol. 2017;14:749-762.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 4684]  [Cited by in RCA: 4287]  [Article Influence: 476.3]  [Reference Citation Analysis (13)]
6.  Aggarwal R, Sounderajah V, Martin G, Ting DSW, Karthikesalingam A, King D, Ashrafian H, Darzi A. Diagnostic accuracy of deep learning in medical imaging: a systematic review and meta-analysis. NPJ Digit Med. 2021;4:65.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 748]  [Cited by in RCA: 506]  [Article Influence: 101.2]  [Reference Citation Analysis (4)]
7.  Saraiva MM, González-Haba M, Widmer J, Mendes F, Gonda T, Agudo B, Ribeiro T, Costa A, Fazel Y, Lera ME, Horneaux de Moura E, Ferreira de Carvalho M, Bestetti A, Afonso J, Martins M, Almeida MJ, Vilas-Boas F, Moutinho-Ribeiro P, Lopes S, Fernandes J, Ferreira J, Macedo G. Deep Learning and Automatic Differentiation of Pancreatic Lesions in Endoscopic Ultrasound: A Transatlantic Study. Clin Transl Gastroenterol. 2024;15:e00771.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 8]  [Cited by in RCA: 11]  [Article Influence: 5.5]  [Reference Citation Analysis (0)]
8.  Sharma P, Hassan C. Artificial Intelligence and Deep Learning for Upper Gastrointestinal Neoplasia. Gastroenterology. 2022;162:1056-1066.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 74]  [Cited by in RCA: 67]  [Article Influence: 16.8]  [Reference Citation Analysis (5)]
9.  van der Sommen F, de Groof J, Struyvenberg M, van der Putten J, Boers T, Fockens K, Schoon EJ, Curvers W, de With P, Mori Y, Byrne M, Bergman JJGHM. Machine learning in GI endoscopy: practical guidance in how to interpret a novel field. Gut. 2020;69:2035-2045.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 111]  [Cited by in RCA: 106]  [Article Influence: 17.7]  [Reference Citation Analysis (5)]
10.  European Study Group on Cystic Tumours of the Pancreas. European evidence-based guidelines on pancreatic cystic neoplasms. Gut. 2018;67:789-804.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 1177]  [Cited by in RCA: 1056]  [Article Influence: 132.0]  [Reference Citation Analysis (14)]
11.  Huang Y, Zhu T, Zhang X, Li W, Zheng X, Cheng M, Ji F, Zhang L, Yang C, Wu Z, Ye G, Lin Y, Wang K. Longitudinal MRI-based fusion novel model predicts pathological complete response in breast cancer treated with neoadjuvant chemotherapy: a multicenter, retrospective study. EClinicalMedicine. 2023;58:101899.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 136]  [Reference Citation Analysis (2)]
12.  Pușcaș IM, Gâta A, Roman A, Albu S, Gâta VA, Irimie A. Integrating Radiomics and Deep-Learning for Prognostic Evaluation in Nasopharyngeal Carcinoma. Medicina (Kaunas). 2025;61:1310.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 2]  [Cited by in RCA: 3]  [Article Influence: 3.0]  [Reference Citation Analysis (0)]
13.  Marra A, Morganti S, Pareja F, Campanella G, Bibeau F, Fuchs T, Loda M, Parwani A, Scarpa A, Reis-Filho JS, Curigliano G, Marchiò C, Kather JN. Artificial intelligence entering the pathology arena in oncology: current applications and future perspectives. Ann Oncol. 2025;36:712-725.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 1]  [Cited by in RCA: 60]  [Article Influence: 60.0]  [Reference Citation Analysis (0)]
14.  Kitano M, Yoshida T, Itonaga M, Tamura T, Hatamaru K, Yamashita Y. Impact of endoscopic ultrasonography on diagnosis of pancreatic cancer. J Gastroenterol. 2019;54:19-32.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 265]  [Cited by in RCA: 236]  [Article Influence: 33.7]  [Reference Citation Analysis (13)]
15.  Lu X, Zhang S, Ma C, Peng C, Lv Y, Zou X. The diagnostic value of EUS in pancreatic cystic neoplasms compared with CT and MRI. Endosc Ultrasound. 2015;4:324-329.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 43]  [Cited by in RCA: 45]  [Article Influence: 4.1]  [Reference Citation Analysis (0)]
16.  Chen D, Hu F, Nian G, Yang T. Deep Residual Learning for Nonlinear Regression. Entropy (Basel). 2020;22:193.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 24]  [Cited by in RCA: 29]  [Article Influence: 4.8]  [Reference Citation Analysis (0)]
17.  Zech JR, Badgeley MA, Liu M, Costa AB, Titano JJ, Oermann EK. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study. PLoS Med. 2018;15:e1002683.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 1154]  [Cited by in RCA: 855]  [Article Influence: 106.9]  [Reference Citation Analysis (6)]
18.  Chiu SH, Li HC, Chang WC, Wu CC, Lin HH, Lo CH, Chang PY. Improving the prediction of patient survival with the aid of residual convolutional neural network (ResNet) in colorectal cancer with unresectable liver metastases treated with bevacizumab-based chemotherapy. Cancer Imaging. 2024;24:165.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 6]  [Reference Citation Analysis (0)]
19.  Ali RR, Yaacob NM, Alqaryouti MH, Sadeq AE, Doheir M, Iqtait M, Rachmawanto EH, Sari CA, Yaacob SS. Learning Architecture for Brain Tumor Classification Based on Deep Convolutional Neural Network: Classic and ResNet50. Diagnostics (Basel). 2025;15:624.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 5]  [Reference Citation Analysis (0)]
20.  Kong C, Yan D, Liu K, Yin Y, Ma C. Multiple deep learning models based on MRI images in discriminating glioblastoma from solitary brain metastases: a multicentre study. BMC Med Imaging. 2025;25:171.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 6]  [Reference Citation Analysis (0)]
21.  Deng K, Wang L, Liu Y, Li X, Hou Q, Cao M, Ng NN, Wang H, Chen H, Yeom KW, Zhao M, Wu N, Gao P, Shi J, Liu Z, Li W, Tian J, Song J. A deep learning-based system for survival benefit prediction of tyrosine kinase inhibitors and immune checkpoint inhibitors in stage IV non-small cell lung cancer patients: A multicenter, prognostic study. EClinicalMedicine. 2022;51:101541.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 3]  [Cited by in RCA: 34]  [Article Influence: 8.5]  [Reference Citation Analysis (4)]
22.  Yang Y, Yang J, Shen L, Chen J, Xia L, Ni B, Ge L, Wang Y, Lu S. A multi-omics-based serial deep learning approach to predict clinical outcomes of single-agent anti-PD-1/PD-L1 immunotherapy in advanced stage non-small-cell lung cancer. Am J Transl Res. 2021;13:743-756.  [PubMed]  [DOI]
23.  Yousaf MN, Chaudhary FS, Ehsan A, Suarez AL, Muniraj T, Jamidar P, Aslanian HR, Farrell JJ. Endoscopic ultrasound (EUS) and the management of pancreatic cancer. BMJ Open Gastroenterol. 2020;7:e000408.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 9]  [Cited by in RCA: 59]  [Article Influence: 11.8]  [Reference Citation Analysis (2)]
24.  Ishii Y, Serikawa M, Tsuboi T, Kawamura R, Tsushima K, Nakamura S, Hirano T, Fukiage A, Mori T, Ikemoto J, Kiyoshita Y, Saeki S, Tamura Y, Miyamoto S, Chayama K. Role of Endoscopic Ultrasonography and Endoscopic Retrograde Cholangiopancreatography in the Diagnosis of Pancreatic Cancer. Diagnostics (Basel). 2021;11:238.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 12]  [Cited by in RCA: 11]  [Article Influence: 2.2]  [Reference Citation Analysis (0)]
25.  M MM, T R M, V VK, Guluwadi S. Enhancing brain tumor detection in MRI images through explainable AI using Grad-CAM with Resnet 50. BMC Med Imaging. 2024;24:107.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 156]  [Cited by in RCA: 34]  [Article Influence: 17.0]  [Reference Citation Analysis (0)]
26.  Klauschen F, Dippel J, Keyl P, Jurmeister P, Bockmayr M, Mock A, Buchstab O, Alber M, Ruff L, Montavon G, Müller KR. Toward Explainable Artificial Intelligence for Precision Pathology. Annu Rev Pathol. 2024;19:541-570.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 84]  [Cited by in RCA: 76]  [Article Influence: 38.0]  [Reference Citation Analysis (0)]
27.  Yao J, Cao K, Hou Y, Zhou J, Xia Y, Nogues I, Song Q, Jiang H, Ye X, Lu J, Jin G, Lu H, Xie C, Zhang R, Xiao J, Liu Z, Gao F, Qi Y, Li X, Zheng Y, Lu L, Shi Y, Zhang L. Deep Learning for Fully Automated Prediction of Overall Survival in Patients Undergoing Resection for Pancreatic Cancer: A Retrospective Multicenter Study. Ann Surg. 2023;278:e68-e79.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 34]  [Cited by in RCA: 24]  [Article Influence: 8.0]  [Reference Citation Analysis (1)]
28.  Dudas D, Saghand PG, Dilling TJ, Perez BA, Rosenberg SA, El Naqa I. Deep Learning-Guided Dosimetry for Mitigating Local Failure of Patients With Non-Small Cell Lung Cancer Receiving Stereotactic Body Radiation Therapy. Int J Radiat Oncol Biol Phys. 2024;119:990-1000.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 9]  [Reference Citation Analysis (0)]
29.  Bontempi D, Nuernberg L, Pai S, Krishnaswamy D, Thiriveedhi V, Hosny A, Mak RH, Farahani K, Kikinis R, Fedorov A, Aerts HJWL. End-to-end reproducible AI pipelines in radiology using the cloud. Nat Commun. 2024;15:6931.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 15]  [Cited by in RCA: 8]  [Article Influence: 4.0]  [Reference Citation Analysis (0)]
30.  Elhaddad M, Hamam S. AI-Driven Clinical Decision Support Systems: An Ongoing Pursuit of Potential. Cureus. 2024;16:e57728.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 120]  [Reference Citation Analysis (0)]
31.  Wichtmann BD, Albert S, Zhao W, Maurer A, Rödel C, Hofheinz RD, Hesser J, Zöllner FG, Attenberger UI. Are We There Yet? The Value of Deep Learning in a Multicenter Setting for Response Prediction of Locally Advanced Rectal Cancer to Neoadjuvant Chemoradiotherapy. Diagnostics (Basel). 2022;12:1601.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 4]  [Cited by in RCA: 11]  [Article Influence: 2.8]  [Reference Citation Analysis (0)]
32.  Wu D, Smith D, VanBerlo B, Roshankar A, Lee H, Li B, Ali F, Rahman M, Basmaji J, Tschirhart J, Ford A, VanBerlo B, Durvasula A, Vannelli C, Dave C, Deglint J, Ho J, Chaudhary R, Clausdorff H, Prager R, Millington S, Shah S, Buchanan B, Arntfield R. Improving the Generalizability and Performance of an Ultrasound Deep Learning Model Using Limited Multicenter Data for Lung Sliding Artifact Identification. Diagnostics (Basel). 2024;14:1081.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 9]  [Reference Citation Analysis (0)]
33.  Ma J, He Y, Li F, Han L, You C, Wang B. Segment anything in medical images. Nat Commun. 2024;15:654.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 2885]  [Cited by in RCA: 946]  [Article Influence: 473.0]  [Reference Citation Analysis (1)]
34.  Oh S, Kim YJ, Park YT, Kim KG. Automatic Pancreatic Cyst Lesion Segmentation on EUS Images Using a Deep-Learning Approach. Sensors (Basel). 2021;22:245.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 7]  [Cited by in RCA: 22]  [Article Influence: 4.4]  [Reference Citation Analysis (0)]
35.  Yang H, Yang M, Chen J, Yao G, Zou Q, Jia L. Multimodal deep learning approaches for precision oncology: a comprehensive review. Brief Bioinform. 2024;26:bbae699.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 55]  [Reference Citation Analysis (1)]
36.  Steyaert S, Qiu YL, Zheng Y, Mukherjee P, Vogel H, Gevaert O. Multimodal deep learning to predict prognosis in adult and pediatric brain tumors. Commun Med (Lond). 2023;3:44.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 86]  [Cited by in RCA: 59]  [Article Influence: 19.7]  [Reference Citation Analysis (0)]
37.  Xu Y, Hosny A, Zeleznik R, Parmar C, Coroller T, Franco I, Mak RH, Aerts HJWL. Deep Learning Predicts Lung Cancer Treatment Response from Serial Medical Imaging. Clin Cancer Res. 2019;25:3266-3275.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 196]  [Cited by in RCA: 371]  [Article Influence: 53.0]  [Reference Citation Analysis (4)]
38.  Jin C, Yu H, Ke J, Ding P, Yi Y, Jiang X, Duan X, Tang J, Chang DT, Wu X, Gao F, Li R. Predicting treatment response from longitudinal images using multi-task deep learning. Nat Commun. 2021;12:1851.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 22]  [Cited by in RCA: 157]  [Article Influence: 31.4]  [Reference Citation Analysis (1)]
Footnotes

Peer review: Externally peer reviewed.

Peer-review model: Single blind

Specialty type: Gastroenterology and hepatology

Country of origin: China

Peer-review report’s classification

Scientific quality: Grade B, Grade B, Grade C, Grade C, Grade C

Novelty: Grade B, Grade B, Grade B, Grade B, Grade B

Creativity or innovation: Grade B, Grade B, Grade B, Grade C, Grade C

Scientific significance: Grade B, Grade B, Grade B, Grade C, Grade C

P-Reviewer: Cao L, Chief, Chief Physician, PhD, Professor, China; Goyal O, DM, MD, Professor, India; Rodrigues de Bastos DR, Researcher, Paraguay S-Editor: Bai Y L-Editor: Filipodia P-Editor: Wang WB

Write to the Help Desk