BPG is committed to discovery and dissemination of knowledge
Retrospective Study Open Access
Copyright: ©Author(s) 2026. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution-NonCommercial (CC BY-NC 4.0) license. No commercial re-use. See permissions. Published by Baishideng Publishing Group Inc.
World J Gastrointest Oncol. Sep 15, 2026; 18(9): 119889
Published online Sep 15, 2026. doi: 10.4251/wjgo.119889
Multicenter deep learning model for pancreatic cancer detection using endoscopic ultrasound
Xin-Ying Yu, Qiang He, Department of Gastroenterology, Beijing Tiantan Hospital, Capital Medical University, Beijing 100071, China
Jun-Qiang Ye, Electronic Information and Communication, Huazhong University of Science and Technology, Wuhan 430074, Hubei Province, China
Zhen He, Department of Gastroenterology, Beijing Friendship Hospital, Capital Medical University, Beijing 100050, China
ORCID number: Qiang He (0000-0001-5419-3360).
Co-corresponding authors: Zhen He and Qiang He.
Author contributions: Yu XY and He Z designed the research study and revised the manuscript; Ye JQ developed the machine learning algorithms and wrote the initial manuscript draft; He Q collected the data, performed the research, and contributed to the drafting and revision of the manuscript; all authors have read and approve the final manuscript. Our study was a multicenter collaborative effort involving two major centers. To appropriately reflect the contributions and responsibilities of each participating site, we designated one corresponding author per center. Specifically, He Q from the Department of Gastroenterology, Beijing Tiantan Hospital, Capital Medical University, serves as the corresponding author for the first center. The other co-corresponding author (He Z) is from the Department of Gastroenterology, Beijing Friendship Hospital, Capital Medical University, representing the second center. This arrangement ensures that each center has a dedicated point of contact for scientific inquiries, data verification, and administrative matters. It also acknowledges the equal intellectual and logistical input from both sites, which is a common and transparent practice in multicenter studies. We believe that co-corresponding authorship accurately represents the collaborative nature of our work and facilitates efficient communication with the research community.
AI contribution statement: The initial draft of the manuscript (including all scientific content such as abstract, introduction, materials and methods, results, discussion, and conclusion) was entirely written by the author without the use of any artificial intelligence text generation tools (such as ChatGPT). The core scientific content, research design, data interpretation, or conclusions are not generated by artificial intelligence. During the revision and submission preparation process, we only use AI assisted language polishing tools to improve grammar, spelling, and readability, similar to using professional editing services. This is a common practice to ensure language clarity. The manuscript did not use artificial intelligence tools to generate any numbers, images, or other visual elements. There are no artificial intelligence tools involved in research design, data analysis, or result interpretation.
Institutional review board statement: The study was reviewed and approved by the IRB of Beijing Tiantan Hospital, Capital Medical University (Approval No. KY 2020-089-02).
Informed consent statement: All study participants provided informed written consent prior to study enrollment.
Conflict-of-interest statement: The authors declare that they have no conflict of interest.
Data sharing statement: No additional data are available.
Corresponding author: Qiang He, Department of Gastroenterology, Beijing Tiantan Hospital, Capital Medical University, No. 119 South Fourth Ring Road West, Fengtai District, Beijing 100071, China. 229476289@qq.com
Received: February 10, 2026
Revised: March 29, 2026
Accepted: May 12, 2026
Published online: September 15, 2026
Processing time: 212 Days and 3.1 Hours

Abstract
BACKGROUND

The early diagnosis of pancreatic cancer via endoscopic ultrasound (EUS) is challenging. Existing deep learning models mostly rely on single-center data and often experience significant generalization decline across different centers due to variations in imaging equipment, settings, and populations.

AIM

To improve pancreatic cancer detection by developing a robust deep learning model capable of handling data heterogeneity across multiple EUS centers.

METHODS

We constructed a multicenter dataset (383 patients, 2362 images) from two hospitals. After performing frequency-domain analysis of cross-center heterogeneity to quantify imaging variations and distribution shifts between the centers, we developed a multicenter data-oriented deep learning classification network named MCEUS-C2Net, integrating local and global features with channel attention. The model was trained and validated on one center's data (182 training, 60 validation, 60 test cases) and independently evaluated on the second (81 cases). Patient-level partitioning ensured independent cross-center evaluation.

RESULTS

MCEUS-C2Net achieved accuracy, sensitivity, and F1 score of 94.51%, 97.09%, and 95.24% in the internal test, outperforming ResNet-50, Swin transformer, and MedViTV2. In external validation, it maintained robust performance with 90.59% accuracy, 90.50% sensitivity, and an area under the curve of 0.9704, showing the smallest performance decline among all models. Confusion matrix analysis confirmed fewer false positives and negatives. Frequency-domain analysis identified significant cross-center distribution differences. Attention heatmaps demonstrated that MCEUS-C2Net stably and autonomously focused on pancreatic lesion sites with high clinical consistency, whereas comparison models exhibited dispersed attention. These results confirm the model’s superior cross-center robustness and interpretability for pancreatic cancer identification.

CONCLUSION

Our approach achieves high-performance pancreatic cancer classification across multicenter data. It effectively addresses imaging heterogeneity, providing a robust, clinically viable solution for standardized EUS diagnosis in real-world settings.

Key Words: Pancreatic cancer lesions; Endoscopic ultrasound; Deep learning; Multicenter study; Computer-aided diagnosis; Attention mechanism

Core Tip: This study introduces the first multicenter endoscopic ultrasound dataset for pancreatic lesions. We also propose a novel deep learning model, MCEUS-C2Net. It integrates local and global features with channel attention mechanisms. This design effectively overcomes cross-center imaging heterogeneity. The model accurately differentiates cancerous from noncancerous pancreatic lesions. During external validation, it demonstrated exceptional generalization. Ultimately, this research provides a valuable benchmark dataset and a robust deep learning network. This network holds the potential to enhance the accuracy of clinical decision-making in real-world settings.



INTRODUCTION

Pancreatic cancer is difficult to diagnose early and has a very poor prognosis for patients because of its insidious early symptoms, the limited sensitivity of conventional imaging tests such as enhanced computed tomography and magnetic resonance imaging to small lesions, and the insufficient specificity of serum markers[1]. In this context, endoscopic ultrasound (EUS) has become the most sensitive imaging method for detecting pancreatic tumors, especially early lesions with diameters of less than 2 cm. This high sensitivity is due to the insertion of a high-frequency ultrasound probe into the digestive tract to scan adjacent to the pancreas, which provides a much higher spatial resolution than conventional in vitro imaging[2,3]. Not only has it been recognized by studies at home and abroad as one of the most reliable and efficient tests for diagnosing small pancreatic lesions, but its guided fine-needle aspiration (EUS-FNA/B) technique is also the “gold standard” for differentiating between benign and malignant tumors[4,5]. To further improve diagnostic performance of EUS, enhanced techniques such as contrast-enhanced ultrasound (CH-EUS) and elastography (EUS-EG) have been applied in clinical practice; these approaches can effectively assess the blood flow and hardness characteristics of microlesions[6,7].

However, traditional methods based on manual feature extraction processes often rely on human experience in medical image classification tasks, so they have a limited ability to recognize patterns. Consequently, they have difficulty fully characterizing the complex morphological and textural differences between pancreatic cancer and non-cancerous lesions, which often leads to insufficient robustness in practical applications.

With the rapid development of deep learning techniques, end-to-end feature learning paradigms have shown significant advantages in a variety of visual tasks. Studies have shown that compared with traditional methods, deep learning methods can automatically learn discriminative features in the field of medical image classification and achieve better performance. The fusion analysis technique of utilizing artificial intelligence (AI) with EUS images has become an important trend[8,9], and deep learning models have shown comparable potential to that of experienced physicians for differentiating pancreatic lesions[10-13], providing a new path to overcome the reliance on individual experience in diagnosis tasks and achieve standardization. Although some recent studies have explored advanced modalities like AI-assisted hyperspectral imaging to enhance gastrointestinal tissue differentiation[14-16], standard EUS remains the mainstream diagnostic tool for pancreatic lesions, making it crucial to unlock the full potential of robust AI for EUS.

However, most existing studies have constructed models based on single-center datasets. Differences among the imaging equipment, acquisition protocols, and patient populations of different medical centers lead to significant heterogeneity in EUS images, commonly known as domain shift. Consequently, models trained on single-center data often show severe performance degradation in cross-center external validation scenarios. From a clinical perspective, this poor generalization is not merely a technical limitation but a major deployment barrier and a direct risk to patients; an unreliable model could lead to missed diagnoses of malignant tumors or trigger unnecessary invasive biopsies for benign lesions, ultimately misleading medical decisions. Furthermore, while collecting and annotating multicenter EUS data is costly and public datasets are scarce, a clear research gap remains: There is a continued need for multicenter datasets, network architectures tailored for multicenter classification, and external validation. To address these critical gaps, in this paper, a multicenter EUS pancreatic disease dataset is constructed from two medical centers, and a novel multicenter data-oriented deep learning classification network (MCEUS-C2Net) is proposed to robustly classify pancreatic cancer and noncancerous lesions.

MATERIALS AND METHODS
Patient information

We constructed a multicenter EUS pancreatic disease imaging dataset, named MEUS-PCBL, for the classification of pancreatic cancer and benign lesions. The dataset was collected from 559 patients aged 45 years to 82 years who underwent EUS-guided fine-needle aspiration/biopsy (EUS-FNA/B) for pancreatic lesions at the Beijing Tiantan Hospital and Beijing Friendship Hospital between January 2020 and June 2025. The tumor lesions were pathologically diagnosed using tissue specimens obtained from EUS-guided fine-needle aspiration/biopsy (EUS-FNA/B) or surgical resection. To ensure strict temporal correlation between the imaging and the reference standard, all EUS-FNA/B procedures were performed concurrently during the same EUS examination. For patients undergoing surgical treatment, the interval between the EUS examination and surgical pathology ranged from 7 days to 14 days, with all cases strictly within 30 days, thereby ensuring the morphological consistency of the lesions and preventing diagnostic discrepancies caused by disease progression over time. The final diagnosis was divided into two major categories: Cancerous [pancreatic ductal adenocarcinoma (PDAC) and pancreatic acinar cell carcinoma (PDCC)] and noncancerous [neuroendocrine neoplasms (NEN), autoimmune pancreatitis (AIP), and solid pseudopapillaryoma (SPN)]. Patients were included if they underwent a complete EUS examination covering the entire pancreas and had a definitive diagnosis confirmed by surgical pathology, EUS-FNA/B pathology, or long-term clinical follow-up of at least 6 months. Furthermore, the inclusion criteria required high-quality EUS images clearly displaying the lesions without significant motion artifacts or gastrointestinal gas obstruction, as well as complete and traceable clinical, pathological, and follow-up data. The exclusion criteria were cases with no video images before conducting EUS-FNA/B; no CH-EUS endoscopy; and poor image quality due to bubbles, blurriness, defocusing, or artifacts. All the images were relabeled, and the degree of conformance was checked by 10 clinically experienced ultrasound physicians. To ensure the absolute objectivity and accuracy of the ground truth labels, a strict double-blind design was implemented. The physicians responsible for extracting and labeling the EUS image frames were completely blinded to the patients' pathological diagnoses throughout the process, while the pathologists were strictly blinded to the EUS images and any AI model predictions during their evaluations, effectively eliminating confirmation bias from the source. A total of 176 patients (158 without video images and 18 with poor image quality) were excluded according to the exclusion criteria, and ultimately, 383 patients were included in this study. As shown in Figure 1, the Beijing Friendship Hospital development cohort included 302 patients who were randomly divided into training, validation, and test cohorts (182 in the training cohort, 60 in the validation cohort, and 60 in the test cohort). The Beijing Tiantan Hospital, as an external test cohort, included 81 patients. The final dataset included a total of 2362 EUS images, with 379 images derived from the Beijing Tiantan Hospital and 1983 images acquired from the Beijing Friendship Hospital. Static frames were manually extracted by experienced endoscopists, with a median of 6 images per patient, in order to adequately capture the characteristic features of the lesions from different angles and in various morphological presentations. The EUS images obtained from different centers varied in their resolutions and imaging conditions, reflecting the diversity of the utilized equipment and acquisition processes in real clinical settings.

Figure 1
Figure 1 Flowchart of the study. EUS: Endoscopic ultrasound; AI: Artificial intelligence.

To further improve data transparency and allow readers to comprehensively assess the datase’s representativeness, the baseline demographic and clinical characteristics of the 383 included patients are summarized in Table 1. The mean age of the total cohort was 66.7 ± 12.5 years, with comparable age distributions between the Beijing Friendship Hospital (67.2 ± 12.2 years) and the Beijing Tiantan Hospital (64.8 ± 13.1 years). The gender distribution was generally balanced, comprising 202 males (52.7%) and 181 females (47.3%). Regarding clinical characteristics, the dataset maintains a consistent class distribution across both centers. Overall, cancerous lesions accounted for 58.7% (n = 225) of the total cohort, predominantly consisting of (PDAC, n = 219) and a small subset of (PACC, n = 6). Noncancerous lesions made up the remaining 41.3% (n = 158), which included SPN, n = 72, AIP, n = 55, and NENs, n = 31. Importantly, the proportional distributions of these disease subtypes in the Beijing Tiantan Hospital external test cohort (n = 81) closely mirrored those in the Beijing Friendship Hospital development cohort (n = 302), ensuring the representativeness and structural balance of the multicenter data.

Table 1 Baseline demographic and clinical characteristics of the study population, n (%).
Characteristics
Total cohort (n = 383)
Beijing Friendship Hospital (n = 302)
Beijing Tiantan Hospital (n = 81)
Age (years), mean ± SD66.7 ± 12.567.2 ± 12.264.8 ± 13.1
Gender
    Male202 (52.7)162 (53.6)40 (49.4)
    Female181 (47.3)140 (46.4)41 (50.6)
Class distribution
Cancerous225 (58.7)181 (59.9)44 (54.3)
Noncancerous158 (41.3)121 (40.1)37 (45.7)
Lesion subtypes
Cancerous group
    Pancreatic ductal adenocarcinoma219 (57.2)176 (58.3)43 (53.1)
    Pancreatic acinar cell carcinoma6 (1.6)5 (1.7)1 (1.2)
Noncancerous group
    Solid pseudopapillaryoma72 (18.8)56 (18.5)16 (19.8)
    Autoimmune pancreatitis55 (14.4)41 (13.6)14 (17.3)
    Neuroendocrine neoplasms31 (8.1)24 (7.9)7 (8.6)

To our knowledge, MEUS-PCBL is currently the first publicly reported multicenter EUS dataset consisting of pancreatic disease images, providing clinically representative benchmark data for the study and validation of EUS-assisted diagnosis algorithms. Some of the MEUS-PCBL data samples are shown in Figure 2.

Figure 2
Figure 2 The MEUS-PCBL dataset. A: Benign lesion case from Beijing Tiantan Hospital; B: Benign lesion case from Beijing Friendship Hospital; C: Malignant tumor lesion case from Beijing Tiantan Hospital; D: Malignant tumor lesion case from Beijing Friendship Hospital.
Data collection and multicenter analysis of frequency-domain distribution differences

EUS was performed using an Olympus EU-ME2 ultrasound system equipped with an Olympus GF-UCT260 or GF-UCT240 curved linear EUS or a Pentax EG-38-J10UT curved linear EUS. During each patient’s EUS examination, video recordings of the EUS images were made from the insertion point until the removal of the endoscope.

To analyze the differences among the frequency-domain distributions of the EUS images acquired at different centers, this paper used a two-dimensional Fourier transform to map the spatial-domain images to the frequency domain to obtain their spectral energy distribution characteristics. The Fourier transform can effectively represent the structural and texture information of an image at different scales, where the low-frequency components mainly reflect the overall structure and intensity distribution, whereas the high-frequency components correspond to edge details, texture variations, and noise characteristics. On this basis, a single image was used as the statistical unit, and the ratio of high-frequency energy to low-frequency energy was calculated as the frequency-domain feature description to quantify the influences of different imaging centers and imaging conditions on the spectral characteristics of each image. A box plot was subsequently constructed to visualize the distribution of the high-frequency energy proportions determined from different centers and the cancer and noncancer categories to depict the concentration tendencies and dispersion degrees of the samples within different centers. The high/Low-frequency energy ratio distributions of the EUS images derived from different centers are shown in Figure 3, and the corresponding statistical results are summarized in Table 2. As shown in Table 2, significant distribution differences were observed in the proportions of high-frequency energy in the frequency domain among the different imaging centers. Specifically, the overall proportion of high-frequency energy in the images of Center A (Center A_0 and Center A_1) was significantly greater than that in Center B (Center B_0 and Center B_1), with mean values of 0.061 and 0.048, respectively. The corresponding values for Center B were only 0.021 and 0.024. Moreover, the data distribution at Center A was wider, with larger standard deviations and higher maxima (up to 0.56), indicating that the central image was more volatile in terms of its high-frequency details and noise components. In contrast, the proportion of high-frequency energy in Center B was generally lower, and the corresponding distribution was more concentrated, reflecting higher consistency and stability among the spectral characteristics exhibited under imaging conditions.

Figure 3
Figure 3 High/Low-frequency energy ratio distributions of endoscopic ultrasound images derived from the two centers. Center A represents the training data from the Beijing Friendship Hospital, and Center B represents the external validation data from the Beijing Tiantan Hospital. The subscript 0 indicates the cancerous category (Center A_0, Center B_0), while the subscript 1 indicates the noncancerous category (Center A_1, Center B_1).
Table 2 Dual-center energy statistics for the two hospitals.
Center
Mean
SD
Min
Max
Center A_00.0614740.0562340.0087410.559571
Center A_10.0478960.0400440.0091320.320019
Center B_00.0212910.0092370.0075140.047186
Center B_10.0236280.0172170.0093670.120883

Notably, even within the same imaging center, differences between the proportions of high-frequency energy for the cancer (class 0) and noncancer (class 1) samples still existed. For example, in Center A, the mean values of Center A_0 and Center A_1 were 0.061 and 0.048, respectively, and similar interclass shifts were observed in Center B. This suggests that the differences between disease categories can also affect the distribution of frequency-domain features in an image and may lead to the partial overlap of different categories in the frequency-domain space.

In summary, the proportion of high-frequency energy in the frequency domain not only revealed significant distribution shifts between the centers but also reflected potential confounding factors between the categories contained within the same center. This spectral distribution inconsistency, caused by both imaging center differences and disease features, may have induced the model to learn noncausal high-frequency features related to the center, thereby weakening its generalization ability on external datasets. This phenomenon further suggests that in the EUS image analysis task, it is necessary to introduce modeling strategies that can suppress center biases and enhance the robustness of discriminative features.

Data preprocessing

During the data preprocessing phase, we first labeled and organized all EUS images according to their lesion types. The final diagnosis was divided into two major categories based on clinicopathological results: Cancerous (including PDAC and PDCC) and noncancerous (including NENs, AIP, and solid pseudopapillary neoplasms). This classification was in line with the actual clinical diagnostic process and helped the model learn the imaging differences between cancer and common benign lesions. Due to the multicenter data coming from different hospitals, certain differences were observed in their EUS equipment models, imaging parameters and operating habits, resulting in inconsistencies among the original images in terms of their spatial resolutions, luminance distributions and contrast levels. To reduce the impacts of these factors on the model training process and improve the generalization ability of the model for use with multicenter data, all EUS images needed to undergo uniform data preprocessing schemes before being input into the network. The steps were as follows.

Size uniformity: Considering the resolution differences among the images collected by different devices, to ensure consistency for the input data at the spatial scale and to accommodate the network structure and memory limitations, all EUS images were uniformly scaled to a spatial resolution of 640 × 480 after the noncorrelated areas were cropped out. This operation helped maintain the overall structural information of each image while mitigating the problem of feature shifts caused by size differences.

Data normalization: To mitigate the impacts of brightness and contrast variations on the model training process under different scanning conditions, the maximum-intensity normalization method was used to linearly map the image pixel values to the[0,1] interval. Through normalization processing, the numerical distributions of different samples could be kept consistent, thereby improving the stability of the model training process and accelerating its convergence.

Data augmentation: To simulate the uncertainties exhibited by the probe angles and imaging positions during real clinical operations and effectively expand the number of training samples, multiple data augmentation strategies, including random rotation, random scaling, and random horizontal flipping, were implemented on the images during the training phase. Data augmentation could increase the robustness of the model to scale variations and spatial transformations, thereby reducing the risk of overfitting and enhancing the generalizability of the model. In addition, we removed nine data samples with poor imaging quality levels.

To ensure the objectivity of the model evaluation and prevent potential data leakage, the partitioning of the dataset into training, validation, and internal testing sets was strictly performed at the patient level rather than the image level. Specifically, all EUS images belonging to a single patient were assigned as a unified group to only one of the subsets. This strategy ensures that the model is tested on entirely unseen cases, thereby providing a rigorous assessment of its diagnostic generalization performance across different individuals.

Network model architecture and training process

Model architecture: This study proposes a deep learning network for classifying pancreatic cancer and noncancerous lesions in multicenter EUS images, termed MCEUS-C2Net (Multicenter EUS Image Classification Network for Cancerous vs Noncancerous Lesions), as illustrated in Figure 4. The proposed architecture is developed based on MedViT V2[17] with a hierarchical feature extraction design consisting of four stages. Given an input EUS image of size 640 × 480 × 3, stage 1 first performs initial encoding using the local feature extraction (LFE) module and downsamples the feature map to 1/4 resolution (approximately 160 × 120). This stage primarily captures low-level discriminative features such as textures, edges, and local anatomical structures.

Figure 4
Figure 4 Architecture of the MCEUS-C2Net model. CSA: Channel self-attention; DiNA: Dilated neighborhood attention; E-MHSA: Efficient multi-head self-attention; GFE: Global feature extraction; KAN: Kolmogorov-Arnold network; LFE: Local feature extraction; MHCA: Multi-head convolutional attention; LFFN: Locally feed-forward network.

To enhance local feature representation, a channel self-attention (CSA) mechanism is introduced after the dilated neighborhood attention (DiNA) block within the LFE module. Specifically, the CSA module adopts a projection strategy based on group convolution with the number of groups equal to the channel dimension[18]. The input feature is projected into channel-wise query Q, key K, and value V representations[19,20].

Channel-wise attention is then computed via matrix multiplication followed by a Softmax operation, enabling the modeling of global inter-channel dependencies. This design allows the network to adaptively recalibrate channel responses by suppressing redundant or irrelevant features (e.g., speckle noise induced by different imaging devices across centers) while emphasizing discriminative texture patterns associated with pancreatic lesions.

The CSA mechanism operates jointly with the sparse global attention mechanism embedded in the DiNA block. While sparse attention expands the effective receptive field within local windows under low computational cost, CSA refines channel-wise feature importance. This combination enables the LFE module to achieve both efficient local detail modeling and enhanced context awareness. Finally, the LFE module concludes with a locally feed-forward network, which further processes the recalibrated spatial features to enhance local representation.

Building upon the enhanced LFE module, stages 2-4 adopt a hierarchical hybrid strategy by progressively stacking LFE blocks with global feature extraction (GFE) modules in an alternating manner. Specifically, the network comprises a total of 40 feature extraction blocks distributed across the four stages (with block depths of[3,4,30,3], respectively). The overall architecture contains approximately 57.77 M parameters, striking an optimal balance between robust hierarchical feature extraction capability and computational efficiency. As the network depth increases, the spatial resolution is progressively reduced to 1/8 (80 × 60), 1/16 (40 × 30), and 1/32 (20 × 15), while the number of feature channels increases accordingly (i.e., 128, 256, and 512) to enhance high-level semantic representation.

The GFE module is built upon an efficient multi-head self-attention mechanism, following the design principle of MedViT V2. It is responsible for modeling long-range dependencies and capturing global contextual relationships across spatial regions. By aggregating cross-region semantic information, the GFE module compensates for the limitations of purely local modeling approaches, particularly in scenarios requiring global structural understanding of pancreatic lesions. Through the stepwise abstraction and fusion of CSA-enhanced local features and globally-aware semantic representations, the network progressively constructs high-level discriminative features. Following this global abstraction, a multi-head convolutional attention module is utilized to further refine these high-level semantic features by leveraging convolutional inductive biases. These features are then fed into the classification head, which utilizes a Kolmogorov-Arnold Network for adaptive feature transformation and aggregation, leading to the final prediction of pancreatic cancer vs noncancerous lesions.

Due to significant variations in imaging devices, acquisition protocols, and patient distributions across centers, multicenter EUS data exhibit notable domain shifts. The proposed architecture addresses this challenge by strengthening global context modeling and enabling effective interaction between local and global features. In particular, the integration of CSA enhances the robustness of local representations against center-specific noise and artifacts, while the hierarchical local-global fusion improves generalization across heterogeneous data distributions. This design ensures stable and reliable classification performance in real-world multicenter clinical settings. To protect institutional intellectual property and the proprietary nature of the clinical deployment algorithms, the source code and pre-trained weights of MCEUS-C2Net are not publicly available at this time.

Model training: To complete the binary task of classifying pancreatic cancer and noncancerous lesions in EUS, supervised learning is used in this paper to train the proposed network in an end-to-end manner. The model feeds EUS images and their corresponding labels via batch processing with a batch size of 16 during the training phase and updates the network parameters through the backpropagation algorithm. To enhance the model's robustness against multicenter imaging variations, comprehensive data augmentation is applied during training, including random rotation (range ± 15°), random horizontal flipping, and random zoom (0.9 × to 1.1 ×). Furthermore, architectural regularization techniques, including Dropout and DropPath (stochastic depth), were embedded within the network layers to further mitigate overfitting. The dataset maintains a relatively balanced class distribution at both the patient and image levels (with an approximate cancerous-to-noncancerous ratio of 1.4:1). Therefore, no additional class-rebalancing strategies were required. In each round of training, the model is first placed in the training mode to perform forward inference on the current batch of samples, and the loss function values are calculated on the basis of the differences between the predictions of the model and the real labels. In this paper, the cross-entropy loss function is used as the optimization objective, and the training loss is gradually minimized through gradient backpropagation.

The average training loss induced for each epoch is recorded during the model training process to monitor the convergence of the model. To evaluate the classification performance of the model on unseen data, the model is validated in a round-by-round manner during training using the internal validation set. During the validation phase, the model is switched to the evaluation mode, and forward inference is performed on the validation set samples without gradient updates. The predictions produced for all the validation samples are aggregated, processed via a Softmax activation to obtain class probabilities, which were then used to calculate the overall classification performance metrics of the model. The model uses the Adam optimizer for parameter updates, with an initial learning rate set at 1 × 10-5 and a maximum training round count of 800 epochs. This upper limit of 800 epochs was empirically set to provide sufficient optimization space for complete convergence. During training, the model’s performance was continuously monitored on the internal validation set. Instead of early stopping, a dynamic checkpointing strategy was utilized to prevent overfitting: Whenever the model achieved a new optimal area under the curve (AUC) on the validation set, the corresponding model parameters were saved. After completing the full 800 epochs, the checkpoint with the highest historical validation AUC was selected for all subsequent experimental analyses. All the experiments are conducted on a PC platform running Ubuntu 20.04, and the model training and inference processes are implemented on the PyTorch (version 2.1) deep learning framework in the Python language. The experimental hardware environment includes an NVIDIA GeForce RTX 3090 GPU with 24 GB of video memory, an Intel E5-2667 2.4 GHz processor, and 64 GB of system memory.

Model evaluation and visualization

To fully evaluate the classification performance of the model in differentiating pancreatic cancer lesions from noncancerous lesions in EUS images, multiple evaluation metrics, including precision, sensitivity, accuracy, specificity, and the F1 score, are used in this study[21]. These indicators reflect the discriminative ability of the model in clinically relevant scenarios from different perspectives and thus allow a more objective and comprehensive analysis of the performance of the model.

Sensitivity is used to measure the ability of the model to detect cases of pancreatic cancer, that is, the proportion of real pancreatic cancer samples that are correctly identified. In clinical applications, a missed diagnosis of pancreatic cancer can lead to delayed treatment for a patient; thus, increased sensitivity is important for assisting clinical decision-making results. Specificity is used to assess the ability of the model to correctly identify noncancerous lesions such as NENs and AIP, reflecting its ability to reduce the numbers of misdiagnoses and overdiagnoses, and is equally crucial for avoiding unnecessary invasive tests and treatments.

Precision measures the proportion of samples predicted as pancreatic cancer by the model that are actually cancer and is used to reflect the reliability of the prediction results produced by the model. Especially in the cancer screening scenario, high precision helps to reduce the false-positive rate. The accuracy, which reflects the overall classification accuracy achieved for the model over all samples, can visually assess the overall discriminative performance of the model, but certain limitations may be encountered when this metric alone is used in cases where the category distribution is unbalanced.

In addition, the F1 score, which strikes a balance between precision and sensitivity, provides a more comprehensive reflection of the overall performance of the model in the pancreatic cancer identification task. By combining the above multiple evaluation metrics, the performance of the model is assessed from multiple clinically relevant dimensions in this paper to ensure the stability and reliability of the model on multicenter EUS data.

To rigorously evaluate the classification performance and reliability of the models, 95% confidence interval (CI) for all quantitative evaluation metrics were calculated using the bootstrapping method with 1000 resamples. Furthermore, to assess the statistical significance of the performance differences between the proposed MCEUS-C2Net and the comparative baseline models, McNemar’s test was employed on the paired prediction results. A two-sided P < 0.05 was considered to indicate statistical significance. All statistical analyses were performed using Python.

To further analyze the interpretability of the decision-making process of the model and to validate the rationality of the regions the model focuses on, we conduct a qualitative visualization analysis of the attention heatmap generated by the model. Specifically, a representative EUS image is selected from the test set to show its original image, the attention heatmap generated by the model, and the fusion results of both outputs. In the attention heatmap, colors ranging from blue to red indicate the attention paid by the model to different spatial regions from low to high.

RESULTS

In accordance with the exclusion criteria, 176 patients (158 without video images and 18 with poor image quality) are excluded; ultimately, 383 patients are included in this study. As part of the development cohort, 302 patients are included and randomly divided into a training cohort, a validation cohort and a test cohort (with 182 patients in the training cohort, 60 in the validation cohort and 60 in the test cohort). In the external test cohort, 81 patients are included.

To fully evaluate how well MCEUS-C2Net distinguishes between pancreatic cancer lesions from noncancerous lesions in EUS images and to validate its stability and robustness in real-world clinical applications, both internal and external test experiments are conducted. The internal test is based on a data distribution that is consistent with the source of the training data and is used to evaluate the discriminative ability of the model under known distribution conditions; the external tests use data from derived different centers as independent.

Validation sets to simulate the data distribution differences that may be encountered in a real clinical deployment environment, thereby further evaluating the generalization performance of the model. By comparing the performances achieved by the model on the internal and external test sets, a more comprehensive reflection of its suitability for multicenter EUS data can be achieved.

In addition to quantitative evaluation metrics, visual analysis methods are introduced in this paper to enhance the interpretability of the prediction results yielded by the model. Specifically, attention heatmaps are used to visualize the key regions that the network focuses on during the classification process. The attention heatmaps visually reflect the image regions that the model focuses on when making classification decisions by projecting the intermediate feature mapping of the model back into the original image space. This method helps to analyze whether the model focuses on the anatomical structures and lesion areas that are associated with pancreatic lesions, thereby verifying whether the decision-making basis of the model is in line with clinical imaging cognition to a certain extent.

In addition, a confusion matrix is used to analyze the classification results produced by the model on the test set. The confusion matrix can visually show the predictive relationship between pancreatic cancer and noncancerous lesions, including the specific cases of correct classifications and misclassifications. Through a visual analysis of the confusion matrix, it is possible to further identify the strengths and weaknesses of the model for different categories and clarify the main sources of misdiagnoses and missed diagnoses.

Test results and analysis

After the model training process is completed, we select four representative classification models for conducting comparative experiments, including ResNet-50[22], a classic deep convolutional neural network that effectively alleviates vanishing gradients by introducing residual connections and is widely used in various medical image classification tasks. The Swin transformer[23] is a transformer network based on a hierarchical window self-attention mechanism that can model the global and local context information of images while maintaining computational efficiency. MedViTV2[23] is a lightweight visual transformer model that is optimized for medical imaging tasks and combines convolution and self-attention mechanisms to achieve excellent classification performance while maintaining low computational overhead. MCEUS-C2Net is the approach proposed in this paper. To comprehensively evaluate the classification performance of the models and their generalizability under multicenter data conditions, the above methods are systematically compared on internal validation sets and external independent datasets, and the quantitative results are summarized in Tables 3 and 4, respectively. In the internal validation experiments, the test data are derived from the same source as that of the training data, and the classification performance of the comparison models is relatively ideal (Table 2). Among them, MCEUS-C2Net achieves the best results in terms of precision, sensitivity, accuracy, specificity and the F1 score. By analyzing the 95%CI, we observe that MCEUS-C2Net exhibits the narrowest interval ranges and the highest lower bounds across all metrics, indicating highly stable and reliable predictive ability. To control for the false discovery rate (FDR) during multiple comparisons, the Benjamini-Hochberg correction was applied to the P values of McNemar’s tests. On the internal dataset, MCEUS-C2Net showed statistically significant improvements over ResNet-50 (P = 0.021, P < 0.05), Swin transformer (P = 0.024, P < 0.05), and MedViTV2 (P = 0.043, P < 0.05).

Table 3 Internal validation results.
Methods
Precision (%) (95%CI)↑
Sensitivity (%) (95%CI)↑
Accuracy (%) (95%CI)↑
Specificity (%) (95%CI)↑
F1 score (%) (95%CI)↑
ResNet-50a86.96 (82.17-90.95)94.14 (91.47-97.49)89.12 (86.41-92.05)82.76 (77.43-88.30)90.50 (87.76-93.16)
Swin_transformera89.29 (83.64-92.31)94.34 (90.48-96.97)90.53 (86.92-92.82)85.71 (79.89-90.64)91.74 (88.02-93.65)
MedViTV2a90.09 (85.32-93.24)95.24 (91.54-97.66)91.50 (87.95-93.85)86.75 (81.71-91.33)92.59 (89.20-94.48)
MCEUS-C2Net (ours)93.46 (89.08-95.97)97.09 (95.48-99.51)94.51 (92.31-96.67)91.14 (86.70-94.92)95.24 (93.01-97.01)
Table 4 External validation results.
Methods
Precision (%) (95%CI)↑
Sensitivity (%) (95%CI)↑
Accuracy (%) (95%CI)↑
Specificity (%) (95%CI)↑
F1 score (%) (95%CI)↑
ResNet-50b84.57 (79.00-89.88)74.00 (67.88-80.49)78.65 (74.32-82.97)84.12 (78.31-89.51)78.93 (74.13-83.42)
Swin_transformerb86.63 (81.87-91.24)81.00 (75.26-86.26)82.97 (78.92-86.49)85.29 (79.88-90.30)83.72 (79.49-87.35)
MedViTV2a86.57 (82.03-91.35)87.00 (82.16-91.67)85.68 (82.16-89.19)84.12 (78.33-89.54)86.78 (83.24-90.10)
MCEUS-C2Net (ours)91.88 (87.75-95.26)90.50 (86.47-94.23)90.59 (87.84-93.51)90.59 (85.98-94.51)91.18 (88.32-93.77)

Specifically, in the external validation experiments, the test data come from independent datasets that do not participate in the model training process, and they exhibit significant imaging equipment, scanning parameter, and data distribution differences relative to those of the training data (Table 3). Under this more challenging setup, the classification performance of some of the comparative methods declines significantly. Specifically, the sensitivity, accuracy and F1 score of ResNet-50 decrease by 20.14%, 10.47% and 11.57%, respectively. The metrics of the Swin transformer decrease by 13.34%, 7.56%, and 8.02%, respectively. The performance of MedViTV2 also decreases, but to a lesser extent, with the corresponding metrics decreasing by 8.24%, 5.82%, and 5.81%, respectively. Furthermore, the 95%CI of these baseline models widened significantly, reflecting their instability and high uncertainty when facing unknown data distributions. In contrast, compared with the competing models, MCEUS-C2Net maintains high and stable classification performance in the external validation, with significantly smaller declines in terms of all the metrics. On the external dataset, MCEUS-C2Net achieved a significant performance improvement over MedViTV2 (P = 0.005, P < 0.05), and demonstrated highly significant superiority over ResNet-50 (P = 0.003, P < 0.01) and Swin transformer (P = 0.005, P < 0.01).

To further analyze the misclassification rate of each method on the external datasets, the corresponding confusion matrix is shown in Figure 5, where T represents the pancreatic cancer category and F represents the noncancerous lesion category. ResNet-50 has the greatest number of misclassified samples, with 52 and 27 false-positive and false-negative samples, respectively, suggesting that this method has a relatively high risk of misdiagnoses and missed diagnoses in external data. The Swin transformer improves these results in terms of false positives but still yields 38 and 25 false-positive and false-negative samples, respectively. The number of false positives yielded by MedViTV2 is further reduced to 26, but the number of false negatives remains 27, and the overall number of misclassified samples is still relatively high. In contrast, MCEUS-C2Net has the fewest misclassified samples, with only 19 false positives and 16 false negatives, indicating a more balanced performance in terms of reducing misdiagnoses and missed diagnoses and a more stable and reliable discriminative ability on external independent datasets. To go beyond descriptive statistics and understand the clinical implications, an error pattern analysis was conducted on the misclassified samples based on our specific pathological categorizations. A detailed review revealed two primary error modes. First, the false-positive cases (noncancerous lesions misclassified as cancer) were predominantly associated with AIP and atypical NENs. AIP, in particular, often presents as a focal hypoechoic, heterogeneous mass mimicking the morphological characteristics of PDAC, which frequently confused the baseline models. Second, the false-negative cases (cancers misclassified as noncancerous) mainly involved well-circumscribed PDCC or early-stage PDACs lacking classic invasive textural features. The well-defined borders of PACC can mimic the appearance of SPN or NENs, leading to algorithmic misjudgments. Notably, MCEUS-C2Net effectively mitigated these specific clinical error patterns. To quantitatively evaluate the classification performance of MCEUS-C2Net, a further analysis is conducted to determine its ability to distinguish pancreatic cancer lesions from noncancerous lesions in EUS images, and AUC is calculated. The AUC, training loss curve, and ROC curve produced for the external validation set of the model are shown in Figure 6, with the corresponding AUC values reaching 0.9704. The AUC values of the baseline models (ResNet-50, Swin transformer, and MedViTV2) on the external dataset were all lower than that of MCEUS-C2Net. The calibration curves and decision curve analysis (DCA) results of the models on the external validation set are shown in Figure 7. As observed in Figure 7A, the prediction curve of MCEUS-C2Net aligns most closely with the ideal diagonal line and exhibits the highest linearity, indicating optimal calibration among the compared models. Furthermore, the DCA curves in Figure 7B reveal that across a wide range of clinical risk thresholds, MCEUS-C2Net consistently yields the highest net clinical benefit, demonstrating relatively optimal clinical utility compared to the traditional baseline models. The above results fully validate the superior performance and good generalization ability of MCEUS-C2Net on this task.

Figure 5
Figure 5 Confusion matrix diagrams produced for the four algorithms in the external validation experiment. A: ResNet-50; B: Swin transformer; C: MedViTV2; D: MCEUS-C2Net. T: Pancreatic cancer category; F: Noncancerous lesion category.
Figure 6
Figure 6 Performance curves of MCEUS-C2Net during training and external validation. A: Area under the curve variation curve during training; B: Loss curve during training; C: Receiver operating characteristic curve for the external validation dataset. AUC: Area under the curve.
Figure 7
Figure 7 Clinical utility evaluation of MCEUS-C2Net and baseline models on the external validation dataset. A: Calibration curves comparing the predicted probabilities of pancreatic cancer with the actual observed frequencies; B: Decision curve analysis evaluating the net clinical benefit across different threshold probabilities.
Ablation study

To systematically validate the independent contributions of the pure GFE, pure LFE, and the CSA mechanism, and to explicitly distinguish our proposed architecture from the baseline MedViTV2, we conducted comprehensive ablation studies by decoupling the network.

On the internal validation set (Table 5), the pure GFE and pure LFE modules achieved baseline accuracies of 86.41% and 88.46%, respectively. Notably, the integration of the CSA module with LFE (pure LFE + CSA) independently improved the accuracy to 89.49%. This demonstrates the efficacy of the CSA mechanism in refining local spatial features and suppressing noise, even in the absence of global context. Ultimately, our complete MCEUS-C2Net achieved the highest accuracy of 94.51%, outperforming the baseline LFE + GFE (91.50%). Furthermore, MCEUS-C2Net exhibited the narrowest 95%CI among all configurations, indicating highly stable and reliable predictive capability.

Table 5 Ablation study results on the internal validation set.
Methods
Precision (%) (95%CI)↑
Sensitivity (%) (95%CI)↑
Accuracy (%) (95%CI)↑
Specificity (%) (95%CI)↑
F1 score (%) (95%CI)↑
GFEb86.85 (82.25-91.39)88.10 (83.76-92.56)86.41 (82.82-89.74)84.44 (79.28-89.64)87.47 (83.99-90.70)
LFEb88.73 (84.46-92.72)90.00 (85.92-93.97)88.46 (85.13-91.54)86.67 (81.82-91.33)89.36 (86.15-92.38)
LFE + CSAb90.05 (85.86-93.91)90.48 (86.57-94.45)89.49 (86.41-92.56)88.33 (83.60-92.90)90.26 (87.12-93.15)
LFE + GFEa90.09 (85.32-93.24)95.24 (91.54-97.66)91.50 (87.95-93.85)86.75 (81.71-91.33)92.59 (89.20-94.48)
MCEUS-C2Net (ours)93.46 (89.08-95.97)97.09 (95.48-99.51)94.51 (92.31-96.67)91.14 (86.70-94.92)95.24 (93.01-97.01)

The necessity and complementary nature of each module are most pronounced in the external cohort (Table 6), where severe speckle noise and domain shifts are present. Relying solely on either the pure GFE or pure LFE led to drastic performance degradation, yielding external accuracies of only 78.11% and 80.27%, respectively, accompanied by noticeably wider 95%CI that reflect high uncertainty on unseen data distributions. While the baseline MedViTV2 (LFE + GFE) improved the accuracy to 85.68%, it remained vulnerable to cross-center heterogeneity. By incorporating the CSA module prior to the global fusion, our final MCEUS-C2Net dramatically boosted the external accuracy to 90.59%. Crucially, McNemar’s tests, with P-values adjusted using the Benjamini-Hochberg FDR correction for multiple comparisons, confirmed that the proposed complete model yielded significant improvements over all ablated configurations (P < 0.05 or P < 0.01). These results conclusively substantiate that the LFE, GFE, and CSA mechanisms are highly complementary and indispensable for robust generalization.

Table 6 Ablation study results on the external validation set.
Methods
Precision (%) (95%CI)↑
Sensitivity (%) (95%CI)↑
Accuracy (%) (95%CI)↑
Specificity (%) (95%CI)↑
F1 score (%) (95%CI)↑
GFEb83.62 (78.15-89.10)74.00 (67.88-80.49)78.11 (73.78-82.43)82.94 (77.06-88.54)78.51 (73.63-82.95)
LFEb86.71 (81.39-92.03)75.00 (68.93-81.42)80.27 (75.95-84.32)86.47 (81.03-91.47)80.43 (75.80-84.85)
LFE + CSAb87.10 (82.18-92.19)81.00 (75.50-86.46)83.24 (79.18-87.30)85.88 (80.46-91.14)83.94 (79.77-87.77)
LFE + GFEa86.57 (82.03-91.35)87.00 (82.16-91.67)85.68 (82.16-89.19)84.12 (78.33-89.54)86.78 (83.24-90.10)
MCEUS-C2Net (ours)91.88 (87.75-95.26)90.50 (86.47-94.23)90.59 (87.84-93.51)90.59 (85.98-94.51)91.18 (88.32-93.77)
Attention visualization analysis

The original images, attention heatmaps and fusion results of ResNet-50, the Swin transformer, MedViTV2 and the MCEUS-C2Net model proposed in this paper are shown in Figure 8. As shown in the figure, the high-response regions of MCEUS-C2Net are concentrated mainly at the locations of pancreatic lesions, which are highly consistent with the areas that clinicians focus on during the actual diagnosis process. Notably, this pattern of concern does not rely on any artificially labeled lesion areas but is rather autonomously learned by the model during end-to-end training. These findings suggest that the proposed approach can effectively focus on the key regions that are closely related to the discrimination between pancreatic cancer and noncancerous lesions, providing good interpretability support for the decision-making process of the model. This phenomenon is attributed mainly to the channel attention mechanism introduced in the model. By adaptively weighting feature responses in the channel dimension, this mechanism enhances the discriminative feature expressions that are related to pancreatic lesions while suppressing interference from redundant or background information, thereby improving the model’s perceptions of critical structures in complex EUS imaging environments. In contrast, ResNet-50, the Swin transformer and MedViTV2 show varying shifts in their attention distributions on the external test datasets, with the high-response regions often being more dispersed and some areas of concern having lower consistency with the actual lesion locations. This suggests that the comparison models are more vulnerable in terms of their feature attention stability when faced with external data possessing changing imaging conditions and data distributions. In contrast, MCEUS-C2Net is able to maintain stable and reasonable attention distributions on external datasets, further verifying its stronger robustness in multicenter data scenarios and its potential clinical application value.

Figure 8
Figure 8 Visual comparison of attention heatmaps generated by different models. A: Original endoscopic ultrasound images; B: Attention heatmaps produced by ResNet-50, Swin transformer, MedViTV2, and MCEUS-C2Net; C: Fusion of original images and attention heatmaps.
DISCUSSION

With the development of AI, especially deep learning[23], an increasing number of works have demonstrated the outstanding performance of AI in the field of digestive diseases and other medical fields[24-26].

In this study, a multicenter dataset with significant distribution differences based on EUS images acquired from two hospitals, different imaging devices, and different patient sources was constructed, and the generalization performance of deep learning models in the task of automatically classifying pancreatic cancer and noncancerous lesions was systematically evaluated. The experimental results revealed that the proposed MCEUS-C2Net approach achieved stable and excellent classification performance in both an internal validation and an external validation conducted on an independent dataset, especially in the latter case, and its accuracy remained above 90%, demonstrating good cross-center adaptability.

Previous studies have also used deep learning architectures to develop AI systems for diagnosing pancreatic diseases in EUS images[27,28], but none of these systems have studied cross-center data, and multicenter EUS data in actual clinical settings possess significant distribution differences. An energy spectrum analysis of multicenter pancreatic EUS images revealed significant spectral distribution and imaging characteristic differences among different centers, mainly due to acquisition equipment difference, different scanning parameter settings, and changes in the compositions of the patient population. Such distribution differences often lead to higher model performance on single-center data but a significant decline in performance on external datasets, as evidenced in previous studies and the comparative experiments conducted in this study. For example, in the external validation, both ResNet-50 and the Swin transformer resulted in significant decreases in their sensitivity and F1 score metrics, whereas MedViTV2 performed relatively stably but was still affected by data distribution changes.

In contrast, the declines exhibited by the various performance indicators of MCEUS-C2Net on the external validation set were significantly smaller than those of the comparison methods (Table 3). This suggests that the model has potential to help standardize the interpretation of EUS images in real endoscopy suites, which is particularly valuable for less-experienced endoscopists when facing complex clinical scenarios. These results indicate that the proposed method can maintain a more stable discriminative ability when addressing unseen data distributions. This advantage is due mainly to the structural design of the model, which combines both LFE and global context modeling capabilities and introduces a channel attention mechanism for the adaptive reweighting of feature representations. By strengthening the discriminative features that are related to pancreatic lesions in the channel dimension and suppressing redundant or irrelevant information, the model can respond more effectively to the distribution variations caused by multicenter data, thereby improving its generalization performance.

This study shows that MCEUS-C2Net does not overly rely on the feature patterns acquired at specific centers or under specific imaging conditions but rather learns more robust and discriminative feature representations in complex multicenter data environments. This trait gives it greater application potential in real clinical scenarios, especially for cross-device, cross-center EUS-assisted diagnostic tasks, providing a practically feasible technical solution for the early identification of pancreatic cancer and clinical decision support.

First, although multicenter data acquired from two hospitals were introduced in this study, the external validation remains limited as it relies on only a single external center. Additionally, the overall sample size was still relatively limited. The acquisition and standardized labeling processes employed for multicenter EUS data are difficult to execute in actual clinical practice, and the sample size limitation may have affected the stability of the obtained statistical results to some extent. Future studies will further expand the data size through the accumulation of data over longer time spans and multi-institutional collaboration to obtain more robust and representative assessment results. Another limitation relates to potential selection bias and spectrum bias. Currently, achieving high classification accuracy across multicenter data while simultaneously accounting for suboptimal, low-quality images from actual clinical practice remains a significant challenge. Therefore, in the future, we will attempt to include more images that present diagnostic difficulties arising from routine operator handling. Furthermore, the variety of noncancerous lesions included in the current study is somewhat limited. Future studies will incorporate a broader spectrum of noncancerous disease types with granular annotations to systematically evaluate the model’s performance across specific subtypes and mitigate potential spectrum bias. Regarding model transparency, although the regions of interest (ROIs) were visualized through attention heatmaps (e.g., Gradient-weighted Class Activation Mapping), providing qualitative support for the model’s decision-making process, this analysis remains primarily descriptive. We acknowledge the necessity of “quantitative verification” to further enhance the persuasiveness of the model’s interpretability. However, as this multicenter retrospective study involves a large-scale dataset, obtaining precise pixel-level manual masks or bounding boxes for all images currently poses a significant logistical bottleneck. Consequently, the current study did not involve strict quantitative comparisons with expert-labeled lesion sites or specific localization metrics (such as IoU or pointing accuracy). Future studies will incorporate expert-drawn ROIs on representative subsets to systematically verify the “rationality” of the model’s focus, providing a more rigorous alignment between AI-identified features and clinical diagnostic gold standards. Notably, another major limitation of the current study is the lack of a head-to-head performance comparison between the proposed AI model and human clinical experts (such as experienced endosonographers and radiologists). To fully establish its clinical utility, future prospective studies must involve a multi-reader, multi-case clinical evaluation to systematically compare the diagnostic accuracy and efficiency of the AI model with those of human experts of varying experience levels.

Furthermore, while MCEUS-C2Net demonstrates promising retrospective performance and robust cross-center generalization, transitioning to real-world clinical deployment faces practical challenges. First, EUS is inherently operator-dependent; inconsistencies in scanning methods, probe angles, and equipment settings introduce significant image heterogeneity. In fact, addressing this exact clinical pain point was the core motivation behind our multicenter design, and our model has indeed proven its effectiveness in mitigating such cross-center variance at the algorithmic level. However, for optimal real-world deployment, relying solely on algorithmic generalization is insufficient. Establishing standardized EUS scanning protocols across different institutions will be a crucial practical prerequisite to minimize extreme operator-induced noise and maximize the AI’s clinical utility. Second, real-time inference is absolutely essential. Since EUS is a dynamic, continuous procedure, a clinically viable AI system must achieve a high processing frame rate with minimal latency, providing instantaneous feedback to guide endoscopists rather than acting as an offline post-processing tool. Finally, seamless integration into the existing endoscopic workflow is critical. Ideally, the proposed model should be deployed as a concurrent computer-aided diagnosis plug-in. It could display real-time diagnostic probabilities on a secondary monitor (or via picture-in-picture) without disrupting the physician's primary field of view or prolonging the procedure time. Addressing these hardware-software integration and standardization hurdles will be the primary focus of our future prospective clinical trials.

CONCLUSION

In this study, the first multicenter EUS pancreatic disease imaging dataset, MEUS-PCBL, was constructed, and a classification network, MCEUS-C2Net, was proposed for multicenter data. The network effectively mitigated the generalization performance degradation caused by the distribution differences exhibited by multicenter data. Furthermore, in cross-center tests, it demonstrated excellent robustness as a potentially applicable tool compared with existing AI models. This model could help standardize EUS reading across different hospitals and may assist clinicians in making more accurate diagnoses. Future work will further expand the data scale and integrate pixel-level annotation to enhance the quantitative interpretability validation of the model.

References
1.  He R, Jiang W, Wang C, Li X, Zhou W. Global burden of pancreatic cancer attributable to metabolic risks from 1990 to 2019, with projections of mortality to 2030. BMC Public Health. 2024;24:456.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 33]  [Cited by in RCA: 30]  [Article Influence: 15.0]  [Reference Citation Analysis (0)]
2.  Almasri B, Ali A. Role of endoscopic ultrasound elastography in differential diagnosis of pancreatic solid masses. Qatar Med J. 2021;2021:40.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 1]  [Cited by in RCA: 2]  [Article Influence: 0.4]  [Reference Citation Analysis (0)]
3.  Vitali F, Zundler S, Jesper D, Wildner D, Strobel D, Frulloni L, Neurath MF. Diagnostic Endoscopic Ultrasound in Pancreatology: Focus on Normal Variants and Pancreatic Masses. Visc Med. 2023;39:121-130.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 6]  [Cited by in RCA: 6]  [Article Influence: 2.0]  [Reference Citation Analysis (0)]
4.  Zhang S, Ni M, Wang P, Zheng J, Sun Q, Xu G, Peng C, Shen S, Zhang W, Huang S, Wang L, Zou X, Lv Y. Diagnostic value of endoscopic ultrasound-guided fine needle aspiration with rapid on-site evaluation performed by endoscopists in solid pancreatic lesions: A prospective, randomized controlled trial. J Gastroenterol Hepatol. 2022;37:1975-1982.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 21]  [Cited by in RCA: 20]  [Article Influence: 5.0]  [Reference Citation Analysis (0)]
5.  Gheorghiu M, Sparchez Z, Rusu I, Bolboacă SD, Seicean R, Pojoga C, Seicean A. Direct Comparison of Elastography Endoscopic Ultrasound Fine-Needle Aspiration and B-Mode Endoscopic Ultrasound Fine-Needle Aspiration in Diagnosing Solid Pancreatic Lesions. Int J Environ Res Public Health. 2022;19:1302.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 10]  [Cited by in RCA: 10]  [Article Influence: 2.5]  [Reference Citation Analysis (1)]
6.  Puga-Tejada M, Del Valle R, Oleas R, Egas-Izquierdo M, Arevalo-Mora M, Baquerizo-Burgos J, Ospina J, Soria-Alcivar M, Pitanga-Lukashok H, Robles-Medranda C. Endoscopic ultrasound elastography for malignant pancreatic masses and associated lymph nodes: Critical evaluation of strain ratio cutoff value. World J Gastrointest Endosc. 2022;14:524-535.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in CrossRef: 6]  [Cited by in RCA: 6]  [Article Influence: 1.5]  [Reference Citation Analysis (0)]
7.  Conti CB, Mulinacci G, Salerno R, Dinelli ME, Grassia R. Applications of endoscopic ultrasound elastography in pancreatic diseases: From literature to real life. World J Gastroenterol. 2022;28:909-917.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in CrossRef: 14]  [Cited by in RCA: 10]  [Article Influence: 2.5]  [Reference Citation Analysis (0)]
8.  Kuwahara T, Hara K, Mizuno N, Haba S, Okuno N, Fukui T, Urata M, Yamamoto Y. Current status of artificial intelligence analysis for the treatment of pancreaticobiliary diseases using endoscopic ultrasonography and endoscopic retrograde cholangiopancreatography. DEN Open. 2024;4:e267.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 17]  [Cited by in RCA: 14]  [Article Influence: 7.0]  [Reference Citation Analysis (0)]
9.  Udriștoiu AL, Cazacu IM, Gruionu LG, Gruionu G, Iacob AV, Burtea DE, Ungureanu BS, Costache MI, Constantin A, Popescu CF, Udriștoiu Ș, Săftoiu A. Real-time computer-aided diagnosis of focal pancreatic masses from endoscopic ultrasound imaging based on a hybrid convolutional and long short-term memory neural network model. PLoS One. 2021;16:e0251701.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 53]  [Cited by in RCA: 46]  [Article Influence: 9.2]  [Reference Citation Analysis (0)]
10.  Kenner B, Chari ST, Kelsen D, Klimstra DS, Pandol SJ, Rosenthal M, Rustgi AK, Taylor JA, Yala A, Abul-Husn N, Andersen DK, Bernstein D, Brunak S, Canto MI, Eldar YC, Fishman EK, Fleshman J, Go VLW, Holt JM, Field B, Goldberg A, Hoos W, Iacobuzio-Donahue C, Li D, Lidgard G, Maitra A, Matrisian LM, Poblete S, Rothschild L, Sander C, Schwartz LH, Shalit U, Srivastava S, Wolpin B. Artificial Intelligence and Early Detection of Pancreatic Cancer: 2020 Summative Review. Pancreas. 2021;50:251-279.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 137]  [Cited by in RCA: 95]  [Article Influence: 19.0]  [Reference Citation Analysis (0)]
11.  Daher H, Punchayil SA, Ismail AAE, Fernandes RR, Jacob J, Algazzar MH, Mansour M. Advancements in Pancreatic Cancer Detection: Integrating Biomarkers, Imaging Technologies, and Machine Learning for Early Diagnosis. Cureus. 2024;16:e56583.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 19]  [Cited by in RCA: 8]  [Article Influence: 4.0]  [Reference Citation Analysis (0)]
12.  Sijithra PC, Santhi N, Ramasamy N. A review study on early detection of pancreatic ductal adenocarcinoma using artificial intelligence assisted diagnostic methods. Eur J Radiol. 2023;166:110972.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 15]  [Cited by in RCA: 11]  [Article Influence: 3.7]  [Reference Citation Analysis (0)]
13.  Cui H, Zhao Y, Xiong S, Feng Y, Li P, Lv Y, Chen Q, Wang R, Xie P, Luo Z, Cheng S, Wang W, Li X, Xiong D, Cao X, Bai S, Yang A, Cheng B. Diagnosing Solid Lesions in the Pancreas With Multimodal Artificial Intelligence: A Randomized Crossover Trial. JAMA Netw Open. 2024;7:e2422454.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 33]  [Cited by in RCA: 40]  [Article Influence: 20.0]  [Reference Citation Analysis (1)]
14.  Fang YJ, Lee KH, Karmakar R, Mukundan A, Nagisetti Y, Huang CW, Wang HC. Transforming Endoscopic Image Classification with Spectrum-Aided Vision for Early and Accurate Cancer Identification. Diagnostics (Basel). 2025;15:2732.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 2]  [Cited by in RCA: 5]  [Article Influence: 5.0]  [Reference Citation Analysis (0)]
15.  Chou CK, Lee KH, Karmakar R, Mukundan A, Gade PC, Gupta D, Su CC, Chen TH, Ko CY, Wang HC. Emulating Hyperspectral and Narrow-Band Imaging for Deep-Learning-Driven Gastrointestinal Disorder Detection in Wireless Capsule Endoscopy. Bioengineering (Basel). 2025;12:953.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 1]  [Cited by in RCA: 7]  [Article Influence: 7.0]  [Reference Citation Analysis (0)]
16.  Weng WC, Huang CW, Su CC, Mukundan A, Karmakar R, Chen TH, Avhad AR, Chou CK, Wang HC. Optimizing Esophageal Cancer Diagnosis with Computer-Aided Detection by YOLO Models Combined with Hyperspectral Imaging. Diagnostics (Basel). 2025;15:1686.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 8]  [Cited by in RCA: 14]  [Article Influence: 14.0]  [Reference Citation Analysis (0)]
17.  Manzari ON, Ahmadabadi H, Kashiani H, Shokouhi SB, Ayatollahi A. MedViT: A robust vision transformer for generalized medical image classification. Comput Biol Med. 2023;157:106791.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 376]  [Cited by in RCA: 137]  [Article Influence: 45.7]  [Reference Citation Analysis (0)]
18.  Nejati Manzari O, Asgariandehkordi H, Koleilat T, Xiao Y, Rivaz H. Medical image classification with KAN-integrated transformers and dilated neighborhood attention. Appl Soft Comput. 2026;186:114045.  [PubMed]  [DOI]  [Full Text]
19.  He K, Gan C, Li Z, Rekik I, Yin Z, Ji W, Gao Y, Wang Q, Zhang J, Shen D. Transformers in medical image analysis. Intelligent Medicine. 2023;3:59-78.  [PubMed]  [DOI]  [Full Text]
20.  Zhao H, Gou Y, Li B, Peng D, Lv J, Peng X.   Comprehensive and Delicate: An Efficient Transformer for Image Restoration. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE, 2023: 14122-14132.  [PubMed]  [DOI]  [Full Text]
21.  Yue Y, Li Z.   Medmamba: Vision mamba for medical image classification. 2024 Preprint. Available from: arXiv:2403.03849.  [PubMed]  [DOI]  [Full Text]
22.  He K, Zhang X, Ren S, Sun J.   Deep Residual Learning for Image Recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas: IEEE, 2016: 770-778.  [PubMed]  [DOI]  [Full Text]
23.  Balikov DA, Hu K, Liu CJ, Betz BL, Chinnaiyan AM, Devisetty LV, Venneti S, Tomlins SA, Cani AK, Rao RC. Comparative Molecular Analysis of Primary Central Nervous System Lymphomas and Matched Vitreoretinal Lymphomas by Vitreous Liquid Biopsy. Int J Mol Sci. 2021;22:9992.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 13]  [Cited by in RCA: 16]  [Article Influence: 3.2]  [Reference Citation Analysis (0)]
24.  Kuwahara T, Hara K, Mizuno N, Haba S, Okuno N, Koda H, Miyano A, Fumihara D. Current status of artificial intelligence analysis for endoscopic ultrasonography. Dig Endosc. 2021;33:298-305.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 45]  [Cited by in RCA: 40]  [Article Influence: 8.0]  [Reference Citation Analysis (2)]
25.  Wang H, Ni D, Wang Y. Recursive Deformable Pyramid Network for Unsupervised Medical Image Registration. IEEE Trans Med Imaging. 2024;43:2229-2240.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 92]  [Cited by in RCA: 33]  [Article Influence: 16.5]  [Reference Citation Analysis (0)]
26.  Zhou Y, Chia MA, Wagner SK, Ayhan MS, Williamson DJ, Struyven RR, Liu T, Xu M, Lozano MG, Woodward-Court P, Kihara Y; UK Biobank Eye & Vision Consortium, Altmann A, Lee AY, Topol EJ, Denniston AK, Alexander DC, Keane PA. A foundation model for generalizable disease detection from retinal images. Nature. 2023;622:156-163.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 701]  [Cited by in RCA: 534]  [Article Influence: 178.0]  [Reference Citation Analysis (2)]
27.  Tonozuka R, Itoi T, Nagata N, Kojima H, Sofuni A, Tsuchiya T, Ishii K, Tanaka R, Nagakawa Y, Mukai S. Deep learning analysis for the detection of pancreatic cancer on endosonographic images: a pilot study. J Hepatobiliary Pancreat Sci. 2021;28:95-104.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 112]  [Cited by in RCA: 89]  [Article Influence: 17.8]  [Reference Citation Analysis (1)]
28.  Zhang J, Zhu L, Yao L, Ding X, Chen D, Wu H, Lu Z, Zhou W, Zhang L, An P, Xu B, Tan W, Hu S, Cheng F, Yu H. Deep learning-based pancreas segmentation and station recognition system in EUS: development and validation of a useful training tool (with video). Gastrointest Endosc. 2020;92:874-885.e3.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 94]  [Cited by in RCA: 83]  [Article Influence: 13.8]  [Reference Citation Analysis (3)]
Footnotes

Peer review: Externally peer reviewed.

Peer-review model: Single blind

Specialty type: Oncology

Country of origin: China

Peer-review report’s classification

Scientific quality: Grade B, Grade B, Grade C, Grade D

Novelty: Grade B, Grade B, Grade C, Grade C

Creativity or innovation: Grade B, Grade B, Grade C, Grade C

Scientific significance: Grade B, Grade B, Grade C, Grade D

P-Reviewer: Mukundan A, Adjunct Professor, Editor, Postdoctoral Fellow, Research Dean, Taiwan; Othman AA, Lecturer, MD, PhD, Egypt; Tan HS, PhD, Professor, China S-Editor: Qu XL L-Editor: A P-Editor: Zhao YQ

Write to the Help Desk