Published online Sep 15, 2026. doi: 10.4251/wjgo.119889
Revised: March 29, 2026
Accepted: May 12, 2026
Published online: September 15, 2026
Processing time: 212 Days and 3.1 Hours
The early diagnosis of pancreatic cancer via endoscopic ultrasound (EUS) is cha
To improve pancreatic cancer detection by developing a robust deep learning model capable of handling data heterogeneity across multiple EUS centers.
We constructed a multicenter dataset (383 patients, 2362 images) from two hos
MCEUS-C2Net achieved accuracy, sensitivity, and F1 score of 94.51%, 97.09%, and 95.24% in the internal test, outperforming ResNet-50, Swin transformer, and MedViTV2. In external validation, it maintained robust per
Our approach achieves high-performance pancreatic cancer classification across multicenter data. It effectively addresses imaging heterogeneity, providing a robust, clinically viable solution for standardized EUS diagnosis in real-world settings.
Core Tip: This study introduces the first multicenter endoscopic ultrasound dataset for pancreatic lesions. We also propose a novel deep learning model, MCEUS-C2Net. It integrates local and global features with channel attention mechanisms. This design effectively overcomes cross-center imaging heterogeneity. The model accurately differentiates cancerous from non
- Citation: Yu XY, Ye JQ, He Z, He Q. Multicenter deep learning model for pancreatic cancer detection using endoscopic ultrasound. World J Gastrointest Oncol 2026; 18(9): 119889
- URL: https://www.wjgnet.com/1948-5204/full/v18/i9/119889.htm
- DOI: https://dx.doi.org/10.4251/wjgo.119889
Pancreatic cancer is difficult to diagnose early and has a very poor prognosis for patients because of its insidious early symptoms, the limited sensitivity of conventional imaging tests such as enhanced computed tomography and magnetic resonance imaging to small lesions, and the insufficient specificity of serum markers[1]. In this context, endoscopic ultrasound (EUS) has become the most sensitive imaging method for detecting pancreatic tumors, especially early lesions with diameters of less than 2 cm. This high sensitivity is due to the insertion of a high-frequency ultrasound probe into the digestive tract to scan adjacent to the pancreas, which provides a much higher spatial resolution than conventional in vitro imaging[2,3]. Not only has it been recognized by studies at home and abroad as one of the most reliable and efficient tests for diagnosing small pancreatic lesions, but its guided fine-needle aspiration (EUS-FNA/B) technique is also the “gold standard” for differentiating between benign and malignant tumors[4,5]. To further improve diagnostic per
However, traditional methods based on manual feature extraction processes often rely on human experience in medical image classification tasks, so they have a limited ability to recognize patterns. Consequently, they have difficulty fully characterizing the complex morphological and textural differences between pancreatic cancer and non-cancerous lesions, which often leads to insufficient robustness in practical applications.
With the rapid development of deep learning techniques, end-to-end feature learning paradigms have shown sig
However, most existing studies have constructed models based on single-center datasets. Differences among the imaging equipment, acquisition protocols, and patient populations of different medical centers lead to significant heterogeneity in EUS images, commonly known as domain shift. Consequently, models trained on single-center data often show severe performance degradation in cross-center external validation scenarios. From a clinical perspective, this poor generalization is not merely a technical limitation but a major deployment barrier and a direct risk to patients; an un
We constructed a multicenter EUS pancreatic disease imaging dataset, named MEUS-PCBL, for the classification of pancreatic cancer and benign lesions. The dataset was collected from 559 patients aged 45 years to 82 years who underwent EUS-guided fine-needle aspiration/biopsy (EUS-FNA/B) for pancreatic lesions at the Beijing Tiantan Hospital and Beijing Friendship Hospital between January 2020 and June 2025. The tumor lesions were pathologically diagnosed using tissue specimens obtained from EUS-guided fine-needle aspiration/biopsy (EUS-FNA/B) or surgical resection. To ensure strict temporal correlation between the imaging and the reference standard, all EUS-FNA/B procedures were performed concurrently during the same EUS examination. For patients undergoing surgical treatment, the interval between the EUS examination and surgical pathology ranged from 7 days to 14 days, with all cases strictly within 30 days, thereby ensuring the morphological consistency of the lesions and preventing diagnostic discrepancies caused by disease progression over time. The final diagnosis was divided into two major categories: Cancerous [pancreatic ductal adenocarcinoma (PDAC) and pancreatic acinar cell carcinoma (PDCC)] and noncancerous [neuroendocrine neoplasms (NEN), autoimmune pancreatitis (AIP), and solid pseudopapillaryoma (SPN)]. Patients were included if they underwent a complete EUS examination covering the entire pancreas and had a definitive diagnosis confirmed by surgical patho
To further improve data transparency and allow readers to comprehensively assess the datase’s representativeness, the baseline demographic and clinical characteristics of the 383 included patients are summarized in Table 1. The mean age of the total cohort was 66.7 ± 12.5 years, with comparable age distributions between the Beijing Friendship Hospital (67.2 ± 12.2 years) and the Beijing Tiantan Hospital (64.8 ± 13.1 years). The gender distribution was generally balanced, comprising 202 males (52.7%) and 181 females (47.3%). Regarding clinical characteristics, the dataset maintains a consistent class distribution across both centers. Overall, cancerous lesions accounted for 58.7% (n = 225) of the total cohort, predominantly consisting of (PDAC, n = 219) and a small subset of (PACC, n = 6). Noncancerous lesions made up the remaining 41.3% (n = 158), which included SPN, n = 72, AIP, n = 55, and NENs, n = 31. Importantly, the proportional distributions of these disease subtypes in the Beijing Tiantan Hospital external test cohort (n = 81) closely mirrored those in the Beijing Friendship Hospital development cohort (n = 302), ensuring the representativeness and structural balance of the multicenter data.
| Characteristics | Total cohort (n = 383) | Beijing Friendship Hospital (n = 302) | Beijing Tiantan Hospital (n = 81) |
| Age (years), mean ± SD | 66.7 ± 12.5 | 67.2 ± 12.2 | 64.8 ± 13.1 |
| Gender | |||
| Male | 202 (52.7) | 162 (53.6) | 40 (49.4) |
| Female | 181 (47.3) | 140 (46.4) | 41 (50.6) |
| Class distribution | |||
| Cancerous | 225 (58.7) | 181 (59.9) | 44 (54.3) |
| Noncancerous | 158 (41.3) | 121 (40.1) | 37 (45.7) |
| Lesion subtypes | |||
| Cancerous group | |||
| Pancreatic ductal adenocarcinoma | 219 (57.2) | 176 (58.3) | 43 (53.1) |
| Pancreatic acinar cell carcinoma | 6 (1.6) | 5 (1.7) | 1 (1.2) |
| Noncancerous group | |||
| Solid pseudopapillaryoma | 72 (18.8) | 56 (18.5) | 16 (19.8) |
| Autoimmune pancreatitis | 55 (14.4) | 41 (13.6) | 14 (17.3) |
| Neuroendocrine neoplasms | 31 (8.1) | 24 (7.9) | 7 (8.6) |
To our knowledge, MEUS-PCBL is currently the first publicly reported multicenter EUS dataset consisting of pan
EUS was performed using an Olympus EU-ME2 ultrasound system equipped with an Olympus GF-UCT260 or GF-UCT240 curved linear EUS or a Pentax EG-38-J10UT curved linear EUS. During each patient’s EUS examination, video recordings of the EUS images were made from the insertion point until the removal of the endoscope.
To analyze the differences among the frequency-domain distributions of the EUS images acquired at different centers, this paper used a two-dimensional Fourier transform to map the spatial-domain images to the frequency domain to obtain their spectral energy distribution characteristics. The Fourier transform can effectively represent the structural and texture information of an image at different scales, where the low-frequency components mainly reflect the overall structure and intensity distribution, whereas the high-frequency components correspond to edge details, texture variations, and noise characteristics. On this basis, a single image was used as the statistical unit, and the ratio of high-frequency energy to low-frequency energy was calculated as the frequency-domain feature description to quantify the influences of different imaging centers and imaging conditions on the spectral characteristics of each image. A box plot was subsequently constructed to visualize the distribution of the high-frequency energy proportions determined from different centers and the cancer and noncancer categories to depict the concentration tendencies and dispersion degrees of the samples within different centers. The high/Low-frequency energy ratio distributions of the EUS images derived from different centers are shown in Figure 3, and the corresponding statistical results are summarized in Table 2. As shown in Table 2, significant distribution differences were observed in the proportions of high-frequency energy in the frequency domain among the different imaging centers. Specifically, the overall proportion of high-frequency energy in the images of Center A (Center A_0 and Center A_1) was significantly greater than that in Center B (Center B_0 and Center B_1), with mean values of 0.061 and 0.048, respectively. The corresponding values for Center B were only 0.021 and 0.024. Moreover, the data distribution at Center A was wider, with larger standard deviations and higher maxima (up to 0.56), indicating that the central image was more volatile in terms of its high-frequency details and noise components. In contrast, the proportion of high-frequency energy in Center B was generally lower, and the corresponding distribution was more concentrated, reflecting higher consistency and stability among the spectral characteristics exhibited under imaging conditions.
| Center | Mean | SD | Min | Max |
| Center A_0 | 0.061474 | 0.056234 | 0.008741 | 0.559571 |
| Center A_1 | 0.047896 | 0.040044 | 0.009132 | 0.320019 |
| Center B_0 | 0.021291 | 0.009237 | 0.007514 | 0.047186 |
| Center B_1 | 0.023628 | 0.017217 | 0.009367 | 0.120883 |
Notably, even within the same imaging center, differences between the proportions of high-frequency energy for the cancer (class 0) and noncancer (class 1) samples still existed. For example, in Center A, the mean values of Center A_0 and Center A_1 were 0.061 and 0.048, respectively, and similar interclass shifts were observed in Center B. This suggests that the differences between disease categories can also affect the distribution of frequency-domain features in an image and may lead to the partial overlap of different categories in the frequency-domain space.
In summary, the proportion of high-frequency energy in the frequency domain not only revealed significant distribution shifts between the centers but also reflected potential confounding factors between the categories contained within the same center. This spectral distribution inconsistency, caused by both imaging center differences and disease features, may have induced the model to learn noncausal high-frequency features related to the center, thereby weakening its generalization ability on external datasets. This phenomenon further suggests that in the EUS image analysis task, it is necessary to introduce modeling strategies that can suppress center biases and enhance the robustness of discriminative features.
During the data preprocessing phase, we first labeled and organized all EUS images according to their lesion types. The final diagnosis was divided into two major categories based on clinicopathological results: Cancerous (including PDAC and PDCC) and noncancerous (including NENs, AIP, and solid pseudopapillary neoplasms). This classification was in line with the actual clinical diagnostic process and helped the model learn the imaging differences between cancer and common benign lesions. Due to the multicenter data coming from different hospitals, certain differences were observed in their EUS equipment models, imaging parameters and operating habits, resulting in inconsistencies among the original images in terms of their spatial resolutions, luminance distributions and contrast levels. To reduce the impacts of these factors on the model training process and improve the generalization ability of the model for use with multicenter data, all EUS images needed to undergo uniform data preprocessing schemes before being input into the network. The steps were as follows.
Size uniformity: Considering the resolution differences among the images collected by different devices, to ensure consistency for the input data at the spatial scale and to accommodate the network structure and memory limitations, all EUS images were uniformly scaled to a spatial resolution of 640 × 480 after the noncorrelated areas were cropped out. This operation helped maintain the overall structural information of each image while mitigating the problem of feature shifts caused by size differences.
Data normalization: To mitigate the impacts of brightness and contrast variations on the model training process under different scanning conditions, the maximum-intensity normalization method was used to linearly map the image pixel values to the[0,1] interval. Through normalization processing, the numerical distributions of different samples could be kept consistent, thereby improving the stability of the model training process and accelerating its convergence.
Data augmentation: To simulate the uncertainties exhibited by the probe angles and imaging positions during real clinical operations and effectively expand the number of training samples, multiple data augmentation strategies, including random rotation, random scaling, and random horizontal flipping, were implemented on the images during the training phase. Data augmentation could increase the robustness of the model to scale variations and spatial transformations, thereby reducing the risk of overfitting and enhancing the generalizability of the model. In addition, we removed nine data samples with poor imaging quality levels.
To ensure the objectivity of the model evaluation and prevent potential data leakage, the partitioning of the dataset into training, validation, and internal testing sets was strictly performed at the patient level rather than the image level. Specifically, all EUS images belonging to a single patient were assigned as a unified group to only one of the subsets. This strategy ensures that the model is tested on entirely unseen cases, thereby providing a rigorous assessment of its diagnostic generalization performance across different individuals.
Model architecture: This study proposes a deep learning network for classifying pancreatic cancer and noncancerous lesions in multicenter EUS images, termed MCEUS-C2Net (Multicenter EUS Image Classification Network for Cancerous vs Noncancerous Lesions), as illustrated in Figure 4. The proposed architecture is developed based on MedViT V2[17] with a hierarchical feature extraction design consisting of four stages. Given an input EUS image of size 640 × 480 × 3, stage 1 first performs initial encoding using the local feature extraction (LFE) module and downsamples the feature map to 1/4 resolution (approximately 160 × 120). This stage primarily captures low-level discriminative features such as textures, edges, and local anatomical structures.
To enhance local feature representation, a channel self-attention (CSA) mechanism is introduced after the dilated neighborhood attention (DiNA) block within the LFE module. Specifically, the CSA module adopts a projection strategy based on group convolution with the number of groups equal to the channel dimension[18]. The input feature is projected into channel-wise query Q, key K, and value V representations[19,20].
Channel-wise attention is then computed via matrix multiplication followed by a Softmax operation, enabling the modeling of global inter-channel dependencies. This design allows the network to adaptively recalibrate channel responses by suppressing redundant or irrelevant features (e.g., speckle noise induced by different imaging devices across centers) while emphasizing discriminative texture patterns associated with pancreatic lesions.
The CSA mechanism operates jointly with the sparse global attention mechanism embedded in the DiNA block. While sparse attention expands the effective receptive field within local windows under low computational cost, CSA refines channel-wise feature importance. This combination enables the LFE module to achieve both efficient local detail modeling and enhanced context awareness. Finally, the LFE module concludes with a locally feed-forward network, which further processes the recalibrated spatial features to enhance local representation.
Building upon the enhanced LFE module, stages 2-4 adopt a hierarchical hybrid strategy by progressively stacking LFE blocks with global feature extraction (GFE) modules in an alternating manner. Specifically, the network comprises a total of 40 feature extraction blocks distributed across the four stages (with block depths of[3,4,30,3], respectively). The overall architecture contains approximately 57.77 M parameters, striking an optimal balance between robust hierarchical feature extraction capability and computational efficiency. As the network depth increases, the spatial resolution is progressively reduced to 1/8 (80 × 60), 1/16 (40 × 30), and 1/32 (20 × 15), while the number of feature channels increases accordingly (i.e., 128, 256, and 512) to enhance high-level semantic representation.
The GFE module is built upon an efficient multi-head self-attention mechanism, following the design principle of MedViT V2. It is responsible for modeling long-range dependencies and capturing global contextual relationships across spatial regions. By aggregating cross-region semantic information, the GFE module compensates for the limitations of purely local modeling approaches, particularly in scenarios requiring global structural understanding of pancreatic lesions. Through the stepwise abstraction and fusion of CSA-enhanced local features and globally-aware semantic representations, the network progressively constructs high-level discriminative features. Following this global abstraction, a multi-head convolutional attention module is utilized to further refine these high-level semantic features by leveraging convolutional inductive biases. These features are then fed into the classification head, which utilizes a Kolmogorov-Arnold Network for adaptive feature transformation and aggregation, leading to the final prediction of pancreatic cancer vs noncancerous lesions.
Due to significant variations in imaging devices, acquisition protocols, and patient distributions across centers, mul
Model training: To complete the binary task of classifying pancreatic cancer and noncancerous lesions in EUS, su
The average training loss induced for each epoch is recorded during the model training process to monitor the convergence of the model. To evaluate the classification performance of the model on unseen data, the model is validated in a round-by-round manner during training using the internal validation set. During the validation phase, the model is switched to the evaluation mode, and forward inference is performed on the validation set samples without gradient updates. The predictions produced for all the validation samples are aggregated, processed via a Softmax activation to obtain class probabilities, which were then used to calculate the overall classification performance metrics of the model. The model uses the Adam optimizer for parameter updates, with an initial learning rate set at 1 × 10-5 and a maximum training round count of 800 epochs. This upper limit of 800 epochs was empirically set to provide sufficient optimization space for complete convergence. During training, the model’s performance was continuously monitored on the internal validation set. Instead of early stopping, a dynamic checkpointing strategy was utilized to prevent overfitting: Whenever the model achieved a new optimal area under the curve (AUC) on the validation set, the corresponding model para
To fully evaluate the classification performance of the model in differentiating pancreatic cancer lesions from noncancerous lesions in EUS images, multiple evaluation metrics, including precision, sensitivity, accuracy, specificity, and the F1 score, are used in this study[21]. These indicators reflect the discriminative ability of the model in clinically relevant scenarios from different perspectives and thus allow a more objective and comprehensive analysis of the performance of the model.
Sensitivity is used to measure the ability of the model to detect cases of pancreatic cancer, that is, the proportion of real pancreatic cancer samples that are correctly identified. In clinical applications, a missed diagnosis of pancreatic cancer can lead to delayed treatment for a patient; thus, increased sensitivity is important for assisting clinical decision-making results. Specificity is used to assess the ability of the model to correctly identify noncancerous lesions such as NENs and AIP, reflecting its ability to reduce the numbers of misdiagnoses and overdiagnoses, and is equally crucial for avoiding unnecessary invasive tests and treatments.
Precision measures the proportion of samples predicted as pancreatic cancer by the model that are actually cancer and is used to reflect the reliability of the prediction results produced by the model. Especially in the cancer screening scenario, high precision helps to reduce the false-positive rate. The accuracy, which reflects the overall classification accuracy achieved for the model over all samples, can visually assess the overall discriminative performance of the model, but certain limitations may be encountered when this metric alone is used in cases where the category distribution is unbalanced.
In addition, the F1 score, which strikes a balance between precision and sensitivity, provides a more comprehensive reflection of the overall performance of the model in the pancreatic cancer identification task. By combining the above multiple evaluation metrics, the performance of the model is assessed from multiple clinically relevant dimensions in this paper to ensure the stability and reliability of the model on multicenter EUS data.
To rigorously evaluate the classification performance and reliability of the models, 95% confidence interval (CI) for all quantitative evaluation metrics were calculated using the bootstrapping method with 1000 resamples. Furthermore, to assess the statistical significance of the performance differences between the proposed MCEUS-C2Net and the com
To further analyze the interpretability of the decision-making process of the model and to validate the rationality of the regions the model focuses on, we conduct a qualitative visualization analysis of the attention heatmap generated by the model. Specifically, a representative EUS image is selected from the test set to show its original image, the attention heatmap generated by the model, and the fusion results of both outputs. In the attention heatmap, colors ranging from blue to red indicate the attention paid by the model to different spatial regions from low to high.
In accordance with the exclusion criteria, 176 patients (158 without video images and 18 with poor image quality) are excluded; ultimately, 383 patients are included in this study. As part of the development cohort, 302 patients are included and randomly divided into a training cohort, a validation cohort and a test cohort (with 182 patients in the training cohort, 60 in the validation cohort and 60 in the test cohort). In the external test cohort, 81 patients are included.
To fully evaluate how well MCEUS-C2Net distinguishes between pancreatic cancer lesions from noncancerous lesions in EUS images and to validate its stability and robustness in real-world clinical applications, both internal and external test experiments are conducted. The internal test is based on a data distribution that is consistent with the source of the training data and is used to evaluate the discriminative ability of the model under known distribution conditions; the external tests use data from derived different centers as independent.
Validation sets to simulate the data distribution differences that may be encountered in a real clinical deployment environment, thereby further evaluating the generalization performance of the model. By comparing the performances achieved by the model on the internal and external test sets, a more comprehensive reflection of its suitability for multicenter EUS data can be achieved.
In addition to quantitative evaluation metrics, visual analysis methods are introduced in this paper to enhance the interpretability of the prediction results yielded by the model. Specifically, attention heatmaps are used to visualize the key regions that the network focuses on during the classification process. The attention heatmaps visually reflect the image regions that the model focuses on when making classification decisions by projecting the intermediate feature mapping of the model back into the original image space. This method helps to analyze whether the model focuses on the anatomical structures and lesion areas that are associated with pancreatic lesions, thereby verifying whether the decision-making basis of the model is in line with clinical imaging cognition to a certain extent.
In addition, a confusion matrix is used to analyze the classification results produced by the model on the test set. The confusion matrix can visually show the predictive relationship between pancreatic cancer and noncancerous lesions, including the specific cases of correct classifications and misclassifications. Through a visual analysis of the confusion matrix, it is possible to further identify the strengths and weaknesses of the model for different categories and clarify the main sources of misdiagnoses and missed diagnoses.
After the model training process is completed, we select four representative classification models for conducting comparative experiments, including ResNet-50[22], a classic deep convolutional neural network that effectively alleviates vanishing gradients by introducing residual connections and is widely used in various medical image classification tasks. The Swin transformer[23] is a transformer network based on a hierarchical window self-attention mechanism that can model the global and local context information of images while maintaining computational efficiency. MedViTV2[23] is a lightweight visual transformer model that is optimized for medical imaging tasks and combines convolution and self-attention mechanisms to achieve excellent classification performance while maintaining low computational overhead. MCEUS-C2Net is the approach proposed in this paper. To comprehensively evaluate the classification performance of the models and their generalizability under multicenter data conditions, the above methods are systematically compared on internal validation sets and external independent datasets, and the quantitative results are summarized in Tables 3 and 4, respectively. In the internal validation experiments, the test data are derived from the same source as that of the training data, and the classification performance of the comparison models is relatively ideal (Table 2). Among them, MCEUS-C2Net achieves the best results in terms of precision, sensitivity, accuracy, specificity and the F1 score. By analyzing the 95%CI, we observe that MCEUS-C2Net exhibits the narrowest interval ranges and the highest lower bounds across all metrics, indicating highly stable and reliable predictive ability. To control for the false discovery rate (FDR) during multiple comparisons, the Benjamini-Hochberg correction was applied to the P values of McNemar’s tests. On the internal dataset, MCEUS-C2Net showed statistically significant improvements over ResNet-50 (P = 0.021, P < 0.05), Swin transformer (P = 0.024, P < 0.05), and MedViTV2 (P = 0.043, P < 0.05).
| Methods | Precision (%) (95%CI)↑ | Sensitivity (%) (95%CI)↑ | Accuracy (%) (95%CI)↑ | Specificity (%) (95%CI)↑ | F1 score (%) (95%CI)↑ |
| ResNet-50a | 86.96 (82.17-90.95) | 94.14 (91.47-97.49) | 89.12 (86.41-92.05) | 82.76 (77.43-88.30) | 90.50 (87.76-93.16) |
| Swin_transformera | 89.29 (83.64-92.31) | 94.34 (90.48-96.97) | 90.53 (86.92-92.82) | 85.71 (79.89-90.64) | 91.74 (88.02-93.65) |
| MedViTV2a | 90.09 (85.32-93.24) | 95.24 (91.54-97.66) | 91.50 (87.95-93.85) | 86.75 (81.71-91.33) | 92.59 (89.20-94.48) |
| MCEUS-C2Net (ours) | 93.46 (89.08-95.97) | 97.09 (95.48-99.51) | 94.51 (92.31-96.67) | 91.14 (86.70-94.92) | 95.24 (93.01-97.01) |
| Methods | Precision (%) (95%CI)↑ | Sensitivity (%) (95%CI)↑ | Accuracy (%) (95%CI)↑ | Specificity (%) (95%CI)↑ | F1 score (%) (95%CI)↑ |
| ResNet-50b | 84.57 (79.00-89.88) | 74.00 (67.88-80.49) | 78.65 (74.32-82.97) | 84.12 (78.31-89.51) | 78.93 (74.13-83.42) |
| Swin_transformerb | 86.63 (81.87-91.24) | 81.00 (75.26-86.26) | 82.97 (78.92-86.49) | 85.29 (79.88-90.30) | 83.72 (79.49-87.35) |
| MedViTV2a | 86.57 (82.03-91.35) | 87.00 (82.16-91.67) | 85.68 (82.16-89.19) | 84.12 (78.33-89.54) | 86.78 (83.24-90.10) |
| MCEUS-C2Net (ours) | 91.88 (87.75-95.26) | 90.50 (86.47-94.23) | 90.59 (87.84-93.51) | 90.59 (85.98-94.51) | 91.18 (88.32-93.77) |
Specifically, in the external validation experiments, the test data come from independent datasets that do not par
To further analyze the misclassification rate of each method on the external datasets, the corresponding confusion matrix is shown in Figure 5, where T represents the pancreatic cancer category and F represents the noncancerous lesion category. ResNet-50 has the greatest number of misclassified samples, with 52 and 27 false-positive and false-negative samples, respectively, suggesting that this method has a relatively high risk of misdiagnoses and missed diagnoses in external data. The Swin transformer improves these results in terms of false positives but still yields 38 and 25 false-positive and false-negative samples, respectively. The number of false positives yielded by MedViTV2 is further reduced to 26, but the number of false negatives remains 27, and the overall number of misclassified samples is still relatively high. In contrast, MCEUS-C2Net has the fewest misclassified samples, with only 19 false positives and 16 false negatives, indicating a more balanced performance in terms of reducing misdiagnoses and missed diagnoses and a more stable and reliable discriminative ability on external independent datasets. To go beyond descriptive statistics and understand the clinical implications, an error pattern analysis was conducted on the misclassified samples based on our specific pathological categorizations. A detailed review revealed two primary error modes. First, the false-positive cases (noncancerous lesions misclassified as cancer) were predominantly associated with AIP and atypical NENs. AIP, in particular, often presents as a focal hypoechoic, heterogeneous mass mimicking the morphological characteristics of PDAC, which frequently confused the baseline models. Second, the false-negative cases (cancers misclassified as noncancerous) mainly involved well-circumscribed PDCC or early-stage PDACs lacking classic invasive textural features. The well-defined borders of PACC can mimic the appearance of SPN or NENs, leading to algorithmic misjudgments. Notably, MCEUS-C2Net effectively mitigated these specific clinical error patterns. To quantitatively evaluate the classification performance of MCEUS-C2Net, a further analysis is conducted to determine its ability to distinguish pancreatic cancer lesions from noncancerous lesions in EUS images, and AUC is calculated. The AUC, training loss curve, and ROC curve produced for the external validation set of the model are shown in Figure 6, with the corresponding AUC values reaching 0.9704. The AUC values of the baseline models (ResNet-50, Swin transformer, and MedViTV2) on the external dataset were all lower than that of MCEUS-C2Net. The calibration curves and decision curve analysis (DCA) results of the models on the external validation set are shown in Figure 7. As observed in Figure 7A, the prediction curve of MCEUS-C2Net aligns most closely with the ideal diagonal line and exhibits the highest linearity, indicating optimal calibration among the compared models. Furthermore, the DCA curves in Figure 7B reveal that across a wide range of clinical risk thresholds, MCEUS-C2Net consistently yields the highest net clinical benefit, demonstrating relatively optimal clinical utility compared to the traditional baseline models. The above results fully validate the superior performance and good generalization ability of MCEUS-C2Net on this task.
To systematically validate the independent contributions of the pure GFE, pure LFE, and the CSA mechanism, and to explicitly distinguish our proposed architecture from the baseline MedViTV2, we conducted comprehensive ablation studies by decoupling the network.
On the internal validation set (Table 5), the pure GFE and pure LFE modules achieved baseline accuracies of 86.41% and 88.46%, respectively. Notably, the integration of the CSA module with LFE (pure LFE + CSA) independently improved the accuracy to 89.49%. This demonstrates the efficacy of the CSA mechanism in refining local spatial features and suppressing noise, even in the absence of global context. Ultimately, our complete MCEUS-C2Net achieved the highest accuracy of 94.51%, outperforming the baseline LFE + GFE (91.50%). Furthermore, MCEUS-C2Net exhibited the narrowest 95%CI among all configurations, indicating highly stable and reliable predictive capability.
| Methods | Precision (%) (95%CI)↑ | Sensitivity (%) (95%CI)↑ | Accuracy (%) (95%CI)↑ | Specificity (%) (95%CI)↑ | F1 score (%) (95%CI)↑ |
| GFEb | 86.85 (82.25-91.39) | 88.10 (83.76-92.56) | 86.41 (82.82-89.74) | 84.44 (79.28-89.64) | 87.47 (83.99-90.70) |
| LFEb | 88.73 (84.46-92.72) | 90.00 (85.92-93.97) | 88.46 (85.13-91.54) | 86.67 (81.82-91.33) | 89.36 (86.15-92.38) |
| LFE + CSAb | 90.05 (85.86-93.91) | 90.48 (86.57-94.45) | 89.49 (86.41-92.56) | 88.33 (83.60-92.90) | 90.26 (87.12-93.15) |
| LFE + GFEa | 90.09 (85.32-93.24) | 95.24 (91.54-97.66) | 91.50 (87.95-93.85) | 86.75 (81.71-91.33) | 92.59 (89.20-94.48) |
| MCEUS-C2Net (ours) | 93.46 (89.08-95.97) | 97.09 (95.48-99.51) | 94.51 (92.31-96.67) | 91.14 (86.70-94.92) | 95.24 (93.01-97.01) |
The necessity and complementary nature of each module are most pronounced in the external cohort (Table 6), where severe speckle noise and domain shifts are present. Relying solely on either the pure GFE or pure LFE led to drastic performance degradation, yielding external accuracies of only 78.11% and 80.27%, respectively, accompanied by no
| Methods | Precision (%) (95%CI)↑ | Sensitivity (%) (95%CI)↑ | Accuracy (%) (95%CI)↑ | Specificity (%) (95%CI)↑ | F1 score (%) (95%CI)↑ |
| GFEb | 83.62 (78.15-89.10) | 74.00 (67.88-80.49) | 78.11 (73.78-82.43) | 82.94 (77.06-88.54) | 78.51 (73.63-82.95) |
| LFEb | 86.71 (81.39-92.03) | 75.00 (68.93-81.42) | 80.27 (75.95-84.32) | 86.47 (81.03-91.47) | 80.43 (75.80-84.85) |
| LFE + CSAb | 87.10 (82.18-92.19) | 81.00 (75.50-86.46) | 83.24 (79.18-87.30) | 85.88 (80.46-91.14) | 83.94 (79.77-87.77) |
| LFE + GFEa | 86.57 (82.03-91.35) | 87.00 (82.16-91.67) | 85.68 (82.16-89.19) | 84.12 (78.33-89.54) | 86.78 (83.24-90.10) |
| MCEUS-C2Net (ours) | 91.88 (87.75-95.26) | 90.50 (86.47-94.23) | 90.59 (87.84-93.51) | 90.59 (85.98-94.51) | 91.18 (88.32-93.77) |
The original images, attention heatmaps and fusion results of ResNet-50, the Swin transformer, MedViTV2 and the MCEUS-C2Net model proposed in this paper are shown in Figure 8. As shown in the figure, the high-response regions of MCEUS-C2Net are concentrated mainly at the locations of pancreatic lesions, which are highly consistent with the areas that clinicians focus on during the actual diagnosis process. Notably, this pattern of concern does not rely on any artificially labeled lesion areas but is rather autonomously learned by the model during end-to-end training. These findings suggest that the proposed approach can effectively focus on the key regions that are closely related to the discrimination between pancreatic cancer and noncancerous lesions, providing good interpretability support for the decision-making process of the model. This phenomenon is attributed mainly to the channel attention mechanism introduced in the model. By adaptively weighting feature responses in the channel dimension, this mechanism enhances the discriminative feature expressions that are related to pancreatic lesions while suppressing interference from redundant or background in
With the development of AI, especially deep learning[23], an increasing number of works have demonstrated the out
In this study, a multicenter dataset with significant distribution differences based on EUS images acquired from two hospitals, different imaging devices, and different patient sources was constructed, and the generalization performance of deep learning models in the task of automatically classifying pancreatic cancer and noncancerous lesions was systematically evaluated. The experimental results revealed that the proposed MCEUS-C2Net approach achieved stable and excellent classification performance in both an internal validation and an external validation conducted on an in
Previous studies have also used deep learning architectures to develop AI systems for diagnosing pancreatic diseases in EUS images[27,28], but none of these systems have studied cross-center data, and multicenter EUS data in actual clinical settings possess significant distribution differences. An energy spectrum analysis of multicenter pancreatic EUS images revealed significant spectral distribution and imaging characteristic differences among different centers, mainly due to acquisition equipment difference, different scanning parameter settings, and changes in the compositions of the patient population. Such distribution differences often lead to higher model performance on single-center data but a significant decline in performance on external datasets, as evidenced in previous studies and the comparative ex
In contrast, the declines exhibited by the various performance indicators of MCEUS-C2Net on the external validation set were significantly smaller than those of the comparison methods (Table 3). This suggests that the model has potential to help standardize the interpretation of EUS images in real endoscopy suites, which is particularly valuable for less-experienced endoscopists when facing complex clinical scenarios. These results indicate that the proposed method can maintain a more stable discriminative ability when addressing unseen data distributions. This advantage is due mainly to the structural design of the model, which combines both LFE and global context modeling capabilities and introduces a channel attention mechanism for the adaptive reweighting of feature representations. By strengthening the discriminative features that are related to pancreatic lesions in the channel dimension and suppressing redundant or irrelevant in
This study shows that MCEUS-C2Net does not overly rely on the feature patterns acquired at specific centers or under specific imaging conditions but rather learns more robust and discriminative feature representations in complex mul
First, although multicenter data acquired from two hospitals were introduced in this study, the external validation remains limited as it relies on only a single external center. Additionally, the overall sample size was still relatively limited. The acquisition and standardized labeling processes employed for multicenter EUS data are difficult to execute in actual clinical practice, and the sample size limitation may have affected the stability of the obtained statistical results to some extent. Future studies will further expand the data size through the accumulation of data over longer time spans and multi-institutional collaboration to obtain more robust and representative assessment results. Another limitation relates to potential selection bias and spectrum bias. Currently, achieving high classification accuracy across multicenter data while simultaneously accounting for suboptimal, low-quality images from actual clinical practice remains a sig
Furthermore, while MCEUS-C2Net demonstrates promising retrospective performance and robust cross-center generalization, transitioning to real-world clinical deployment faces practical challenges. First, EUS is inherently operator-dependent; inconsistencies in scanning methods, probe angles, and equipment settings introduce significant image heterogeneity. In fact, addressing this exact clinical pain point was the core motivation behind our multicenter design, and our model has indeed proven its effectiveness in mitigating such cross-center variance at the algorithmic level. However, for optimal real-world deployment, relying solely on algorithmic generalization is insufficient. Esta
In this study, the first multicenter EUS pancreatic disease imaging dataset, MEUS-PCBL, was constructed, and a classification network, MCEUS-C2Net, was proposed for multicenter data. The network effectively mitigated the generalization performance degradation caused by the distribution differences exhibited by multicenter data. Furthermore, in cross-center tests, it demonstrated excellent robustness as a potentially applicable tool compared with existing AI models. This model could help standardize EUS reading across different hospitals and may assist clinicians in making more accurate diagnoses. Future work will further expand the data scale and integrate pixel-level annotation to enhance the quantitative interpretability validation of the model.
| 1. | He R, Jiang W, Wang C, Li X, Zhou W. Global burden of pancreatic cancer attributable to metabolic risks from 1990 to 2019, with projections of mortality to 2030. BMC Public Health. 2024;24:456. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 33] [Cited by in RCA: 30] [Article Influence: 15.0] [Reference Citation Analysis (0)] |
| 2. | Almasri B, Ali A. Role of endoscopic ultrasound elastography in differential diagnosis of pancreatic solid masses. Qatar Med J. 2021;2021:40. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 1] [Cited by in RCA: 2] [Article Influence: 0.4] [Reference Citation Analysis (0)] |
| 3. | Vitali F, Zundler S, Jesper D, Wildner D, Strobel D, Frulloni L, Neurath MF. Diagnostic Endoscopic Ultrasound in Pancreatology: Focus on Normal Variants and Pancreatic Masses. Visc Med. 2023;39:121-130. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 6] [Cited by in RCA: 6] [Article Influence: 2.0] [Reference Citation Analysis (0)] |
| 4. | Zhang S, Ni M, Wang P, Zheng J, Sun Q, Xu G, Peng C, Shen S, Zhang W, Huang S, Wang L, Zou X, Lv Y. Diagnostic value of endoscopic ultrasound-guided fine needle aspiration with rapid on-site evaluation performed by endoscopists in solid pancreatic lesions: A prospective, randomized controlled trial. J Gastroenterol Hepatol. 2022;37:1975-1982. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 21] [Cited by in RCA: 20] [Article Influence: 5.0] [Reference Citation Analysis (0)] |
| 5. | Gheorghiu M, Sparchez Z, Rusu I, Bolboacă SD, Seicean R, Pojoga C, Seicean A. Direct Comparison of Elastography Endoscopic Ultrasound Fine-Needle Aspiration and B-Mode Endoscopic Ultrasound Fine-Needle Aspiration in Diagnosing Solid Pancreatic Lesions. Int J Environ Res Public Health. 2022;19:1302. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 10] [Cited by in RCA: 10] [Article Influence: 2.5] [Reference Citation Analysis (1)] |
| 6. | Puga-Tejada M, Del Valle R, Oleas R, Egas-Izquierdo M, Arevalo-Mora M, Baquerizo-Burgos J, Ospina J, Soria-Alcivar M, Pitanga-Lukashok H, Robles-Medranda C. Endoscopic ultrasound elastography for malignant pancreatic masses and associated lymph nodes: Critical evaluation of strain ratio cutoff value. World J Gastrointest Endosc. 2022;14:524-535. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in CrossRef: 6] [Cited by in RCA: 6] [Article Influence: 1.5] [Reference Citation Analysis (0)] |
| 7. | Conti CB, Mulinacci G, Salerno R, Dinelli ME, Grassia R. Applications of endoscopic ultrasound elastography in pancreatic diseases: From literature to real life. World J Gastroenterol. 2022;28:909-917. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in CrossRef: 14] [Cited by in RCA: 10] [Article Influence: 2.5] [Reference Citation Analysis (0)] |
| 8. | Kuwahara T, Hara K, Mizuno N, Haba S, Okuno N, Fukui T, Urata M, Yamamoto Y. Current status of artificial intelligence analysis for the treatment of pancreaticobiliary diseases using endoscopic ultrasonography and endoscopic retrograde cholangiopancreatography. DEN Open. 2024;4:e267. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 17] [Cited by in RCA: 14] [Article Influence: 7.0] [Reference Citation Analysis (0)] |
| 9. | Udriștoiu AL, Cazacu IM, Gruionu LG, Gruionu G, Iacob AV, Burtea DE, Ungureanu BS, Costache MI, Constantin A, Popescu CF, Udriștoiu Ș, Săftoiu A. Real-time computer-aided diagnosis of focal pancreatic masses from endoscopic ultrasound imaging based on a hybrid convolutional and long short-term memory neural network model. PLoS One. 2021;16:e0251701. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 53] [Cited by in RCA: 46] [Article Influence: 9.2] [Reference Citation Analysis (0)] |
| 10. | Kenner B, Chari ST, Kelsen D, Klimstra DS, Pandol SJ, Rosenthal M, Rustgi AK, Taylor JA, Yala A, Abul-Husn N, Andersen DK, Bernstein D, Brunak S, Canto MI, Eldar YC, Fishman EK, Fleshman J, Go VLW, Holt JM, Field B, Goldberg A, Hoos W, Iacobuzio-Donahue C, Li D, Lidgard G, Maitra A, Matrisian LM, Poblete S, Rothschild L, Sander C, Schwartz LH, Shalit U, Srivastava S, Wolpin B. Artificial Intelligence and Early Detection of Pancreatic Cancer: 2020 Summative Review. Pancreas. 2021;50:251-279. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 137] [Cited by in RCA: 95] [Article Influence: 19.0] [Reference Citation Analysis (0)] |
| 11. | Daher H, Punchayil SA, Ismail AAE, Fernandes RR, Jacob J, Algazzar MH, Mansour M. Advancements in Pancreatic Cancer Detection: Integrating Biomarkers, Imaging Technologies, and Machine Learning for Early Diagnosis. Cureus. 2024;16:e56583. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 19] [Cited by in RCA: 8] [Article Influence: 4.0] [Reference Citation Analysis (0)] |
| 12. | Sijithra PC, Santhi N, Ramasamy N. A review study on early detection of pancreatic ductal adenocarcinoma using artificial intelligence assisted diagnostic methods. Eur J Radiol. 2023;166:110972. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 15] [Cited by in RCA: 11] [Article Influence: 3.7] [Reference Citation Analysis (0)] |
| 13. | Cui H, Zhao Y, Xiong S, Feng Y, Li P, Lv Y, Chen Q, Wang R, Xie P, Luo Z, Cheng S, Wang W, Li X, Xiong D, Cao X, Bai S, Yang A, Cheng B. Diagnosing Solid Lesions in the Pancreas With Multimodal Artificial Intelligence: A Randomized Crossover Trial. JAMA Netw Open. 2024;7:e2422454. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 33] [Cited by in RCA: 40] [Article Influence: 20.0] [Reference Citation Analysis (1)] |
| 14. | Fang YJ, Lee KH, Karmakar R, Mukundan A, Nagisetti Y, Huang CW, Wang HC. Transforming Endoscopic Image Classification with Spectrum-Aided Vision for Early and Accurate Cancer Identification. Diagnostics (Basel). 2025;15:2732. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 2] [Cited by in RCA: 5] [Article Influence: 5.0] [Reference Citation Analysis (0)] |
| 15. | Chou CK, Lee KH, Karmakar R, Mukundan A, Gade PC, Gupta D, Su CC, Chen TH, Ko CY, Wang HC. Emulating Hyperspectral and Narrow-Band Imaging for Deep-Learning-Driven Gastrointestinal Disorder Detection in Wireless Capsule Endoscopy. Bioengineering (Basel). 2025;12:953. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 1] [Cited by in RCA: 7] [Article Influence: 7.0] [Reference Citation Analysis (0)] |
| 16. | Weng WC, Huang CW, Su CC, Mukundan A, Karmakar R, Chen TH, Avhad AR, Chou CK, Wang HC. Optimizing Esophageal Cancer Diagnosis with Computer-Aided Detection by YOLO Models Combined with Hyperspectral Imaging. Diagnostics (Basel). 2025;15:1686. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 8] [Cited by in RCA: 14] [Article Influence: 14.0] [Reference Citation Analysis (0)] |
| 17. | Manzari ON, Ahmadabadi H, Kashiani H, Shokouhi SB, Ayatollahi A. MedViT: A robust vision transformer for generalized medical image classification. Comput Biol Med. 2023;157:106791. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 376] [Cited by in RCA: 137] [Article Influence: 45.7] [Reference Citation Analysis (0)] |
| 18. | Nejati Manzari O, Asgariandehkordi H, Koleilat T, Xiao Y, Rivaz H. Medical image classification with KAN-integrated transformers and dilated neighborhood attention. Appl Soft Comput. 2026;186:114045. [DOI] [Full Text] |
| 19. | He K, Gan C, Li Z, Rekik I, Yin Z, Ji W, Gao Y, Wang Q, Zhang J, Shen D. Transformers in medical image analysis. Intelligent Medicine. 2023;3:59-78. [DOI] [Full Text] |
| 20. | Zhao H, Gou Y, Li B, Peng D, Lv J, Peng X. Comprehensive and Delicate: An Efficient Transformer for Image Restoration. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE, 2023: 14122-14132. [DOI] [Full Text] |
| 21. | Yue Y, Li Z. Medmamba: Vision mamba for medical image classification. 2024 Preprint. Available from: arXiv:2403.03849. [DOI] [Full Text] |
| 22. | He K, Zhang X, Ren S, Sun J. Deep Residual Learning for Image Recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas: IEEE, 2016: 770-778. [DOI] [Full Text] |
| 23. | Balikov DA, Hu K, Liu CJ, Betz BL, Chinnaiyan AM, Devisetty LV, Venneti S, Tomlins SA, Cani AK, Rao RC. Comparative Molecular Analysis of Primary Central Nervous System Lymphomas and Matched Vitreoretinal Lymphomas by Vitreous Liquid Biopsy. Int J Mol Sci. 2021;22:9992. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 13] [Cited by in RCA: 16] [Article Influence: 3.2] [Reference Citation Analysis (0)] |
| 24. | Kuwahara T, Hara K, Mizuno N, Haba S, Okuno N, Koda H, Miyano A, Fumihara D. Current status of artificial intelligence analysis for endoscopic ultrasonography. Dig Endosc. 2021;33:298-305. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 45] [Cited by in RCA: 40] [Article Influence: 8.0] [Reference Citation Analysis (2)] |
| 25. | Wang H, Ni D, Wang Y. Recursive Deformable Pyramid Network for Unsupervised Medical Image Registration. IEEE Trans Med Imaging. 2024;43:2229-2240. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 92] [Cited by in RCA: 33] [Article Influence: 16.5] [Reference Citation Analysis (0)] |
| 26. | Zhou Y, Chia MA, Wagner SK, Ayhan MS, Williamson DJ, Struyven RR, Liu T, Xu M, Lozano MG, Woodward-Court P, Kihara Y; UK Biobank Eye & Vision Consortium, Altmann A, Lee AY, Topol EJ, Denniston AK, Alexander DC, Keane PA. A foundation model for generalizable disease detection from retinal images. Nature. 2023;622:156-163. [RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)] [Cited by in Crossref: 701] [Cited by in RCA: 534] [Article Influence: 178.0] [Reference Citation Analysis (2)] |
| 27. | Tonozuka R, Itoi T, Nagata N, Kojima H, Sofuni A, Tsuchiya T, Ishii K, Tanaka R, Nagakawa Y, Mukai S. Deep learning analysis for the detection of pancreatic cancer on endosonographic images: a pilot study. J Hepatobiliary Pancreat Sci. 2021;28:95-104. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 112] [Cited by in RCA: 89] [Article Influence: 17.8] [Reference Citation Analysis (1)] |
| 28. | Zhang J, Zhu L, Yao L, Ding X, Chen D, Wu H, Lu Z, Zhou W, Zhang L, An P, Xu B, Tan W, Hu S, Cheng F, Yu H. Deep learning-based pancreas segmentation and station recognition system in EUS: development and validation of a useful training tool (with video). Gastrointest Endosc. 2020;92:874-885.e3. [RCA] [PubMed] [DOI] [Full Text] [Cited by in Crossref: 94] [Cited by in RCA: 83] [Article Influence: 13.8] [Reference Citation Analysis (3)] |