BPG is committed to discovery and dissemination of knowledge
Retrospective Study Open Access
Copyright: ©Author(s) 2026. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution-NonCommercial (CC BY-NC 4.0) license. No commercial re-use. See permissions. Published by Baishideng Publishing Group Inc.
World J Gastroenterol. Aug 28, 2026; 32(32): 120382
Published online Aug 28, 2026. doi: 10.3748/wjg.120382
Automatic recognition of tumour-infiltrating lymphocytes in pathological biopsy images of the gastric mucosa
Yu Fan, Department of Pathology, Shaanxi Provincial Hospital of Traditional Chinese Medicine, Xi’an 710028, Shaanxi Province, China
Su-Nan Wang, Ying-Ying Li, Shenzhen Polytechnic University, Shenzhen 518055, Guangdong Province, China
Bo Jiang, Department of Pathology, People’s Liberation Army Joint Logistic Support Force 990th Hospital, Zhumadian 463000, Henan Province, China
Chao-Ya Zhu, Department of Pathology, Third Affiliated Hospital, Zhengzhou University, Zhengzhou 450052, Henan Province, China
Xing-Hai Liao, Department of Surgery, Southern Medical University Shenzhen Hospital, Shenzhen 518110, Guangdong Province, China
Fa-Shun Zhang, Department of Pathology, Xuchang Central Hospital, Xuchang 461099, Henan Province, China
Yang-Kun Wang, Department of Pathology, The Fourth People’s Hospital of Longgang District, Shenzhen 518123, Guangdong Province, China
ORCID number: Yu Fan (0009-0003-3137-9017); Su-Nan Wang (0000-0002-1836-1360); Bo Jiang (0000-0001-9301-5567); Ying-Ying Li (0000-0002-1971-9985); Chao-Ya Zhu (0000-0002-4421-7514); Xing-Hai Liao (0009-0001-2566-1978); Fa-Shun Zhang (0009-0009-6769-4848); Yang-Kun Wang (0000-0002-9970-3093).
Co-first authors: Yu Fan and Su-Nan Wang.
Author contributions: Wang YK conceived and designed the study; Fan Y ,Liao XH and Jiang B collected data, sorted pathological samples, and drafted the manuscript; Wang SN and Fan Y constructed the model, optimized the algorithm, implemented ablation experiments, and analyzed data; Jiang B and Fan Y performed pathological image annotation, established ground truth, and verified interobserver consistency; Li YY and Fan Y conducted image preprocessing including regions of interest extraction and colour deconvolution, and constructed the dataset; Zhu CY, Fan Y and Zhang FS collated clinical follow-up data and performed survival analysis of early gastric cancer patients; Liao XH , Fan Y and Zhang FS collected multicentre samples and confirmed clinical information; Zhang FS and Fan Y performed pathological diagnosis of gastric mucosal lesions and conducted double-blind verification of sample types; Wang YK acquired funding, revised the manuscript, and gave final approval of the version to be published. Fan Y and Wang SN contributed equally to this work as co-first authors.
Supported by Shenzhen Basic Research Special Natural Science Foundation Project, No. JCYJ202506044185911015.
Institutional review board statement: Approval from the Ethics Committee of Shaanxi Provincial Academy of Traditional Chinese Medicine and Shaanxi Provincial Hospital of Traditional Chinese Medicine, Exemption Approval No. (2023) Lunshenmian No. (15), date: January 10, 2023. The study complied with the Declaration of Helsinki and China’s Measures for Ethical Review of Life Science and Medical Research Involving Human Subjects.
Informed consent statement: Written informed consent was obtained from all participants.
Conflict-of-interest statement: The authors declare no conflicts of interest.
Data sharing statement: All data generated or analyzed during this study are included in this published article.
Corresponding author: Yang-Kun Wang, Department of Pathology, The Fourth People’s Hospital of Longgang District, No. 2 Jinjian Road, Nanwan Subdistrict, Longgang District, Shenzhen 518123, Guangdong Province, China. dr.wyk@163.com
Received: February 26, 2026
Revised: March 12, 2026
Accepted: April 21, 2026
Published online: August 28, 2026
Processing time: 160 Days and 17.5 Hours

Abstract
BACKGROUND

The immune landscape of the tumour microenvironment, particularly tumour-infiltrating lymphocytes (TILs), is pivotal in the progression of gastric mucosal lesions. However, conventional manual assessment of TILs is hindered by interobserver variability, poor reproducibility, and labour-intensive quantification, restricting its clinical utility.

AIM

To develop an artificial intelligence (AI)-based metric, the gastric-AI-TIL (G-AI-TIL) index, for precise lesion grading and prognostic stratification in early gastric cancer (EGC).

METHODS

We retrospectively collected 320 whole-slide images of gastric mucosal biopsy samples from three medical centres. A multiscale two-stage convolutional neural network (CNN), comprising a gastric-CNN for lesion segmentation and a G-TIL-CNN for TIL enumeration, was constructed. Image preprocessing involved region of interest extraction and Ruifrok-Johnston colour deconvolution to normalize staining variations. Two senior pathologists performed the annotations in a double-blind manner to establish the ground truth. The prognostic value of the G-AI-TIL was evaluated using Kaplan-Meier and multivariate Cox regression analyses.

RESULTS

In the independent test set (n = 96), the G-TIL-CNN model achieved an accuracy of 99.2% (kappa = 0.98), significantly outperforming manual evaluation by pathologists (86.3% accuracy; P < 0.001). A progressive increase in the G-AI-TIL index correlated with lesion malignancy, as follows: Chronic atrophic gastritis < intestinal metaplasia < high-grade intraepithelial neoplasia < EGC (P < 0.001). Furthermore, a high G-AI-TIL (≥ 28.5%) was identified as an independent protective factor for both 3-year disease-free survival [hazard ratio (HR) = 0.58] and overall survival (HR = 0.55; P = 0.009) in patients with EGC.

CONCLUSION

The proposed AI model enables objective, high-throughput quantification of TILs. The G-AI-TIL index serves as a robust biomarker for grading gastric mucosal lesions and stratifying EGC risk, providing a quantitative basis for personalized treatment strategies, such as the decision to perform endoscopic resection and surgery.

Key Words: Artificial intelligence; Gastric cancer; Tumour-infiltrating lymphocytes; Deep learning; Prognosis; Digital pathology

Core Tip: This study retrospectively collected 320 whole-slide images of gastric mucosal biopsies and constructed a two-stage convolutional neural network model, which underwent multi-step preprocessing and annotated training. The verification results showed the model achieved a 99.2% accuracy in tumour-infiltrating lymphocytes (TIL) identification. It was found that the gastric-artificial intelligence-TIL (G-AI-TIL) index increases with the elevated malignancy of gastric mucosal lesions, and a high G-AI-TIL index acts as an independent protective factor for the survival of early gastric cancer patients. This research provides a novel objective tool for the precise diagnosis and treatment of gastric cancer.



INTRODUCTION

The malignant progression of gastric mucosal lesions is a process regulated in multiple stages by multiple factors that follow the classic pathway of “chronic atrophic gastritis (CAG) → intestinal metaplasia (IM) → intraepithelial neoplasia → gastric cancer”. Among them, the immune status of the tumour microenvironment is the core link in regulating the malignant transformation of lesions[1]. As key indicators of the intensity of the local antitumour immune response, tumour-infiltrating lymphocytes (TILs) are closely related to the activity and prognosis of gastric mucosal lesions. In CAG, mild infiltration of TILs indicates that inflammation is controllable, but if it persists, it may promote the progression of mucosal atrophy; in high-grade intraepithelial neoplasia (HGIN), an increase in TIL density is often accompanied by the activation of immune surveillance of abnormal cells, which may delay the transformation of lesions into cancer, and in early-stage gastric cancer, high TIL infiltration is often associated with inhibited tumour cell proliferation and reduced recurrence risk[2,3]. However, traditional TIL evaluation relies on subjective assessment of hematoxylin and eosin (H&E)-stained sections by pathologists on the basis of the three-level semiquantitative standard of “none/nonactive/active”, which has significant limitations. On the one hand, the interobserver consistency is low (with a Cohen’s kappa of only 0.63-0.72). In particular, when TILs are differentiated between IM and low-grade intraepithelial neoplasia, misjudgment is likely to occur because of the similar morphology of goblet cells and scattered TILs; however, accurate quantification of the TIL density is not possible, which makes clinical risk stratification of gastric mucosal lesions difficult. For example, the ability of TIL evaluation to distinguish HGIN from early-stage invasive cancer may directly affect the choice of treatment options (endoscopic submucosal dissection vs radical surgery)[4,5].

With the integration of digital pathology and artificial intelligence (AI) technology, convolutional neural networks (CNNs) have provided a technological breakthrough for solving the standardization problem of TIL evaluation[6,7]. The digitization of whole slide images (WSIs) has made the automated analysis of massive amounts of pathological image data possible. Multiscale CNN models can integrate image features at different magnifications to simultaneously capture the spatial distribution patterns of TILs and detailed features of cell nuclei. Colour deconvolution technology can correct the intensity heterogeneity of different batches of H&E-stained tissue, increasing the robustness of the model for the analysis of samples from multiple centres. In addition, the CNN-based TIL quantification index can achieve an integrated analysis of “lesion recognition-density calculation–prognosis correlation”, providing an objective and reproducible evaluation tool for pathological diagnosis of gastric mucosa[8,9]. In this study, on the basis of a multicentre retrospective cohort, a multiscale two-stage CNN model was constructed to achieve fully automated recognition and quantification of TILs in gastric mucosa WSIs, and its association with lesion grade and prognosis in early gastric cancer (EGC) was explored. The aim was to address the subjective limitations of traditional TIL evaluation methods and provide a new technological approach for the precise diagnosis of gastric mucosa lesions. The overall research process is shown in Figure 1, covering all aspects of data collection, image preprocessing, model training, and clinical application.

Figure 1
Figure 1 Overall workflow of the multiscale two-stage convolutional neural network model for automatic recognition of gastric mucosal tumour-infiltrating lymphocytes and prognostic analysis. This workflow shows four core stages of automated tumour-infiltrating lymphocyte (TIL) analysis and prognostic evaluation in gastric mucosal biopsy images. A: A total of 320 whole slide images (WSIs) from three centres were divided into training and independent test sets (7:3); B: WSIs underwent regions of interest extraction, Ruifrok-Johnston colour deconvolution (to eliminate staining heterogeneity) and 299 pixel × 299 pixel patch generation; C: A two-stage convolutional neural network (CNN) was constructed-a gastric CNN (based on pretrained Inception-ResNet-v2) for lesion segmentation/grading and a gastric artificial intelligence-TIL (G-AI-TIL) with three multiscale branches (10 × /20 × /40 ×) and attention fusion for TIL enumeration; D: The G-AI-TIL index was calculated to analyse its correlation with lesion grade; Kaplan-Meier and Cox regression verified its prognostic value in early gastric cancer, and a prognostic nomogram was constructed. WSI: Whole slide image; ROI: Region of interest; H/E: Hematoxylin and eosin; CNN: Convolutional neural network; CAG: Chronic atrophic gastritis; IM: Intestinal metaplasia; HGIN: High-grade intraepithelial neoplasia; EGC: Early gastric cancer; TIL: Tumor-infiltrating lymphocytes; G-AI-TIL: Gastric artificial intelligence-based tumor-infiltrating lymphocytes.
MATERIALS AND METHODS
Research subjects and sample selection

Gastric mucosal pathological biopsy samples collected between January 2018 and December 2023 from three tertiary-grade A hospitals were included. The inclusion criteria were as follows: (1) Double-blind diagnosis by two senior pathologists with ≥ 10 years of experience and lesion types that included CAG, IM, HGIN, and EGC; (2) Samples fixed in 10% neutral formalin and embedded in paraffin and H&E staining quality that was sufficient (no slide detachment, over- or understaining, or tissue damage); and (3) Complete clinical data, including patient age, sex, and Helicobacter pylori infection status, and a follow-up time of ≥ 12 months for EGC patients [including disease-free survival (DFS) and overall survival (OS) data]. The exclusion criteria were as follows: (1) Previous gastric surgery, radiotherapy, chemotherapy, or immunotherapy; (2) Artificial damage or staining artefacts (such as bubbles or wrinkles); and (3) Missing clinical follow-up data.

A total of 320 samples meeting the inclusion criteria were collected, including 80 cases of CAG (45 males and 35 females, with a median age of 56 years), 80 cases of IM (43 males and 37 females, with a median age of 58 years), 80 cases of HGIN (47 males and 33 females, with a median age of 60 years), and 80 cases of EGC (49 males and 31 females, with a median age of 62 years). This was a retrospective multicentre study. The use of all gastric mucosal biopsy specimens and clinical data was exempted by the ethics committees of the participating institutions. Approval from the Ethics Committee of Shaanxi Provincial Academy of Traditional Chinese Medicine and Shaanxi Provincial Hospital of Traditional Chinese Medicine, Exemption Approval No. (2023) Lunshenmian No. (15), date: January 10, 2023. The study complied with the Declaration of Helsinki and China’s Measures for Ethical Review of Life Science and Medical Research Involving Human Subjects. All the samples were double-anonymized to protect patient privacy. Informed consent was waived, as the ethics committee confirmed that there was no privacy or rights infringement.

WSI digitization and dataset construction

The paraffin sections were digitized using a Leica Aperio AT2 whole-slide scanner (Leica Biosystems, Germany). The magnification of the scanning objective was 40 ×, the pixel size was 0.24 μm, the output image format was SVS, and the image resolution was 8000 pixels × 10000 pixels. The SVS-formatted images were read using OpenSlide software (v3.4.1). The important histopathological fields of view were selected and segmented into nonoverlapping image patches of 299 pixels × 299 pixels (to match the input size of the Inception-ResNet-v2 model). Image patches with a background proportion > 50% were excluded (the background area was defined as pixels with a grayscale value > 220 in the H channel accounting for more than 50%).

The 320 samples were divided into a training set and an independent test set at a ratio of 7:3. The training set consisted of 224 cases (56 cases of CAG, 56 cases of IM, 56 cases of HGIN, and 56 cases of EGC) and contained 58240 effective image patches. The test set included 96 cases (24 cases of CAG, 24 cases of IM, 24 cases of HGIN, and 24 cases of EGC), with 31104 effective image patches.

Image preprocessing

The WSI processing workflow was standardized to ensure model robustness. First, regions of interest (ROIs) containing mucosal tissue were automatically extracted to reduce background computation. To mitigate the impact of staining variations across different batches and centres, the Ruifrok-Johnston colour deconvolution algorithm was applied to separate the H&E-stained channels. The normalized ROIs were subsequently tessellated into fixed-size image patches to serve as inputs for the deep learning network. This study introduces an ablation experiment on input channels, comparing the performance of models with 2-channel H&E (obtained via Ruifrok-Johnston colour deconvolution) vs 3-channel red green blue (RGB) inputs. The experiments are conducted on the same training and testing datasets, with all other model parameters held constant.

Image annotation and ground truth definition

To establish a high-quality ground truth, image annotation was conducted by two experienced gastrointestinal pathologists in a double-blind manner using the HALO digital imaging analysis platform (v3.6.4134; Indica Labs, United States). The annotation protocol followed a strict three-step procedure: (1) Lesion delineation: The gastric mucosal lesion areas were precisely outlined, and normal mucosa, fibromuscular stroma, and background artefacts were rigorously excluded; (2) Cellular segmentation: Within the defined lesion areas, TILs were annotated on the basis of the following morphological criteria: Lymphocytes exhibiting a high nuclear-to-cytoplasmic ratio, dense chromatin, and focal or diffuse infiltration patterns. Non-TIL areas, including epithelial cells, goblet cells, and stromal components, were excluded; and (3) Consensus and adjudication: Any discordant annotations were adjudicated by a third senior pathologist to reach a consensus. Interobserver agreement was assessed using Cohen’s kappa coefficient, which was 0.92, indicating excellent consistency and high-quality annotation standards.

Construction and training of the multiscale two-stage CNN model

First stage: Gastric-CNN (gastric mucosal lesion area recognition network): The pretrained Inception-ResNet-v2 model, which is pretrained on the ImageNet dataset and has a strong ability to extract general image features, is used. For the task of gastric mucosa lesion recognition, the following adjustments are made to the model: Input layer: H-E-deconvolved image patches of 299 pixels × 299 pixels with 2 channels (H channel and E channel).

Channel dimension adaptation for the pretrained model: To address the compatibility issue between the 2-channel (H/E) input and the 3-channel weights of the pretrained Inception-ResNet-v2 model, this study performed channel dimension mapping optimization on the input layer of the model: The 2-channel (H/E) feature map is dimensionally expanded using 1 × 1 convolution kernels, mapping the number of feature channels from 2 to 3. During the mapping process, the convolution kernel weights are initialized using the Xavier normal distribution, and this layer is set to a trainable state. The weights of the first 10 convolutional layers (including the inception modules and ResNet residual connections) of the original pretrained model remain frozen, and only the parameters of the mapping layer and subsequent unfrozen layers are updated. This operation not only retains the general features such as edges and textures learned by the pretrained model but also realizes the effective input of 2-channel H/E images, solving the technical problem of channel dimension mismatch.

Feature extraction layer: Freeze the first 10 layers of the pretrained model (including the inception module and ResNet residual connection) to retain the learned general features such as edges and textures. The learning rate of the unfrozen layers is set to 0.0001 to prevent the pretrained features from being damaged.

Classification layer: Replace the 1000-class output layer of the original model with a 4-class layer (corresponding to CAG/IM/HGIN/EGC). A new fully connected layer (with 1024 neurons and a ReLU activation function) is added, and the softmax activation function is used to output the probabilities of each lesion type.

Second stage: G-TIL-CNN (TIL recognition network): Using the lesion area image patches output by the Gastric-CNN as input, a multiscale TIL recognition network is constructed. The specific architecture is shown in Figure 2.

Figure 2
Figure 2 Schematic workflow and multiscale convolutional neural network architecture for tumour-infiltrating lymphocyte recognition in gastric mucosa. The pipeline consists of four primary stages: (1) Dataset preparation: The study cohort was divided into a training set and an independent test set; (2) Image preprocessing: Original hematoxylin and eosin images were decomposed into 2-channel (H and E) components via color deconvolution to resolve channel dimension mismatch and enhance staining robustness; (3) Model core: A multiscale convolutional neural network architecture featuring parallel branches and attention mechanisms was employed to extract multi-resolution features; and (4) Output and visualization: The model generates attention heatmaps that precisely highlight tumour-infiltrating lymphocyte nuclear regions (red) while effectively suppressing interference from background structures such as goblet cells. H/E: Hematoxylin and eosin.

Multiscale feature extraction branches: The low-scale branch (10 × magnification) uses a 5 × 5 convolution kernel (stride 1, padding 2) to extract the spatial distribution features of the TILs within the lesion area; the medium-scale branch (20 × magnification) combines the local binary pattern (LBP8,1riu2) and variance (VAR8,1) to construct joint texture features; and the high-scale branch (40 × magnification) uses a 3 × 3 convolution kernel (stride 1, padding 1) to extract the detailed features of the TIL cell nuclei.

Feature fusion layer: The channel attention mechanism is adopted to perform weighted fusion on the three-branch features, and the channel weights of each branch feature are calculated (in this study, the fusion weights for the high-, medium-, and low-scale branches are the optimal values dynamically learned during model training, rather than manually fixed hyperparameters. At the initial stage of model training, the weight of each branch was set to 1/3. After 30 training epochs, the channel attention mechanism automatically adjusted the weights on the basis of the contribution of the features, eventually converging to 0.4 (high-scale), 0.3 (medium-scale), and 0.3 (low-scale). This weight distribution reflects the core role of high-scale nuclear detail features in TIL recognition.) to enhance the contribution of key features.

Classification layer: A binary classification (TIL/non-TIL) is set, and the sigmoid activation function is used to output the classification probability. The loss function is the cross-entropy loss with class weights (the weight of the TIL samples is 3, and the weight of the non-TIL samples is 1) to address class imbalance caused by the low proportion (approximately 25%) of the TIL samples in the training set.

Model training parameters: The optimizer used is Adam with an initial learning rate of 0.0005, which decays by 10% every 5 epochs. The batch size is 16, and the number of training epochs is 30. Data augmentation (random rotation from -90° to 90°, scaling from 0.8 times to 1.2 times, horizontal flipping, and random adjustment of pixel values by ± 30) is adopted to improve generalizability. The training environment is an NVIDIA A100 GPU (40 GB video memory), and it is implemented on the PyTorch 2.0 framework. The total training time is approximately 11.2 hours (including Gastric-CNN and G-TIL-CNN). The model weights have been uploaded to the Zenodo platform for easy verification and reuse.

Image postprocessing: The TIL classification results output by the G-TIL-CNN are postprocessed to eliminate misjudgment of image patches (such as misclassifying stromal cells as TILs) and noise interference, ultimately improving the accuracy of TIL recognition (the false positive rate is reduced to less than 0.2% after processing).

To quantify the intensity of immune infiltration, we defined the “gastric mucosa TIL density index” (G-AI-TIL). For each whole-slide image (WSI) sample, the G-AI-TIL is calculated as follows: G-AI-TIL = × 100%, where NTIL represents the number of image patches identified as TIL-positive by the model and Ntotal represents the total number of image patches identified as the region of interest (ROI) in this case. This index reflects the spatial infiltration density of lymphocytes within the lesion area, with values ranging from 0% to 100%.

Quantitative indicators

Quantitative indicators of TILs: The gastric mucosa TIL density index (G-AI-TIL) was defined as follows: G-AI-TIL = (NTIL/total) × 100%. Here, NTIL represents the number of image patches classified as TIL by G-TIL-CNN, and Nnon-TIL represents the number of image patches classified as non-TIL. The G-AI-TIL values range from 0%-100%. A higher value indicates more significant TIL infiltration in the lesion area.

Model performance evaluation metrics: Classification performance metrics: Accuracy (number of correctly classified image patches/total number of image patches × 100%), specificity (number of correctly classified non-TIL image patches/actual number of non-TIL image patches × 100%), sensitivity (number of correctly classified TIL image patches/actual number of TIL image patches × 100%), Cohen’s kappa coefficient (used as a measure of the consistency between the model and manual annotation, kappa > 0.8 indicates excellent consistency), and F1 score [harmonic mean of precision and recall, F1 = (2 × precision × recall)/(precision + recall)].

Quantitative consistency: The Pearson correlation coefficient between the TIL density (cells/mm2) determined manually by pathologists and that quantified by G-AI-TIL was calculated to evaluate the consistency between the model’s quantification results and the gold standard.

Statistical analysis

Correlation analysis of lesion grade: The Kruskal-Wallis test was used to compare the differences in the G-AI-TIL among the four types of gastric mucosal lesions, and the Tukey method was used for pairwise comparisons as a post hoc test.

Prognostic association analysis: Follow-up was performed for a total of 96 patients with EGC. DFS was defined as the time from diagnosis to tumour recurrence or death from any cause, and OS was defined as the time from diagnosis to death from any cause. The Kaplan-Meier method was used to plot survival curves, and the log-rank test was used to compare the survival differences between the high/Low TIL groups (using the median G-AI-TIL, 28.5%, as the cut-off). This threshold was determined by “maximizing prognostic significance”. After the 25th, 30th, and 35th percentile thresholds were validated, the log-rank χ2 value corresponding to 28.5% was the highest. A Cox proportional hazards regression model was constructed with age, sex, tumour invasion depth (submucosa vs mucosa), and ulcer status (present vs absent) as covariates to evaluate the independent effect of the G-AI-TIL on the prognosis of patients with EGC.

Comparison of model performance: The log-likelihood ratio test was used to compare the goodness of fit between the “prognostic model incorporating G-AI-TIL” and the “model incorporating traditional TIL grading”.

RESULTS
Performance of the G-TIL-CNN model

Overall performance: The G-TIL-CNN demonstrated excellent classification ability in the independent test set (a total of 31104 image patches). As shown in Table 1, the model achieved an accuracy of 99.2%, a specificity of 100.0%, and a sensitivity of 98.7%. The F1 score reached 0.99, and Cohen’s kappa coefficient was 0.98, indicating high consistency between the model’s prediction results and the gold-standard annotations of pathologists. In contrast, the accuracy of traditional manual evaluation (performed by intermediate pathologists) was only 86.3%, and the kappa coefficient was 0.72. Statistical analysis confirmed that the model’s performance was significantly better than that of the manual evaluation (P < 0.001).

Table 1 Performance indicators of the gastric-tumour-infiltrating lymphocytes-convolutional neural network model in the training set and test set (%).
Dataset
Accuracy (95%CI)
Specificity (95%CI)
Sensitivity (95%CI)
Cohen’s Kappa (95%CI)
F1 score (95%CI)
Compared with manual assessment (P value)
Training99.5 (99.2-99.8)99.8 (99.6-100.0)99.3 (98.9-99.7)0.99 (0.98-1.00)99.5 (99.2-99.8)< 0.001
Testing99.2 (98.8-99.6)100.0 (99.9-100.0)98.7 (98.1-99.3)0.98 (0.97-0.99)99.3 (98.9-99.7)

The confusion matrix of the test set is shown in Table 2. The model has no false positives (non-TILs misjudged as TILs), and the missed detection rate is only 1.3%, further verifying its high specificity and sensitivity.

Table 2 Confusion matrix of the gastric-tumour-infiltrating lymphocytes-convolutional neural network model on the test set, n (%).
Actual label/predicted label
TIL
Non-TIL
Total
Recall (%)
Miss rate (%)
TNR (%)
TIL7632 (24.5)98 (0.3)7730 (24.8)98.71.3-
Non-TIL0 (0.0)23374 (75.2)23374 (75.2)100.00.0100.0
Total7632 (24.5)23472 (75.5)31104 (100)---
Precision (%)100.099.6----
False discovery rate (%)0.00.4----

TIL recognition performance of each lesion type: In different gastric mucosal lesions, there were slight differences in the TIL recognition accuracy of the G-TIL-CNN (Table 3). The recognition accuracy was the highest for EGC (99.5%). In EGC, TILs mostly showed a clustered distribution, with a clear boundary from tumour epithelial cells and significant nuclear atypia, making them easy to distinguish from other cells. The recognition accuracy was the lowest for IM (98.6%). This occurred mainly because the goblet cells in the IM samples (with cytoplasmic vacuoles and small, round nuclei) were morphologically similar to scattered TILs, resulting in a small number of misjudgments. However, the model still maintained a specificity of 99.2%, and no non-TIL regions were misjudged as TILs, meeting the clinical diagnostic requirements.

Table 3 Tumour-infiltrating lymphocyte recognition performance of the gastric-tumour-infiltrating lymphocytes-convolutional neural network for different gastric mucosal lesions (%).
Lesion type
Accuracy (95%CI)
Specificity (95%CI)
Sensitivity (95%CI)
Chronic atrophic gastritis99.1 (98.5-99.7)99.5 (99.1-99.9)98.8 (98.0-99.6)
Intestinal metaplasia98.6 (97.8-99.4)99.2 (98.6-99.8)97.9 (96.8-99.0)
High-grade intraepithelial neoplasia99.3 (98.7-99.9)99.7 (99.3-100.0)99.0 (98.2-99.8)
Early gastric cancer99.5 (99.0-100.0)100.0 (99.8-100.0)99.1 (98.4-99.8)

Interpretability analysis of the model was further performed with quantitative evaluation and clinical physician validation to address the “black box” issue of deep learning. For the attention heatmaps generated by the model, the intersection over union (IoU) and Dice similarity coefficient were used to quantify the overlap between the model’s attention regions and the regions manually annotated with TILs by pathologists. In the independent test set (n = 96), the average IoU reached 0.89, and the Dice similarity coefficient was 0.94, indicating a high degree of consistency between the model’s attention focus and the core regions of pathological diagnosis. Furthermore, a blind validation was conducted by 5 senior gastrointestinal pathologists with more than 10 years of diagnostic experience (from different hospital pathology departments). Pathologists were required to score the rationality of the attention heatmaps on a 1-5 scale (5 for complete consistency with clinical diagnostic logic) and evaluate the auxiliary value of the heatmaps for TIL recognition. The results revealed that the average rationality score of the heatmaps was 4.8, and 92% of the pathologists believed that the heatmaps could reduce the time of TIL recognition by more than 30% and effectively avoid the misjudgment of goblet cells, which confirmed the clinical recognition of the attention mechanism of the model.

An interpretability analysis of the model is shown in Figure 3. As seen from the attention heatmap, when the model identifies TILs, it focuses mainly on the lymphocyte nuclear region rather than the background or interfering cells (such as goblet cells), which confirms the rationality of the model’s judgement and alleviates the “black box” problem of AI.

Figure 3
Figure 3 Performance evaluation of the gastric tumour-infiltrating lymphocyte-convolutional neural network model for tumour-infiltrating lymphocyte recognition. Model performance was evaluated on 31104 test set patches. A: The receiver operating characteristic curve shows excellent discriminative ability for tumour-infiltrating lymphocyte (TILs)/non-TILs; B: The confusion matrix reveals 100.0% specificity (no false positives) and 98.7% sensitivity (1.3% miss rate); C: The scatter plot shows a strong positive correlation (Pearson r = 0.93, P < 0.001) between the gastric artificial intelligence-based TIL and the pathologist-manual TIL density, confirming quantitative consistency; D: The bar chart shows that the model outperforms the manual evaluation in terms of accuracy (99.2% vs 86.3%) and kappa coefficient (0.98 vs 0.72, P < 0.001), indicating that the interobserver variability is reduced. ROC: Receiver operating characteristic; AUC: Area under the curve; TIL: Tumour-infiltrating lymphocyte; CNN: Convolutional neural network; G-AI-TIL: Gastric artificial intelligence-based tumor-infiltrating lymphocytes.

To verify that the epithelial layer-identified TILs are true positives rather than misidentifications of epithelial cell nuclei or goblet cells, this study conducted the following dual validation experiments: (1) Morphological quantitative validation: Morphological features, including the nuclear-to-cytoplasmic ratio, nuclear area, chromatin density, and 12 other indicators, were extracted from the TIL regions identified in the epithelial layer shown in Figure 4. These features were compared with those of pathologist-annotated true TILs, goblet cells, and epithelial cell nuclei. The results revealed that the morphological features of the epithelial layer-identified regions were 98.9% similar to those of true TILs, while the similarity to goblet cells and epithelial cell nuclei was only 32.5% and 28.7%, respectively; and (2) Immunohistochemical gold standard validation: Thirty randomly selected samples containing epithelial layer TIL identifications (including 10 IM samples, 10 HGIN samples, and 10 EGC samples) were subjected to CD3/CD8 immunohistochemical staining. The results revealed that the epithelial layer TIL regions identified by the model were specifically labelled by CD3/CD8 antibodies, with a positive concordance rate of 100%, confirming the absence of false-positive annotations. These results indicate that the TILs in the epithelial layer identified by the model are true intraepithelial lymphocytes and are not misidentified, thereby ensuring the accuracy of the G-AI-TIL index calculation.

Figure 4
Figure 4 Model interpretability analysis and immunohistochemistry validation of tumour-infiltrating lymphocyte recognition. Representative images (scale bar = 50 μm) of chronic atrophic gastritis, intestinal metaplasia, high-grade intraepithelial neoplasia and early gastric cancer compared with original hematoxylin and eosin staining, pathologists’ double-blind manual annotation, artificial intelligence (AI) attention heatmaps and CD3/CD8 immunohistochemistry (IHC) validation. Red regions in heatmaps indicate the model’s focus on tumour-infiltrating lymphocyte (TIL) nuclear regions, avoiding interference from goblet cells (IM). AI-recognized epithelial TIL regions show 100% positive concordance with CD3/CD8 IHC staining (black arrows), confirming that there are no false positives. The attention regions of the model are highly consistent with the manual annotations (IoU = 0.89; Dice coefficient = 0.94). CAG: Chronic atrophic gastritis; IM: Intestinal metaplasia; HGIN: High-grade intraepithelial neoplasia; EGC: Early gastric cancer; IHC: Immunohistochemistry; AI: Artificial intelligence; H/E: Hematoxylin and eosin.

Ablation experiment of feature fusion weights: To verify the effectiveness of this weighting, an ablation study was added: Using equal-weight fusion (0.33/0.33/0.33) as the control group, the model performance of the two fusion methods was compared on an independent test set (n = 96). The results revealed that the dynamic learning weight group achieved an accuracy of 99.2%, with kappa = 0.98, whereas the equal-weight control group achieved 96.5%, with kappa = 0.89 (P < 0.001); moreover, the dynamic learning weight group demonstrated a significantly greater accuracy (98.6%) than the equal-weight control group (94.2%) in distinguishing TILs from goblet cells in IM samples did, confirming the rationality of the enhanced high-scale branch weight shown in Figure 5A.

Figure 5
Figure 5 Ablation experiment results of the gastric tumour-infiltrating lymphocyte-convolutional neural network model. A: Compared with equal weight fusion, dynamic learning weight fusion achieves higher accuracy (99.2% vs 96.5%) and kappa (0.98 vs 0.89, P < 0.001), with better intestinal metaplasia tumour-infiltrating lymphocyte recognition (98.6% vs 94.2%); B: 2-channel haematoxylin/eosin input outperforms 3-channel red green blue input in terms of accuracy (99.2% vs 94.7%), specificity (100.0% vs 92.3%) and sensitivity (98.7% vs 95.1%, P < 0.001), eliminating staining batch interference and improving multicentre sample robustness. H/E: Haematoxylin/eosin; RGB: Red green blue.

Ablation experiment of the input channel types: The results revealed that the 2-channel H/E input group achieved an accuracy of 99.2%, a specificity of 100.0%, and a sensitivity of 98.7% on the independent test set; the 3-channel RGB input group had an accuracy of 94.7%, a specificity of 92.3%, and a sensitivity of 95.1% (P < 0.001) shown in Figure 5B. The main reason for the performance degradation of the RGB input group is the spectral interference caused by batch differences in H&E staining. In contrast, the 2-channel H/E input, by separating the staining channels, eliminates the effect of dye overlap and significantly improves the robustness of the model to multicentre samples. Although the 2-channel input loses part of the RGB spectral information, it retains most of the core nuclear (H) and cytoplasmic (E) features in pathological diagnosis through colour deconvolution, and their contribution to TIL recognition far outweighs the loss of spectral information (Table 4).

Table 4 Ablation experiment performance indicators of the gastric-tumour-infiltrating lymphocytes-convolutional neural network model (independent test set).
Ablation experiment type
Experimental group
Accuracy (95%CI)
Specificity (95%CI)
Sensitivity (95%CI)
Cohen’s Kappa
P value (vs optimal group)
Feature fusion weightDynamic learning weight (0.4/0.3/0.3)99.2 (98.8-99.6)100.0 (99.9-100.0)98.7 (98.1-99.3)0.98< 0.001
Feature fusion weightEqual weight (0.33/0.33/0.33)96.5 (95.8-97.2)97.8 (97.1-98.5)95.1 (94.2-96.0)0.89
Input channel type2-channel H/E (Colour Deconvolution)99.2 (98.8~99.6)100.0 (99.9~100.0)98.7 (98.1-99.3)0.98< 0.001
Input channel type3-channel RGB94.7 (93.9-95.5)92.3 (91.2-93.4)95.1 (94.0-96.2)0.87
Association between the G-AI-TIL and the grade of gastric mucosal lesions

The results of the Kruskal-Wallis test revealed significant differences in the G-AI-TIL among the four types of gastric mucosal lesions (H = 52.36, P < 0.001), with a trend towards “increasing malignancy of lesions → gradual increase in the G-AI-TIL” (Figure 6). The post hoc Tukey test results revealed that the median G-AI-TIL was significantly greater for EGC (28.5%) than for HGIN (15.2%, P < 0.001), that for HGIN was significantly greater than that for IM (8.7%, P < 0.001), and that for IM was significantly greater than that for CAG (5.3%, P < 0.01). Additionally, the G-AI-TIL was significantly positively correlated with the TIL density (cells/mm2) manually determined by pathologists (Pearson correlation coefficient = 0.93; P < 0.001), indicating high consistency between the model quantification results and the gold standard.

Figure 6
Figure 6 Correlations between the gastric artificial intelligence-based tumor-infiltrating lymphocytes and pathological grade of gastric mucosal lesions. The box plot shows the gastric artificial intelligence-based tumor-infiltrating lymphocytes (G-AI-TIL) distribution across the four lesion types (24 cases each) in the test set, with the x-axis ordered by increasing malignancy [chronic atrophic gastritis (CAG)→intestinal metaplasia (IM)→high-grade intraepithelial neoplasia (HGIN)→early gastric cancer (EGC)]. The results of the Kruskal-Wallis test and the Tukey post hoc test confirmed a stepwise increase in the median G-AI-TIL: CAG (5.3%) < IM (8.7%, P < 0.01) < HGIN (15.2%, P < 0.001) < EGC (28.5%, P < 0.001). The G-AI-TIL is positively correlated with lesion malignancy and serves as a quantitative index for grading gastric mucosal lesions. CAG: Chronic atrophic gastritis; IM: Intestinal metaplasia; HGIN: High-grade intraepithelial neoplasia; EGC: Early gastric cancer; G-AI-TIL: Gastric artificial intelligence-based tumor-infiltrating lymphocytes.
Association between the G-AI-TIL and the prognosis of patients with EGC

Follow-up results: The median follow-up time for 96 patients with EGC was 42 months (range: 12-60 months). During this period, recurrence occurred in 28 patients (29.2%), and 24 patients (25.0%) died. Among them, 21 patients (21.9%) died from tumour-related causes, and 3 patients (3.1%) died from non-tumour-related causes.

Kaplan-Meier survival analysis: Using the median G-AI-TIL (28.5%) as the cut-off, 96 patients with EGC were divided into a high-TIL group (G-AI-TIL ≥ 28.5%, n = 48) and a low-TIL group (G-AI-TIL < 28.5%, n = 48). The Kaplan-Meier survival curves (Figure 7A and B) revealed that the 3-year DFS rate in the high-TIL group was 82.1% (95%CI: 70.3%-93.9%), which was significantly greater than that in the low-TIL group (63.5%, 95%CI: 50.1%-76.9%, log-rank χ2 = 6.89; P = 0.009), and the 3-year OS rate in the high-TIL group was 85.7% (95%CI: 75.1%-96.3%), which was significantly greater than that in the low-TIL group (67.2%, 95%CI: 54.3%-80.1%, log-rank χ2 = 7.53; P = 0.006).

Figure 7
Figure 7 Prognostic value of the gastric artificial intelligence-based tumor-infiltrating lymphocytes index in early gastric cancer. A and B: Kaplan-Meier curves for disease-free survival and overall survival (OS) stratified by the median gastric artificial intelligence-based tumor-infiltrating lymphocytes (G-AI-TIL) cut-off of 28.5% (log-rank test). The high G-AI-TIL group exhibited significantly better survival outcomes; C: Forest plot of multivariate Cox regression analysis for OS, identifying high G-AI-TIL ( ≥ 28.5%) as an independent protective factor alongside risk factors such as age and submucosal invasion; D: A prognostic nomogram predicting 3- and 5-year OS probabilities by integrating G-AI-TIL, age, and invasion depth, with matching calibration curves demonstrating high accuracy (C-index = 0.81). EGC: Early gastric cancer; HR: Hazard ratio; CI: Confidence interval; G-AI-TIL: Gastric artificial intelligence-based tumor-infiltrating lymphocytes.

Multivariate Cox regression and prognostic prediction model: The variables that were significant in the univariate analysis were included in the multivariate Cox proportional hazards regression model. The results revealed that a G-AI-TIL ≥ 28.5% was an independent protective factor for both DFS [hazard ratio (HR) = 0.58; 95%CI: 0.37-0.91; P = 0.018] and OS (HR = 0.55; 95%CI: 0.35-0.86; P = 0.009) in patients with EGC. This implies that a high immune infiltration status was associated with a 42% reduction in recurrence risk and a 45% reduction in death risk. In contrast, tumour invasion depth reaching the submucosa was a significant risk factor (OS HR = 1.93) (Figure 7C and D).

To improve clinical utility, we integrated the G-AI-TIL index, patient age, and tumour invasion depth to construct a nomogram. This tool can intuitively calculate the total score for each patient, thereby predicting their 3-year and 5-year survival probabilities. The calibration curve shows a high degree of consistency between the predicted values and the actual observed values (C-index = 0.81), indicating that the model has good accuracy in individualized prognosis assessment (Table 5).

Table 5 Multivariate Cox proportional hazards regression analysis of prognostic factors in patients with early gastric cancer.
Prognostic Index
Variable
HR
95%CI
P value
DFSG-AI-TIL ≥ 28.5% (vs < 28.5%)0.580.37-0.910.018
Age (per 1-year increase)1.041.02-1.06< 0.001
Tumour invasion depth (submucosa vs mucosa)1.871.12-3.120.016
Ulcers (present vs absent)1.320.81-2.150.257
OSG-AI-TIL ≥ 28.5% (vs < 28.5%)0.550.35-0.860.009
Age (per 1-year increase)1.051.03-1.07< 0.001
Tumour invasion depth (submucosa vs mucosa)1.931.15-3.240.013
Ulcers (present vs absent)1.410.85-2.340.180
DISCUSSION
Technical advantages and clinical value of the model

(1) Multiscale feature fusion addresses the core challenges in TIL recognition: The identification of gastric mucosal TILs faces the challenges of “high tissue heterogeneity and diverse target morphologies”. For example, scattered TILs in IM samples resemble goblet cells in shape, and the boundaries between clustered TILs and tumour cells in EGC are blurred[10,11]. This model captures the spatial distribution, texture differences, and nuclear details of TILs through low-, medium-, and high-scale branches, respectively. By integrating the attention mechanism to weight and fuse key features, it effectively distinguishes TILs from interfering cells in different lesion types. The Cohen’s kappa coefficient of the test set reached 0.98, significantly outperforming traditional manual evaluation and resolving the clinical issue of “large interobserver variability”[12]; (2) Colour deconvolution enhances model robustness: Batch differences in H&E staining are important factors affecting the generalizability of pathological image analysis models. Differences in staining reagents and procedures across different hospitals and batches may lead to varying staining intensities for the same lesion[13,14]. In this study, the Ruifrok-Johnston colour deconvolution algorithm is employed to decompose RGB images into H-E single channels, eliminating the interference of dye overlap. This enables the model to maintain stable performance for multicentre samples from three hospitals, laying a foundation for clinical translation[15,16]; and (3) G-AI-TIL enables integrated “diagnosis–prognosis” assessment: Traditional TIL evaluation can provide only semiquantitative data and cannot be directly correlated with clinical prognosis[15,16]. The G-AI-TIL index defined in this study can not only be used to quantitatively distinguish four types of lesions (CAG, IM, HGIN, and EGC, with a trend of “increasing malignancy → increasing G-AI–TIL”) but also serve as an independent prognostic protective factor for EGC (HR = 0.55; P = 0.009). These results suggest that high TIL infiltration may improve the prognosis of EGC patients by enhancing the local antitumour immune response (e.g., CD8+ T cells killing tumour cells), providing an objective basis for clinical risk stratification. For example, EGC patients with high TIL infiltration may be considered for shorter follow-up intervals or endoscopic minimally invasive treatment instead of radical surgery.

In addition, the performance of the model in this study is significantly superior to that of previous AI studies related to gastric mucosal TILs and has the advantage of full automation, as follows: (1) The multisequence correlation coefficient for TIL recognition in various cancers (including melanoma) is 0.82, while the Cohen’s kappa coefficient of the G-TIL-CNN in this study reaches 0.98, indicating high consistency; (2) The area under the curve in previous studies was 0.954, but manual delineation of the tumour area was needed. In contrast, this study requires no human intervention, and the accuracy rate (99.2%) is better than the approximately 95% accuracy rate reported in previous studies[17]; and (3) In terms of clinical value, the processing time for a single WSI is only 2.3 minutes (including automatic lesion recognition and TIL enumeration), which is 3.7 times more efficient than traditional manual evaluation (8.5 minutes), making it suitable for large-scale screening of gastric mucosal lesions.

Biological mechanisms and clinical implications

This study revealed that the G-AI-TIL index increases with increasing lesion malignancy (CAG→EGC) and that high TIL infiltration in EGC patients indicates a favourable prognosis. This phenomenon reflects the dynamic response of the body’s immune system to tumorigenesis. In the early stages of HGIN and EGC, the release of tumour neoantigens activates local immune surveillance, recruiting many lymphocytes (especially CD8+ T cells) to infiltrate and eliminate abnormal cells, which explains why the TIL density is greater in lesions with greater malignancy[18,19]. However, for patients already diagnosed with EGC, a high TIL density means that the host retains a strong antitumour immune capacity, which can effectively inhibit the metastasis and recurrence of minimal residual disease, thus resulting in better survival benefits[20]. These findings have direct clinical guiding significance: EGC patients with a low G-AI-TIL (< 28.5%) are suggested to be in an “immune desert” state with a high risk of recurrence. Even if the tumour is confined to the mucosal layer, a more aggressive follow-up strategy is recommended, or additional adjuvant therapy should be considered rather than simply endoscopic resection.

Comparison with traditional TIL assessment methods

Traditional manual evaluation of TILs relies on the subjective experience of pathologists and has significant limitations: (1) Low consistency: In this study, the kappa coefficient of manual evaluation by intermediate-level pathologists was only 0.72. Misjudgements are particularly prone to occur when IM and HGIN are differentiated; (2) Insufficient quantification accuracy: Samples can only be graded as “none/mild/moderate/severe” without the determination of specific density values; and (3) Low efficiency: The average time for manual evaluation of a single WSI is 8.5 minutes, while this model takes only 2.3 minutes to process a single WSI, significantly improving diagnostic efficiency. This model effectively compensates for the deficiencies of traditional methods through automated and quantitative analysis: (1) High consistency (kappa = 0.98), avoiding subjective biases; (2) Precise quantification (the G-AI-TIL can be accurate to 0.1%), supporting the identification of subtle pathological differences; and (3) High efficiency and suitability for large-scale screening and follow-up. In addition, the prognostic predictive value of the model is superior to that of traditional manual grading, further highlighting its clinical application potential.

Research limitations and future directions

Sample size and limited disease stage: The sample size was relatively limited (320 cases, including 96 cases of EGC), and advanced gastric cancer cases were not included. In the future, it will be necessary to expand the multicentre sample size (covering different regions and different pathological centres) to further validate the performance of the model in advanced gastric cancer and other gastric mucosal lesions (such as low-grade intraepithelial neoplasia). In view of the limitation that the cohort of this study included only Chinese individuals, the research team formulated a multicentre international external validation plan: It is intended to cooperate with the pathology departments of hospitals in Japan, South Korea, Europe, the United States and other regions to collect a total of 500 gastric mucosal lesion samples (including those of CAG, IM, HGIN, EGC and advanced gastric cancer). The samples cover different ethnic groups in East Asia, Europe and the United States, and different H&E staining reagents and scanner models (Leica, Hamamatsu, Aperio) are used in each centre. The validation contents include the following: (1) The consistency of the G-AI-TIL index among different ethnic groups and different detection platforms; (2) The applicability of the 28.5% threshold and the correction of ethnic-specific thresholds when necessary; and (3) The performance of the model in TIL identification of advanced gastric cancer and the analysis of its prognostic value. Moreover, standardized records of the staining protocols and scanning parameters of multicentre samples will be obtained, a standardized database of pathological images will be constructed, and the cross-platform adaptability of the model will be further improved through data augmentation (simulation of different staining intensities and scanning resolutions).

Insufficient model interpretability: Although the attention heatmaps of the model have been quantitatively evaluated and clinically validated (with an average IoU of 0.89 and a clinical rationality score of 4.8), the “black-box” characteristic of CNNs still makes it impossible to fully clarify the specific feature dimensions (e.g., nuclear texture and spatial distribution) that drive TIL classification. In addition, the model has potential failure modes: When the samples have severe staining inhomogeneity (excessively deep/shallow haematoxylin staining) or tissue folding, the attention heatmaps may have local offsets, which may lead to a slight decrease in recognition accuracy. In the future, the key feature dimensions of TIL classification will be further mined through feature visualization technology and a staining quality assessment module will be added to the model postprocessing step to automatically mark unqualified samples and prompt manual review. Moreover, the key regions recognized by the model (such as TIL cell nuclei and intercellular spaces) are more intuitively visualized through high-resolution attention heatmaps to further increase clinician trust in the model.

Failure to integrate multimodal data: This study was based solely on H&E-stained images and did not incorporate immunohistochemical data (e.g., staining of CD3 and CD8) to identify TIL subtypes. In the future, multimodal data (H&E staining + immunohistochemistry) can be integrated to construct a combined evaluation model of “TIL density–TIL subtype”, further refining the analysis of the immune microenvironment and providing a more accurate basis for the selection of immunotherapy regimens.

Population specificity of the threshold: The G-AI-TIL threshold (28.5%) is based on the cohort from three hospitals in this study, which may be affected by geographical location and pathological staining procedures. In the future, samples from different regions and pathological centres should be included to validate the universality of the threshold.

Dependence on manually annotated data: Model training relies on double-blind annotation by pathologists. If there are biases in the annotation (even though kappa = 0.92 in this study), it may affect the model performance. Immunohistochemistry (e.g., CD3 and CD8 staining) can subsequently be used as the “gold standard” to reduce manual annotation errors.

CONCLUSION

The multiscale two-stage CNN model constructed in this study achieves automated and high-precision identification of TILs in gastric mucosal biopsy images. The developed gastric mucosal G-AI-TIL effectively quantifies TILs and distinguishes gastric carcinogenesis stages, including CAG, IM, HGIN, and EGC. Furthermore, the G-AI-TIL serves as an independent protective factor for both DFS and OS in EGC patients. By overcoming the subjectivity and inefficiency of traditional assessment methods, this model provides an objective tool for precise diagnosis, risk stratification, and personalized treatment planning, demonstrating significant potential for clinical translation.

References
1.  Sung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, Bray F. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J Clin. 2021;71:209-249.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 76817]  [Cited by in RCA: 70831]  [Article Influence: 14166.2]  [Reference Citation Analysis (67)]
2.  Zullo A, Rago A, Felici S, Licci S, Ridola L, Caravita di Toritto T. Onset and Progression of Precancerous Lesions on Gastric Mucosa of Patients Treated for Gastric Lymphoma. J Gastrointestin Liver Dis. 2020;29:27-31.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 6]  [Cited by in RCA: 12]  [Article Influence: 2.0]  [Reference Citation Analysis (0)]
3.  Wang Y, Liu H, Zhang M, Xu J, Zheng L, Liu P, Chen J, Liu H, Chen C. Epigenetic reprogramming in gastrointestinal cancer: biology and translational perspectives. MedComm (2020). 2024;5:e670.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 9]  [Cited by in RCA: 9]  [Article Influence: 4.5]  [Reference Citation Analysis (0)]
4.  Zavros Y, Merchant JL. The immune microenvironment in gastric adenocarcinoma. Nat Rev Gastroenterol Hepatol. 2022;19:451-467.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 164]  [Cited by in RCA: 137]  [Article Influence: 34.3]  [Reference Citation Analysis (0)]
5.  Ma H, Srivastava S, Ho SWT, Xu C, Lian BSX, Ong X, Tay ST, Sheng T, Lum HYJ, Abdul Ghani SAB, Chu Y, Huang KK, Goh YT, Lee M, Hagihara T, Ng CSY, Tan ALK, Zhang Y, Ding Z, Zhu F, Ng MSW, Joseph CRC, Chen H, Li Z, Zhao JJ, Rha SY, Teh M, Yeong J, Yong WP, So JB, Sundar R, Tan P. Spatially Resolved Tumor Ecosystems and Cell States in Gastric Adenocarcinoma Progression and Evolution. Cancer Discov. 2025;15:767-792.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 49]  [Cited by in RCA: 37]  [Article Influence: 37.0]  [Reference Citation Analysis (0)]
6.  Ugolini F, De Logu F, Iannone LF, Brutti F, Simi S, Maio V, de Giorgi V, Maria di Giacomo A, Miracco C, Federico F, Peris K, Palmieri G, Cossu A, Mandalà M, Massi D, Laurino M. Tumor-Infiltrating Lymphocyte Recognition in Primary Melanoma by Deep Learning Convolutional Neural Network. Am J Pathol. 2023;193:2099-2110.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 17]  [Reference Citation Analysis (0)]
7.  Xia S, Xia Y, Liu T, Luo Y, Pang PC. Application of deep learning models in gastric cancer pathology image analysis: a systematic scoping review. BMC Cancer. 2025;25:1257.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 6]  [Cited by in RCA: 6]  [Article Influence: 6.0]  [Reference Citation Analysis (0)]
8.  Chu ML, Ge XM, Eastham J, Nguyen T, Fuji RN, Sullivan R, Ruderman D. Assessment of Color Reproducibility and Mitigation of Color Variation in Whole Slide Image Scanners for Toxicologic Pathology. Toxicol Pathol. 2023;51:313-328.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 5]  [Reference Citation Analysis (0)]
9.  Tsou P, Wu CJ. Classifying driver mutations of papillary thyroid carcinoma on whole slide image: an automated workflow applying deep convolutional neural network. Front Endocrinol (Lausanne). 2024;15:1395979.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 2]  [Reference Citation Analysis (0)]
10.  Zhang X, Liu K, Zhang K, Li X, Sun Z, Wei B. SAMS-Net: Fusion of attention mechanism and multi-scale features network for tumor infiltrating lymphocytes segmentation. Math Biosci Eng. 2023;20:2964-2979.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 9]  [Cited by in RCA: 8]  [Article Influence: 2.7]  [Reference Citation Analysis (0)]
11.  Khan R, Alzaben N, Daradkeh YI, Lee MY, Ullah I. Bilateral collaborative streams with multi-modal attention network for accurate polyp segmentation. Sci Rep. 2025;15:34182.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 3]  [Reference Citation Analysis (0)]
12.  Huo X, Tian S, Yu L, Zhang W, Li A, Yang Q, Song J. MM-HiFuse: multi-modal multi-task hierarchical feature fusion for esophagus cancer staging and differentiation classification. Complex Intell Syst. 2025;11:113.  [PubMed]  [DOI]  [Full Text]
13.  Xu Y, Yang S, Zhu Y, Yao S, Li Y, Ye H, Ye Y, Li Z, Wu L, Zhao K, Huang L, Liu Z. Artificial intelligence for quantifying Crohn's-like lymphoid reaction and tumor-infiltrating lymphocytes in colorectal cancer. Comput Struct Biotechnol J. 2022;20:5586-5594.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 5]  [Reference Citation Analysis (0)]
14.  Yang J, Ye H, Fan X, Li Y, Wu X, Zhao M, Hu Q, Ye Y, Wu L, Li Z, Zhang X, Liang C, Wang Y, Xu Y, Li Q, Yao S, You D, Zhao K, Liu Z. Artificial intelligence for quantifying immune infiltrates interacting with stroma in colorectal cancer. J Transl Med. 2022;20:451.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 15]  [Cited by in RCA: 20]  [Article Influence: 5.0]  [Reference Citation Analysis (0)]
15.  Li R, Li J, Wang Y, Liu X, Xu W, Sun R, Xue B, Zhang X, Ai Y, Du Y, Jiang J. The artificial intelligence revolution in gastric cancer management: clinical applications. Cancer Cell Int. 2025;25:111.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 12]  [Cited by in RCA: 11]  [Article Influence: 11.0]  [Reference Citation Analysis (1)]
16.  Zeng J, Song D, Li K, Cao F, Zheng Y. Deep learning model for predicting postoperative survival of patients with gastric cancer. Front Oncol. 2024;14:1329983.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 10]  [Reference Citation Analysis (2)]
17.  Liu DHW, Kim YW, Sefcovicova N, Laye JP, Hewitt LC, Irvine AF, Vromen V, Janssen Y, Davarzani N, Fazzi GE, Jolani S, Melotte V, Magee DR, Kook MC, Kim H, Langer R, Cheong JH, Grabsch HI. Tumour infiltrating lymphocytes and survival after adjuvant chemotherapy in patients with gastric cancer: post-hoc analysis of the CLASSIC trial. Br J Cancer. 2023;128:2318-2325.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 2]  [Cited by in RCA: 16]  [Article Influence: 5.3]  [Reference Citation Analysis (0)]
18.  Ge J, Xiao X, Zhou H, Tang M, Bai J, Zou X, Zhang C, Huang C, Feng X, Liu T, Yi X, Xia X, Liu H, Chen Z. Single-cell profiling reveals tumour cell heterogeneity accompanying a pre-malignant and immunosuppressive microenvironment in gastric adenocarcinoma. Clin Transl Med. 2023;13:e1490.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 6]  [Reference Citation Analysis (0)]
19.  Chen C, Mejbel HA, Pathak T, Krasinskas A, Reid M, Corredor G, Fu P, Willis JE, Madabhushi A. Artificial intelligence defines spatial patterns of tumor-infiltrating lymphocytes highly associated with outcome - a pan-GI cancer study. ESMO Open. 2025;10:105757.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 8]  [Cited by in RCA: 5]  [Article Influence: 5.0]  [Reference Citation Analysis (0)]
20.  Huang W, Wang X, Zhong R, Li Z, Zhou K, Lyu Q, Han JE, Chen T, Islam MT, Yuan Q, Ahmad MU, Chen S, Chen C, Huang J, Xie J, Shen Y, Xiong W, Shen L, Xu Y, Yang F, Xu Z, Li G, Jiang Y. Multimodal radiopathomics signature for prediction of response to immunotherapy-based combination therapy in gastric cancer using interpretable machine learning. Cancer Lett. 2025;631:217930.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 10]  [Reference Citation Analysis (0)]
Footnotes

Peer review: Externally peer reviewed.

Peer-review model: Single blind

Specialty type: Gastroenterology and hepatology

Country of origin: China

Peer-review report’s classification

Scientific quality: Grade B, Grade B

Novelty: Grade B, Grade C

Creativity or innovation: Grade B, Grade C

Scientific significance: Grade B, Grade B

P-Reviewer: Byeon H, Associate Professor, Director, PhD, Research Dean, South Korea; Wang CL, MD, PhD, China S-Editor: Qu XL L-Editor: A P-Editor: Wang CH

Write to the Help Desk