©The Author(s) 2025.
World J Gastroenterol. Dec 21, 2025; 31(47): 112921
Published online Dec 21, 2025. doi: 10.3748/wjg.v31.i47.112921
Published online Dec 21, 2025. doi: 10.3748/wjg.v31.i47.112921
Table 2 Summary of key studies of large language models in the field of gastrointestinal cancer
| Ref. | Year | Models | Objectives | Datasets | Performance | Evaluation |
| Syed et al[29] | 2022 | BERTi | Developed fine-tuned BERTi for integrated colonoscopy reports | 34165 reports | F1-scores of 91.76%, 92.25%, 88.55% for colonoscopy, pathology, and radiology | Manual chart review by 4 expert-guided reviewer |
| Lahat et al[30] | 2023 | GPT | Assessed GPT performance in addressing 110 real-world gastrointestinal inquiries | 110 real-life questions | Moderate accuracy (3.4-3.9/5) for treatment and diagnostic queries | Assessed by three gastroenterologists using a 1-5 scale for accuracy etc. |
| Lee et al[31] | 2023 | GPT-3.5 | Examined GPT-3.5’s responses to eight frequently asked colonoscopy questions | 8 colonoscopy-related questions | GPT answers had extremely low text similarity (0%-16%) | Four gastroenterologists rated the answers on a 7-point Likert scale |
| Emile et al[32] | 2023 | GPT-3.5 | Analyzed GPT-3.5’s ability to generate appropriate responses to CRC questions | 38 CRC questions | 86.8% deemed appropriately, with 95% concordance on 2022 ASCRS guidelines | Three surgery experts assessed answers using ASCRS guidelines |
| Moazzam et al[33] | 2023 | GPT | Investigated the quality of GPT’s responses to pancreatic cancer-related questions | 30 pancreatic cancer-questions | 80% responses were “very good” or “excellent” | Responses were graded by 20 experts against a clinical benchmark |
| Yeo et al[34] | 2023 | GPT | Assessed GPT’s performance in answering questions regarding cirrhosis and HCC | 164 questions about cirrhosis and HCC | 79.1% correctness for cirrhosis and 74% for HCC, but only 47.3% comprehensiveness | Responses were reviewed by two hepatologists and resolved by a 3rd reviewer |
| Cao et al[35] | 2023 | GPT-3.5 | Examined GPT-3.5’s capacity to answer on liver cancer screening and diagnosis | 20 questions | 48% answers were accurate, with frequent errors in LI-RADS categories | Six fellowship-trained physicians from three centers assessed answers |
| Gorelik et al[36] | 2024 | GPT-4 | Evaluated GPT-4’s ability to provide guideline-aligned recommendations | 275 colonoscopy reports | Aligned with experts in 87% of scenarios, showing no significant accuracy gap | Advice assessed by consensus review with multiple experts |
| Gorelik et al[37] | 2023 | GPT-4 | Analyzed GPT-4’s effectiveness in post-colonoscopy management guidance | 20 clinical scenarios | 90% followed guidelines, with 85% correctness and strong agreement (κ = 0.84) | Assessed by two senior gastroenterologists for guideline compliance |
| Zhou et al[38] | 2023 | GPT-3.5 and GPT-4 | Developed a gastric cancer consultation system and automated report generator | 23 medical knowledge questions | 91.3% appropriate gastric cancer advice (GPT-4), 73.9% for GPT-3.5 | The evaluation was conducted by reviewers with medical standards |
| Yang et al[39] | 2025 | RECOVER (LLM) | Designed a LLM-based remote patient monitoring system for postoperative care | 7 design sessions, 5 interviews | Six major design strategies for integrating clinical guidelines and information | Clinical staff reviewed and provided feedback on the design and functionality |
| Kerbage et al[40] | 2024 | GPT-4 | Evaluated GPT-4’s accuracy in responding to IBS, IBD, and CRC screening | 65 questions (45 patients, 20 doctors) | 84% of answers were accurate | Assessed independently by three senior gastroenterologists |
| Tariq et al[41] | 2024 | GPT-3.5, GPT-4, and Bard | Compared the efficacy of GPT-3.5, GPT 4, and Bard (July 2023 version) in answering 47 common colonoscopy patient queries | 47 queries | GPT 4 outperformed GPT-3.5 and Bard, with 91.4% fully accurate responses vs 6.4% and 14.9%, respectively | Responses were scored by two specialists on a 0-2 point scale and resolved by a 3rd reviewer |
| Maida et al[42] | 2025 | GPT-4 | Evaluated GPT-4’s suitability in addressing screening, diagnostic, therapeutic inquiries | 15 CRC screening inquiries | 4.8/6 for CRC screening accuracy, 2.1/3 for completeness scored | Assessment involved 20 experts and 20 non-experts rating the answers |
| Atarere et al[43] | 2024 | BingChat, GPT, YouChat | Tested the appropriateness of GPT, BingChat, and YouChat in patient education and patient-physician communication | 20 questions (15 on CRC screening and 5 patient-related) | GPT and YouChat provided more reliable answers than BingChat, but all models had occasional inaccuracies | Two board-certified physicians and one Gastroenterologist graded the responses |
| Chang et al[44] | 2024 | GPT-4 | Compared GPT-4’s accuracy, reliability, and alignment of colonoscopy recommendations | 505 colonoscopy reports | 85.7% of cases matched USMSTF guidelines | Assessment was conducted by an expert panel under USMSTF guidelines |
| Lim et al[45] | 2024 | GPT-4 | Compared a contextualized GPT model with standard GPT in colonoscopy screening | 62 example use cases | Contextualized GPT-4 outperformed standard GPT-4 | Compare the GPT4 against a model with relevant screening guidelines |
| Munir et al[46] | 2024 | GPT | Evaluated the quality and utility of responses for three GI surgeries | 24 research questions | Modest quality and vary significantly based on the type of procedure | Responses were graded by 45 expert surgeons |
| Truhn et al[47] | 2024 | GPT-4 | Created a structured data parsing module with GPT-4 for clinical text processing | 100 CRC reports | 99% accuracy for T-stage extraction, 96% for N-stage, and 94% for M-stage | Accuracy of GPT-4 was compared with manually extracted data by experts |
| Choo et al[48] | 2024 | GPT | Designed a clinical decision-support system to generate personalized management plans | 30 stage III recurrent CRC patients | 86.7% agree with tumor board decisions, 100% for second-line therapies | The recommendations were compared with the decision plans made by the MDT |
| Huo et al[49] | 2024 | GPT, BingChat, Bard, Claude 2 | Established a multi-AI platform framework to optimize CRC screening recommendations | Responses for 3 patient cases | GPT aligned with guidelines in 66.7% of cases, while other AIs showed greater divergence | Clinician and patient advice was compared to guidelines |
| Pereyra et al[50] | 2024 | GPT-3.5 | Optimized GPT-3.5 for personalized CRC screening recommendations | 238 physicians | GPT scored 4.57/10 for CRC screening, vs 7.72/10 for physicians | Answers were compared against a group of surgeons |
| Peng et al[51] | 2024 | GPT-3.5 | Built a GPT-3.5-powered system for answering CRC-related queries | 131 CRC questions | 63.01 mean accuracy, but low comprehensiveness scores (0.73-0.83) | Two physicians reviewed each response, with a third consulted for discrepancies |
| Ma et al[52] | 2024 | GPT-3.5 | Established GPT-3.5-based quality control for post-esophageal ESD procedures | 165 esophageal ESD cases | 92.5%-100% accuracy across post-esophageal ESD quality metrics | Two QC members and a senior supervisor conducted assessment |
| Cohen et al[53] | 2025 | LLaMA-2, Mistral-v0.1 | Explored the ability of LLMs to extract PD-L1 biomarker details for research purposes | 232 EHRs from 10 cancer types | Fine-tuned LLMs outperformed LSTM trained on > 10000 examples | Assessed by 3 clinical experts against manually curated answers |
| Scherbakov et al[54] | 2025 | Mixtral 8 × 7 B | Assessed LLM to extract stressful events from social history of clinical notes | 109556 patients, 375334 notes | Arrest or incarceration (OR = 0.26, 95%CI: 0.06-0.77) | One human reviewer assessed the precision and recall of extracted events |
| Chatziisaak et al[55] | 2025 | GPT-4 | Evaluated the concordance of therapeutic recommendations generated by GPT | 100 consecutive CRC patients | 72.5% complete concordance, 10.2% partial concordance, and 17.3% discordance | Three reviewers independently assessed concordance with MDT |
| Saraiva et al[56] | 2025 | GPT-4 | Assessed GPT-4’s performance in interpreting images in gastroenterology | 740 images | Capsule endoscopy: Accuracies 50.0%-90.0% (AUCs 0.50-0.90) | Three experts reviewed and labeled images for CE |
| Siu et al[57] | 2025 | GPT-4 | Evaluated the efficacy, quality, and readability of GPT-4’s responses | 8 patient-style questions | Accurate (40), safe (4.25), appropriate (4.00), actionable (4.00), effective (4.00) | Evaluated by 8 colorectal surgeons |
| Horesh et al[58] | 2025 | GPT-3.5 | Evaluated management recommendations of GPT in clinical settings | 15 colorectal or anal cancer patients | Rating 48 for GPT recommendations, 4.11 for decision justification | Evaluated by 3 experienced colorectal surgeons |
| Ellison et al[59] | 2025 | GPT-3.5, Perplexity | Compared readability using different prompts | 52 colorectal surgery materials | Average 7.0-9.8, Ease 53.1-65.0, Modified 9.6-11.5 | Compared mean scores between baseline and documents generated by AI |
| Ramchandani et al[60] | 2025 | GPT-4 | Validated the use of GPT-4 for identifying articles discussing perioperative and preoperative risk factors for esophagectomy | 1967 studies for title and abstract screening | Perioperative: Agreement rate = 85.58%, AUC = 0.87. Preoperative: Agreement rate = 78.75%, AUC = 0.75 | Decisions were compared with those of three independent human reviewers |
| Zhang et al[61] | 2025 | GPT-4, DeepSeek, GLM-4, Qwen, LLaMa3 | To evaluate the consistency of LLMs in generating diagnostic records for hepatobiliary cases using the HepatoAudit dataset | 684 medical records covering 20 hepatobiliary diseases | Precision: GPT-4 reached a maximum of 93.42%. Recall: Generally below 70%, with some diseases below 40% | Professional physicians manually verified and corrected all the data |
| Spitzl et al[62] | 2025 | Claude-3.5, GPT-4o, DeepSeekV3, Gemini 2 | Assessed the capability of state-of-the-art LLMs to classify liver lesions based solely on textual descriptions from MRI reports | 88 fictitious MRI reports designed to resemble real clinical documentation | Micro F1-score and macro F1-score: Claude 3.5 Sonnet 0.91 and 0.78, GPT-4o 0.76 and 0.63, DeepSeekV3 0.84 and 0.70, Gemini 2.0 Flash 0.69 and 0.55 | Model performance was assessed using micro and macro F1-scores benchmarked against ground truth labels |
| Sheng et al[63] | 2025 | GPT-4o and Gemini | Investigated the diagnostic accuracies for focal liver lesions | 228 adult patients with CT/MRI reports | Two-step GPT-4o, single-step GPT-4o and single-step Gemini (78.9%, 68.0%, 73.2%) | Six radiologists reviewed the images and clinical information in two rounds (alone, with LLM) |
| Williams et al[64] | 2025 | GPT-4-32K | Determined LLM extract reasons for a lack of follow-up colonoscopy | 846 patients' clinical notes | Overall accuracy: 89.3%, reasons: Refused/not interested (35.2%) | A physician reviewer checked 10% of LLM-generated labels |
| Lu et al[65] | 2025 | MoE-HRS | Used a novel MoE combined with LLMs for risk prediction and personalized healthcare recommendations | SNPs, medical and lifestyle data from United Kingdom Biobank | MoE-HRS outperformed state-of-the-art cancer risk prediction models in terms of ROC-AUC, precision, recall, and F1 score | LLMs-generated advice were validated by clinical medical staff |
| Yang et al[66] | 2025 | GPT-4 | Explored the use of LLMs to enhance doctor-patient communication | 698 pathology reports of tumors | Average communication time decreased by over 70%, from 35 to 10 min (P < 0.001) | Pathologists evaluated the consistency between original and AI reports |
| Jain et al[67] | 2025 | GPT-4, GPT-3.5, Gemini | Studied the performance of LLMs across 20 clinicopathologic scenarios in gastrointestinal pathology | 20 clinicopathologic scenarios in GI | Diagnostic accuracy: Gemini Advanced (95%, P = 0.01), GPT-4 (90%, P = 0.05), GPT-3.5 (65%) | Two fellowship-trained pathologists independently assessed the responses of the models |
| Xu et al[68] | 2025 | GPT-4, GPT-4o, Gemini | Assessed the performance of LLMs in predicting immunotherapy response in unresectable HCC | Multimodal data from 186 patients | Accuracy and sensitivity: GPT-4o (65% and 47%) Gemini-GPT (68% and 58%). Physicians (72% and 70%) | Six physicians (three radiologists and three oncologists) independently assessed the same dataset |
| Deroy et al[69] | 2025 | GPT-3.5 Turbo | Explored the potential of LLMs as a question-answering (QA) tool | 30 training and 50 testing queries | A1: 0.546 (maximum value); A2: 0.881 (maximum value across three runs) | Model-generated answers were compared to the gold standard |
| Ye et al[70] | 2025 | BioBERT-based | Proposed a novel framework that incorporates clinical features to enhance multi-omics clustering for cancer subtyping | Six cancer datasets across three omics levels | Mean survival score of 2.20, significantly higher than other methods | Three independent clinical experts review and validate the clustering results |
- Citation: Shi L, Huang R, Zhao LL, Guo AJ. Foundation models: Insights and implications for gastrointestinal cancer. World J Gastroenterol 2025; 31(47): 112921
- URL: https://www.wjgnet.com/1007-9327/full/v31/i47/112921.htm
- DOI: https://dx.doi.org/10.3748/wjg.v31.i47.112921