BPG is committed to discovery and dissemination of knowledge
Review
©The Author(s) 2025.
World J Gastroenterol. Dec 21, 2025; 31(47): 112921
Published online Dec 21, 2025. doi: 10.3748/wjg.v31.i47.112921
Table 2 Summary of key studies of large language models in the field of gastrointestinal cancer
Ref.
Year
Models
Objectives
Datasets
Performance
Evaluation
Syed et al[29]2022BERTiDeveloped fine-tuned BERTi for integrated colonoscopy reports34165 reportsF1-scores of 91.76%, 92.25%, 88.55% for colonoscopy, pathology, and radiologyManual chart review by 4 expert-guided reviewer
Lahat et al[30]2023GPTAssessed GPT performance in addressing 110 real-world gastrointestinal inquiries110 real-life questions Moderate accuracy (3.4-3.9/5) for treatment and diagnostic queriesAssessed by three gastroenterologists using a 1-5 scale for accuracy etc.
Lee et al[31]2023GPT-3.5Examined GPT-3.5’s responses to eight frequently asked colonoscopy questions8 colonoscopy-related questionsGPT answers had extremely low text similarity (0%-16%)Four gastroenterologists rated the answers on a 7-point Likert scale
Emile et al[32]2023GPT-3.5Analyzed GPT-3.5’s ability to generate appropriate responses to CRC questions38 CRC questions86.8% deemed appropriately, with 95% concordance on 2022 ASCRS guidelinesThree surgery experts assessed answers using ASCRS guidelines
Moazzam et al[33]2023GPTInvestigated the quality of GPT’s responses to pancreatic cancer-related questions30 pancreatic cancer-questions80% responses were “very good” or “excellent”Responses were graded by 20 experts against a clinical benchmark
Yeo et al[34]2023GPTAssessed GPT’s performance in answering questions regarding cirrhosis and HCC164 questions about cirrhosis and HCC79.1% correctness for cirrhosis and 74% for HCC, but only 47.3% comprehensivenessResponses were reviewed by two hepatologists and resolved by a 3rd reviewer
Cao et al[35]2023GPT-3.5Examined GPT-3.5’s capacity to answer on liver cancer screening and diagnosis20 questions48% answers were accurate, with frequent errors in LI-RADS categoriesSix fellowship-trained physicians from three centers assessed answers
Gorelik et al[36]2024GPT-4Evaluated GPT-4’s ability to provide guideline-aligned recommendations275 colonoscopy reportsAligned with experts in 87% of scenarios, showing no significant accuracy gapAdvice assessed by consensus review with multiple experts
Gorelik et al[37]2023GPT-4Analyzed GPT-4’s effectiveness in post-colonoscopy management guidance20 clinical scenarios90% followed guidelines, with 85% correctness and strong agreement (κ = 0.84)Assessed by two senior gastroenterologists for guideline compliance
Zhou et al[38]2023GPT-3.5 and GPT-4Developed a gastric cancer consultation system and automated report generator23 medical knowledge questions91.3% appropriate gastric cancer advice (GPT-4), 73.9% for GPT-3.5The evaluation was conducted by reviewers with medical standards
Yang et al[39]2025RECOVER (LLM)Designed a LLM-based remote patient monitoring system for postoperative care7 design sessions, 5 interviewsSix major design strategies for integrating clinical guidelines and informationClinical staff reviewed and provided feedback on the design and functionality
Kerbage et al[40]2024GPT-4Evaluated GPT-4’s accuracy in responding to IBS, IBD, and CRC screening65 questions (45 patients, 20 doctors)84% of answers were accurateAssessed independently by three senior gastroenterologists
Tariq et al[41]2024GPT-3.5, GPT-4, and BardCompared the efficacy of GPT-3.5, GPT 4, and Bard (July 2023 version) in answering 47 common colonoscopy patient queries47 queriesGPT 4 outperformed GPT-3.5 and Bard, with 91.4% fully accurate responses vs 6.4% and 14.9%, respectivelyResponses were scored by two specialists on a 0-2 point scale and resolved by a 3rd reviewer
Maida et al[42]2025GPT-4Evaluated GPT-4’s suitability in addressing screening, diagnostic, therapeutic inquiries15 CRC screening inquiries4.8/6 for CRC screening accuracy, 2.1/3 for completeness scoredAssessment involved 20 experts and 20 non-experts rating the answers
Atarere et al[43]2024BingChat, GPT, YouChatTested the appropriateness of GPT, BingChat, and YouChat in patient education and patient-physician communication20 questions (15 on CRC screening and 5 patient-related)GPT and YouChat provided more reliable answers than BingChat, but all models had occasional inaccuraciesTwo board-certified physicians and one Gastroenterologist graded the responses
Chang et al[44]2024GPT-4Compared GPT-4’s accuracy, reliability, and alignment of colonoscopy recommendations505 colonoscopy reports85.7% of cases matched USMSTF guidelinesAssessment was conducted by an expert panel under USMSTF guidelines
Lim et al[45]2024GPT-4Compared a contextualized GPT model with standard GPT in colonoscopy screening62 example use casesContextualized GPT-4 outperformed standard GPT-4Compare the GPT4 against a model with relevant screening guidelines
Munir et al[46]2024GPTEvaluated the quality and utility of responses for three GI surgeries24 research questionsModest quality and vary significantly based on the type of procedureResponses were graded by 45 expert surgeons
Truhn et al[47]2024GPT-4Created a structured data parsing module with GPT-4 for clinical text processing100 CRC reports99% accuracy for T-stage extraction, 96% for N-stage, and 94% for M-stageAccuracy of GPT-4 was compared with manually extracted data by experts
Choo et al[48]2024GPTDesigned a clinical decision-support system to generate personalized management plans30 stage III recurrent CRC patients86.7% agree with tumor board decisions, 100% for second-line therapiesThe recommendations were compared with the decision plans made by the MDT
Huo et al[49]2024GPT, BingChat, Bard, Claude 2Established a multi-AI platform framework to optimize CRC screening recommendationsResponses for 3 patient casesGPT aligned with guidelines in 66.7% of cases, while other AIs showed greater divergenceClinician and patient advice was compared to guidelines
Pereyra et al[50]2024GPT-3.5Optimized GPT-3.5 for personalized CRC screening recommendations238 physiciansGPT scored 4.57/10 for CRC screening, vs 7.72/10 for physiciansAnswers were compared against a group of surgeons
Peng et al[51]2024GPT-3.5Built a GPT-3.5-powered system for answering CRC-related queries131 CRC questions63.01 mean accuracy, but low comprehensiveness scores (0.73-0.83)Two physicians reviewed each response, with a third consulted for discrepancies
Ma et al[52]2024GPT-3.5Established GPT-3.5-based quality control for post-esophageal ESD procedures165 esophageal ESD cases92.5%-100% accuracy across post-esophageal ESD quality metricsTwo QC members and a senior supervisor conducted assessment
Cohen et al[53]2025LLaMA-2, Mistral-v0.1Explored the ability of LLMs to extract PD-L1 biomarker details for research purposes232 EHRs from 10 cancer typesFine-tuned LLMs outperformed LSTM trained on > 10000 examplesAssessed by 3 clinical experts against manually curated answers
Scherbakov et al[54]2025Mixtral 8 × 7 BAssessed LLM to extract stressful events from social history of clinical notes109556 patients, 375334 notesArrest or incarceration (OR = 0.26, 95%CI: 0.06-0.77)One human reviewer assessed the precision and recall of extracted events
Chatziisaak et al[55]2025GPT-4Evaluated the concordance of therapeutic recommendations generated by GPT100 consecutive CRC patients72.5% complete concordance, 10.2% partial concordance, and 17.3% discordanceThree reviewers independently assessed concordance with MDT
Saraiva et al[56]2025GPT-4Assessed GPT-4’s performance in interpreting images in gastroenterology740 imagesCapsule endoscopy: Accuracies 50.0%-90.0% (AUCs 0.50-0.90)Three experts reviewed and labeled images for CE
Siu et al[57]2025GPT-4Evaluated the efficacy, quality, and readability of GPT-4’s responses8 patient-style questionsAccurate (40), safe (4.25), appropriate (4.00), actionable (4.00), effective (4.00)Evaluated by 8 colorectal surgeons
Horesh et al[58]2025GPT-3.5Evaluated management recommendations of GPT in clinical settings15 colorectal or anal cancer patientsRating 48 for GPT recommendations, 4.11 for decision justificationEvaluated by 3 experienced colorectal surgeons
Ellison et al[59]2025GPT-3.5, PerplexityCompared readability using different prompts52 colorectal surgery materialsAverage 7.0-9.8, Ease 53.1-65.0, Modified 9.6-11.5Compared mean scores between baseline and documents generated by AI
Ramchandani et al[60]2025GPT-4Validated the use of GPT-4 for identifying articles discussing perioperative and preoperative risk factors for esophagectomy1967 studies for title and abstract screeningPerioperative: Agreement rate = 85.58%, AUC = 0.87. Preoperative: Agreement rate = 78.75%, AUC = 0.75Decisions were compared with those of three independent human reviewers
Zhang et al[61]2025GPT-4, DeepSeek, GLM-4, Qwen, LLaMa3To evaluate the consistency of LLMs in generating diagnostic records for hepatobiliary cases using the HepatoAudit dataset684 medical records covering 20 hepatobiliary diseasesPrecision: GPT-4 reached a maximum of 93.42%. Recall: Generally below 70%, with some diseases below 40%Professional physicians manually verified and corrected all the data
Spitzl et al[62]2025Claude-3.5, GPT-4o, DeepSeekV3, Gemini 2Assessed the capability of state-of-the-art LLMs to classify liver lesions based solely on textual descriptions from MRI reports88 fictitious MRI reports designed to resemble real clinical documentationMicro F1-score and macro F1-score: Claude 3.5 Sonnet 0.91 and 0.78, GPT-4o 0.76 and 0.63, DeepSeekV3 0.84 and 0.70, Gemini 2.0 Flash 0.69 and 0.55Model performance was assessed using micro and macro F1-scores benchmarked against ground truth labels
Sheng et al[63]2025GPT-4o and GeminiInvestigated the diagnostic accuracies for focal liver lesions228 adult patients with CT/MRI reportsTwo-step GPT-4o, single-step GPT-4o and single-step Gemini (78.9%, 68.0%, 73.2%)Six radiologists reviewed the images and clinical information in two rounds (alone, with LLM)
Williams et al[64]2025GPT-4-32KDetermined LLM extract reasons for a lack of follow-up colonoscopy846 patients' clinical notesOverall accuracy: 89.3%, reasons: Refused/not interested (35.2%)A physician reviewer checked 10% of LLM-generated labels
Lu et al[65]2025MoE-HRSUsed a novel MoE combined with LLMs for risk prediction and personalized healthcare recommendationsSNPs, medical and lifestyle data from United Kingdom BiobankMoE-HRS outperformed state-of-the-art cancer risk prediction models in terms of ROC-AUC, precision, recall, and F1 scoreLLMs-generated advice were validated by clinical medical staff
Yang et al[66]2025GPT-4Explored the use of LLMs to enhance doctor-patient communication698 pathology reports of tumorsAverage communication time decreased by over 70%, from 35 to 10 min (P < 0.001)Pathologists evaluated the consistency between original and AI reports
Jain et al[67]2025GPT-4, GPT-3.5, GeminiStudied the performance of LLMs across 20 clinicopathologic scenarios in gastrointestinal pathology20 clinicopathologic scenarios in GIDiagnostic accuracy: Gemini Advanced (95%, P = 0.01), GPT-4 (90%, P = 0.05), GPT-3.5 (65%)Two fellowship-trained pathologists independently assessed the responses of the models
Xu et al[68]2025GPT-4, GPT-4o, GeminiAssessed the performance of LLMs in predicting immunotherapy response in unresectable HCCMultimodal data from 186 patientsAccuracy and sensitivity: GPT-4o (65% and 47%) Gemini-GPT (68% and 58%). Physicians (72% and 70%)Six physicians (three radiologists and three oncologists) independently assessed the same dataset
Deroy et al[69]2025GPT-3.5 TurboExplored the potential of LLMs as a question-answering (QA) tool30 training and 50 testing queriesA1: 0.546 (maximum value); A2: 0.881 (maximum value across three runs)Model-generated answers were compared to the gold standard
Ye et al[70]2025BioBERT-basedProposed a novel framework that incorporates clinical features to enhance multi-omics clustering for cancer subtypingSix cancer datasets across three omics levels Mean survival score of 2.20, significantly higher than other methodsThree independent clinical experts review and validate the clustering results


Write to the Help Desk