Copyright: ©Author(s) 2026. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution-NonCommercial (CC BY-NC 4.0) license. No commercial re-use. See permissions. Published by Baishideng Publishing Group Inc.
World J Gastroenterol. Oct 7, 2026; 32(37): 119857
Published online Oct 7, 2026. doi: 10.3748/wjg.119857
Published online Oct 7, 2026. doi: 10.3748/wjg.119857
Evaluating large language models in Helicobacter pylori-related question answering: From knowledge tests to patient queries
Shi-Ping Sun, Yi Li, Han Min, Department of Gastroenterology, The Affiliated Suzhou Hospital of Nanjing Medical University, Suzhou Municipal Hospital, Gusu School, Nanjing Medical University, Suzhou 215000, Jiangsu Province, China
Dan-Ye Niu, Department of Clinical Nutrition, The Affiliated Suzhou Hospital of Nanjing Medical University, Suzhou Municipal Hospital, Gusu School, Nanjing Medical University, Suzhou 215000, Jiangsu Province, China
Ming-Kai Yuan, Department of Hepatobiliary and Pancreatic Surgery, The Affiliated Suzhou Hospital of Nanjing Medical University, Suzhou Municipal Hospital, Gusu School, Nanjing Medical University, Suzhou 215000, Jiangsu Province, China
Li Liu, Department of Gastroenterology, The Affiliated Suzhou Hospital of Nanjing Medical University, Suzhou Municipal Hospital, Suzhou Women and Children’s Health Hospital, Suzhou 215000, Jiangsu Province, China
Co-first authors: Shi-Ping Sun and Dan-Ye Niu.
Author contributions: Sun SP, Niu DY, and Min H conceptualized the study, and drafted the manuscript; Sun SP, Niu DY, and Yuan MK developed the methodology; Sun SP, Niu DY, Li Y, and Min H curated the data; Sun SP and Niu DY performed the formal analysis, and contributed equally as co-first authors; Sun SP, Niu DY, Li Y, and Min H conducted the investigation; Sun SP, Liu L, and Min H reviewed and edited the manuscript; Yuan MK, Liu L and Min H supervised the study. All authors approved the final version to publish.
AI contribution statement: No AI tools (including ChatGPT, Grammarly, DeepL, or any other AI-based tools) were used in the preparation of this manuscript. No part of the main text (including the Abstract, Introduction, Materials and Methods, Results, Discussion, or Conclusion) was generated by AI. AI tools were not used for language polishing, translation, data analysis, or writing assistance. AI tools did not participate in the study design, data analysis, or interpretation of results. No images in the manuscript were generated by AI. We confirm that the manuscript was entirely prepared by the authors.
Supported by Suzhou Major Disease Multicenter Clinical Research Project, No. DZXYJ202508.
Institutional review board statement: This study was approved by the Ethical Committee of the Suzhou Hospital Affiliated to Nanjing Medical University, No. K-2025-162-K01.
Informed consent statement: All participants provided informed consent prior to participation.
Conflict-of-interest statement: All the authors report no relevant conflicts of interest for this article.
STROBE statement: The authors have read the STROBE Statement-checklist of items, and the manuscript was prepared and revised according to the STROBE Statement-checklist of items.
Data sharing statement: All data are included in the article and its Supplementary material.
Corresponding author: Han Min, MD, Full Professor, Department of Gastroenterology, The Affiliated Suzhou Hospital of Nanjing Medical University, Suzhou Municipal Hospital, Gusu School, Nanjing Medical University, No. 26 Daoqian Street, Gusu District, Suzhou 215000, Jiangsu Province, China. minhan1981@njmu.edu.cn
Received: February 9, 2026
Revised: March 20, 2026
Accepted: May 28, 2026
Published online: October 7, 2026
Processing time: 205 Days and 19.6 Hours
Revised: March 20, 2026
Accepted: May 28, 2026
Published online: October 7, 2026
Processing time: 205 Days and 19.6 Hours
Core Tip
Core Tip: This three-phase study evaluated large language models for Chinese Helicobacter pylori health-information support, progressing from multiple-choice testing to guideline-based questions and real-world patient queries. Performance was clearly task dependent: The model that performed best in structured assessments was not necessarily the most effective in patient interaction. In phase 3, DeepSeek showed better correctness, completeness, perceived readability, and patient satisfaction. These findings support multidimensional evaluation and clinician oversight before patient-facing deployment.