BPG is committed to discovery and dissemination of knowledge
Observational Study Open Access
Copyright: ©Author(s) 2026. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution-NonCommercial (CC BY-NC 4.0) license. No commercial re-use. See permissions. Published by Baishideng Publishing Group Inc.
World J Gastroenterol. Oct 7, 2026; 32(37): 119857
Published online Oct 7, 2026. doi: 10.3748/wjg.119857
Evaluating large language models in Helicobacter pylori-related question answering: From knowledge tests to patient queries
Shi-Ping Sun, Yi Li, Han Min, Department of Gastroenterology, The Affiliated Suzhou Hospital of Nanjing Medical University, Suzhou Municipal Hospital, Gusu School, Nanjing Medical University, Suzhou 215000, Jiangsu Province, China
Dan-Ye Niu, Department of Clinical Nutrition, The Affiliated Suzhou Hospital of Nanjing Medical University, Suzhou Municipal Hospital, Gusu School, Nanjing Medical University, Suzhou 215000, Jiangsu Province, China
Ming-Kai Yuan, Department of Hepatobiliary and Pancreatic Surgery, The Affiliated Suzhou Hospital of Nanjing Medical University, Suzhou Municipal Hospital, Gusu School, Nanjing Medical University, Suzhou 215000, Jiangsu Province, China
Li Liu, Department of Gastroenterology, The Affiliated Suzhou Hospital of Nanjing Medical University, Suzhou Municipal Hospital, Suzhou Women and Children’s Health Hospital, Suzhou 215000, Jiangsu Province, China
ORCID number: Han Min (0009-0005-8576-2243).
Co-first authors: Shi-Ping Sun and Dan-Ye Niu.
Author contributions: Sun SP, Niu DY, and Min H conceptualized the study, and drafted the manuscript; Sun SP, Niu DY, and Yuan MK developed the methodology; Sun SP, Niu DY, Li Y, and Min H curated the data; Sun SP and Niu DY performed the formal analysis, and contributed equally as co-first authors; Sun SP, Niu DY, Li Y, and Min H conducted the investigation; Sun SP, Liu L, and Min H reviewed and edited the manuscript; Yuan MK, Liu L and Min H supervised the study. All authors approved the final version to publish.
AI contribution statement: No AI tools (including ChatGPT, Grammarly, DeepL, or any other AI-based tools) were used in the preparation of this manuscript. No part of the main text (including the Abstract, Introduction, Materials and Methods, Results, Discussion, or Conclusion) was generated by AI. AI tools were not used for language polishing, translation, data analysis, or writing assistance. AI tools did not participate in the study design, data analysis, or interpretation of results. No images in the manuscript were generated by AI. We confirm that the manuscript was entirely prepared by the authors.
Supported by Suzhou Major Disease Multicenter Clinical Research Project, No. DZXYJ202508.
Institutional review board statement: This study was approved by the Ethical Committee of the Suzhou Hospital Affiliated to Nanjing Medical University, No. K-2025-162-K01.
Informed consent statement: All participants provided informed consent prior to participation.
Conflict-of-interest statement: All the authors report no relevant conflicts of interest for this article.
STROBE statement: The authors have read the STROBE Statement-checklist of items, and the manuscript was prepared and revised according to the STROBE Statement-checklist of items.
Data sharing statement: All data are included in the article and its Supplementary material.
Corresponding author: Han Min, MD, Full Professor, Department of Gastroenterology, The Affiliated Suzhou Hospital of Nanjing Medical University, Suzhou Municipal Hospital, Gusu School, Nanjing Medical University, No. 26 Daoqian Street, Gusu District, Suzhou 215000, Jiangsu Province, China. minhan1981@njmu.edu.cn
Received: February 9, 2026
Revised: March 20, 2026
Accepted: May 28, 2026
Published online: October 7, 2026
Processing time: 205 Days and 14.3 Hours

Abstract
BACKGROUND

Large language models (LLMs) are increasingly used for patient education, but their reliability for Chinese Helicobacter pylori (H. pylori)-related counseling remains unclear.

AIM

To evaluate the performance of LLMs in Chinese H. pylori-related question answering across structured tests, guideline-based questions, and real-world patient queries.

METHODS

This three-phase comparative study evaluated four LLMs. In phase 1, models answered 112 H. pylori single-best-answer questions. In phase 2, they answered 30 guideline-based clinical questions, and three blinded senior gastroenterologists rated correctness, completeness, readability, helpfulness, and safety. Readability was also assessed using the language difficulty understanding tool for general public. Based on performance in phases 1 and 2, two models were selected for phase 3, in which 40 patients provided 120 real-world questions. Clinicians rated responses, and patients rated satisfaction and perceived readability.

RESULTS

In phase 1, accuracy ranged from 75.0% to 85.0%. In phase 2, ChatGPT5 achieved the highest scores across all domains, including correctness (4.63), completeness (4.89), readability (4.88), helpfulness (4.86), and safety (4.59). In phase 3, DeepSeek outperformed ChatGPT5 in correctness (4.80 vs 3.85, P < 0.001), completeness (4.80 vs 4.05, P = 0.008), perceived readability (4.65 vs 4.13, P = 0.006), and patient satisfaction (4.73 vs 4.05, P < 0.001).

CONCLUSION

LLM performance in H. pylori-related health-information support was task dependent. Strong performance in structured assessments did not necessarily predict better real-world patient interaction. Multidimensional evaluation and clinician oversight are needed before patient-facing use.

Key Words: Helicobacter pylori; Large language models; Patient education; Readability; Patient satisfaction; Clinical decision support systems; Artificial intelligence

Core Tip: This three-phase study evaluated large language models for Chinese Helicobacter pylori health-information support, progressing from multiple-choice testing to guideline-based questions and real-world patient queries. Performance was clearly task dependent: The model that performed best in structured assessments was not necessarily the most effective in patient interaction. In phase 3, DeepSeek showed better correctness, completeness, perceived readability, and patient satisfaction. These findings support multidimensional evaluation and clinician oversight before patient-facing deployment.



INTRODUCTION

Helicobacter pylori (H. pylori) infection is one of the most prevalent chronic bacterial infections worldwide. A recent systematic review and meta-analysis including 1748 studies from 111 countries reported that the pooled crude prevalence of H. pylori infection during 2015-2022 was approximately 43.9% (95% confidence interval: 42.3%-45.5%) among adults and remained as high as 35.1% (95% confidence interval: 30.5%-40.1%) among children and adolescents, with marked variation across World Health Organization regions, highlighting its ongoing public health significance[1]. H. pylori infection is a major etiological factor for gastric cancer; community-based screening and eradication in high-risk areas can reduce the incidence of gastric cancer[2].

In recent years, the expanding evidence base in evidence-based medicine has driven continuous updates of international and national guidelines/consensus statements on H. pylori diagnosis and treatment. For example, the Maastricht VI/Florence consensus systematically evaluated and issued recommendations regarding indications, diagnostic approaches, treatment strategies, gastric cancer prevention, and the relationship between H. pylori infection and the gut microbiome[3]. Nevertheless, H. pylori management remains challenging in real-world practice. Patients have substantial information needs regarding testing and eradication workflows, adverse drug reactions, timing of post-treatment confirmation, management of household contacts, and lifestyle interventions. Meanwhile, increasing antibiotic resistance has made the selection of appropriate eradication regimens and the improvement of adherence and treatment effectiveness key unresolved issues in routine care[4]. Accordingly, providing patients with accurate, clear, and actionable health information is an essential component of standardized H. pylori management and improved eradication success. Beyond its high prevalence, H. pylori remains a major yet under-recognized global health problem. As a group 1 carcinogen, persistent H. pylori infection contributes to the burden of gastric cancer and peptic ulcer disease, while many infected individuals remain undiagnosed until symptoms or complications prompt medical attention. Delayed recognition and management may also increase avoidable healthcare utilization and costs. In this context, scalable tools that can deliver early, accurate, and comprehensible patient education may help narrow the gap between guideline recommendations and patient awareness, thereby supporting timely testing, treatment, and follow-up[5-7].

Concurrently, advances in generative artificial intelligence have brought large language models (LLMs) into the public arena for health information seeking. LLMs can engage in natural-language dialogue, summarize and explain medical knowledge, and generate patient education materials, and they have shown broad potential in medical education and clinical information processing[8]. For patients with H. pylori infection, LLMs may facilitate rapid understanding of key issues such as why eradication is needed, how to take medications, how to handle adverse effects, when to perform post-treatment testing, and how to prevent reinfection, thereby potentially improving health literacy and care efficiency.

However, the reliability and safety of LLMs for medical question answering remain controversial. Prior studies and reviews have noted that LLMs may exhibit outdated knowledge, factual errors, or “hallucinated” content, with substantial variability across clinical topics, which constrains their direct use in real-world healthcare settings[9]. In the context of H. pylori infection, where guidelines are frequently updated, regimen selection depends on local resistance patterns, and patient concerns span the entire pathway from diagnosis and treatment to confirmation of eradication and prevention, there is a clear need for systematic, multidimensional evaluation of LLM responses in terms of accuracy, completeness, readability, helpfulness, and safety to inform appropriate use by clinicians and patients.

Therefore, we compared multiple currently available LLMs in H. pylori-related question answering and developed a stepwise evaluation framework: First, assessing knowledge accuracy using structured single-best-answer questions; second, conducting blinded expert ratings across multiple quality domains for guideline-based key clinical questions while quantifying readability; and third, evaluating real-world utility in patient-generated questions by integrating clinician ratings and patient-reported experience measures. This study aims to provide evidence to guide the selection and responsible use of LLMs for health information support in patients with H. pylori infection.

MATERIALS AND METHODS
Study design

We conducted a three-phase, stepwise comparative study to systematically evaluate the performance of currently available LLMs in answering H. pylori-related questions in Chinese (Figure 1). The workflow comprised: (1) A multiple-choice test to quantify objective accuracy on structured medical knowledge; (2) A guideline-oriented clinical question assessment integrating blinded expert ratings and quantitative Chinese readability evaluation; and (3) A real-world patient question assessment combining clinician ratings and patient-reported experience measures to determine practical utility for clinical information support.

Figure 1
Figure 1 Overall study design and workflow of the three-phase evaluation framework. MCQ: Multiple-choice question; LLM: Large language model; LDU-TGP: Ludong University Text Grading Platform.

In phase 1, 112 H. pylori-related single-best-answer multiple-choice questions (MCQs) were administered to four LLMs. To ensure consistency in adjudication, each model was instructed to provide a single explicit option for each item. Model-selected options were compared with reference answers to calculate overall accuracy and category-specific accuracy across five domains (definition, diagnosis, clinical manifestations, treatment, follow-up and prevention).

In phase 2, gastroenterologists developed 30 H. pylori-related clinical questions based on authoritative clinical guidelines/consensus statements and submitted them to the same four LLMs. Each question was asked three times to each model using identical wording (all three runs completed on the same day), and each run was conducted in a new, independent chat session to eliminate potential carryover from conversational memory. All responses were exported as plain text and stripped of any features that might reveal model identity. For blinded evaluation, all responses were randomly shuffled prior to rating; additionally, for each model-question pair, each rater received only one response randomly selected from the three repetitions to reduce recognition bias. Three senior gastroenterologists (associate chief physician level or above) independently rated each response using a 5-point Likert scale across five dimensions: Correctness, completeness, readability, helpfulness, and safety. For each response and each dimension, the final score was defined as the median of the three raters’ scores. In parallel, response readability was quantified using a Chinese text grading platform.

Based on overall performance in phases 1 and 2, the two best-performing LLMs were selected for phase 3. In phase 3, we collected 120 real-world H. pylori-related clinical questions from 40 patients (three questions per patient). Patient questions were randomly assigned to the two selected models, and each question was asked once. Responses were similarly converted to plain text and de-identified prior to evaluation. The same three gastroenterologists rated the responses using the same Likert-scale framework, with the median of the three ratings used as the final score. Patients also rated satisfaction and perceived readability of the model responses using 5-point scales.

Across all phases, standardized Chinese prompts were used for all LLMs, instructing the model to respond as a “professional gastroenterologist” to either guideline-based clinical questions or patient-generated questions. All interactions were conducted in independent sessions to enhance comparability and reproducibility. Model names/versions and access dates for each study phase are provided in the Supplementary material.

Participants

In phase 3, we enrolled 40 patients presenting for clinical care with H. pylori-related conditions. Each patient submitted three clinical questions related to their own H. pylori infection, yielding a total of 120 real-world patient questions. Before model prompting and subsequent evaluation, all questions were de-identified to ensure removal of any personally identifiable information.

Inclusion criteria: (1) Age ≥ 18 years; (2) Clinically confirmed H. pylori infection, defined by at least one positive test result among 13C/14C urea breath test, stool H. pylori antigen test, or endoscopic biopsy–based testing; (3) Ability to communicate in Chinese and complete the question submission procedure; and (4) Voluntary participation in the study.

Exclusion criteria: (1) Significant communication barriers or inability to complete the study procedures; (2) Critical illness or any condition deemed by the investigators to make participation inappropriate; and (3) Submitted questions unrelated to H. pylori infection or containing substantial third-party information that precluded valid evaluation.

Question submission and allocation procedure. Under the guidance of study personnel, patients submitted open-ended questions reflecting their real concerns regarding H. pylori diagnosis, eradication treatment, post-treatment confirmation, and prevention. Each question was asked once in phase 3. Patient questions were randomly assigned to the two LLMs selected for phase 3 to minimize potential imbalance in question difficulty between models. Model outputs were converted to plain text and stripped of model-identifying features before proceeding to clinician rating and patient evaluation.

Question banks

We used three sources of H. pylori-related questions corresponding to the three study phases: A structured MCQ bank (phase 1), a guideline-oriented clinical question bank (phase 2), and a real-world patient question bank (phase 3).

Phase 1: MCQ bank. The phase 1 question bank was derived from physician test questions used in the Department of Gastroenterology, Suzhou Hospital Affiliated to Nanjing Medical University. A total of 112 H. pylori-related single-best-answer MCQs were selected. To enable stratified analyses, three gastroenterologists categorized the questions into five domains based on the assessed content: Definition, diagnosis, clinical manifestations, treatment, and follow-up and prevention. Each MCQ had a reference answer, and models were instructed to output a single option for accuracy calculations.

Phase 2: Guideline-oriented clinical question bank. The phase 2 question set was developed by three gastroenterologists based on authoritative clinical guidelines/consensus statements, primarily referencing the Maastricht consensus, American College of Gastroenterology guidelines, and relevant Chinese guidelines/consensus documents. The questions were formulated to cover key decision points in H. pylori management (e.g., diagnostic strategies, selection of eradication regimens, medication precautions and management of adverse effects, post-treatment confirmation and prevention of reinfection, and special-population management). A total of 30 clinical questions were finalized through group discussion. To assess response stability and reduce random variation, each question was posed to each model three times using identical wording in independent new chat sessions; the resulting responses were used for expert ratings and readability analyses.

Phase 3: Real-world patient question bank. Phase 3 questions were collected from patients with H. pylori-related conditions in clinical practice. Forty patients each submitted three questions relevant to their own disease management, yielding 120 clinical questions. Questions were open-ended natural-language queries without predefined categorization. To reflect real-world use, each question was asked once to each of the two LLMs selected for phase 3, and the responses proceeded to clinician ratings and patient evaluations.

Outcomes

Model responses generated in phases 2 and 3 were independently evaluated by three senior gastroenterologists (associate chief physician level or above). To minimize recognition bias and reduce the risk of information leakage, all model outputs were converted to plain text before assessment, and any elements that could reveal model identity or system-specific features (e.g., distinctive formatting, signature phrasing, or numbering styles) were removed. All responses were then randomly shuffled and distributed to raters in a blinded manner.

In phase 2, each model produced three independent responses per question. To further reduce the likelihood that raters could identify a specific model and to mitigate learning effects from repeated answers, each rater evaluated only one response randomly selected from the three repetitions for each model-question pair. In phase 3, each patient question generated a single model response; responses underwent the same plain-text conversion, de-identification, and random shuffling procedures prior to blinded rating.

Clinicians used a 5-point Likert scale to assess the quality of each response across five domains: Correctness, completeness, readability, helpfulness, and safety. For the correctness domain, the anchor definitions were as follows: 1 = completely incorrect; 2 = more incorrect than correct; 3 = equally incorrect and correct; 4 = more correct than incorrect; and 5 = completely correct.

All other domains were rated on a 1-5 scale (higher scores indicating better performance) to capture overall quality in terms of information coverage, clarity of expression, helpfulness for patient understanding and decision-making, and risk mitigation. For each response and each domain, the final score was defined as the median of the three clinicians’ ratings. Phase 2 and phase 3 domain scores were derived using the same procedure and were used for between-model comparisons.

In phase 3, in addition to clinician ratings, patients evaluated satisfaction and perceived readability of the model responses using 5-point scales (higher scores indicating greater satisfaction or better readability). Patient-reported evaluations were used to complement clinician assessments by capturing acceptability and user experience in real-world communication settings. The original model outputs and the corresponding expert ratings are provided in the Supplementary material (phase 2: Supplementary Tables 1-4; phase 3: Supplementary Tables 5 and 6).

To quantify the Chinese readability of LLM-generated responses, we used the Ludong University Text Grading Platform (LDU-TGP) to automatically evaluate responses from phases 2 and 3[10]. Before readability analysis, all responses were converted to plain text and underwent de-identification and removal of model-identifying features. Each response was then entered into the platform as an independent text, and two outputs were recorded as readability metrics: The reading difficulty score and the recommended reading age. A higher reading difficulty score indicates greater comprehension difficulty and poorer readability (i.e., a lower level of understandability).

Statistical analysis

All statistical analyses and figures were generated using IBM SPSS Statistics (version 29.0.1.0) and GraphPad Prism (version 10). Data distributions were summarized descriptively. Baseline characteristics were summarized as mean ± SD for continuous variables and n (%) for categorical variables; between-group comparisons (ChatGPT vs DeepSeek) were performed using an independent-samples t test for age and Fisher’s exact test for categorical variables (sex and education level; Freeman-Halton extension for the 2 × 3 table), as appropriate. Because Likert-scale ratings are ordinal, domain scores and readability metrics were primarily summarized as the median with interquartile range, and between-group comparisons were performed using nonparametric tests.

For phase 1 (MCQs), overall and domain-specific accuracy rates were calculated for each model and compared descriptively. For phase 2 (guideline-oriented questions), differences among the four models in each rating domain and readability metric were assessed using the Friedman test. When the overall test was significant, post hoc pairwise comparisons were conducted with Bonferroni adjustment for multiple testing. For phase 3 (real-world patient questions), differences between the two models selected for phase 3 in each rating domain and readability metric were compared using the Mann-Whitney U test (nonparametric test for independent samples). All tests were two-sided, and statistical significance was defined as P < 0.05.

RESULTS
Phase 1 performance

Across H. pylori-related MCQs, all four LLMs achieved generally high accuracy across knowledge domains, although performance varied by domain (Figure 2). Accuracy was 100% for all models in the clinical manifestations domain (n = 4) and the follow-up and prevention domain (n = 11). Accuracy was also high in the definition domain (n = 50), ranging from 86.0% to 90.0%. In contrast, accuracy was lower in the diagnosis domain (n = 20), ranging from 75.0% to 85.0%; DeepSeek achieved the highest accuracy (85.0%), whereas ChatGPT5 had the lowest (75.0%).

Figure 2
Figure 2 Performance comparison of large language models on multiple-choice questions. H. pylori: Helicobacter pylori; MCQ: Multiple-choice question.

Greater between-model differences were observed in the treatment domain (n = 28). Grok achieved the highest accuracy (96.4%), followed by DeepSeek (92.9%), while ChatGPT5 (85.7%) and ChatGPT4o (82.1%) performed relatively lower. Overall, DeepSeek and Grok outperformed the other models in most domains, with the most pronounced advantage observed in treatment-related questions (Figure 2).

Phase 2 clinician-rated evaluation

In phase 2 (30 guideline-oriented clinical questions), expert ratings differed significantly among the four LLMs across five domains, correctness, completeness, readability, helpfulness, and safety (Figure 3). Overall, ChatGPT5 achieved the highest scores across all domains, with mean scores of 4.63 for correctness, 4.89 for completeness, 4.88 for readability, 4.86 for helpfulness, and 4.59 for safety. DeepSeek ranked second overall (correctness 4.45, completeness 4.43, readability 4.50, helpfulness 4.54, safety 4.40), whereas ChatGPT4o showed intermediate performance across domains (correctness 4.34, completeness 4.16, readability 4.22, helpfulness 4.24, safety 4.25). In contrast, Grok received the lowest scores in multiple domains other than correctness, particularly for completeness (3.75), readability (3.66), and helpfulness (3.72), and also showed a relatively lower safety score (3.86) (Figure 3).

Figure 3
Figure 3 Evaluation results of large language models on clinical questions. aP < 0.05, bP < 0.01, cP < 0.001, and dP < 0.00001.

Post hoc pairwise comparisons with Bonferroni adjustment indicated that ChatGPT5 differed significantly from the other models in most domains (P ≤ 0.05). In addition, Grok differed significantly from the other models in completeness, readability, and helpfulness (Figure 3; Supplementary Tables 1-4).

Automated Chinese readability evaluation using the LDU-TGP platform showed significant between-model differences in both the recommended reading age and the reading difficulty score (Figure 4). For recommended reading age, DeepSeek yielded the highest mean recommended age (15.33 years), followed by Grok (14.91 years), whereas ChatGPT5 (14.02 years) and ChatGPT4o (13.99 years) were lower (Figure 4). Bonferroni-adjusted pairwise comparisons indicated no significant difference between ChatGPT5 and ChatGPT4o (P = 1.000); however, both were significantly lower than DeepSeek and Grok (all P < 0.001). The difference between DeepSeek and Grok was not significant (P = 1.000).

Figure 4
Figure 4  Multidimensional assessment of large language models on patient-generated questions.

For the reading difficulty score (higher scores indicating greater comprehension difficulty and poorer readability), ChatGPT5 had the lowest mean score (13.11), followed by ChatGPT4o (13.92), whereas Grok and DeepSeek scored 16.00 and 17.19, respectively (Figure 4). Bonferroni-adjusted comparisons showed no significant difference between ChatGPT5 and ChatGPT4o (P = 1.000), whereas both differed significantly from DeepSeek and Grok (all P < 0.001). The difference between DeepSeek and Grok was not significant (P = 1.000). Overall, based on LDU-TGP objective metrics, ChatGPT5 and ChatGPT4o were associated with lower recommended reading ages and lower reading difficulty scores, whereas DeepSeek showed the highest objective linguistic difficulty.

Phase 3 patient-centered evaluation

In phase 3, 40 patients with H. pylori-related conditions were enrolled. Baseline characteristics of phase 3 participants are summarized in Table 1. The ChatGPT5 and DeepSeek groups were comparable in age (40.4 ± 9.8 years vs 40.8 ± 10.1 years; P = 0.899), sex distribution (male: 45.0% vs 45.0%; P = 1.000), and education level (P = 0.805). Each patient contributed three real-world clinical questions, yielding a total of 120 questions. The two selected models each answered questions from 20 patients (60 questions per model), followed by clinician ratings and patient evaluations (Figure 5; Supplementary Tables 5 and 6).

Figure 5
Figure 5 Radar chart summarizing the performance of different models across multiple evaluation dimensions. aP < 0.05, bP < 0.01.
Table 1 Baseline characteristics of phase 3 participants, n (%).
Characteristic

ChatGPT5 (n = 20)
DeepSeek (n = 20)
Overall (n = 40)
P value
Age (years), mean ± SD40.4 ± 9.840.8 ± 10.140.6 ± 9.80.899
Sex1.000
Male9 (45.0)9 (45.0)18 (45.0)
Female11 (55.0)11 (55.0)22 (55.0)
Education level0.805
Junior high school or below1 (5.0)2 (10.0)3 (7.5)
Senior high school/vocational high school8 (40.0)9 (45.0)17 (42.5)
College or above11 (55.0)9 (45.0)20 (50.0)

Based on 5-point Likert ratings by clinicians, DeepSeek scored significantly higher than ChatGPT5 in correctness and completeness: Correctness, 4.80 vs 3.85 (Z = 3.834, P < 0.001), and completeness, 4.80 vs 4.05 (Z = 2.635, P = 0.008; Figure 5). No statistically significant differences were observed between the two models for clinician-rated readability (4.30 vs 4.25, P = 0.796), helpfulness (4.60 vs 4.25, P = 0.090), or safety (4.00 vs 3.95, P = 0.564; Figure 5).

Patient-reported evaluations showed that DeepSeek achieved significantly higher perceived readability and satisfaction than ChatGPT5: Perceived readability, 4.65 vs 4.13 (Z = 2.722, P = 0.006), and satisfaction, 4.73 vs 4.05 (P < 0.001; Figure 5). Potentially problematic outputs (any expert rating ≤ 3.5 in correctness, completeness, or safety) were catalogued in Supplementary material.

DISCUSSION

This study compared several currently available LLMs in answering H. pylori-related questions. In phase 1, overall accuracy on MCQs was high across models, yet domain-level performance diverged: All models achieved perfect accuracy in the clinical manifestations and follow-up/prevention domains, whereas accuracy was lower for diagnosis and the greatest between-model variability was observed for treatment. These findings suggest that model performance is more likely to diverge as questions shift from standardized knowledge recall to content requiring greater clinical decision-making. Phase 2, which evaluated guideline-oriented clinical questions, further supported this observation. Expert Likert ratings revealed significant differences among the four models across correctness, completeness, readability, helpfulness, and safety; ChatGPT5 achieved the highest overall ratings, whereas Grok scored relatively lower in completeness, readability, and helpfulness. In phase 3, using real-world patient-generated questions, the relative strengths of the two selected models were not fully concordant: DeepSeek outperformed ChatGPT5 in clinician-rated correctness and completeness, and patients also reported higher perceived readability and satisfaction for DeepSeek, while no significant differences were observed between models in clinician-rated readability, helpfulness, or safety. Notably, objective linguistic readability metrics (recommended reading age and reading difficulty score) did not fully align with clinicians’ subjective readability ratings, indicating that medical health-information evaluation should consider both language complexity and information quality. Overall, our findings suggest that LLMs have potential value for H. pylori-related health information support; however, performance is highly dependent on task type and evaluation domain, underscoring the importance of multidimensional, context-specific assessment and cautious model selection.

Compared with the generally high accuracy observed in phase 1 MCQs, differences among models were more pronounced in the clinical question-answering tasks of phases 2 and 3, largely because the two task types impose different levels of complexity. MCQs provide fixed stems and constrained answer options, requiring the model to select from a limited set of alternatives and thus often yielding higher apparent accuracy. In contrast, clinical question answering requires the model to articulate the problem in natural language, cover key elements comprehensively, provide actionable recommendations, and, when appropriate, flag risks or advise timely medical attention. These demands place greater emphasis on information organization and the completeness of clinical reasoning, thereby amplifying performance gaps.

This divergence was particularly evident for diagnosis- and treatment-related content. Diagnostic questions commonly involve the choice of testing modality, medication-related considerations and timing windows before testing, and the appropriate timing of post-treatment confirmation, whereas treatment questions typically require simultaneous consideration of regimen selection, treatment duration, management of adverse effects, and follow-up strategies[11]. Omission or inadequate explanation of any critical step can differentiate model performance across domains such as completeness, helpfulness, and safety, even when the overall direction of the response is broadly correct. Moreover, real-world patient questions are often phrased colloquially and may lack key contextual information, further increasing response difficulty[12,13]. Taken together, while multiple-choice performance reflects foundational knowledge, clinical question answering more directly captures the practical utility of LLMs in real-world communication settings.

In addition, our findings indicate that the relative strengths of different models are not consistent even within the same disease domain, and that these differences are most evident in whether responses are sufficiently comprehensive, clearly articulated, and readily usable for clinical communication. In the guideline-oriented questions of phase 2, ChatGPT5 achieved the highest overall scores across correctness, completeness, readability, helpfulness, and safety, suggesting a greater tendency toward structured response organization, for example, providing point-by-point explanations, stating necessary assumptions, and offering follow-up guidance or precautions, thereby aligning more closely with experts’ expectations of a “clinically usable” answer. By contrast, Grok received lower ratings for completeness, readability, and helpfulness, implying that its responses may have been more concise or insufficiently covered key steps; thus, even when the central conclusion was partly correct, the absence of actionable guidance or omission of important caveats likely contributed to less favorable clinician appraisal.

Notably, readability metrics revealed a different dimension of between-model variation. LDU-TGP results demonstrated significant differences in recommended reading age and reading difficulty score, indicating that models vary substantially in linguistic features such as word choice, syntactic complexity, and information density. At the same time, objective linguistic readability did not fully align with clinicians’ subjective readability ratings, underscoring that “easy to read” is not synonymous with “clinically clear” or “unlikely to be misunderstood”. In clinical health-information contexts, oversimplification may lead to inadequate content coverage, whereas overly technical language may raise the comprehension barrier; different models’ trade-offs between these two tendencies may partly explain the observed discrepancies between readability metrics and expert ratings[14,15].

In the real-world patient questions of phase 3, DeepSeek outperformed ChatGPT5 in clinician-rated correctness and completeness, and patients also reported higher satisfaction and perceived readability for DeepSeek. These findings suggest that DeepSeek’s responses may have been better aligned with patients’ ways of asking questions and may have more directly addressed the issues of greatest concern to patients, thereby improving user experience. Conversely, no significant differences were observed between the two models in clinician-rated readability, helpfulness, or safety, indicating that in real-world patient scenarios, between-model differences may be more concentrated in whether key medical information is provided accurately and comprehensively rather than across all domains. Collectively, these results imply that models may differ in information coverage, communication style, and risk-mitigation cues; therefore, model selection for patient education should be guided by the specific use case and evaluation domains rather than relying on a single metric.

An additional consideration is that model performance in this study should be interpreted within a language-specific and culture-specific context. LLMs are not linguistically neutral; their outputs are shaped by uneven training distributions across languages, which may lead to differences in accuracy, safety, and communicative adequacy between Chinese and English or among different Chinese-language settings. In patient-facing communication, seemingly small variations in how gastric discomfort, medication intolerance, or follow-up concerns are expressed may alter how well a model captures intent and prioritizes clinically relevant information. Accordingly, adaptation of LLMs for broader clinical use requires more than literal translation; it may also require alignment with regional clinical guidelines, local health-literacy expectations, and the pragmatic features of patient language in different settings[16,17].

Regarding readability, we assessed “readability” through two complementary approaches: (1) Clinicians’ subjective readability ratings on the Likert scale; and (2) Objective linguistic metrics generated by LDU-TGP (recommended reading age and reading difficulty score). The two approaches were not fully concordant, suggesting that “readability” in medical contexts comprises at least two layers. Objective metrics primarily capture intrinsic linguistic difficulty (e.g., lexical/syntactic complexity and information density), whereas clinicians’ notion of readability often also reflects communication quality, such as whether explanations are clear, whether key points are covered in a clinically coherent manner, whether actionable steps are provided, and whether ambiguity is avoided. Patient-reported perceived readability in phase 3 further indicates that whether patients can “understand and find the answer helpful” depends not only on linguistic difficulty but also on whether the response directly addresses their concerns. Therefore, when evaluating LLMs for patient-facing health information support, objective linguistic metrics should be considered alongside clinician- and patient-reported experience, and no single metric should be used as a proxy for overall usability[18].

From a clinical perspective, our findings suggest that LLMs may serve as a supplementary tool for H. pylori patient education and pre-/post-visit counseling, supporting explanation of basic concepts, helping patients understand testing and post-treatment confirmation workflows, improving medication adherence, and providing general advice on lifestyle modification and prevention of reinfection. However, in practice, it is essential to emphasize that LLMs should be positioned as a source of health information support rather than a substitute for clinical diagnosis and treatment, and that clear boundaries should be established for high-risk content[19]. For questions involving eradication regimen selection, contraindications and drug-drug interactions, special populations, prior treatment failure, and post-treatment confirmation strategies, LLM outputs should be treated as informational references and ultimately verified by clinicians to avoid inappropriate recommendations arising from missing patient-specific details[20].

To enhance safety and usability, standardized implementation strategies may be adopted in clinical settings: First, explicitly communicate to patients that model-generated content cannot replace medical advice and encourage them to bring the responses to clinic for verification; second, preferentially instruct the model to respond in a structured format to reduce omissions; third, incorporate fixed safety prompts when addressing treatment- or risk-related topics; and finally, tailor responses through layered presentation according to patients’ comprehension levels to balance informational completeness with readability[21]. Overall, with clear boundaries, structured outputs, and human oversight, LLMs may improve the efficiency and experience of H. pylori-related health information access, while caution remains warranted for individualized clinical decision-making.

In addition, this study has several limitations. First, LLMs evolve rapidly; we evaluated publicly available versions during the study period, and performance may change with model updates, training data, or inference strategies, rendering our findings time sensitive. Accordingly, we provide the access dates and interface-displayed model versions in Supplementary Table 7. Second, although the phase 2 guideline-oriented questions were derived from authoritative guidelines/consensus statements and covered key clinical decision points, the composition and difficulty distribution of the questions may still have been influenced by investigator selection. Likewise, the phase 1 multiple-choice items were drawn from a single-center internal test bank and may be subject to biases in item type and difficulty. Third, the phase 3 patient sample was obtained from a single clinical source and was relatively small; while real-world patient questions were collected, we did not stratify analyses by question type, which may have limited more granular interpretation of model differences across specific topics. Because all patient questions were collected from a single Chinese clinical setting, we were also unable to assess whether regional dialects, local expressions, or other Chinese-language communication patterns would alter model performance in real-world use. Although overall ratings were favorable, we observed occasional potentially unsafe, incomplete, or inaccurate outputs. These mainly involved non-standard regimen descriptions and overconfident, non-verifiable quantitative statements, which may affect patient-facing decision making (Supplementary material). Representative examples are provided in Supplementary material (e.g., S3 | P2-Q1-Run1; S3 | P2-Q27-Run2; S5 | P3-P17-Q3). For example, one response described bismuth quadruple therapy using an atypical antibiotic combination (S3 | P2-Q1-Run1), another included specific resistance-rate statements without clear sourcing (S3 | P2-Q27-Run2), and a patient-facing testing-preparation answer was judged potentially unsafe by experts (S5 | P3-P17-Q3) (Supplementary material).

The observed error patterns likely reflect not only incomplete domain knowledge, but also a structural limitation of current autoregressive LLMs, which optimize next-token prediction rather than grounded clinical state reasoning. As a result, responses may appear fluent and persuasive while still omitting key decision points, overstating certainty, or generating clinically unsafe details. From this perspective, future progress may depend not only on larger models or prompt refinement, but also on advances in reasoning-oriented medical artificial intelligence. From this perspective, future progress may depend not only on larger models or prompt refinement, but also on advances in reasoning-oriented medical artificial intelligence. Concept-centered reasoning frameworks, including emerging large concept models and related pathway-aware architectures, may help preserve conceptual integrity by anchoring outputs to verified medical ontologies, clinical pathways, and higher-level reasoning structures rather than relying primarily on surface statistical associations. Although these frameworks remain at an early stage, they may offer a potentially more robust route to reducing hallucinations and improving diagnostic reliability, especially when combined with human-in-the-loop review and explicit safety constraints[13,22].

Future studies should validate these findings in larger samples, across multiple centers, and in more diverse populations, while adopting more refined evaluation frameworks. For example, model performance could be compared by question topic and clinical scenario to better identify weaknesses at “high-risk decision points” (e.g., regimen selection, management of adverse effects, and care of special populations)[23]. In addition, incorporating more rigorous inter-rater reliability assessments and more explicit safety endpoints would strengthen the robustness of conclusions. Finally, patient-facing output optimization strategies, such as layered presentation, structured templates, and standardized risk-warning modules, should be explored and evaluated for their impact on real-world outcomes including patient understanding, adherence, and healthcare-seeking behavior, thereby providing stronger evidence to guide the responsible use of LLMs in H. pylori patient education.

CONCLUSION

Currently available LLMs showed generally good performance in H. pylori-related health-information support, but their strengths were clearly task dependent. Importantly, the model that performed best in structured knowledge and guideline-oriented assessments was not necessarily the most effective in real-world patient interaction, where correctness, completeness, and patient-perceived readability had greater practical relevance. Therefore, performance on multiple-choice or exam-style tasks alone should not be used as a surrogate for clinical utility. Multidimensional evaluation, clearly defined use boundaries, and clinician oversight remain essential before LLMs are deployed for patient-facing H. pylori education.

References
1.  Chen YC, Malfertheiner P, Yu HT, Kuo CL, Chang YY, Meng FT, Wu YX, Hsiao JL, Chen MJ, Lin KP, Wu CY, Lin JT, O'Morain C, Megraud F, Lee WC, El-Omar EM, Wu MS, Liou JM. Global Prevalence of Helicobacter pylori Infection and Incidence of Gastric Cancer Between 1980 and 2022. Gastroenterology. 2024;166:605-619.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 385]  [Cited by in RCA: 376]  [Article Influence: 188.0]  [Reference Citation Analysis (1)]
2.  Pan KF, Li WQ, Zhang L, Liu WD, Ma JL, Zhang Y, Ulm K, Wang JX, Zhang L, Bajbouj M, Zhang LF, Li M, Vieth M, Quante M, Wang LH, Suchanek S, Mejías-Luque R, Xu HM, Fan XH, Han X, Liu ZC, Zhou T, Guan WX, Schmid RM, Gerhard M, Classen M, You WC. Gastric cancer prevention by community eradication of Helicobacter pylori: a cluster-randomized controlled trial. Nat Med. 2024;30:3250-3260.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 63]  [Cited by in RCA: 80]  [Article Influence: 40.0]  [Reference Citation Analysis (1)]
3.  Malfertheiner P, Megraud F, Rokkas T, Gisbert JP, Liou JM, Schulz C, Gasbarrini A, Hunt RH, Leja M, O'Morain C, Rugge M, Suerbaum S, Tilg H, Sugano K, El-Omar EM; European Helicobacter and Microbiota Study group. Management of Helicobacter pylori infection: the Maastricht VI/Florence consensus report. Gut. 2022;gutjnl-2022.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 1155]  [Cited by in RCA: 1006]  [Article Influence: 251.5]  [Reference Citation Analysis (13)]
4.  Du RC, Lv NH, Hu Y. [Interpretation of the Maastricht VI/Florence consensus report on management of Helicobacter pylori infection and its updates]. Zhonghua Xiaohua Zazhi. 2023;43 352-356 [DOI:10.3760/cma.j.cn311367-20221014.  [PubMed]  [DOI]
5.  Li Y, Choi H, Leung K, Jiang F, Graham DY, Leung WK. Global prevalence of Helicobacter pylori infection between 1980 and 2022: a systematic review and meta-analysis. Lancet Gastroenterol Hepatol. 2023;8:553-564.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 392]  [Cited by in RCA: 378]  [Article Influence: 126.0]  [Reference Citation Analysis (15)]
6.  Shah S, Cappell K, Sedgley R, Pelletier C, Jacob R, Bonafede M, Yadlapati R. Healthcare costs among patients with newly diagnosed helicobacter pylori infection in the United States: a linked claims-EHR study. J Med Econ. 2023;26:1227-1236.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 2]  [Reference Citation Analysis (0)]
7.  Morgan E, Clifford G, Park JY.   The global epidemiology of gastric cancer and Helicobacter pylori: current and future perspectives for prevention. In: Park JY, editor. Population-Based Helicobacter pylori Screen-and-Treat Strategies for Gastric Cancer Prevention: Guidance on Implementation. Lyon (FR): International Agency for Research on Cancer, 2025.  [PubMed]  [DOI]
8.  Tian S, Jin Q, Yeganova L, Lai PT, Zhu Q, Chen X, Yang Y, Chen Q, Kim W, Comeau DC, Islamaj R, Kapoor A, Gao X, Lu Z. Opportunities and challenges for ChatGPT and large language models in biomedicine and health. Brief Bioinform. 2023;25:bbad493.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 278]  [Cited by in RCA: 179]  [Article Influence: 59.7]  [Reference Citation Analysis (0)]
9.  Wang D, Ye J, Li J, Liang J, Zhang Q, Hu Q, Pan C, Wang D, Liu Z, Shi W, Guo M, Li F, Du W, Zheng YF. Enhancing Large Language Models for Improved Accuracy and Safety in Medical Question Answering: Comparative Study. JMIR Med Educ. 2025;11:e70190.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 4]  [Reference Citation Analysis (0)]
10.  Cheng Y, Xu DK, Dong J. [A study on the analysis of key factors of text reading difficulty grading and the readability formula based on a corpus of language teaching materials]. Yuyan Wenzi Yingyong. 2020;132-143.  [PubMed]  [DOI]
11.  Shah SC, Iyer PG, Moss SF. AGA Clinical Practice Update on the Management of Refractory Helicobacter pylori Infection: Expert Review. Gastroenterology. 2021;160:1831-1841.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 197]  [Cited by in RCA: 183]  [Article Influence: 36.6]  [Reference Citation Analysis (4)]
12.  Abanmy NO, Al-Ghreimil N, Alsabhan JF, Al-Baity H, Aljadeed R. Evaluating the accuracy of ChatGPT in delivering patient instructions for medications: an exploratory case study. Front Artif Intell. 2025;8:1550591.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 7]  [Reference Citation Analysis (0)]
13.  Wu J, Wu X, Zheng Y, Yang J. Clinical pathway-aware large language models for reliable and transparent medical dialogue. J Biomed Inform. 2025;172:104942.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 3]  [Reference Citation Analysis (0)]
14.  Picton B, Andalib S, Spina A, Camp B, Solomon SS, Liang J, Chen PM, Chen JW, Hsu FP, Oh MY. Assessing AI Simplification of Medical Texts: Readability and Content Fidelity. Int J Med Inform. 2025;195:105743.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 6]  [Cited by in RCA: 18]  [Article Influence: 18.0]  [Reference Citation Analysis (0)]
15.  Okuhara T, Furukawa E, Okada H, Yokota R, Kiuchi T. Readability of written information for patients across 30 years: A systematic review of systematic reviews. Patient Educ Couns. 2025;135:108656.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 5]  [Cited by in RCA: 42]  [Article Influence: 42.0]  [Reference Citation Analysis (0)]
16.  Li P, Xu Y, Liu X, Shen Z, Wang Y, Lv X, Lu Z, Wu H, Zhuang J, Chen Y. Large Language Models in Patient Health Communication for Atherosclerotic Cardiovascular Disease: Pilot Cross-Sectional Comparative Analysis. JMIR Med Inform. 2026;14:e81422.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 2]  [Cited by in RCA: 3]  [Article Influence: 3.0]  [Reference Citation Analysis (0)]
17.  Yao Z, Duan L, Xu S, Chi L, Sheng D. Performance of Large Language Models in the Non-English Context: Qualitative Study of Models Trained on Different Languages in Chinese Medical Examinations. JMIR Med Inform. 2025;13:e69485.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 15]  [Reference Citation Analysis (0)]
18.  Bedi S, Liu Y, Orr-Ewing L, Dash D, Koyejo S, Callahan A, Fries JA, Wornow M, Swaminathan A, Lehmann LS, Hong HJ, Kashyap M, Chaurasia AR, Shah NR, Singh K, Tazbaz T, Milstein A, Pfeffer MA, Shah NH. Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review. JAMA. 2025;333:319-328.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 53]  [Cited by in RCA: 354]  [Article Influence: 354.0]  [Reference Citation Analysis (0)]
19.  Hager P, Jungmann F, Holland R, Bhagat K, Hubrecht I, Knauer M, Vielhauer J, Makowski M, Braren R, Kaissis G, Rueckert D. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat Med. 2024;30:2613-2622.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 462]  [Cited by in RCA: 348]  [Article Influence: 174.0]  [Reference Citation Analysis (0)]
20.  Howell MD. Generative artificial intelligence, patient safety and healthcare quality: a review. BMJ Qual Saf. 2024;33:748-754.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 15]  [Cited by in RCA: 33]  [Article Influence: 16.5]  [Reference Citation Analysis (0)]
21.  Hua Y, Xia W, Bates D, Hartstein GL, Kim HT, Li M, Nelson BW, Stromeyer Iv C, King D, Suh J, Zhou L, Torous J. Standardizing and Scaffolding Health Care AI-Chatbot Evaluation: Systematic Review. JMIR AI. 2025;4:e69006.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 11]  [Reference Citation Analysis (0)]
22.  Merchant SA, Merchant N, Varghese SL, Shaikh MJS. Large language models and large concept models in radiology: Present challenges, future directions, and critical perspectives. World J Radiol. 2025;17:114754.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 3]  [Reference Citation Analysis (10)]
23.  CHART Collaborative. Reporting guidelines for chatbot health advice studies: explanation and elaboration for the Chatbot Assessment Reporting Tool (CHART). BMJ. 2025;390:e083305.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 6]  [Cited by in RCA: 15]  [Article Influence: 15.0]  [Reference Citation Analysis (0)]
Footnotes

Peer review: Externally peer reviewed.

Peer-review model: Single blind

Specialty type: Gastroenterology and hepatology

Country of origin: China

Peer-review report’s classification

Scientific quality: Grade A, Grade B, Grade B

Novelty: Grade B, Grade B, Grade B

Creativity or innovation: Grade A, Grade B, Grade B

Scientific significance: Grade A, Grade B, Grade C

P-Reviewer: Merchant SA, Chairman, Dean, Emeritus Professor, India; Romanchuk OP, DM, Full Professor, MD, PhD, Ukraine S-Editor: Wu S L-Editor: A P-Editor: Wang CH

Write to the Help Desk