Copyright: ©Author(s) 2026.
World J Methodol. Sep 20, 2026; 16(3): 116022
Published online Sep 20, 2026. doi: 10.5662/wjm.v16.i3.116022
Published online Sep 20, 2026. doi: 10.5662/wjm.v16.i3.116022
Table 6 Patient assessment of large language model responses
| Characteristics | ChatGPT-5a | Gemini-2.5a | Claude-4a | P value | ||
| A vs B | A vs C | B vs C | ||||
| Overall | ||||||
| Comprehensiveness | 3.90 ± 0.29 | 3.94 ± 0.23 | 3.95 ± 0.21 | 0.5018 | 0.3859 | 0.8416 |
| Actionability | 3.41 ± 0.62 | 3.81 ± 0.40 | 3.86 ± 0.34 | 0.0001 | 0.0002 | 0.5538 |
| Empathy | 3.74 ± 0.46 | 3.89 ± 0.31 | 3.86 ± 0.35 | 0.0954 | 0.1941 | 0.6850 |
| Domain wise | ||||||
| Domain 1: General understanding | ||||||
| Comprehensiveness | 3.84 ± 0.36 | 3.95 ± 0.22 | 3.97 ± 0.18 | 0.1076 | 0.0472 | 0.6616 |
| Actionability | 3.83 ± 0.39 | 3.92 ± 0.27 | 3.75 ± 0.44 | 0.2397 | 0.3982 | 0.0432 |
| Empathy | 3.56 ± 0.55 | 3.87 ± 0.35 | 3.86 ± 0.35 | 0.0040 | 0.0053 | 0.8999 |
| Domain 2: Symptoms and diagnosis | ||||||
| Comprehensiveness | 3.93 ± 0.26 | 3.93 ± 0.25 | 3.94 ± 0.23 | 1.0000 | 0.8577 | 0.8546 |
| Actionability | 3.56 ± 0.61 | 3.88 ± 0.32 | 3.85 ± 0.36 | 0.0049 | 0.0126 | 0.6984 |
| Empathy | 3.75 ± 0.45 | 3.88 ± 0.32 | 3.86 ± 0.34 | 0.1456 | 0.2270 | 0.7898 |
| Domain 3: Causes and risk factors | ||||||
| Comprehensiveness | 3.89 ± 0.31 | 3.91 ± 0.29 | 3.95 ± 0.22 | 0.7694 | 0.3274 | 0.4946 |
| Actionability | 3.51 ± 0.61 | 3.81 ± 0.39 | 3.85 ± 0.21 | 0.0116 | 0.0015 | 0.5744 |
| Empathy | 3.77 ± 0.43 | 3.91 ± 0.29 | 3.90 ± 0.31 | 0.0960 | 0.1298 | 0.8834 |
| Domain 4: Dietary considerations | ||||||
| Comprehensiveness | 3.90 ± 0.31 | 3.97 ± 0.19 | 3.93 ± 0.27 | 0.2330 | 0.6499 | 0.4516 |
| Actionability | 3.13 ± 0.65 | 3.97 ± 0.16 | 3.89 ± 0.32 | 0.0001 | 0.0001 | 0.1667 |
| Empathy | 3.77 ± 0.46 | 3.92 ± 0.29 | 3.79 ± 0.44 | 0.0890 | 0.8450 | 0.1276 |
| Domain 5: Treatment and management | ||||||
| Comprehensiveness | 3.92 ± 0.26 | 3.95 ± 0.24 | 3.96 ± 0.18 | 0.5980 | 0.4320 | 0.8357 |
| Actionability | 3.25 ± 0.62 | 3.62 ± 0.49 | 3.86 ± 0.34 | 0.0046 | 0.0001 | 0.0141 |
| Empathy | 3.79 ± 0.43 | 3.88 ± 0.32 | 3.82 ± 0.41 | 0.2977 | 0.7534 | 0.4735 |
| Domain 6: Lifestyle factors | ||||||
| Comprehensiveness | 3.93 ± 0.26 | 3.96 ± 0.19 | 3.99 ± 0.17 | 0.5624 | 0.2315 | 0.4647 |
| Actionability | 3.22 ± 0.61 | 3.79 ± 0.41 | 3.84 ± 0.37 | 0.0001 | 0.0001 | 0.1788 |
| Empathy | 3.78 ± 0.43 | 3.91 ± 0.29 | 3.89 ± 0.31 | 0.1217 | 0.1989 | 0.7694 |
- Citation: Goyal K, Goyal MK, Taranikanti V, Wander P, Chowdhary R, Kalra S, Prashar G, Vuthaluru AR, Goyal O. Do bots provide correct and adequate guidance regarding acidity: A blinded comparison rated by patients and physicians. World J Methodol 2026; 16(3): 116022
- URL: https://www.wjgnet.com/2222-0682/full/v16/i3/116022.htm
- DOI: https://dx.doi.org/10.5662/wjm.v16.i3.116022