BPG is committed to discovery and dissemination of knowledge
Observational Study
Copyright: ©Author(s) 2026.
World J Methodol. Sep 20, 2026; 16(3): 116022
Published online Sep 20, 2026. doi: 10.5662/wjm.v16.i3.116022
Table 5 Physician categorical assessment of responses generated by ChatGPT-5, Gemini-2.5, and Claude-4 across four key domains: Comprehensiveness, accuracy, actionability, and empathy, n (%)
Characteristics
ChatGPT-5a
Gemini-2.5a
Claude-4a
P value
Comprehensiveness
Minimal2 (5.12)1 (2.6)1 (2.6)0.029
Moderately32 (82.1)27(69.2)20 (51.3)
Highly5 (12.8)11(28.2)18 (46.2)
Accuracy
Partially5 (12.8)1 (2.6)1 (2.6)0.013
Mostly32 (82.1)27(69.2)26 (66.7)
Entirely accurate and evidence-based2 (5.1)11 (28.2)12 (30.8)
Actionability
Slightly actionable10 (25.5)1 (2.6)1 (2.6)< 0.00001
Moderately actionable26 (66.7)14 (35.8)20 (51.2)
Actionable 2 (5.2)17 (43.7)17 (43.6)
Highly actionable1 (2.6)7 (17.9)1 (2.6)
Empathy
Minimally empathetic17 (43.65)1 (2.6)1 (2.6)0.00001
Moderately empathetic20 (51.3)26 (66.7)25 (64.1)
Highly empathetic2 (5.1)12 (30.8)13 (33.3)


Write to the Help Desk