BPG is committed to discovery and dissemination of knowledge
Observational Study
Copyright: ©Author(s) 2026.
World J Methodol. Sep 20, 2026; 16(3): 116022
Published online Sep 20, 2026. doi: 10.5662/wjm.v16.i3.116022
Figure 2
Figure 2 Patient-rated comprehensiveness, actionability, and empathy of responses generated by large language models. A: Patient assessment of comprehensiveness of responses across large language models. The stacked bar chart represents the distribution of patient-perceived comprehensiveness (minimal, moderately, and highly comprehensive) in responses generated by ChatGPT-5, Gemini-2.5, and Claude-4. Percentages correspond to values shown in the accompanying table, with ChatGPT-5 responses being mostly rated as “moderately comprehensive”, while Gemini-2.5 and Claude-4 demonstrated a higher proportion of “highly comprehensive” responses (P < 0.05); B: Patient assessment of actionability of responses across large language models. The stacked bar chart shows the distribution of patient-rated actionability of responses generated by ChatGPT-5, Gemini-2.5, and Claude-4. Ratings ranged from “slight actionable” to “highly actionable”. ChatGPT-5 responses were most frequently rated as “moderately actionable”, whereas Gemini-2.5 and Claude-4 demonstrated higher proportions of “actionable” and “highly actionable” responses. Differences across models were statistically significant (P < 0.05); C: Patient assessment of empathy of responses generated by ChatGPT-5, Gemini-2.5, and Claude-4. Responses were categorized as minimally empathetic, moderately empathetic, or highly empathetic. ChatGPT-5 responses were more often rated as minimally empathetic (38.5%), whereas Gemini-2.5 (38.5%) and Claude-4 (35.9%) produced a greater proportion of highly empathetic responses. Differences across models were statistically significant (P < 0.05).


Write to the Help Desk