BPG is committed to discovery and dissemination of knowledge
Retrospective Study
Copyright: ©Author(s) 2026.
World J Gastroenterol. Jun 28, 2026; 32(24): 118690
Published online Jun 28, 2026. doi: 10.3748/wjg.118690
Table 2 Overall performance comparison of five multimodal large language models in phase A

Gemini-2.5-Pro
Grok-4
GPT-4o
GPT-5
Qwen-VL-Max
Accuracy0.5020.4170.5940.7170.353
Precision0.6390.5520.6350.7310.488
Recall0.5020.4170.5940.7170.353
F1 score (95%CI)0.480 (0.452-0.574)0.415 (0.369-0.481)0.602 (0.522-0.648)0.720 (0.665-0.773)0.338 (0.263-0.382)
Cohen’s κ0.3430.2390.4490.6080.133
MAE0.6110.7700.4520.2970.767
MSE0.8661.1940.5440.3251.021
RMSE0.9301.0930.7380.5701.011
SD0.7810.9180.6970.5640.953
CV0.7130.8380.6370.5150.870
r value10.681 0.590 0.777 0.852 0.478


Write to the Help Desk