BPG is committed to discovery and dissemination of knowledge
Retrospective Study
Copyright: ©Author(s) 2026.
World J Gastroenterol. Jun 28, 2026; 32(24): 118690
Published online Jun 28, 2026. doi: 10.3748/wjg.118690
Table 3 Overall performance comparison of five multimodal large language models in phase B

Gemini-2.5-Pro
Grok-4
GPT-4o
GPT-5
Qwen-VL-Max
Accuracy0.5410.4200.5830.7100.406
Precision0.6210.4800.5900.7240.515
Recall0.5410.4200.5830.7100.406
F1 score (95%CI)0.546 (0.490-0.611)0.426 (0.363-0.481)0.584 (0.502-0.618)0.715 (0.664-0.768)0.408 (0.313-0.439)
Cohen’s κ0.3810.2290.4250.5960.190
MAE0.5650.7530.4730.3000.707
MSE0.7991.1410.5870.3220.940
RMSE0.8941.0680.7660.5670.969
SD0.8430.9730.7550.5660.951
CV0.7690.8880.6890.5170.868
r value10.6440.5730.7490.8500.499


Write to the Help Desk