Copyright: ©Author(s) 2026.
World J Gastroenterol. Jun 28, 2026; 32(24): 118690
Published online Jun 28, 2026. doi: 10.3748/wjg.118690
Published online Jun 28, 2026. doi: 10.3748/wjg.118690
Table 2 Overall performance comparison of five multimodal large language models in phase A
| Gemini-2.5-Pro | Grok-4 | GPT-4o | GPT-5 | Qwen-VL-Max | |
| Accuracy | 0.502 | 0.417 | 0.594 | 0.717 | 0.353 |
| Precision | 0.639 | 0.552 | 0.635 | 0.731 | 0.488 |
| Recall | 0.502 | 0.417 | 0.594 | 0.717 | 0.353 |
| F1 score (95%CI) | 0.480 (0.452-0.574) | 0.415 (0.369-0.481) | 0.602 (0.522-0.648) | 0.720 (0.665-0.773) | 0.338 (0.263-0.382) |
| Cohen’s κ | 0.343 | 0.239 | 0.449 | 0.608 | 0.133 |
| MAE | 0.611 | 0.770 | 0.452 | 0.297 | 0.767 |
| MSE | 0.866 | 1.194 | 0.544 | 0.325 | 1.021 |
| RMSE | 0.930 | 1.093 | 0.738 | 0.570 | 1.011 |
| SD | 0.781 | 0.918 | 0.697 | 0.564 | 0.953 |
| CV | 0.713 | 0.838 | 0.637 | 0.515 | 0.870 |
| r value1 | 0.681 | 0.590 | 0.777 | 0.852 | 0.478 |
- Citation: Zhao XY, Shen Y, He JJ, Zhou QY, Jiang LS, Zhou ZR, An FM, Zhan Q, Sun J, Feng W. Comparative evaluation of multimodal large language models for Mayo Endoscopic Subscore grading in ulcerative colitis. World J Gastroenterol 2026; 32(24): 118690
- URL: https://www.wjgnet.com/1007-9327/full/v32/i24/118690.htm
- DOI: https://dx.doi.org/10.3748/wjg.118690