Copyright: ©Author(s) 2026.
World J Gastroenterol. Jun 28, 2026; 32(24): 118690
Published online Jun 28, 2026. doi: 10.3748/wjg.118690
Published online Jun 28, 2026. doi: 10.3748/wjg.118690
Table 5 Diagnostic accuracy by Mayo Endoscopic Subscore grade among physicians and multimodal large language models
| MES grade | 0 | 1 | 2 | 3 |
| Expert 1 | 100 | 53.4 | 81.8 | 82.9 |
| Expert 2 | 98.2 | 67.1 | 66.7 | 62.9 |
| GPT-5 | 86.5 | 56.3 | 62.4 | 87.1 |
| GPT-4o | 85.3 | 45.2 | 53.8 | 52.2 |
| Gemini-2.5-Pro | 93.1 | 38.9 | 45.4 | 60.0 |
| Grok-4 | 88.9 | 30.1 | 35.1 | 40.3 |
| Qwen-VL-Max | 76.7 | 26.3 | 33.3 | 38.5 |
- Citation: Zhao XY, Shen Y, He JJ, Zhou QY, Jiang LS, Zhou ZR, An FM, Zhan Q, Sun J, Feng W. Comparative evaluation of multimodal large language models for Mayo Endoscopic Subscore grading in ulcerative colitis. World J Gastroenterol 2026; 32(24): 118690
- URL: https://www.wjgnet.com/1007-9327/full/v32/i24/118690.htm
- DOI: https://dx.doi.org/10.3748/wjg.118690