Copyright: ©Author(s) 2026.
World J Gastroenterol. Jun 28, 2026; 32(24): 118690
Published online Jun 28, 2026. doi: 10.3748/wjg.118690
Published online Jun 28, 2026. doi: 10.3748/wjg.118690
Figure 5 Confusion matrices of multimodal large language models in phase A.
A-J: Confusion matrices illustrating the Mayo Endoscopic Subscore classification outputs of five multimodal large language models: GPT-5 (A and B), GPT-4o (C and D), Gemini-2.5-Pro (E and F), Grok-4 (G and H), and Qwen-VL-Max (I and J). The horizontal axis denotes the true Mayo Endoscopic Subscore grades, and the vertical axis represents the predicted grades generated by each model. Color intensity corresponds to the number or proportion of images in each classification cell with deeper red tones indicating higher counts or frequencies.
- Citation: Zhao XY, Shen Y, He JJ, Zhou QY, Jiang LS, Zhou ZR, An FM, Zhan Q, Sun J, Feng W. Comparative evaluation of multimodal large language models for Mayo Endoscopic Subscore grading in ulcerative colitis. World J Gastroenterol 2026; 32(24): 118690
- URL: https://www.wjgnet.com/1007-9327/full/v32/i24/118690.htm
- DOI: https://dx.doi.org/10.3748/wjg.118690