Copyright: ©Author(s) 2026.
World J Gastroenterol. Jun 28, 2026; 32(24): 118690
Published online Jun 28, 2026. doi: 10.3748/wjg.118690
Published online Jun 28, 2026. doi: 10.3748/wjg.118690
Figure 3 Segment-wise performance evaluation of multimodal large language models in phase B.
A: Overall ranking of five multimodal large language models (Gemini-2.5-Pro, Grok-4, GPT-4o, GPT-5, and Qwen-VL-Max) based on weighted composite scores; B: Summary of key performance metrics (accuracy, F1 score, κ, precision, and recall); C and D: Heatmaps illustrating segment-specific accuracy and mean absolute error across six colonic regions; E: Bar chart showing average accuracy by segment with pairwise comparisons assessed using the McNemar test (P < 0.05); F: Variability analysis of accuracy across segments (lower values indicate greater stability; G and H: Radar plots depicting segment-level and overall performance profiles across the five models. MAE: Mean absolute error; ACC: Accuracy; Corr: Correlation.
- Citation: Zhao XY, Shen Y, He JJ, Zhou QY, Jiang LS, Zhou ZR, An FM, Zhan Q, Sun J, Feng W. Comparative evaluation of multimodal large language models for Mayo Endoscopic Subscore grading in ulcerative colitis. World J Gastroenterol 2026; 32(24): 118690
- URL: https://www.wjgnet.com/1007-9327/full/v32/i24/118690.htm
- DOI: https://dx.doi.org/10.3748/wjg.118690