Copyright: ©Author(s) 2026.
World J Gastroenterol. Jun 28, 2026; 32(24): 118690
Published online Jun 28, 2026. doi: 10.3748/wjg.118690
Published online Jun 28, 2026. doi: 10.3748/wjg.118690
Figure 2 Comprehensive performance analysis of five multimodal large language models.
A: Overall ranking of five multimodal large language models based on composite performance scores; B: Comparison of core evaluation metrics (accuracy, F1 score, κ, precision, recall). Statistical tests: McNemar for accuracy, precision, recall, and F1 score; paired Wilcoxon for κ (P < 0.05 considered significant); C: Radar chart illustrating performance across five key dimensions; D: Heatmap summarizing six quantitative indicators (accuracy, precision, recall, F1 score, κ, and mean absolute error); E: Precision-recall scatter plot demonstrating the trade-off between precision and recall; F: Comparison of error metrics (mean absolute error and root mean squared error) where lower values indicate superior performance. MAE: Mean absolute error; RMSE: Root mean squared error.
- Citation: Zhao XY, Shen Y, He JJ, Zhou QY, Jiang LS, Zhou ZR, An FM, Zhan Q, Sun J, Feng W. Comparative evaluation of multimodal large language models for Mayo Endoscopic Subscore grading in ulcerative colitis. World J Gastroenterol 2026; 32(24): 118690
- URL: https://www.wjgnet.com/1007-9327/full/v32/i24/118690.htm
- DOI: https://dx.doi.org/10.3748/wjg.118690