BPG is committed to discovery and dissemination of knowledge
Retrospective Study
Copyright: ©Author(s) 2026.
World J Gastroenterol. Jun 28, 2026; 32(24): 118690
Published online Jun 28, 2026. doi: 10.3748/wjg.118690
Figure 3
Figure 3 Segment-wise performance evaluation of multimodal large language models in phase B. A: Overall ranking of five multimodal large language models (Gemini-2.5-Pro, Grok-4, GPT-4o, GPT-5, and Qwen-VL-Max) based on weighted composite scores; B: Summary of key performance metrics (accuracy, F1 score, κ, precision, and recall); C and D: Heatmaps illustrating segment-specific accuracy and mean absolute error across six colonic regions; E: Bar chart showing average accuracy by segment with pairwise comparisons assessed using the McNemar test (P < 0.05); F: Variability analysis of accuracy across segments (lower values indicate greater stability; G and H: Radar plots depicting segment-level and overall performance profiles across the five models. MAE: Mean absolute error; ACC: Accuracy; Corr: Correlation.


Write to the Help Desk