©The Author(s) 2026.
World J Gastroenterol. Feb 28, 2026; 32(8): 115297
Published online Feb 28, 2026. doi: 10.3748/wjg.v32.i8.115297
Published online Feb 28, 2026. doi: 10.3748/wjg.v32.i8.115297
Table 1 Comparative model summary
| Model name | Architecture family | Pretraining | Mean IoU, % | Mean Dice | PET score, % | GRR, % | Inference latency (milliseconds/frame) | Model size | Notes |
| U-Net | CNN | N/A | 74.88 | (Self 79.37 + EDD2020 67.63)/2 = 73.50% | 45.74 | 65.41 | 3.64 | 31.46 M, approximately 126 MB | Classic encoder-decoder with skip connections; excellent throughput with very high FPS but lower PET; simple architecture may limit generalization |
| ResNet + U-Net | CNN | ResNet encoder typically ImageNet-1K | 80.63 | (83.59 + 73.97)/2 = 78.78% | 82.58 | 67.03 | 5.18 | 32.52 M, approximately 130 MB | Residual backbone improves accuracy and robustness with modest compute; good PET |
| ConvNeXt + UPerNet | CNN | Typically ImageNet-1K | 82.69 | (85.65 + 76.90)/2 = 81.27% | 84.70 | 68.79 | 5.65 | 41.37 M, approximately 165 MB | Modern CNN with ViT-inspired design; strong accuracy and good throughput |
| M2SNet | CNN | N/A | 80.87 | (84.24 + 74.81)/2 = 79.53% | 72.33 | 67.64 | 14.86 | 29.89 M, approximately 120 MB | Multi-scale subtraction units for improved feature complementarity and edge clarity; slower inference among CNNs |
| Dilated SegNet | CNN | ResNet50 backbone | 80.55 | (83.35 + 73.64)/2 = 78.49% | 74.88 | 68.08 | 9.23 | 18.111 M, approximately 72 MB | Dilated convolutions for real-time polyp segmentation; good trade-off, moderate speed |
| PraNet | CNN | N/A | 73.74 | (74.15 + 61.12)/2 = 67.63% | 48.81 | 65.79 | 12.16 | 32.56 M, approximately 130 MB | Parallel partial decoder + reverse attention for boundary refinement; decent accuracy but lower PET and generalization |
| SwinV2 + UPerNet | Transformer | SwinV2 typically ImageNet-pretrained | 82.74 | (85.59 + 76.97)/2 = 81.28% | 78.41 | 68.12 | 12.19 | 41.91 M, approximately 168 MB | SwinV2 hierarchical transformer backbone with UPerNet decoder; strong accuracy with moderate compute |
| SegFormer | Transformer | Typically ImageNet-1K | 82.86 | (93.14 + 77.20)/2 = 85.17% | 92.02 | 70.11 | 9.68 | 24.73 M, approximately 99 MB | Excellent performance-efficiency balance; low FLOPs (4.23 GFLOPs), good generalization: Recommended for real-time clinical use per paper |
| SETR-MLA | Transformer | ViT backbone | 77.42 | (82.14 + 71.48)/2 = 76.81% | 52.45 | 69.67 | 5.55 | 90.77 M, approximately 363 MB | Segmentation transformer with multi-level aggregation; large parameter count but relatively fast inference in this setup |
| TransUNet | Hybrid (CNN + transformer) | ResNet50 + ViT | 74.81 | (77.35 + 65.06)/2 = 71.21% | 26.39 | 67.14 | 13.03 | 105.00 M, approximately 420 MB | Combines CNN encoder and ViT: Strong representational power but heavy compute and lower PET |
| PVTV2 + EMCAD | Transformer | PVTV2 usually ImageNet-pretrained | 82.91 | (85.81 + 77.07)/2 = 81.44% | 88.14 | 71.38 | 12.56 | 26.77 M, approximately 107 MB | Pyramid vision transformer v2 + efficient multi-scale decoding; strong generalization and good PET |
| FCBFormer | Transformer | PVTV2 backbone | 82.00 | (85.18 + 76.03)/2 = 80.61% | 61.89 | 71.52 | 21.43 | 33.09 M, approximately 132 MB | Polyp-specialized transformer variant; strong generalization (highest GRR) but highest inference time among many models, limiting real-time use at high resolution |
| Swin-UMamba | Hybrid (Mamba based) | N/A | 79.45 | (81.23 + 71.12)/2 = 76.18% | 53.31 | 70.93 | 13.00 | 59.89 M, approximately 240 MB | Mamba hybrid leveraging visual-state-space-model; good generalization but relatively large and slower training/inference cost |
| Swin-UMamba-D | Hybrid (Mamba based) | N/A | 83.29 | (86.15 + 77.53)/2 = 81.84% | 88.39 | 69.36 | 12.97 | 27.50 M, approximately 110 MB | Best segmentation performance (average IoU) among study but relatively high training and inference cost; strong segmentation accuracy but moderate generalization |
| UMamba-Bot | Hybrid (Mamba based) | N/A | 71.61 | (75.03 + 61.79)/2 = 68.41% | 39.45 | 64.78 | 6.27 | 28.77 M, approximately 115 MB | Lightweight mamba variant; good FPS but weaker accuracy and generalization |
| UMamba-Enc | Hybrid (Mamba based) | N/A | 71.28 | (75.24 + 61.82)/2 = 68.53% | 37.15 | 65.33 | 7.28 | 27.56 M, approximately 110 MB | Encoder-focused mamba variant; similar trade-offs to UMamba-Bot: Faster but lower accuracy |
| VM-UNETV2 | Hybrid (Mamba like) | N/A | 81.63 | (84.36 + 74.89)/2 = 79.63% | 83.48 | 69.49 | 12.90 | 22.77 M, approximately 91 MB | VM-UNET encoder variant; strong PET and competitive accuracy, GPU-focused design |
- Citation: Yang YH. Bridging innovation and clinical reality: Interpreting the comparative study of deep learning models for multi-class upper gastrointestinal disease segmentation. World J Gastroenterol 2026; 32(8): 115297
- URL: https://www.wjgnet.com/1007-9327/full/v32/i8/115297.htm
- DOI: https://dx.doi.org/10.3748/wjg.v32.i8.115297