©The Author(s) 2025.
World J Gastroenterol. Dec 21, 2025; 31(47): 112921
Published online Dec 21, 2025. doi: 10.3748/wjg.v31.i47.112921
Published online Dec 21, 2025. doi: 10.3748/wjg.v31.i47.112921
Table 3 Summary of key studies of vision foundation models-assisted endoscopy in the field of gastrointestinal cancer
| Model | Year | Architecture | Training algorithm | Parameters | Datasets | Disease studied | Model type | Source code link |
| Surgical-DINO[76] | 2023 | DINOv2 | LoRA layers added to DINOv2, optimizing the LoRA layers | 86.72M | SCARED, Hamlyn | Endoscopic Surgery | Vision | https://github.com/BeileiCui/SurgicalDINO |
| ProMISe[77] | 2023 | SAM (ViT-B) | APM and IPS modules are trained while keeping SAM frozen | 1.3-45.6M | EndoScene, ColonDB etc. | Polyps, Skin Cancer | Vision | NA |
| Polyp-SAM[78] | 2023 | SAM | Strategy as pretrain only the mask decoder while freezing all encoders | NA | CVC-ColonDB Kvasir etc. | Colon Polyps | Vision | https://github.com/ricklisz/Polyp-SAM |
| Endo-FM[79] | 2023 | ViT B/16 | Pretrained using a self-supervised teacher-student framework, and fine-tuned on downstream tasks | 121M | Colonoscopic, LDPolyp etc. | Polyps, erosion, etc. | Vision | https://github.com/med-air/Endo-FM |
| ColonGPT[80] | 2024 | SigLIP-SO, Phi1.5 | Pre-alignment with image-caption pairs, followed by supervised fine-tuning using LoRA | 0.4-1.3B | ColonINST (30k+ images) | Colorectal polyps | Vision | https://github.com/ColonGPT/ColonGPT |
| DeepCPD[81] | 2024 | ViT | Hyperparameters are optimized for colonoscopy datasets, including Adam optimizer | NA | PolypsSet, CP-CHILD-A etc. | CRC | Vision | https://github.com/Zhang-CV/DeepCPD |
| OneSLAM[82] | 2024 | Transformer (CoTracker) | Zero-shot adaptation using TAP + Local Bundle Adjustment | NA | SAGE-SLAM, C3VD etc. | Laparoscopy, Colon | Vision | https://github.com/arcadelab/OneSLAM |
| EIVS[83] | 2024 | Vision Mamba, CLIP | Unsupervised Cycle‑Consistency | 63.41M | 613 WLE, 637 images | Gastrointestinal | Vision | NA |
| APT[84] | 2024 | SAM | Parameter-efficient fine-tuning | NA | Kvasir-SEG, EndoTect etc. | CRC | Vision | NA |
| FCSAM[85] | 2024 | SAM | LayerNorm LoRA fine-tuning strategy | 1.2M | Gastric cancer (630 pairs) etc. | GC, Colon Polyps | Vision | NA |
| DuaPSNet[86] | 2024 | PVTv2-B3 | Transfer learning with pre-trained PVTv2-B3 on ImageNet | NA | LaribPolypDB, ColonDB etc. | CRC | Vision | https://github.com/Zachary-Hwang/Dua-PSNet |
| EndoDINO[87] | 2025 | ViT (B, L, g) | DINOv2 methodology, hyperparameters tuning | 86M to 1B | HyperKvasir, LIMUC | GI Endoscopy | Vision | https://github.com/ZHANGBowen0208/EndoDINO/ |
| PolypSegTrack[88] | 2025 | DINOv2 | One-step fine-tuning on colonoscopic videos without first pre-training | NA | ETIS, CVC-ColonDB etc. | Colon polyps | Vision | NA |
| AiLES[89] | 2025 | RF-Net | Not fine-tuned from external model | NA | 100 GC patients | Gastric cancer | Vision | https://github.com/CalvinSMU/AiLES |
| PPSAM[90] | 2025 | SAM | Fine-tuning with variable bounding box prompt perturbations | NA | EndoScene, ColonDB etc. | Investigated in Ref. | Vision | https://github.com/SLDGroup/PP-SAM |
| SPHINX-Co[91] | 2024 | LLaMA-2 + SPHINX-X | Fine-tuned SPHINX-X on CoPESD with cosine learning rate scheduler | 7B, 13B | CoPESD | Gastric cancer | Multimodal | https://github.com/gkw0010/CoPESD |
| LLaVA-Co[91] | 2024 | LLaVA-1.5 (CLIP-ViT-L) | Fine-tuned LLaVA-1.5 on CoPESD with cosine learning rate scheduler | 7B, 13B | CoPESD | Gastric cancer | Multimodal | https://github.com/gkw0010/CoPESD |
| ColonCLIP[92] | 2025 | CLIP | Prompt tuning with frozen CLIP, then encoder fine-tuning with frozen prompts | 57M, 86M | OpenColonDB | CRC | Multimodal | https://github.com/Zoe-TAN/ColonCLIP-OpenColonDB |
| PSDM[93] | 2025 | Stable Diffusion + CLIP | Continual learning with prompt replay to incrementally train on multiple datasets | NA | PolypGen, ColonDB, Polyplus etc. | CRC | Vision, Generative | The original paper reported a GitHub link for this model, but it is currently unavailable |
| PathoPolypDiff[94] | 2025 | Stable Diffusion v1-4 | Fine-tuned Stable Diffusion v1-4 and locked first U-Net block, fine-tuned remaining blocks | NA | ISIT-UMR Colonoscopy Dataset | CRC | Generative | https://github.com/Vanshali/PathoPolyp-Diff |
- Citation: Shi L, Huang R, Zhao LL, Guo AJ. Foundation models: Insights and implications for gastrointestinal cancer. World J Gastroenterol 2025; 31(47): 112921
- URL: https://www.wjgnet.com/1007-9327/full/v31/i47/112921.htm
- DOI: https://dx.doi.org/10.3748/wjg.v31.i47.112921