©The Author(s) 2025.
World J Gastroenterol. Dec 21, 2025; 31(47): 112921
Published online Dec 21, 2025. doi: 10.3748/wjg.v31.i47.112921
Published online Dec 21, 2025. doi: 10.3748/wjg.v31.i47.112921
Table 4 Summary of key studies of vision foundation models-assisted radiology in the field of gastrointestinal cancer
| Model | Year | Architecture | Training algorithm | Parameters | Datasets | Disease studied | Model type | Source code link |
| PubMedCLIP[98] | 2021 | CLIP | Fine-tuned on ROCO dataset for 50 epochs with Adam optimizer | NA | ROCO, VQA-RAD, SLAKE | Abdomen samples | Multimodal | https://github.com/sarahESL/PubMedCLIP |
| RadFM[97] | 2023 | MedLLaMA-13B | Pre-trained on MedMD and fine-tuned on RadMD | 14B | MedMD, RadMD etc. | Over 5000 diseases | Multimodal | https://github.com/chaoyi-wu/RadFM |
| Merlin[99] | 2024 | I3D-ResNet152 | Multi-task learning with EHR and radiology reports and fine-tuning for specific tasks | NA | 6M images, 6M codes and reports | Multiple diseases, Abdominal | Multimodal | NA |
| MedGemini[100] | 2024 | Gemini | Fine-tuning Gemini 1.0/1.5 on medical QA, multimodal and long-context corpora | 1.5B | MedQA, NEJM, GeneTuring | Various | Multimodal | https://github.com/Google-Health/med-gemini-medqa-relabelling |
| HAIDEF[101] | 2024 | VideoCoCa | Fine-tuning on downstream tasks with limited labeled data | NA | CT volumes and reports | Various | Vision | https://huggingface.co/collections/google/ |
| CTFM[102] | 2024 | Vision Model1 | Trained using a self-supervised learning strategy, employing a SegResNet encoder for the pre-training phase | NA | 26298 CT scans | CT scans (stomach, colon) | Vision | https://aim.hms.harvard.edu/ct-fm |
| MedVersa[103] | 2024 | Vision Model1 | Trained from scratch on the MedInterp dataset and adapted to various medical imaging tasks | NA | MedInterp | Various | Vision | https://github.com/3clyp50/MedVersa_Internal |
| iMD4GC[104] | 2024 | Transformer-based2 | A novel multimodal fusion architecture with cross-modal interaction and knowledge distillation | NA | GastricRes/Sur, TCGA etc. | Gastric cancer | Multimodal | https://github.com/FT-ZHOU-ZZZ/iMD4GC/ |
| Yasaka et al[105] | 2025 | BLIP-2 | LORA with specific fine-tuning of the fc1 layer in the vision and q-former models | NA | 5777 CT scans | Esophageal cancer via chest CT | Multimodal | NA |
- Citation: Shi L, Huang R, Zhao LL, Guo AJ. Foundation models: Insights and implications for gastrointestinal cancer. World J Gastroenterol 2025; 31(47): 112921
- URL: https://www.wjgnet.com/1007-9327/full/v31/i47/112921.htm
- DOI: https://dx.doi.org/10.3748/wjg.v31.i47.112921