©The Author(s) 2025.
World J Gastroenterol. Dec 21, 2025; 31(47): 112921
Published online Dec 21, 2025. doi: 10.3748/wjg.v31.i47.112921
Published online Dec 21, 2025. doi: 10.3748/wjg.v31.i47.112921
Table 1 Summary of common general-purpose foundation models used in gastrointestinal cancer
| Name | Type | Creator | Year | Architecture | Parameters | Modality | OSS | GI cancer applications |
| BERT | LLM | 2018 | Encoder-only transformer | 110M (base), 340M (large) | Text | Yes | NLP, Radio, MLLM | |
| GPT-3 | LLM | OpenAI | 2020 | Decoder-only transformer | 175B | Text | No | NLP |
| ViT | Vision | 2020 | Encoder-only transformer | 86M (base), 307M (large), 632M (huge) | Image | Yes | Endo, Radio, PA, MLLM | |
| DINOv1 | Vision | Meta | 2021 | Encoder-only transformer | 22M, 86M | Image | Yes | Endo, PA |
| CLIP | MM | OpenAI | 2021 | Encoder-encoder | 120-580M | Text, Image | Yes | Endo, Radio, MLLM, directly1 |
| GLM-130B | LLM | Tsinghua | 2022 | Encoder-decoder | 130B | Text | Yes | NLP |
| Stable Diffusion | MM | Stability AI | 2022 | Diffusion model | 1.45B | Text, Image | Yes | NLP, Endo, MLLM, directly |
| BLIP | MM | Salesforce | 2022 | Encoder-decoder | 120M (base), 340M (large) | Text, Image | Yes | Radio, MLLM, directly |
| YouChat | LLM | You.com | 2022 | Fine-tuned LLMs | Unknown | Text | No | NLP |
| Bard | MM | 2023 | Based on PaLM 2 | 340B estimated | Text, Image, Audio, Code | No | NLP | |
| Bing Chat | MM | Microsoft | 2023 | Fine-tuned GPT-4 | Unknown | Text, Image | No | NLP |
| Mixtral 8x7B | LLM | Mistral AI | 2023 | Decoder-only, Mixture-of-Experts (MoE) | 46.7B total (12.9B active per token) | Text | NLP | |
| LLaVA | MM | Microsoft | 2023 | Vision encoder, LLM | 7B, 13B | Text, Image | Yes | PA, MLLM |
| DINOv2 | Vision | Meta | 2023 | Encoder-only transformer | 86M to 1.1B | Image | Yes | Endo, Radio, PA, MLLM, directly |
| Claude 2 | LLM | Anthropic | 2023 | Decoder-only transformer | Unknown | Text | No | NLP |
| GPT-4 | MM | OpenAI | 2023 | Decoder-only transformer | 1.8T (Estimated) | Text, Image | No | NLP, Endo, MLLM, directly |
| LLaMa 2 | LLM | Meta | 2023 | Decoder-only transformer | 7B, 13B, 34B, 70B | Text | Yes | NLP, Endo, MLLM, directly |
| SAM | Vision | Meta | 2023 | Encoder-decoder | 375M, 1.25G, 2.56G | Image | Yes | Endo, directly |
| GPT-4V | MM | OpenAI | 2023 | MM transformer | 1.8T | Text, Image | No | Endo, MLLM |
| Qwen | NLP | Alibaba | 2023 | Decoder-only transformer | 70B, 180B, 720B | Text | Yes | NLP, MLLM |
| GPT-4o | MM | OpenAI | 2024 | MM transformer | Unknown (Larger than GPT-4) | Text, Image, Video | No | NLP |
| LLaMa 3 | LLM | Meta | 2024 | Decoder-only transformer | 8B, 70B, 400B | Text | Yes | NLP, directly |
| Gemini 1.5 | MM | 2024 | MM transformer | 1.6T | Text, Image, Video, Audio | No | NLP, Radio, directly | |
| Claude 3.7 | MM | Anthropic | 2024 | Decoder-only transformer | Unknown | Text, Image | No | NLP, directly |
| YOLOWorld | Vision | IDEA | 2024 | CNN + RepVL-PAN vision-language fusion | 13-110M (depending on scale) | Text, Image | Yes | Endo, directly |
| DeepSeek | LLM | DeepSeek | 2025 | Decoder-only transformer | 671B | Text | Yes | NLP |
| Phi-4 | LLM | Microsoft | 2025 | Decoder-only transformer | 14B (plus), 7B (mini) | Text | Yes | Endo |
- Citation: Shi L, Huang R, Zhao LL, Guo AJ. Foundation models: Insights and implications for gastrointestinal cancer. World J Gastroenterol 2025; 31(47): 112921
- URL: https://www.wjgnet.com/1007-9327/full/v31/i47/112921.htm
- DOI: https://dx.doi.org/10.3748/wjg.v31.i47.112921