BPG is committed to discovery and dissemination of knowledge
Review
©The Author(s) 2025.
World J Gastroenterol. Dec 21, 2025; 31(47): 112921
Published online Dec 21, 2025. doi: 10.3748/wjg.v31.i47.112921
Table 1 Summary of common general-purpose foundation models used in gastrointestinal cancer
Name
Type
Creator
Year
Architecture
Parameters
Modality
OSS
GI cancer applications
BERTLLMGoogle2018Encoder-only transformer110M (base), 340M (large)TextYesNLP, Radio, MLLM
GPT-3LLMOpenAI2020Decoder-only transformer175BTextNoNLP
ViTVisionGoogle2020Encoder-only transformer86M (base), 307M (large), 632M (huge)ImageYesEndo, Radio, PA, MLLM
DINOv1VisionMeta2021Encoder-only transformer22M, 86MImageYesEndo, PA
CLIPMMOpenAI2021Encoder-encoder120-580MText, ImageYesEndo, Radio, MLLM, directly1
GLM-130BLLMTsinghua2022Encoder-decoder130BTextYesNLP
Stable DiffusionMMStability AI2022Diffusion model1.45BText, ImageYesNLP, Endo, MLLM, directly
BLIPMMSalesforce2022Encoder-decoder120M (base), 340M (large)Text, ImageYesRadio, MLLM, directly
YouChatLLMYou.com2022Fine-tuned LLMsUnknownTextNoNLP
BardMMGoogle2023Based on PaLM 2340B estimatedText, Image, Audio, CodeNoNLP
Bing ChatMMMicrosoft2023Fine-tuned GPT-4UnknownText, ImageNoNLP
Mixtral 8x7BLLMMistral AI2023Decoder-only, Mixture-of-Experts (MoE)46.7B total (12.9B active per token)TextNLP
LLaVAMMMicrosoft2023Vision encoder, LLM7B, 13BText, ImageYesPA, MLLM
DINOv2VisionMeta2023Encoder-only transformer86M to 1.1BImageYesEndo, Radio, PA, MLLM, directly
Claude 2LLMAnthropic2023Decoder-only transformerUnknownTextNoNLP
GPT-4MMOpenAI2023Decoder-only transformer1.8T (Estimated)Text, ImageNoNLP, Endo, MLLM, directly
LLaMa 2LLMMeta2023Decoder-only transformer7B, 13B, 34B, 70BTextYesNLP, Endo, MLLM, directly
SAM VisionMeta2023Encoder-decoder375M, 1.25G, 2.56GImageYesEndo, directly
GPT-4VMMOpenAI2023MM transformer1.8T Text, ImageNoEndo, MLLM
QwenNLPAlibaba2023Decoder-only transformer70B, 180B, 720BTextYesNLP, MLLM
GPT-4oMMOpenAI2024MM transformerUnknown (Larger than GPT-4)Text, Image, VideoNoNLP
LLaMa 3LLMMeta2024Decoder-only transformer8B, 70B, 400BTextYesNLP, directly
Gemini 1.5MMGoogle2024MM transformer1.6TText, Image, Video, AudioNoNLP, Radio, directly
Claude 3.7MMAnthropic2024Decoder-only transformerUnknownText, ImageNoNLP, directly
YOLOWorldVisionIDEA2024CNN + RepVL-PAN vision-language fusion13-110M (depending on scale)Text, ImageYes Endo, directly
DeepSeekLLMDeepSeek2025Decoder-only transformer671BTextYesNLP
Phi-4LLMMicrosoft2025Decoder-only transformer14B (plus), 7B (mini)TextYesEndo


Write to the Help Desk