©The Author(s) 2025.
World J Gastroenterol. Dec 21, 2025; 31(47): 112921
Published online Dec 21, 2025. doi: 10.3748/wjg.v31.i47.112921
Published online Dec 21, 2025. doi: 10.3748/wjg.v31.i47.112921
Table 5 Summary of key studies of Vision Foundation Models-assisted pathology in the field of gastrointestinal cancer
| Model | Year | Architecture | Training Algorithm | Paras | WSIs | Tissues | Open source link |
| LUNIT-SSL[110] | 2021 | ViT-S | DINO; full fine-tuning and linear evaluation on downstream tasks | 22M | 3.7K | 32 | https://Lunitio.github.io/research/publications/pathology_ssl |
| CTransPath[111] | 2022 | Swin Transformer | MoCoV3 (SRCL); frozen backbone with linear classifier fine-tuning | 28M | 32K | 32 | https://github.com/Xiyue-Wang/TransPath |
| Phikon[112] | 2023 | ViT-B | iBOT (Masked Image Modeling); fine-tuned with ABMIL/TransMIL on frozen features | 86M | 6K | 16 | https://github.com/owkin/HistoSSLscaling |
| REMEDIS[113] | 2023 | BiT-L (ResNet-152) | SimCLR (contrastive learning); end-to-end fine-tuning on labeled ID/OOD data | 232M | 29K | 32 | https://github.com/google-research/simclr |
| Virchow[114] | 2024 | ViT-H, DINOv2 | DINOv2 (SSL); used frozen embeddings with simple aggregators | 632M | 1.5M | 17 | https://huggingface.co/paige-ai/Virchow |
| Virchow2[115] | 2024 | ViT-H | DINOv2 (SSL); fine-tuned with linear probes or full-tuning on downstream tasks | 632M | 3.1M | 25 | https://huggingface.co/paige-ai/Virchow2 |
| Virchow2G[115] | 2024 | ViT-G | DINOv2 (SSL); fine-tuned with linear probes or full fine-tuning | 1.9B | 3.1M | 25 | https://huggingface.co/paige-ai/Virchow2 |
| Virchow2G mini[115]1 | 2024 | ViT-S, Virchow2G | DINOv2 (SSL); distilled from Virchow2G, then fine-tuned on downstream tasks | 22M | 3.2M | 25 | https://huggingface.co/paige-ai/Virchow2 |
| UNI[9] | 2024 | ViT-L | DINOv2 (SSL); used frozen features with linear probes or few-shot learning | 307M | 100K | 20 | https://github.com/mahmoodlab/UNI |
| Phikon-v2[116] | 2024 | ViT-L | DINOv2 (SSL); frozen ViT and ABMIL ensemble fine-tuning | 307M | 58K | 30 | https://huggingface.co/owkin/phikon-v2 |
| RudolfV[117] | 2024 | ViT-L | DINOv2 (SSL); fine-tuned with optimizing linear classification layer and adapting encoder weights | 304M | 103K | 58 | https://github.com/rudolfv |
| HIBOU-B[118] | 2024 | ViT-B | DINOv2 (SSL); frozen feature extractor, trained linear classifier or attention pooling | 86M | 1.1M | 12 | https://github.com/HistAI/hibou |
| HIBOU-L[118]2 | 2024 | ViT-L | DINOv2 (SSL); frozen feature extractor, trained linear classifier or attention pooling | 307M | 1.1M | 12 | https://github.com/HistAI/hibou |
| H-Optimus-03 | 2024 | ViT-G | DINOv2 (SSL); linear probe and ABMIL on frozen features | 1.1B | > 500K | 32 | https://github.com/bioptimus/releases/ |
| Madeleine[119] | 2024 | CONCH | MAD-MIL; linear probing, prototyping, and full fine-tuning for downstream tasks | 86M | 23K | 2 | https://github.com/mahmoodlab/MADELEINE |
| COBRA[120] | 2024 | Mamba-2 | Self-supervised contrastive pretraining with multiple FMs and Mamba2 architecture | 15M | 3K | 6 | https://github.com/KatherLab/COBRA |
| PLUTO[121] | 2024 | FlexiVit-S | DINOv2; frozen backbone with task-specific heads for fine-tuning | 22M | 158K | 28 | NA |
| HIPT[122] | 2025 | ViT-HIPT | DINO (SSL); fine-tune with gradient accumulation | 10M | 11K | 33 | https://github.com/mahmoodlab/HIPT |
| PathoDuet[123] | 2025 | ViT-B | MoCoV3; fine-tuned using standard supervised learning on labeled downstream task data | 86M | 11K | 32 | https://github.com/openmedlab/PathoDuet |
| Kaiko[124] | 2025 | ViT-L | DINOv2 (SSL); linear probing with frozen encoder on downstream tasks | 303M | 29K | 32 | https://github.com/kaiko-ai/towards_large_pathology_fms |
| PathOrchestra[125] | 2025 | ViT-L | DINOv2; ABMIL, linear probing, weakly supervised classification | 304M | 300K | 20 | https://github.com/yanfang-research/PathOrchestra |
| THREADS[126] | 2025 | ViT-L, CONCHv1.5 | Fine-tune gene encoder, initialize patch encoder randomly | 16M | 47K | 39 | https://github.com/mahmoodlab/trident |
| H0-mini[127] | 2025 | ViT | Using knowledge distillation from H-Optimus-0 | 86M | 6K | 16 | https://huggingface.co/bioptimus/H0-mini |
| TissueConcepts[128] | 2025 | Swin Transformer | Frozen encoder with linear probe for downstream tasks | 27.5M | 7K | 14 | https://github.com/FraunhoferMEVIS/MedicalMultitaskModeling |
| OmniScreen[129] | 2025 | Virchow2 | Attention-aggregated Virchow2 embeddings fine-tuning | 632M | 48K | 27 | https://github.com/OmniScreen |
| BROW[130] | 2025 | ViT-B | DINO (SSL); self-distillation with multi-scale and augmented views | 86M | 11K | 6 | NA |
| BEPH[131] | 2025 | BEiTv2 | BEiTv2 (SSL); supervised fine-tuning on clinical tasks with labeled data | 86M | 11K | 32 | https://github.com/Zhcyoung/BEPH |
| Atlas[132] | 2025 | ViT-H, RudolfV | DINOv2; linear probing with frozen backbone on downstream tasks | 632M | 1.2M | 70 | NA |
- Citation: Shi L, Huang R, Zhao LL, Guo AJ. Foundation models: Insights and implications for gastrointestinal cancer. World J Gastroenterol 2025; 31(47): 112921
- URL: https://www.wjgnet.com/1007-9327/full/v31/i47/112921.htm
- DOI: https://dx.doi.org/10.3748/wjg.v31.i47.112921