Meet Phaet and Mascaret: two pathology foundation models built to generalize across labs
Pathology foundation models are powerful, but they quietly encode which scanner and lab produced a whole slide image, hurting performance across sites. Our model-agnostic fine-tuning recipe removes this fragility, improving robustness and performance together. Two of the models are now available on Hugging Face.
The problem: pathology foundation models that don't travel well
Over the past few years, foundation models have transformed computational pathology and AI-powered digital pathology. Trained on vast collections of digitized slides, they learn general-purpose representations of tissue that can be adapted to a wide range of downstream tasks, from AI biomarker discovery and detection to gene-expression modeling, tumor microenvironment (TME) characterization and survival analysis.
But these models inherit a long-standing weakness of the field. Their representations are entangled with acquisition factors, such as the scanner model or the staining protocol used to digitize a slide. Two images of the same tissue, scanned on different devices or stained in different laboratories, can map to markedly different points in the model's feature space. The result is a domain shift that degrades performance the moment a model is deployed in a laboratory whose acquisition pipeline differs from the one it was trained on. For a field whose promise rests on real-world, multi-site clinical use, that fragility is a fundamental barrier.
Why scaling alone hasn't solved it
The dominant response has been to scale, building ever-larger models trained on more slides. Bigger models are indeed more robust to acquisition shift, but a growing body of work shows that even the most recent foundation models still entangle non-biological factors like scanner, laboratory, and staining into their representations.1,2,3 Dedicated benchmarks now exist specifically to quantify this fragility, and they confirm that perfect robustness has not yet been achieved.4,5
Existing mitigation strategies mostly work around the encoder rather than fixing it: normalizing or augmenting stain appearance at the image level, correcting embeddings after the fact using site statistics, or suppressing site-predictive signals during downstream task training. All of these keep the backbone frozen. Knowledge distillation can help too, but it requires an already-robust teacher model and is computationally expensive. The open question has been how to instill invariance to acquisition factors directly, in a way that works across models.
Our approach: robustifying pathology foundation models
We developed a novel fine-tuning recipe that makes pathology foundation models more robust to acquisition factors, without retraining from scratch and without sacrificing the general quality of the representation. The output is a single, task-agnostic, robust encoder that can be dropped into existing pipelines in place of the original.
Crucially, the method is model-agnostic. We applied the same recipe uniformly to ten different foundation models spanning a range of architectures and pretraining recipes, and it improved every one of them.

The headline result: robustness and performance, together
The striking finding is that there is no trade-off. Across all ten models, fine-tuning improved robustness and downstream performance simultaneously. On average, it raised the PathoROB robustness index by 23% (from 0.72 to 0.87) and increased overall cross-benchmark performance by 43% across three community resources combined: Patho-Bench and THUNDER, two standardized benchmarks for evaluating pathology foundation models, and HEST, a dataset pairing histology with spatial transcriptomics.6,7,8,9
Individual gains reached as high as 72% in robustness (for Phikon-v2, released as Phaet) and 76% in performance (for Midnight-12k, released as Mascaret). After fine-tuning, fine-tuned encoders took every top position on the individual leaderboards, and 8 of the 10 most robust models were fine-tuned versions. All improvements were statistically significant (p< 0.0001).
What the model actually learns
To understand where the gains come from, it helps to look at what the model organizes its feature space around. Before fine-tuning, the leading axis of variation often separates slides by how they were digitized rather than by what the tissue shows. After fine-tuning, that same axis reorganizes around biology.
A clear example comes from the Camelyon dataset: sentinel lymph-node tiles from breast-cancer patients across five Dutch medical centers, digitized on three different whole-slide scanners. In the original Phikon-v2 model, the representation clusters by medical center, not by biology (an Adjusted Rand Index of 0.46 against center versus 0.00 against metastasis status). The three centers sharing one scanner even collapse into a single cloud. After fine-tuning, the centers become intermixed (ARI 0.01) while tumor and normal tiles cleanly separate (ARI 0.67). In effect, fine-tuning re-purposes the dominant directions of the feature space from encoding where a slide was digitized to encoding what the morphology is. Notably, none of these five centers were seen during fine-tuning, so the invariance generalizes to unseen acquisition sources.

Benchmarking foundation models for computational pathology
These results rest on a deliberately broad evaluation: ten foundation models assessed on robustness (PathoROB) and on downstream performance across three established benchmarks covering gene-expression prediction (HEST), tile-level tasks (THUNDER) and slide-level tasks (Patho-Bench). All results, for both base and fine-tuned models, were generated using the official benchmark implementations with default hyperparameters and no additional tuning, so the comparisons are as clean as we could make them.
Two robust models, available for research
We are releasing two of the fine-tuned models for the community to test and benchmark: Phaet, the robustified version of Phikon-v2, and Mascaret, the robustified version of Midnight-12k. The names follow a wave theme drawn from the Waiv brand (a mascaret is a tidal bore, a wave that carries its shape forward) echoing the idea of representations that stay stable as they move from one setting to another.
By making pathology foundation models robust to the scanner and lab that produced a slide, we aim to move AI digital pathology closer to dependable, real-world use producing a foundation for scaling AI precision oncology and precision diagnostics so that no patient is left behind.
If you build or benchmark pathology foundation models, we would love for you to put Phaet and Mascaret through their paces. Request access and download the models on Hugging Face.
For the full method, experiments and results, read the preprint: Robustifying pathology foundation models via fine-tuning (arXiv).
Authors
References
- Jonah Kömen, Edwin D. de Jong, Julius Hense, Hannah Marienwald, Jonas Dippel, Philip Naumann, Eric Marcus, Lukas Ruff, Maximilian Alber, Jonas Teuwen, Frederick Klauschen, and Klaus-Robert Müller. Towards robust foundation models for digital pathology. Nature Communications, 17, 2026. doi:10.1038/s41467-026-73923-2. arXiv:2507.17845.
- Binghao Chai, Jianan Chen, Paul Cool, Fatine Oumlil, Anna Tollitt, David F. Steiner, Tapabrata Chakraborti, and Adrienne M. Flanagan. Impact of tissue staining and scanner variation on the performance of pathology foundation models: a study of sarcomas and their mimics. The Journal of Pathology: Clinical Research, 12(2):e70080, 2026. doi: 10.1002/2056-4538.70080.
- Fredrik K. Gustafsson and Mattias Rantalainen. Evaluating computational pathology foundation models for prostate cancer grading under distribution shifts. arXiv preprint arXiv:2410.06723, 2024
- Jonah Kömen, Edwin D. de Jong, Julius Hense, Hannah Marienwald, Jonas Dippel, Philip Naumann, Eric Marcus, Lukas Ruff, Maximilian Alber, Jonas Teuwen, Frederick Klauschen, and Klaus-Robert Müller. Towards robust foundation models for digital pathology. Nature Communications, 17, 2026. doi: 10.1038/s41467-026-73923-2. arXiv:2507.17845
- Alexandre Filiot, Nicolas Dop, Oussama Tchita, Auriane Riou, Rémy Dubois, Thomas Peeters, Daria Valter, Marin Scalbert, Charlie Saillard, Geneviève Robin, and Antoine Olivier. Distilling foundation models for robust and efficient models in digital pathology. In Medical Image Computing and Computer Assisted Intervention (MICCAI), volume 15966 of Lecture Notes in Computer Science, pages 162–172. Springer, 2025. doi: 10.1007/978-3-032-04981-0_16. arXiv:2501.16239
- PathoROB: https://github.com/bifold-pathomics/PathoROB
- HEST: https://github.com/mahmoodlab/HEST
- THUNDER: https://mics-lab.github.io/thunder/leaderboards/
- Patho-bench: https://github.com/mahmoodlab/patho-bench