AnatCL
Anatomical foundation models for brain MRIs.
Motivation
Weakly supervised contrastive pre-training on brain MRI usually relies on a single attribute, the subject’s age, to decide which scans should be close in representation space. Anatomical measures such as cortical thickness carry more information about the brain, and they can be computed from the MRI itself with standard tools such as FreeSurfer, without extra labels.
Method
AnatCL extends weakly supervised contrastive learning to multiple attributes.
- Anatomical similarity. For two scans, a degree of positiveness is computed from three FreeSurfer measures: cortical thickness, gray matter volume, and surface area, on the Desikan-Killiany (68 regions) or Destrieux (148 regions) atlas.
- Local and global variants. The local variant compares the measures region by region; the global variant compares each measure across the whole brain.
- Objective. The anatomical loss is combined with an age-based loss (y-Aware): L = λ₁ LAnatCL + λ₂ Lage.
- Pre-training. A 3D ResNet-18 trained on the healthy subjects of OpenBHB (3,984 T1-weighted MRIs, VBM preprocessing).
Results
The pre-trained encoder is frozen and evaluated with linear probing on 12 downstream tasks (ADNI, OASIS-3, SchizConnect, ABIDE I) and 10 clinical assessment scores, against SimCLR, brain-age regression (L1), y-Aware, and ExpW.
- AnatCL gives the best average performance across tasks and the lowest brain-age error on OpenBHB (MAE 2.55 years).
- It leads on most schizophrenia tasks, on Asperger’s and PDD-NOS, and on Alzheimer’s disease detection on OASIS-3.
- Ablations show that combining anatomical and age information works better than either one alone.

Independent evaluation: cross-scanner reliability
An independent study by Navarro-González et al. (medRxiv preprint, 2026) measured how stable the embeddings of five brain MRI foundation models are when the same people are scanned on different scanners, using the ON-Harmony travelling-heads dataset (20 participants, 8 scanners, 3 vendors). AnatCL had the highest between-scanner reliability (median ICC 0.97), above the FreeSurfer morphometric baseline (0.93), y-Aware (0.81), and the purely self-supervised models (0.25–0.45). It also had the smallest gap between within-scanner and between-scanner reliability, and always matched the correct subject across scanners.

Code
Code and pre-trained models are available on GitHub.