ICLR 2026 Workshop on Foundation Models for Science (FM4Science) · 2026
Benchmarking Foundation Models for Unsupervised Discovery in Large Multimodal Astrophysical Datasets
We compare AstroPT, AstroCLIP and AION representations for unsupervised discovery in matched Euclid imaging and DESI spectroscopy. A scalable density-estimation pipeline combines per-modality rarity with cross-modal misalignment to reveal model-dependent rankings, instrumental artefacts and physically coherent candidates including AGN and strong gravitational lenses.
- Models
- AstroPT · AION · AstroCLIP
- Data
- Euclid imaging × DESI spectra
- Task
- Unsupervised anomaly discovery
01 · Approach
One dataset, three representation spaces
We encode the same matched galaxy sample with three astronomical foundation models. Lightweight density estimators then rank objects that are rare in an individual modality or unusually misaligned across image and spectrum.
- Extract image, spectral and joint embeddings for each object.
- Estimate rarity inside each representation with normalizing flows.
- Compare rankings and inspect the highest-scoring candidates across models.

Latent geometry differs substantially across models. AstroPT and AION form relatively continuous manifolds, while AstroCLIP is more fragmented and shows a broader image-spectrum alignment distribution.
02 · Findings
The ranking is model-dependent
The most extreme candidates partly agree across models, but the broader rankings diverge. This makes cross-model comparison useful: consensus highlights robust rare systems, while disagreement exposes architecture-specific sensitivity.
Visual inspection surfaces both astrophysical candidates and data-quality issues, including AGN-like spectra, strong-lens morphologies, unusual quiescent systems, diffraction spikes and residual imaging artefacts.

Representative high-ranking candidates pair Euclid image cutouts with DESI spectra, keeping the physical and instrumental interpretation visible together.
Takeaway
Foundation models do not define a single notion of rarity.
Their anomaly rankings reflect different representation geometries and training objectives. Comparing those views is therefore part of the discovery method, not only a benchmarking exercise.
Read the full paper