Papers

ICLR 2026 Workshop on Foundation Models for Science (FM4Science) · 2026

Benchmarking Foundation Models for Unsupervised Discovery in Large Multimodal Astrophysical Datasets

M. Ronceray, M. Huertas-Company, A. Chanson, M. Siudek, A. Preto, M. J. Smith, J. Zoubian, C. Bonini

We compare AstroPT, AstroCLIP and AION representations for unsupervised discovery in matched Euclid imaging and DESI spectroscopy. A scalable density-estimation pipeline combines per-modality rarity with cross-modal misalignment to reveal model-dependent rankings, instrumental artefacts and physically coherent candidates including AGN and strong gravitational lenses.

Models
AstroPT · AION · AstroCLIP
Data
Euclid imaging × DESI spectra
Task
Unsupervised anomaly discovery

We encode the same matched galaxy sample with three astronomical foundation models. Lightweight density estimators then rank objects that are rare in an individual modality or unusually misaligned across image and spectrum.

  1. Extract image, spectral and joint embeddings for each object.
  2. Estimate rarity inside each representation with normalizing flows.
  3. Compare rankings and inspect the highest-scoring candidates across models.
UMAP projections, representative galaxy thumbnails and image-spectrum alignment distributions for AstroPT, AION and AstroCLIP
Figure 01

Latent geometry differs substantially across models. AstroPT and AION form relatively continuous manifolds, while AstroCLIP is more fragmented and shows a broader image-spectrum alignment distribution.

The most extreme candidates partly agree across models, but the broader rankings diverge. This makes cross-model comparison useful: consensus highlights robust rare systems, while disagreement exposes architecture-specific sensitivity.

Visual inspection surfaces both astrophysical candidates and data-quality issues, including AGN-like spectra, strong-lens morphologies, unusual quiescent systems, diffraction spikes and residual imaging artefacts.

Four high-ranking Euclid anomaly candidates with their matched DESI spectra
Figure 02

Representative high-ranking candidates pair Euclid image cutouts with DESI spectra, keeping the physical and instrumental interpretation visible together.

Takeaway

Foundation models do not define a single notion of rarity.

Their anomaly rankings reflect different representation geometries and training objectives. Comparing those views is therefore part of the discovery method, not only a benchmarking exercise.

Read the full paper