Conceptio › Archive › arXiv CS
arXiv CSopen access

Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Back into Plato’s Cave: Examining Cross-modal Representational Convergence at Scale

A. Sophia Koepke1,2,3

arXiv:2604.18572v1 [cs.CV] 20 Apr 2026

1 3

Daniil Zverev2

UC Berkeley

2

Shiry Ginosar4

Alexei A. Efros1

Technical University Munich, MCML

University of Tübingen, Tübingen AI Center

4

Toyota Technical Institute at Chicago

Abstract The Platonic Representation Hypothesis [40] suggests that neural networks trained on different modalities (e.g., text and images) align and eventually converge toward the same representation of reality. If true, this has significant implications for whether modality choice matters at all. We show that the experimental evidence for this hypothesis is fragile and depends critically on the evaluation regime. Alignment is measured using mutual nearest neighbors on small datasets (≈1K samples) and degrades substantially as the dataset is scaled to millions of samples. The alignment that remains between model representations reflects coarse semantic overlap rather than consistent fine-grained structure. Moreover, the evaluations in Huh et al. [40] are done in a one-to-one image-caption setting, a constraint that breaks down in realistic many-to-many settings and further reduces alignment. We also find that the reported trend of stronger language models increasingly aligning with vision does not hold for newer models. Overall, our findings suggest that the current evidence for cross-modal representational convergence is considerably weaker than subsequent works have taken it to be. Models trained on different modalities may learn equally rich representations of the world, just not the same one. Project page: https://akoepke.github.io/cave_umwelten

1

Introduction

The success of Large Language Models (LLMs) is causing much hand-wringing in the computer vision community: do we even need pixels to build machines that understand our world, or is language “all you need”? Several works have demonstrated that models trained only on text data have made progress in solving what were thought to be fundamentally visual problems, such as visual question answering (VQA) [30, 38], visual reasoning [10, 37, 91, 2], or embodied robotics applications [1, 54]. This resonates with the suggestion that text data may make other modalities redundant [83], on the premise that the part of the world that is relevant to humans is manifest in language. On the other hand, it is argued that linguistic data alone cannot yield genuine understanding [8] or allow actual embodiment. After all, there is a reason we visit art museums rather than just read descriptions of paintings in a catalogue. This raises a central question: how do models trained on different modalities represent reality? The Platonic Representation Hypothesis [40] offers a compelling answer: as neural networks grow larger and consume more data, their learned representations will become more and more aligned, no matter which data modality (text, vision, audio, touch, etc.) was used for training. Proponents of language-only learning have interpreted this as validation of their approach: since the choice of modality does not matter as they all lead to the same shared representation, one might as well use Preprint.

Figure 1: Illustration of the mutual nearest neighbor metric used by Huh et al. [40] to measure crossmodal alignment. (a) Sparse regime: given a query image and caption (blue), nearest neighbors (NN) are retrieved independently in image and text embedding spaces. Mutual NN alignment measures whether the NNs are consistent across modalities. (b) Dense regime: as dataset size increases, NNs within each modality get better. The vision model retrieves a car in the same pose, and the language model retrieves a caption of the same car model regardless of pose. At scale, improved withinmodality organization does not translate into cross-modal agreement. language as the most convenient source of data.1 However, the strength of a hypothesis depends on the strength of its evidence, and the experimental protocol underpinning the claim rests on specific methodological choices that have largely gone unexamined in subsequent work. In this paper, we take a closer look at the experimental evidence for the hypothesis and find it to be fragile and to depend critically on the evaluation regime. Huh et al. [40] conducted their analysis on small, sparse datasets with one-to-one correspondences between modalities. However, real-world multi-modal data is large, dense, and inherently many-to-many: one image has many valid descriptions, and a single caption can correspond to many plausible images. These differences fundamentally change what it means for two representations to “align”. In a small dataset, weakly related samples may become nearest neighbors simply because no better alternatives exist (Fig. 1a). Here, two models can agree despite organizing their representations differently. As the dataset grows (i.e. the gallery used for retrieving nearest neighbors gets denser), both models find closer neighbors and cross-modal consistency requires more fine-grained structural alignment (Fig. 1b). A vision model may retrieve an image of a car taken from a similar angle as the query, while the language model retrieves a caption describing the same car model as the query but in a different pose. Both are valid, but inconsistent between modalities, producing a mismatch that gets penalized under the mutual nearest-neighbor metric. This illustrates how mutual nearest-neighbor agreement becomes an increasingly strict measure for alignment in many-to-many regimes. Using a mutual kNN metric with k > 1 is less strict, but does not fundamentally change the conclusion. 1 But analogously, the same argument could be made for vision-only learning [41].

2

In this paper, we examine how cross-modal alignment changes in evaluation settings with large, dense, and non-bijective datasets, and observe the following: Alignment degrades with scale: Increasing the gallery from 1024 to millions of samples causes a sharp drop in cross-modal mutual nearest-neighbor agreement. Coarse agreement persists but fine-grained agreement does not: In controlled settings (e.g., ImageNet), vision and language models reliably retrieve correct-class neighbors but rarely agree on the same instance. Many-to-many correspondence reduces alignment: Allowing multiple valid correspondences per sample leads to drops in agreement, even when retrieved neighbors are semantically sensible. Previously reported trends may not hold for newer models: The claim that stronger language models align better with vision seems to weaken for more recent models. These findings paint a more mixed picture than the small-gallery results from Huh et al. suggest. Models trained on different modalities can learn rich and semantically meaningful structure, yet still organize that structure differently. Low agreement does not imply poor representations, it reflects differences in how information is arranged. Nearly a century ago, von Uexküll [89] argued that every organism inhabits its own perceptual world, or Umwelt, shaped by its senses rather than by an observer-independent reality. The same, we believe, might hold for our models: each constructs its own representational structure, determined by its modality and training data, rather than converging toward a shared model of reality. Though it is still early days, we suspect future evidence will favor von Uexküll over Plato.

2

Related Work

One Platonic Ideal vs many Umwelten. In his “Theory of Forms”, Plato argued that every physical object we perceive is a flawed imitation (a shadow) of some eternal, abstract “ideal” form [71], and only by escaping from the tyranny of our physical senses (leaving the cave of shadows), we can achieve true understanding. But in the 20th century, this argument for a single, unified Platonic Ideal representation has been repeatedly undercut by biologists, psychologists, and philosophers. Biologist von Uexküll argued that every organism inhabits its own perceptual environment, or Umwelt [89]: a tick lives in a world of thermal gradients, a bat in a world of echoes. The different Umwelten might have only little overlap with each other2 . Gibson’s ecological psychology [25] pushed this further, proposing that perception is shaped by what an organism can do in its environment, not by an observerindependent reality. Philosopher Wittgenstein, thinking about language, arrived at a strikingly similar conclusion. He famously argued: “If a lion could speak, we could not understand him” [94], meaning that the lion’s world (goals, instincts, perceived reality) is so utterly different from our own, that even if it spoke English, we would not comprehend the meaning3 . Building on Wittgenstein, psychologist Rosch developed her Prototype Theory of Categorization [75], arguing forcefully against a single platonic ideal as a representation of object categories, proposing a data-driven clustering-based model instead. Representational alignment. The question of representational similarity has been studied extensively in the neurosciences [19, 33, 48]. In machine learning, the parallel question of whether independently trained networks learn similar internal structure has received growing attention. Lenc and Vedaldi [52] investigated the equivalence of representations from different trained models and found that early convolutional layers are more interchangeable than later ones. This task, also referred to as “model stitching”, was later revisited by Bansal et al. [6]. Related to this, Li et al. [53] proposed methods to align neurons across independently trained networks. More recently, Dravid et al. [18] introduced 2 For a tour of von Uexküll’s ideas, see Koenderink’s delightful book [46]. 3 For a great treatment of Wittgenstein’s argument in popular culture, see the episode Darmok of the American TV series

Star Trek: The Next Generation.

3

“Rosetta Neurons,” showing that different vision models share common units corresponding to similar visual concepts across architectures, tasks, and training data. Furthermore, alignment has been linked with shared model capabilities measured by task performances [5, 45, 40, 64, 6]. To directly quantify representational similarities, several metrics have been used to measure correlations between features [64, 36]. Kornblith et al. [47] introduced Central Kernel Alignment (CKA) as a robust measure invariant to orthogonal transformations and isotropic scaling. Huh et al. [40] found the CKA metric to reveal only a “very weak trend of alignment between models” and therefore proposed the use of the mutual kNN metric that measures the overlap of two sets of neighborhoods of size k. Multi-modal alignment. Early efforts to connect images and text utilized human annotations [90]. The curation of large-scale paired image-caption datasets, such as MS-COCO [55] and Visual Genome [49], facilitated the systematic study of cross-modal correspondence and models. The CLIP model [73] by Radford et al. formed a turning point by demonstrating that contrastive learning on web-scale image-text pairs could produce shared embedding spaces. Since, a growing body of work has investigated whether such alignment arises even without explicit joint training. Merullo et al. [62] showed that a simple learned linear transformation could map between frozen vision encoders and LLMs. Moschella et al. [65] use similarities to an anchor set. Maniparambil et al. [59] demonstrated that even unaligned unimodal encoders possess high semantic similarity. Fully unsupervised approaches include blind vision-language matching [77] and unpaired embedding translation via cycle-consistency [42, 98]. Finally, Gupta et al. [31] show that an orthogonal map can map between independently trained multi-modal contrastive models. These results are often seen as evidence for representational convergence. However, they are obtained in restricted settings (e.g. [77] experiments on CIFAR-100 and ImageNet-100) and do not scale to real-world multi-modal data. Our work examines whether alignment survives beyond these constraints, showing that it decreases at scale and reflects coarse categorical agreement rather than shared fine-grained structure. Limits and measurement of emergent cross-modal structure. Several analyses show that alignment between independently trained unimodal encoders depends strongly on data, architecture, and evaluation protocol. Tjandrasuwita et al. [86] find that alignment varies with modality similarity and the balance of shared versus unique information, while Hadgi et al. [32] report weaker alignment for “pure” 3D encoders without careful subspace selection. Zhu et al. [99] further show that video–text alignment depends on temporal richness and text availability. Gröger et al. [29] show that global similarity measures such as CKA are sensitive to network scale and can be altered via null calibration, largely removing evidence of global convergence while leaving local neighborhood similarity (e.g., mutual kNN) more stable, though still evaluated under small-scale and bijective regimes. Beyond similarity metrics, Smith et al. [81] and Kumar et al. [50] show that functional agreement and output behavior can persist even when internal representations are misaligned or entangled, suggesting that behavioral compatibility does not imply shared structure. These caveats echo grounding arguments that text-only learning may be insufficient to recover perceptual structure [7, 51], and motivate multimodal foundation models that integrate perception and language at scale [39, 4, 63, 35].

3

Experimental setup

Mutual kNN metric. To measure alignment between representations from different models, we use the mutual k-nearest-neighbor metric (illustrated in Fig. 1), following Huh et al. [40]. Given a shared gallery set of n datapoints (referred to as mini-batch sampled from the data distribution in [40]) encoded into feature vectors ai ∈ Rd1 and bi ∈ Rd2 by two models with i ∈ {1, · · · , n}, we first L2-normalize each representation. We then retrieve the k nearest neighbors of every query point independently for each model (e.g. image and text query for vision and language encoders): Nka (i) = argtopkj̸=i a⊤ i aj ,

Nkb (i) = argtopkj̸=i b⊤ i bj .

The per-sample score is the number of overlapping samples normalized by k: si =

|Nka (i) ∩ Nkb (i)| , k 4

Query

WIT-1024

DINOv2

LLM

Lighthouse on Schiermonnikoog

Traeth Benllech - the sands at Benllech from the headland above Beach Road

The Weisshorn with Tete de Milon on the farright

The Schubertring section of the Ringstrasse in Vienna

Dutenhofen’s location in Wetzlar

The Carlswerk Smelting Museum at Magdesprung

Query

Chateau de Chalmazel

HildrethLord-Hawley Farmhouse, October 2012

Panoramic view of Liljeholmen

Goose Island brewpub on Clybourn Ave.

Schwarzhausern, canton of Bern

Lighthouse on Eldred Rock island

Lighthouse on Mudou Island in Baisha Township

Old lighthouse Rotes Kliff near Kampen on Sylt

Lighthouse on the Mull of Galloway

WIT-1M

DINOv2

LLM

Lighthouse on Schiermonnikoog

Beach on Schiermonnikoog

Map of Schiermonnikoog NP

Landscape in Schiermonnikoog National Park

North Tower, Schiermonnikoog

Lighthouse on the island Lyran

Satellite image of easternmost point of Schiermonnikoog

Figure 2: Nearest-neighbor quality depends on data density. We show 10 within-modality nearest neighbors for image (DINOv2) and text (LLM) embeddings on a sparse WIT-1024 gallery (top) and a denser WIT-1M gallery (bottom). For text queries, retrieved captions and their corresponding reference images are shown. At smaller scale, nearest neighbors are less semantically precise. Nearestneighbor structure becomes more semantically refined as gallery density increases. and the overall mutual-kNN score is the mean over all samples. A score of 1 means that every point’s k nearest neighbors are identical in both spaces, and a score of 0 means that the k nearest neighbors do not overlap. In the sparse gallery in Fig. 1, the query (blue) retrieves the same neighbor in both image and text spaces, giving a mutual kNN score of 1 for k=1. A score of nk suggests chance-level expected overlap for independent random retrieval, which decreases for growing n. Note that throughout we report raw mutual kNN (as in [40]). Implementation details. For most experiments, we use DINOv2-base [69] as the vision encoder. We refer to this model as DINOv2 in the following. Our primary language model is OpenLlama3b [87, 24] (abbreviated as OpenLlama). Additional models are considered in the supplementary material (Section A.3). For each image and text sample, we extract the representations from all layers of their respective encoders and follow the experimental protocol from [40]. Details about additional models used in Section 4 are provided in Section E.3.2 in the supplementary material. We use Faiss [17] for nearest neighbor computation at scale. Specifically, we use their exact nearest neighbor implementation with IndexFlatL2 which is equivalent to using cosine similarity on normalized vectors.

4

How much do representations align?

In this section, we take a close look at the experimental evidence underpinning the Platonic Representation Hypothesis [40]. The experiments in [40] rest on two foundations that warrant scrutiny: the use of mutual kNN alignment on a small evaluation set of only 1024 samples from the Wikipedia Image-Text (WIT) dataset [82] (WIT-1024), and the use of data with bijective (one-to-one) imagetext correspondences. Typically, these choices are not acknowledged when the hypothesis is cited [60, 76, 57, 9, 13]. The claim is usually invoked in its broad, appealing form rather than in the narrow terms under which experimental support was provided. Here, we analyze how alignment behaves for a finer-grained metric (k=1 instead of k=10), and a denser gallery (million(s of) instead of 1024 samples). We then decompose what mutual kNN alignment actually measures in a controlled setup on ImageNet. This reveals that models individually retrieve correct-class neighbors but rarely agree on which one, suggesting that information is organized differently in each unimodal model. We then turn to the bijective assumption, and examine what 5

(a) Effect of neighborhood sizes k in mutual kNN (both axes log-scaled). Trivially, mutual kNN converges to 1.0 as k approaches the full gallery size. [40] utilized mutual kNN for k = 10.

(b) Alignment between DINOv2 and different LLMs, measured on WIT-1024 and WIT-1M. As observed in [40], alignment (mutual kNN) increases with language performance (measured as 1 − bitsperbyte from [40]) on the 1024-sample set, but this trend breaks with larger gallery size.

Figure 3: Mutual kNN text-image feature alignment when scaling from WIT-1024 to WIT-1M. (a) shows the dependence on neighborhood size k, while (b) examines alignment for different LLMs. The observation from [40], that more capable language models align better with vision largely vanishes at WIT-1M scale. happens when it is relaxed. Finally, we perform a trend check to ask whether the predictions from [40] have held up as models have improved. Sensitivity to k in mutual kNN. Huh et al. [40] reported mutual kNN alignment for k=10. We additionally evaluate at k=1, which requires the two representation spaces to agree on the single nearest neighbor. As shown in Fig. 3a, the metric trivially converges to 1 as k approaches the full gallery size n, since both neighbor sets then contain all samples. Even moderate values of k can inflate scores by capturing broadly similar rather than precisely matching neighbors. In our analyses at larger gallery scales, we perform deduplication to prevent near-duplicate samples from trivially inflating neighborhood overlap (see Section E.1 in the supplementary material for details). 4.1

Alignment across dataset scales

Nearest neighbors in sparse gallery. We now turn to the data, and ask whether the 1024sample gallery used in [40] is too sparse to capture more than coarse structural agreement. As shown in Table 1, the mean cosine similarity between queries and nearest neighbors in terms Table 1: Nearest-neighbor quality across gallery of both image (DINOv2) and text features (Open- sizes. As the gallery grows, nearest neighbors Llama) is significantly lower for the WIT-1024 get closer to the query set in both DINOv2 and gallery compared to WIT-1M (e.g. 0.799 com- OpenLlama embedding spaces, facilitating the pared to 0.906 for DINOv2 at k=1). Note that more fine-grained analysis of cross-modal alignin both cases, we use WIT-1024 as the query set. ment. We visualize nearest neighbors for k=10 for imGallery Model k=1 age and text features on WIT-1024 and WIT-1M WIT-1024 DINOv2 0.799 in Fig. 2. At low density, semantically unrelated WIT-1024 OpenLlama 0.502 samples may end up as nearest neighbors as there WIT-1M DINOv2 0.906 is nothing closer available, meaning that measured WIT-1M OpenLlama 0.757 mutual kNN agreement mainly can reflect the shared lack of alternatives. To get more meaningful insights, we scale the density of the retrieval gallery in the following section.

k=10 0.717 0.400 0.888 0.701

Densification by scaling the gallery size. Having established that the WIT-1024 gallery captures mainly coarse structure, we densify the gallery and test whether alignment persists. We evaluate on up to 1M and 15M gallery samples from the English-text WIT [82] and LAION400M [78] respectively. The best layer pair was determined on the 1024-sample subset of WIT, following [40]. As shown in Fig. 4, alignment scores decrease as gallery size grows for fixed k and query set (WIT-1024). 6

Figure 4: Scaling the gallery size to 1M (WIT) and 15M (LAION) shows a large drop in mutual kNN alignment for k=1 and k=10 for DINOv2 and OpenLlama features. Query

WIT-1024

10K

50K

100K

500K

1M

2004 Volvo S80 S Automatic 2.4 Rear Taken in Warwick

1995 Holden Commodore (VS) Executive station wagon. . .

2003 Kia Carens LX 1.8 Front

2017 Volvo S90 Momentum D4 Automatic 2.0 Rear Taken in Leamington Spa

2005 Volkswagen Touareg V6 Sport Automatic 3.2 Front Taken in Leamington Spa

2004 Volvo S80 S Automatic 2.4 Front Taken in Warwick

2004 Volvo S80 S Automatic 2.4 Front Taken in Warwick

Ondina vitrea (Brusina, 1866); Pyramidellidae

Guraleus fascinus Hedley, 1922; family Mangeliidae

Guraleus fascinus Hedley, 1922; family Mangeliidae

Prothalotia ramburi (Crosse, H., 1864)

Prothalotia ramburi (Crosse, H., 1864)

Talostolida teres (Gmelin, 1791); family Cypraeidae

Pleuroptya violacealis (Syllepte violacealis). . .

Map of the urban area of Novi Sad. . . showing Vidovdansko Naselje

The annexed western quarter of Slovene ethnic territory. . .

Map of municipal area of Zrenjanin - village of Banatski Despotovac. . .

Map of the Baki Petrovac municipality, showing the location of Magli

Map of the Baki Petrovac municipality, showing the location of Magli

Map of the Baki Petrovac municipality, showing the location of Magli

Map of the urban area of Novi Sad. . . showing Bistrica (Novo Naselje)

DINOv2

LLM

DINOv2

LLM

DINOv2

LLM

Figure 5: Nearest-neighbor (k=1) examples with DINOv2 and OpenLlama across gallery scales on WIT-1M. Captions are shown with corresponding images. Mutual kNN matches across modalities are framed green. While the bottom example shows a match at 1M scale, at larger scales each model finds closer but different matches (top three). The mutual kNN alignment scores drop from 0.135 and 0.058 on the 1024-sample gallery to 0.008 and 0.001 on LAION-15M for k = 10 and k = 1 respectively. This confirms that the agreement observed at small scale declines with the transition to finer-grained evaluation at large scale. Nearest neighbors become closer and more semantically similar to the query, placing greater demand on the two representation spaces to agree on subtle distinctions. Interestingly, alignment at k=n/100 remains relatively stable across scales, suggesting that models share some degree of coarse structural agreement. We hypothesize that this amounts to precisely the kind of broad categorical correspondence one would expect from models trained on overlapping internet data, and lacks signal about whether representations are organized in the same way. We also analyze how mutual kNN alignment for various LLMs and DINOv2 behaves at the WIT-1M scale. Reproducing the setting of [40], Fig. 3b shows a clear trend on WIT-1024: stronger language models exhibit higher alignment with visual features. This is a central finding of [40] and one of their most compelling pieces of evidence. However, when we scale the gallery to 1M samples, this trend 7

Query

WIT-1024

10K

100K

1M

10M

15M

A parked Air China Boeing 737-800 at the gate

First Air Boeing 767 at Val-d’Or Airport

N39297 - Boeing 737-824 - United Airlines

TF-NAC - National Airlines Boeing 747-400BCF, SF, BDSF. . .

Screenshot of Air Canada Boeing 747-400 on the ground.

Air China Boeing 737-800 Economy Class BeijingHarbin, China

Air China Boeing 737-800 Economy Class BeijingHarbin, China

The Battle of Nicopolis, as depicted by Turkish miniaturist in 1588

The Martyrdom of Saint Agathius. 16th century work. . .

The Martyrdom of Saint Agathius. 16th century work. . .

Detail of a Venetian Warship from the Mausoleum of Girolamo Michiel. . .

Hippodrome of Constantinople – Procession of the guilds. . .

The Battle of Nicopolis, as depicted by Turkish miniaturist in 1588. . .

The Battle of Nicopolis, as depicted by Turkish miniaturist in 1588. . .

DINOv2

LLM

DINOv2

LLM

Figure 6: Nearest-neighbor (k=1) examples with DINOv2 and OpenLlama across gallery scales on LAION-15M. As the gallery densifies, each model finds closer but different matches (top example). The match at 15M (bottom right) is a near-duplicate that survived our deduplication pipeline. largely vanishes. The gap between LLMs narrows considerably, and the relationship between model capability and alignment weakens. This suggests that the observation in [40] may be a result of the sparse evaluation setting. The nearest-neighbor examples in Figs. 5 and 6 further illustrate this effect. Matches that appear semantically meaningful at small gallery sizes often break down as more candidates are introduced. In vision space, we find better neighbors that deviate from the best text neighbors at larger scale. There are only very few matches found at 1M and 15M data scale, most of which are near-duplicates that our deduplication pipeline did not catch (e.g. a crop shifted by a few pixels). We show additional visualizations in Figs. 21 to 24 in the supplementary material. We additionally test whether the alignment drop with increasing gallery size is merely an artifact of the mutual kNN metric being harder at scale. Specifically, we measure within-modality alignment for two pairs of models: two language models of different scale (OpenLlama-3b and OpenLlama-13b), and, separately, two vision models (DINOv2-base and DINOv2-giant). If mutual kNN alignment collapses for dense galleries regardless of the models being compared, the cross-modal drop observed would be uninformative. If within-modality alignment remains stable, the cross-modal drop is meaningful. As shown in the supplementary material (Fig. 12), unimodal alignment remains much more stable across gallery sizes than the cross-modal alignment reported in the main paper. For the OpenLlama pair, mutual kNN at k=1 stays between [0.59, 0.62], and for the DINOv2 pair between [0.35, 0.45], across all gallery scales. This confirms that mutual kNN does not inherently collapse at scale. Observation 1: Mutual kNN alignment scores decrease for denser galleries for fixed small k, suggesting that mutual kNN is sensitive to gallery sparsity. Furthermore, the reported trend that stronger language models align better with vision weakens substantially at WIT-1M scale, with all models scoring near zero. 4.2

What is captured by cross-modal mutual kNN alignment?

Low mutual kNN alignment at fixed small k could mean two things: the models individually retrieve n poor neighbors, or they each retrieve good neighbors but different ones. The stability at k = 100 hints at the models agreeing at a coarse level but diverging on fine-grained structure. To test this directly, we use the ImageNet [15] validation set, where class labels let us evaluate each model’s retrieval independently. We decompose each query into: (i) whether each model individually retrieves a correct-class neighbor, (ii) whether both do, and (iii) whether they agree on the exact same gallery item (mutual kNN with 8

(a) The query image (left) is matched with galleries of increasing density. As the gallery becomes more dense, DINOv2 and OpenLlama retrieve from the same class, but different instances, illustrating how within-class structure is organized differently across modalities.

(b) Per-modality retrieval accuracy and crossmodal mutual kNN alignment (k=1) as images / captions per class in gallery increase. Modalities individually improve with gallery density, but alignment does not.

Figure 8: Decomposing cross-modal alignment on ImageNet val. (a) shows a qualitative retrieval example where both models find plausible neighbors but disagree on the specific instance. (b) quantifies this: individual class-level retrieval accuracy improves with gallery density, yet strict alignment remains flat, illustrating that models organize within-class structure differently.

k=1). Our query set consists of one image per class (1000 images), and we vary the number of images per class (ipc) in the gallery from 1 to 49. We use detailed image captions (981 words on average) generated by gemini-3-flash-preview [85, 70], making this a favorable setting for alignment (details are provided in Section E.2 in the supplementary material). Fig. 8b reveals that as the gallery densifies, in line with Cover and Hart [12], both models individually improve at retrieving correct-class neighbors. This indicates some degree of shared coarse structure. At larger scale, both models retrieve reasonable neighbors but different ones. We see similar trends when looking at coarser evaluation with k=10 (see Section B.3 in the supplementary material). At 49 images per class in the gallery, DINOv2 succeeds 46.1% of the time and OpenLlama 58.0%. Yet strict alignment on the exact same gallery item remains flat around 11%, even with detailed captions. For reference, alignment with class-name-only captions drops from 0.42 to near zero as ipc increases. The models are individually capable but organize within-class structure differently (Fig. 8a). At ipc=1, strict alignment (23.1%) actually exceeds the rate at which both models retrieve a correct-class neighbor (11.7%), meaning the models often agree on semantically plausible but technically incorrect neighbors (Fig. 7). This reveals what mutual kNN actually captures. It does not measure unimodal representation quality, but agreement on fine-grained structure. Our experiments provide direct evidence that low cross-modal alignment in terms Figure 7: Shared mistake at ipc=1. The query imof mutual kNN is not due to poor representaage (bookstore) is matched by both DINOv2 and tions but rather due to fundamentally different OpenLlama to a library image. The models agree, representational organization within modalities. but on the wrong answer. Both models learn structured, high-quality representations. They simply do not structure them the same way.

Observation 2: As the data gets denser, both models retrieve correct-class neighbors at increasing rates, yet strict cross-modal mutual kNN alignment is flat at 11%. This is not a failure of unimodal representation quality, but of unaligned representational organization across modalities.

9

Figure 9: Effect of relaxing the bijective assumption on text-image alignment, using the CycleReward dataset [3]. We densify one modality by adding more images per caption (left) or more captions per image (right) while keeping the other fixed. Mutual kNN alignment decreases consistently for both k=1 and k=10. 4.3

What happens when the data is not bijective?

In practice, the relationship between modalities, such as image and text, is inherently many-to-many: a single image can be described by countless text descriptions, and a single text caption can correspond to a large set of visually distinct images. More fundamentally, modalities often differ in information content. Specifically, images encode spatial, textural, and perceptual structure that text captures only to a limited extent. On the other hand, text encodes abstraction, negation, and compositional semantics that images do not. One could, in principle, bridge this gap trivially. For instance, one could encode pixel values as text or render captions as images and establish a bijection between those. Those preserve the information, but the inductive structure (the modality-specific properties that make each Figure 10: Illustration of non-bijective (manymodality useful) of each modality is lost. to-many) correspondence between image and To test what happens when bijectivity is relaxed, captions. The nearest neighbor of a text caption we use the CycleReward dataset [3] which pairs for one image (blue) is a caption for a different each real sample with multiple synthetic candi- image (red). However, the nearest image neighdates. The I2T subset contains 11 generated cap- bor for a given image may be another image with tions per real image, and T2I consists of 12 syn- the same caption. thetic images for each text prompt. This directly breaks the bijection, i.e. one-to-one matching, that our earlier analysis assumes. We evaluate mutual kNN by densifying one modality at a time: for T2I we keep the text fixed and increase the number of generated images per prompt, and for I2T we keep the image fixed and add generated captions. For illustration, let us consider the T2I experiments where the closest neighbor in the densified modality is more likely to be a similar image that is associated with the same caption, while in the sparse text space the NN will be a caption for a different image. This creates a scenario where mutual kNN fails (see Fig. 10). We adapt the mutual kNN metric where a match is counted when the retrieved item corresponds to the same source sample, even if it is not the exact same caption or image. In Fig. 9, we see that the mutual kNN scores decrease as the one-to-one assumption is relaxed. Whether this is due to reduced alignment or just a limitation of the metric is an open question. Regardless, the original evidence for convergence depends on an assumption that real-world multi-modal data rarely satisfies. Observation 3: A single image can be described in countless ways, and a single caption can match many visually distinct images. When we progressively relax the one-to-one assumption, mutual kNN alignment drops consistently. The mutual kNN metric cannot distinguish between genuine misalignment and many-to-many correspondence.

10

Figure 11: Testing whether the alignment-LLM performance trend from [40], tested on WIT-1024, holds for recent LLMs on ARC, GSM8K, MMLU, and LogiQA2. Dashed lines show the trend for models from [40] (circles). Recent LLMs (diamonds) do not follow the trend: stronger language models do not seem to be more aligned with DINOv2. 4.4

Trend check: Are the predictions from [40] holding up so far?

Huh et al. [40] predict that as LLMs become stronger, their representations align more with vision representations. This claim is evaluated using three proxies for language performance: HellaSwag [97], GSM8K [11], and (1 − bitsperbyte). In this section, we revisit this trend analysis with an extended set of models and benchmarks. We evaluate 55 LLMs, spanning from BLOOMZ [66] to recently released models. In addition to the evaluations in [40], we use the ARC Challenge [10], MMLU [34], and LogiQA2 [58] benchmarks. These probe arithmetic reasoning, general knowledge, and logical reasoning. We present results for models that surpass Llama-3-70B [27] (the strongest model in [40]) on at least one benchmark, testing whether the alignment-performance trend continues. The full set of 55 models and results on the largely saturated benchmarks are included in Section D in the supplementary material. In Fig. 11, we observe that recent models do not continue on the scaling lines extrapolated from the model set from [40]. Instead, their points do not follow the predicted trend and hint at saturation with respect to DINOv2 features. Furthermore, the R2 averaged across all regression lines ranges from -8.6 to -3.7, confirming that the extrapolated trend is not continued by recent models.

5

Discussion and future work

The Platonic Representation Hypothesis suggests that models trained on different modalities converge toward a shared representation of reality as they scale. Our results indicate a more conditional interpretation of this claim. Mutual kNN agreement is highly sensitive to the evaluation regime: it drops sharply when moving from small galleries to million-scale datasets and degrades further under many-to-many cross-modal correspondences. Moreover, the previously reported trend that stronger language models yield higher alignment does not consistently hold for recent models. These findings do not rule out shared structure across modalities. Overall, our results indicate that small-scale mutual nearest-neighbor evaluations may overstate the degree of convergence by relying on restrictive gallery sizes and one-to-one pairing. Low mutual kNN agreement should not be conflated with weak representations. As our analysis on ImageNet shows, it may instead reflect differences in how fine-grained structure is organized. Rather than converging on a single Platonic representation of reality, modalities appear to inhabit in their own Umwelten, a distinct but coherent representational “cave” where alignment between them is local and partial. 11

Future work: in search of bijection. Prior work on the Platonic Representation Hypothesis [40] evaluates alignment under a one-to-one correspondence assumption between modalities. We have shown that mutual kNN agreement is not reliable once this assumption is relaxed. Likewise, interpretations of the Platonic Representation Hypothesis as evidence that “language is all you need” [92] often rely on evaluation regimes that effectively assume bijective structure between images and text. However, real-world image–text data is fundamentally many-to-many, and the extent to which any approximate bijection exists at the level of representations remains unclear. A key direction for future work is to directly test this assumption, for example by studying whether language can serve as a lossless bottleneck for image reconstruction (i.e. an image-text-image autoencoder). If, as we suspect, this proves illusive for realistic settings (e.g. text bottlenecks of under a thousand words), it would be very interesting to identify and model the part of the joint text-image space forming a bijection (the intersection of the Venn diagram), and disentangle it from the parts that do not. Acknowledgments. This work was in part supported by the BMFTR (FKZ: 16IS24060), the DFG (SFB 1233, project number: 276693517), NSF IIS-2403305, and ONR MURI. This research utilized compute resources at the Tübingen Machine Learning Cloud. The authors thank all Efros group members for valuable discussions that shaped this work, and particularly Tyler Bonnen and Amil Dravid for proofreading the draft. Lastly, we thank Phillip Isola for the thought-provoking hypothesis, and for discussing and engaging openly with our disagreements – a rare kind of intellectual generosity.

References [1] M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, et al. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691, 2022. [2] E. Akyürek, M. Damani, A. Zweiger, L. Qiu, H. Guo, J. Pari, Y. Kim, and J. Andreas. The surprising effectiveness of test-time training for few-shot learning. arXiv preprint arXiv:2411.07279, 2024. [3] H. Bahng, C. Chan, F. Durand, and P. Isola. Cycle consistency as reward: Learning image-text alignment without human preferences. arXiv preprint arXiv:2506.02095, 2025. [4] S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, et al. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631, 2025. [5] R. Balestriero et al. A spline theory of deep learning. In ICML, 2018. [6] Y. Bansal, P. Nakkiran, and B. Barak. Revisiting model stitching to compare neural representations. In NeurIPS, 2021. [7] E. M. Bender and A. Koller. Climbing towards NLU: On meaning, form, and understanding in the age of data. In D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020. [8] J. Browning and Y. LeCun. Ai and the limits of language. Noema Magazine, 2022. [9] W. Chai, E. Song, Y. Du, C. Meng, V. Madhavan, O. Bar-Tal, J.-N. Hwang, S. Xie, and C. D. Manning. Auroracap: Efficient, performant video detailed captioning and a new benchmark. In ICLR, 2025. [10] F. Chollet. On the measure of intelligence. arXiv preprint arXiv:1911.01547, 2019. [11] K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021. [12] T. Cover and P. Hart. Nearest neighbor pattern classification. IEEE transactions on information theory, 1967. [13] G. Dar. mini-vec2vec: Scaling universal geometry alignment with linear transformations. arXiv preprint arXiv:2510.02348, 2025. 12

[14] DeepSeek-AI, D. Guo, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025. [15] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009. [16] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021. [17] M. Douze, A. Guzhva, C. Deng, J. Johnson, G. Szilvasy, P.-E. Mazaré, M. Lomeli, L. Hosseini, and H. Jégou. The faiss library. arXiv preprint arXiv:2401.08281, 2024. [18] A. Dravid, Y. Gandelsman, A. A. Efros, and A. Shocher. Rosetta neurons: Mining the common units in a model zoo. In ICCV, 2023. [19] S. Edelman. Representation is representation of similarities. Behavioral and brain sciences, 1998. [20] L. Gao, J. Tow, B. Abbasi, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, A. Le Noac’h, H. Li, K. McDonell, N. Muennighoff, C. Ociepa, J. Phang, L. Reynolds, H. Schoelkopf, A. Skowron, L. Sutawika, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou. The language model evaluation harness. Zenodo, 07 2024. doi: 10.5281/zenodo.12608602. URL https://zenodo.org/records/12608602. [21] G. D. Gemma Team. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295, 2024. [22] G. D. Gemma Team. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118, 2024. [23] G. D. Gemma Team. Gemma 3 technical report. arXiv preprint arXiv:2503.19786, 2025. [24] X. Geng and H. Liu. Openllama: An open reproduction of llama, 2023. URL https://github. com/openlm-research/open_llama. [25] J. J. Gibson. The Ecological Approach to Visual Perception. Houghton Mifflin, Boston, 1979. ISBN 978-0898593019. [26] A. Gokaslan and V. Cohen. OpenWebTextCorpus, 2019.

Openwebtext corpus.

http://Skylion007.github.io/

[27] A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024. [28] D. Groeneveld, I. Beltagy, P. Walsh, A. Bhagia, R. Kinney, O. Tafjord, A. H. Jha, H. Ivison, I. Magnusson, Y. Wang, S. Arora, D. Atkinson, R. Authur, K. R. Chandu, A. Cohan, J. Dumas, Y. Elazar, Y. Gu, J. Hessel, T. Khot, W. Merrill, J. Morrison, N. Muennighoff, A. Naik, C. Nam, M. E. Peters, V. Pyatkin, A. Ravichander, D. Schwenk, S. Shah, W. Smith, E. Strubell, N. Subramani, M. Wortsman, P. Dasigi, N. Lambert, K. Richardson, L. Zettlemoyer, J. Dodge, K. Lo, L. Soldaini, N. A. Smith, and H. Hajishirzi. Olmo: Accelerating the science of language models. In ACL, 2024. [29] F. Gröger, S. Wen, and M. Brbić. Revisiting the platonic representation hypothesis: An aristotelian view. arXiv preprint arXiv:2602.14486, 2026. [30] S. Gu, C. Clark, and A. Kembhavi. I can’t believe there’s no images! learning visual tasks using only language supervision. In ICCV, 2023. [31] S. Gupta, S. Kansal, S. Jegelka, P. Isola, and V. Garg. Canonicalizing multimodal contrastive representation learning. In ICLR, 2026. 13

[32] S. Hadgi, L. Moschella, A. Santilli, D. Gomez, Q. Huang, E. Rodolà, S. Melzi, and M. Ovsjanikov. Escaping plato’s cave: Towards the alignment of 3d and text latent spaces. In CVPR, 2025. [33] J. V. Haxby, M. I. Gobbini, M. L. Furey, A. Ishai, J. L. Schouten, and P. Pietrini. Distributed and overlapping representations of faces and objects in ventral temporal cortex. Science, 2001. [34] D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding. In ICLR, 2021. [35] W. Hong, W. Yu, X. Gu, G. Wang, G. Gan, H. Tang, J. Cheng, J. Qi, J. Ji, L. Pan, et al. Glm-4.5 v and glm-4.1 v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning. arXiv preprint arXiv:2507.01006, 2025. [36] H. Hotelling. Relations between two sets of variates. In Breakthroughs in statistics: methodology and distribution. 1992. [37] X. Hu, S. Storks, R. L. Lewis, and J. Chai. In-context analogical reasoning with pre-trained language models. In ACL, 2023. [38] Y. Hu, H. Hua, Z. Yang, W. Shi, N. A. Smith, and J. Luo. Promptcap: Prompt-guided image captioning for vqa with gpt-3. In ICCV, 2023. [39] S. Huang, L. Dong, W. Wang, Y. Hao, S. Singhal, S. Ma, T. Lv, L. Cui, O. K. Mohammed, B. Patra, Q. Liu, K. Aggarwal, Z. Chi, J. Bjorck, V. Chaudhary, S. Som, X. Song, and F. Wei. Language is not all you need: Aligning perception with language models. In NeurIPS, 2023. [40] M. Huh, B. Cheung, T. Wang, and P. Isola. The platonic representation hypothesis. In ICML, 2024. [41] P. Isola. Personal communication, 2025. [42] R. Jha, C. Zhang, V. Shmatikov, and J. X. Morris. Harnessing the universal geometry of embeddings. In NeurIPS, 2025. [43] A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. Renard Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. El Sayed. Mistral 7B. arXiv preprint arXiv:2310.06825, 2023. [44] A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand, et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024. [45] J. Jiang, J. Zhou, and Z. Zhu. Tracing representation progression: Analyzing and enhancing layer-wise similarity. arXiv preprint arXiv:2406.14479, 2024. [46] J. J. Koenderink. Sentience. De Clootcrans Press, Trajectum, Netherlands, 2019. [47] S. Kornblith, M. Norouzi, H. Lee, and G. Hinton. Similarity of neural network representations revisited. In ICML, 2019. [48] N. Kriegeskorte, M. Mur, and P. A. Bandettini. Representational similarity analysis-connecting the branches of systems neuroscience. Frontiers in systems neuroscience, 2008. [49] R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. IJCV, 2017. [50] A. Kumar, J. Clune, J. Lehman, and K. O. Stanley. Questioning representational optimism in deep learning: The fractured entangled representation hypothesis. arXiv preprint arXiv:2505.11581, 2025. [51] Y. LeCun et al. A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27. Openreview, 2022. 14

[52] K. Lenc and A. Vedaldi. Understanding image representations by measuring their equivariance and equivalence. In CVPR, 2015. [53] Y. Li, J. Yosinski, J. Clune, H. Lipson, and J. Hopcroft. Convergent learning: Do different neural networks learn the same representations? In ICLR, 2016. [54] J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng. Code as policies: Language model programs for embodied control. In ICRA, 2023. [55] T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. [56] A. H. Liu, S. Subramanian, V. Jouault, A. Sadé, et al. arXiv:2601.08584, 2026.

Ministral 3.

arXiv preprint

[57] D. Liu, S. Zhao, L. Zhuo, W. Lin, Y. Xin, X. Li, Q. Qin, Y. Qiao, H. Li, and P. Gao. Luminamgpt: Illuminate flexible photorealistic text-to-image generation with multimodal generative pretraining. arXiv preprint arXiv:2408.02657, 2025. [58] H. Liu, J. Liu, L. Cui, Z. Teng, N. Duan, M. Zhou, and Y. Zhang. Logiqa2.0: The logicqa dataset for logical reasoning. IEEE Transactions on Audio, Speech, and Language Processing, 2023. [59] M. Maniparambil, R. Akshulakov, Y. A. D. Djilali, S. Narayan, M. E. A. Seddik, K. Mangalam, and N. E. O’Connor. Do vision and language encoders represent the world similarly? In CVPR, 2024. [60] P. Marcos-Manchón and L. Fuentemilla. Shared representations in brains and models reveal a two-route cortical organization during scene perception. arXiv preprint arXiv:2507.13941, 2026. [61] S. Merity, C. Xiong, J. Bradbury, and R. Socher. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843, 2016. [62] J. Merullo, L. Castricato, C. Eickhoff, and E. Pavlick. Linearly mapping from image to text space. In ICLR, 2023. [63] Meta AI. The llama 4 herd: The beginning of a new era of natively multimodal ai innovation, 2025. URL https://ai.meta.com/blog/llama-4-multimodal-intelligence/. [64] A. S. Morcos, M. Raghu, and S. Bengio. Insights on representational similarity in neural networks with canonical correlation. In NeurIPS, 2018. [65] L. Moschella, V. Maiorca, M. Fumero, A. Norelli, F. Locatello, and E. Rodolà. Relative representations enable zero-shot latent space communication. In ICLR, 2023. [66] N. Muennighoff, T. Wang, L. Sutawika, A. Roberts, S. Biderman, T. L. Scao, M. S. Bari, S. Shen, Z.-X. Yong, H. Schoelkopf, X. Tang, D. Radev, A. F. Aji, K. Almubarak, S. Albanie, Z. Alyafeai, A. Webson, E. Raff, and C. Raffel. Crosslingual generalization through multitask finetuning. In ACL, 2023. [67] T. OLMo, P. Walsh, L. Soldaini, D. Groeneveld, K. Lo, S. Arora, A. Bhagia, Y. Gu, S. Huang, M. Jordan, N. Lambert, D. Schwenk, O. Tafjord, T. Anderson, D. Atkinson, F. Brahman, C. Clark, P. Dasigi, N. Dziri, A. Ettinger, M. Guerquin, D. Heineman, H. Ivison, P. W. Koh, J. Liu, S. Malik, W. Merrill, L. J. V. Miranda, J. Morrison, T. Murray, C. Nam, J. Poznanski, V. Pyatkin, A. Rangapur, M. Schmitz, S. Skjonsberg, D. Wadden, C. Wilhelm, M. Wilson, L. Zettlemoyer, A. Farhadi, N. A. Smith, and H. Hajishirzi. 2 olmo 2 furious. arXiv preprint arXiv:2501.00656, 2025. [68] OpenAI. Introducing introducing-gpt-oss/.

gpt-oss,

2025.

URL

https://openai.com/index/

[69] M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y. Huang, S.-W. Li, I. Misra, M. Rabbat, V. Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski. Dinov2: Learning robust visual features without supervision. TMLR, 2024. 15

[70] S. Pichai, D. Hassabis, and K. Kavukcuoglu. A new era of intelligence with Gemini 3. Google Blog (The Keyword), Nov. 2025. URL https://blog.google/ products-and-platforms/products/gemini/gemini-3/. Accessed: 2026-01-01. [71] Plato. Republic. c. 375 BC. [72] A. C. Qwen Team. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. [73] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning transferable visual models from natural language supervision. In ICML, 2021. [74] I. Research. Granite 3.3 8b base, 2025. URL https://huggingface.co/ibm-granite/ granite-3.3-8b-base. [75] E. Rosch. Principles of categorization. In E. Rosch and B. B. Lloyd, editors, Cognition and Categorization, pages 27–48. Lawrence Elbaum Associates, 1978. [76] J. Ruan, A. Abudula, X. Liu, B. Li, Y. Li, C. Wang, Y. Fan, Y. Ge, T. Xiao, and J. Zhu. Ndp: Next distribution prediction as a more broad target. arXiv preprint arXiv:2408.17377, 2024. [77] D. Schnaus, N. Araslanov, and D. Cremers. It’s a (blind) match! towards vision-language correspondence without parallel data. In CVPR, 2025. [78] C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021. [79] A. Singh, R. Hu, V. Goswami, G. Couairon, W. Galuba, M. Rohrbach, and D. Kiela. Public multimodal dataset (PMD). URL https://huggingface.co/datasets/facebook/pmd. [80] A. Singh, R. Hu, V. Goswami, G. Couairon, W. Galuba, M. Rohrbach, and D. Kiela. Flava: A foundational language and vision alignment model. In CVPR, 2022. [81] D. Smith, H. Mannering, and A. Marcu. Functional alignment can mislead: Examining model stitching. In ICML, 2025. [82] K. Srinivasan, K. Raman, J. Chen, M. Bendersky, and M. Najork. Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning. In ACM SIGIR conference on research and development in information retrieval, 2021. [83] I. Sutskever. “the mastermind behind gpt-4 and the future of ai” — eye on a.i. (podcast, season 2 episode 118). https://podcasts.apple.com/us/podcast/ ilya-sutskever-the-mastermind-behind-gpt-4-and/id1438378439?i= 1000604382855, mar 2023. Accessed: 2026-02-28. [84] F.-L. Team. The falcon 3 family of open models, 2024. URL https://huggingface.co/ blog/falcon3. [85] G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023. [86] M. Tjandrasuwita, C. Ekbote, L. Ziyin, and P. P. Liang. Understanding the emergence of multimodal representation alignment. In ICML, 2025. [87] H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. [88] H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023. 16

[89] J. B. Uexküll and G. Kriszat. Streifzuge durch die Umwelten von Tieren und Menschen Ein Bilderbuch unsichtbarer Welten. Springer, 1934. [90] L. Von Ahn, S. Ginosar, M. Kedia, and M. Blum. Improving image search with phetch. In ICASSP, 2007. [91] R. Wang, E. Zelikman, G. Poesia, Y. Pu, N. Haber, and N. D. Goodman. Hypothesis search: Inductive reasoning with language models. arXiv preprint arXiv:2309.05660, 2023. [92] S. L. Wang, P. Isola, and B. Cheung. Words that make language models perceive. arXiv preprint arXiv:2510.02425, 2025. [93] R. Wightman. Pytorch image models, 2019. URL https://github.com/huggingface/ pytorch-image-models. [94] L. Wittgenstein. Philosophical Investigations. Wiley-Blackwell, 1953. [95] A. Young, B. Chen, C. Li, C. Huang, G. Zhang, G. Zhang, G. Wang, H. Li, J. Zhu, J. Chen, et al. Yi: Open foundation models by 01.ai. arXiv preprint arXiv:2403.04652, 2024. [96] C. Zauner. Implementation and benchmarking of perceptual image hash functions. 2010. [97] R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi. Hellaswag: Can a machine really finish your sentence? In ACL, 2019. [98] J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In ICCV, 2017. [99] T. Zhu, T. Han, L. Guibas, V. Pătrăucean, and M. Ovsjanikov. Dynamic reflections: Probing video representations with text alignment. In ICLR, 2026.

17

Supplementary Material Back into Plato’s Cave: Examining Cross-modal Representational Convergence at Scale A

Is the drop in mutual kNN alignment at scale caused by the metric, caption quality, or by specific model choices?

In this section, we provide additional experiments to verify that the main findings in the paper are not due to a confounding variable. A.1

Sanity check: does mutual kNN inherently drop at scale even within modalities?

We test whether the alignment drop with increasing gallery size (Fig. 4) is merely an artifact of the mutual kNN metric being harder at scale. Specifically, we measure within-modality alignment for two pairs of models: two language models of different scale (OpenLlama-3b and OpenLlama-13b), and, separately, two vision models (DINOv2base and DINOv2-giant). If mutual kNN alignment collapses for dense galleries regardless of the models being compared, the cross-modal drop observed in the paper would be uninformative. If within-modality alignment remains stable, the cross-modal drop is meaningful.

(a) OpenLlama-3b and OpenLlama-13b

(b) DINOv2-base and DINOv2-giant

Figure 12: Unimodal mutual kNN alignment as a function of gallery size on WIT-1M. In contrast to cross-modal alignment (Fig. 4), unimodal alignment remains significantly more stable across scales. As shown in Fig. 12, unimodal alignment remains much more stable across gallery sizes than the cross-modal alignment reported in the main paper. For the OpenLlama pair, mutual kNN at k=1 stays between [0.59, 0.62], and for the DINOv2 pair between [0.35, 0.45], across all gallery scales. This confirms that mutual kNN does not inherently collapse at scale, and that the degradation observed for cross-modal pairs reflects an actual property of the representation spaces. A.2

WIT-1M-recap: Is the alignment drop caused by poor captions?

One might hypothesize that the alignment drop at scale is driven by low quality of the WIT captions rather than by a fundamental cross-modal difference. To test this, we recaption WIT-1M using gemini3-flash-preview as described in Section E.2. The resulting WIT-1M-recap dataset contains visually detailed descriptions of around 500 words per image. As shown in Fig. 13, mutual kNN alignment still drops with gallery size. More detailed captions give overall higher mutual kNN scores, but do not prevent the decline in scores. This suggests that caption quality is not the primary driver of the mutual kNN alignment drop. A.3

Mutual cross-modal kNN alignment drops at scale across model pairs

The cross-modal alignment drop reported in Fig. 4 in the paper uses DINOv2-base and OpenLlama-3b. Here, we examine whether similar patterns hold for stronger models. In Fig. 14, we repeat the scaling

Figure 13: Cross-modal mutual kNN alignment on images recaptioned using gemini-3-flash-preview (WIT-1M-recap) as the gallery grows to 1M samples. Detailed captions result in overall higher mutual kNN scores, but do not prevent the drop in scores.

(a) DINOv2-base and OpenLlama-13b

(b) DINOv2-giant and OpenLlama-13b

Figure 14: Cross-modal mutual kNN alignment as gallery grows from WIT-1024 to WIT-1M for additional, stronger model pairs. Replacing DINOv2-base with the stronger DINOv2-giant and OpenLlama-3b (Fig. 4 in the paper) with OpenLlama-13b does not prevent the drop. This suggests that the degradation in mutual kNN alignment was not a result of the limitation of any individual model.

experiment for two additional model pairs: DINOv2-base with OpenLlama-13b, and DINOv2-giant with OpenLlama-13b. Replacing DINOv2-base with the stronger DINOv2-giant and OpenLlama-3b with OpenLlama-13b does not change the pattern. We observe that mutual kNN still drops at scale. This is consistent with Fig. 3b in the paper, which already shows low alignment scores across different LLMs at WIT-1M scale. This confirms that the degradation reported in the paper is not specific to that particular choice of models.

B

Additional ImageNet experiments

The controlled experimental setting on the ImageNet validation set in Sec. 4.2 of the paper provides one of our key findings: models individually retrieve correct-class neighbors at increasing rates as the gallery densifies, yet cross-modal agreement remains flat. Here, we verify that the ImageNet validation set serves as a suitable test bed. Furthermore, we confirm that our observations are not limited to our choice of models or metric settings. B.1

The ImageNet validation set is denser than WIT-1024

A natural question is how the gallery density in our ImageNet experiments compares to the WIT data. As shown in Table 2, nearest-neighbor cosine similarities on the ImageNet validation set are substantially higher than on WIT-1024 and comparable to WIT-1M. Even for only one image per class in the gallery (ipc=1), the ImageNet validation set provides a denser retrieval setting than WIT-1024. 19

Table 2: Nearest-neighbor distances across gallery sizes. ImageNet, even with only one image per class in the gallery (ipc=1), has neighbor distances comparable to WIT-1M, confirming that it operates in a similarly dense retrieval regime. Gallery

Model

Dim

k=1

k=10

WIT-1024 WIT-1024

DINOv2-base OpenLlama3b

768 3200

0.799 0.502

0.717 0.400

WIT-1M WIT-1M

DINOv2-base OpenLlama3b

768 3200

0.906 0.757

0.888 0.701

ImageNet ipc=1 ImageNet ipc=1 ImageNet ipc=1 ImageNet ipc=1

DINOv2-base DINOv2-giant OpenLlama3b LLaMA-65B

768 1536 3200 8192

0.823 0.609 0.928 0.890

0.763 0.496 0.904 0.858

ImageNet ipc=49 ImageNet ipc=49 ImageNet ipc=49 ImageNet ipc=49

DINOv2-base DINOv2-giant OpenLlama3b LLaMA-65B

768 1536 3200 8192

0.887 0.749 0.954 0.926

0.861 0.690 0.944 0.912

This confirms that the ImageNet experiments operate in a denser retrieval regime comparable to WIT-1M, making this a meaningful test bed. B.2

Stronger models do not close the gap for ImageNet

The ImageNet decomposition experiments in Sec. 4.2 in the main paper use DINOv2-base and OpenLlama-3b as its vision and language model respectively. Here, we probe whether stronger models would show better cross-modal agreement that closes the gap to unimodal retrieval accuracy. In Fig. 15, we repeat the experiment with DINOv2-base paired with OpenLlama-65b, and DINOv2giant paired with OpenLlama-65b. The pattern is unchanged: both models individually improve at retrieving correct-class neighbors as the gallery densifies, but strict cross-modal alignment remains flat. Using substantially stronger models on both sides does not close the gap between individual retrieval accuracy and cross-modal agreement.

(a) DINOv2-base and Llama-65b

(b) DINOv2-giant and Llama-65b

Figure 15: ImageNet per-modality retrieval accuracy and cross-modal mutual kNN alignment (k=1) as images / captions per class in gallery increase for different model pairs. Even with substantially stronger models (OpenLlama-65b, DINOv2-giant), individual retrieval improves with gallery density while cross-modal alignment remains flat.

20

B.3

ImageNet ablation shows a similar pattern for k = 10

The main paper reports the ImageNet decomposition experiments with mutual kNN scores for k=1. Here, we verify that the finding is not an artifact of this strict setting. We additionally present how mutual kNN with k=10 evolves when the gallery grows in Fig. 16 for two different model pairs. We again observe that individual retrieval accuracy improves with gallery density while cross-modal alignment, here in terms of mutual kNN with k=10, does not.

(a) DINOv2-base and OpenLlama-3b

(b) DINOv2-giant and OpenLlama-65b

Figure 16: Per-modality retrieval accuracy and cross-modal mutual kNN alignment (k=10) as images / captions per class in gallery increase for two different model pairs. Again, modalities individually improve with gallery density, but mutual kNN alignment, here with k=10, does not.

C

What happens with non-synthetic data that is not bijective?

In the main paper (Sec. 4.3), we use the CycleReward dataset [3] to test alignment when the bijective (one-to-one) assumption is relaxed with synthetic multi-modal correspondences. Here, we complement this analysis using non-synthetic many-to-many correspondences from the WIT dataset [82]. C.1

Non-synthetic dataset with many-to-many correspondences

Natural duplicates in WIT. The WIT dataset naturally contains many-to-many correspondences between images and captions: the same caption can describe many visually distinct images, and the same image is reused across Wikipedia articles with different corresponding text. Specifically, 7.1% of the captions are associated with more than one image, and 24.6% of the images have more than one caption before deduplication (see Section E.1). These naturally occurring one-to-many and manyto-one correspondences provide a complementary test bed for relaxing the bijective (one-to-one) setting without relying on generated images or captions. Within-group deduplication. Grouping by caption text (for Table 3: Natural duplicates in the WIT T2I) or by image (for I2T) can include within-group dupli- dataset after within-group deduplicacates, i.e. a caption group may contain duplicate image, and tion. an image group may contain repeated captions. After withingroup deduplication, the number of qualifying one-to-many T2I I2T samples decreases from 7,844 to 4,975 for T2I and from Unique elements 3.2 M 2.5 M 38,254 to 24,853 for I2T (Table 3). Appearing >1 time

We construct two complementary one-to-many datasets to mirror the experimental setup in Sec. 4.3 of the paper:

7.1%

24.6%

Groups with ≥5 corresp. Before dedup 7,844 38,254 After dedup 4,975 24,853

T2I (text-to-images): We select all 4,975 captions that are associated with at least 5 unique images. For each caption, we take 5 images, yielding a flat dataset of 24,875 image-text pairs. 21

I2T (image-to-texts): We identify 24,853 unique images that are paired with at least 5 distinct captions (by exact string matching). To match the T2I dataset size, we randomly subsample 4,975 images. For each image, we take 5 captions, again yielding 24,875 samples. Mutual kNN also decreases on non-synthetic data when the bijective assumption is relaxed

C.2

We evaluate alignment between DINOv2-base [69] and OpenLlama-3b [24] on the WIT-based T2I and I2T datasets. Fig. 17 shows mutual kNN alignment for k=1 and k=10 as the number of images per caption and vice versa increases from 1 to 5. In both directions, alignment decreases as bijectivity is relaxed. This is consistent with the results on the CycleReward dataset in the main paper (Fig. 10) and reinforces the conclusion that mutual kNN alignment is sensitive to the bijective assumption. When multiple valid correspondences exist for a query, the two modalities are less likely to agree on the same nearest neighbor, even if each individually retrieves a good match. This confirms that the observed drop in alignment for non-bijective setting is not an artifact of synthetic data.

Densifying captions per image k = 10 k=1

0.100 0.050 0.000 1

2

3

Images per caption

4

k = 10 k=1

0.100

Mutual kNN

Mutual kNN

Densifying images per caption 0.075 0.050 0.025 0.000 1

5

(a) T2I: increasing the number of unique images per caption from 1 (bijective) to 5.

2

3

Captions per image

4

5

(b) I2T: increasing the number of unique captions per image from 1 (bijective) to 5.

Figure 17: Effect of relaxing the bijective assumption on mutual kNN alignment using naturally occurring many-to-many correspondences in the WIT data (I2T and T2I subset). Mutual kNN alignment on non-synthetic drops consistently as we increase the number of images per caption and vice versa. This confirms that the pattern observed on CycleReward is not an artifact of synthetic data.

D

Does the alignment vs performance trend predicted by Huh et al. [40] continue with recent LLMs?

To assess whether the alignment vs performance trend predicted by Huh et al. [40] continues with recent language models, we evaluate 55 LLMs (see Section E.3.2 for full list) on six standard benchmarks using the LM Evaluation Harness framework [20]. [40] originally used three benchmarks to measure language capability: HellaSwag [97], GSM8K [11], and (1 − bitsperbyte) on OpenWebText [26]. We replace OpenWebText with Wikitext [61] and extend this analysis to three additional benchmarks that probe different aspects of language understanding: ARC Challenge [10], MMLU [34], and LogiQA2 [58]. D.0.1

Benchmarks and metrics.

Table 4 summarizes the evaluation configuration for each benchmark used in Sec. 4.4 of the paper (we used the default configurations from [20]). D.0.2

Does the alignment vs performance trend hold?

For each benchmark and each DINOv2 variant, we fit a linear regression on the base models used in [40], predicting mutual kNN alignment a from benchmark performance p. We then evaluate how well this trend describes two populations: 22

Table 4: Overview of language model benchmarks used in Sec. 4.4 of the paper. Benchmark

Capability

Few-shot

HellaSwag [97] Wikitext [61] ARC Challenge [10] GSM8K [11] MMLU [34] LogiQA2 [58]

Commonsense reasoning Language modeling Science QA Math reasoning General knowledge Logical reasoning

0 0 4 5 5 5

Metric Accuracy 1 − bits per byte Accuracy Exact match Accuracy Accuracy

Table 5: Average R2 of the linear regression (fitted on the 19 base models from Huh et al. [40]) evaluated on the base models themselves and on the 36 recent models, across all four DINOv2 2 variants. Positive Ravg (new) indicates that the trend from the base models is a good predictor for the new models. Negative values indicate that the regression line is a worse predictor than the mean.

Benchmark

2 Ravg (Huh et al.)

2 Ravg (new)

HellaSwag Wikitext

0.752 0.729

0.297 0.489

ARC GSM8K MMLU LogiQA2

0.702 0.336 0.430 0.431

−0.575 −1.753 −0.662 −1.414

R2 (Huh et al.): The standard coefficient of determination on the data is used to fit the regression, i.e. R2 (Huh et al.) = r2 , where r is the Pearson correlation between mutual kNN alignment and language modelling benchmark score across the 19 base models from Huh et al. [40]. R2 (new models): We apply the line fitted on the base models to the 36 recent models and compute the generalized R2 : P (ai − âi )2 2 R (new) = 1 − P i∈Mnew , 2 i∈Mnew (ai − ānew ) where âi are the linear regression alignment predictions based on language performance pi , ai is the alignment score for the i-th model and ānew is the mean alignment of the new models. When R2 (new) > 0, the relation between alignment and language performance predicted in Huh et al. extrapolates; when R2 (new) < 0, the regression line is a worse predictor than simply predicting the 2 average ānew . Ravg values are reported in Table 5 and Figs. 18 and 19. The results reveal a split across language modelling benchmarks. For HellaSwag and Wikitext, the relation between alignment and language performance observed by Huh et al. partially extends to 2 recent models: the Ravg on new models remains positive (0.297 and 0.489, respectively), indicating that stronger language models according to these benchmarks have higher mutual kNN alignment with DINOv2. Both benchmarks primarily measure next-token prediction quality and commonsense language understanding, which are closely related to the pretraining objective of autoregressive LLMs. In contrast, for the four benchmarks that probe more specialized reasoning abilities: ARC (science QA), GSM8K (arithmetic), MMLU (general knowledge), and LogiQA2 (logical reasoning), the relation between alignment and language performance predicted from the base models Huh et al. does not appear to hold for this set of recent models. 2 Specifically, the Ravg on new models is consistently negative, ranging from −0.575 (ARC) to −1.753 (GSM8K). This means that the linear fit from the base models from Huh et al. [40] is a worse predictor of alignment for recent models than simply predicting the mean. In Figs. 18 and 19, we observe that recent models that are stronger than the best base model (Meta-Llama-3-70B) do not show higher

23

mutual kNN alignment with DINOv2 features. Instead, their alignment scores seem to saturate or decrease. We note that the 36 added (new) models are heterogeneous. They include new models trained on next-token prediction (pre-training), instruction-tuned models, and reasoning-distilled models (e.g. DeepSeek-R1-Distill). We treat them as a single population to test whether the trend extrapolates to recent LLMs. The above results support and extend the finding from Sec. 4.4 of the main paper. The relationship between alignment and language performance from [40] holds for core language modelling benchmarks, but does not seem to generalize to reasoning benchmarks.

E

Experimental setup

In Section E.1, we provide additional details about the deduplication pipeline for the WIT-1M and LAION-15M datasets. We then describe the captioning pipeline used for WIT-1M-recap and for the ImageNet validation set in Section E.2. E.1 E.1.1

WIT-1M and LAION-15M datasets Image deduplication.

We deduplicate the gallery pools from WIT [82] and LAION400M [78] at the image level using perceptual hashing [96]. For each image, a 64-bit hash hi ∈ {0, 1}64 is computed. For this, we first convert the image to grayscale, resize it to 32 × 32, apply a 2D Discrete Cosine Transform, and threshold the top-left 8 × 8 low-frequency coefficients against their median. This produces a binary fingerprint that is robust to minor changes, e.g. due to recompression. We consider images i and j duplicates if their Hamming distance satisfies dH(hi , hj ) ≤ 2, measuring the number of bit positions at which two hashes differ: dH (hi , hj ) =

64 X

1[hi,b ̸= hj,b ].

b=1

We perform deduplication of the gallery against the WIT-1024 images, and within the gallery by keeping the first occurrence in the case of image duplicates. E.1.2

Caption deduplication.

In addition to image deduplication, we do a text deduplication pass to remove gallery samples with captions identical to another gallery sample or to a WIT-1024 query caption. Duplicate captions are undesirable because they allow trivial text-based query-gallery matching, inflating retrieval scores regardless of visual representations. We use exact string matching and remove any gallery samples that match WIT-1024 captions. Among the gallery samples, we discard duplicates and keep only the first occurrence of a duplicate captionsample. E.1.3

WIT-1M.

We obtain a raw pool of 3,582,610 samples from the English-text WIT dataset [82]. To construct the English-only subset of the dataset, we used a subset of [80, 79]. Since [79] only contains the image URLs, we retrieved the corresponding images from [82]. The raw pool undergoes our deduplication pipeline, resulting in 2,389,146 samples. We randomly sampled 1 million samples for the WIT-1M dataset. Deduplication statistics are provided in Table 6 and Table 7. In WIT, 31,845 captions appear repeatedly, accounting for 129,499 samples (5.21%); the most frequent captions are coat of arms (2716×) and Town hall (2361×). 24

Mutual kNN 0.00 0.0

LANGUAGE performance on GSM8K (5 shot) 0.2

0.05

0.4

25

0.5 0.6

dino small dino base dino large

0.60

gemma-3-27b DS-R1-Llama-70b

gemma-2-9b DS-R1-Qwen-14b llama-13b Llama-2-13b-hf OLMo-2-1124-7B gemma-7b Meta-Llama-3-8B Llama-3.1-8B Olmo-3-32Bb Qwen-14b Mistral-7b granite-3.3-8b-base DS-R1-Qwen-32b phi-4 Olmo-3-32b llama-30b Yi-1.5-34B Qwen-32b OLMo-2-1124-13B gemma-2-27b DS-R1-Llama-70b llama-65b Mixtral-8x7b gemma-3-27b aya-expanse-32b Llama-2-70b-hf Llama-3-70b Llama-3.1-70b

Falcon3-10b

openllama13b OLMo-7b DeepSeek-R1-Distill-Llama-8B gemma-3-4b-it Falcon3-7B-Base llama-7b Llama-2-7b-hf Qwen-8b

dino small dino base dino large

phi-4

gemma-3-4b-it Falcon3-7B-Base DeepSeek-R1-Distill-Qwen-7B Yi-1.5-34B Llama-3-70b Olmo-3-32Bb Falcon3-10b gemma-2-9b Llama-3.1-70b Qwen-14b Qwen-4b DS-R1-Qwen-14b aya-expanse-32b gemma-2-27b DS-R1-Qwen-32b Qwen-8b

0.55

OLMo-2-1124-13B Olmo-3-1025-7B

Qwen3-1.7B

0.50

Olmo-3-1025-7B

0.4

DeepSeek-R1-Distill-Llama-8B DeepSeek-R1-Distill-Qwen-1.5B OLMo-2-1124-7B

0.05 gemma-2-2b-it

Mutual kNN 0.05

granite-3.3-8b-base Qwen-32b

Qwen-4b openllama7b gemma-2b

0.3

Mixtral-8x7b

Llama-2-70b-hf

0.45 openllama3b

OLMo-1b bloomz-7b1

DeepSeek-R1-Distill-Qwen-7B Qwen3-1.7B

gemma-3-1b-it

LANGUAGE performance on WIKITEXT 0.2

Meta-Llama-3-8B Llama-3.1-8B gemma-7b

gemma-2-2b-it llama-65b

Mistral-7b

gpt-oss-20b

0.40 bloomz-3b

gpt-oss-20b

0.1

llama-30b

gemma-3-1b-it

LANGUAGE performance on HELLASWAG (0 shot) 0.35

Llama-2-13b-hf

bloomz-1b7

DeepSeek-R1-Distill-Qwen-1.5B

0.0

llama-13b gemma-2b

0.00 bloomz-1b1

gemma-3-270m-it

bloomz-560m

0.00 0.1

Llama-2-7b-hf

OLMo-1b bloomz-1b7 bloomz-3b bloomz-560m bloomz-1b1 openllama3b bloomz-7b1 gemma-3-270m-it openllama7b OLMo-7b Olmo-3-32b openllama13b llama-7b

Mutual kNN et al.) = 0.729 0.20 RR (Huh (new models) = 0.489 2 avg 2 avg

0.15

0.10

dino giant Huh et al., 2024 New models

0.7

et al.) = 0.752 0.20 RR (Huh (new models) = 0.297

2 avg 2 avg

0.15

0.10

dino giant Huh et al., 2024 New models

0.65

et al.) = 0.336 0.20 RR (Huh (new models) = -1.753

2 avg 2 avg

0.15

0.10

dino small dino base dino large dino giant Huh et al., 2024 New models

0.6

0.8

Figure 18: Mutual kNN alignment vs. language benchmark performance for 55 LLMs across four DINOv2 variants, on WikiText, HellaSwag, and GSM8K. Dashed lines show the linear trend fit to the 19 base models from [40]. For WikiText and HellaSwag (top two plots), recent models roughly follow the trend. For GSM8K (bottom plot), the trend is not followed. Llama-3-70b Llama-3.1-70b

Llama-2-70b-hf

llama-65b

Mixtral-8x7b

openllama3b OLMo-7b Qwen-14b DS-R1-Qwen-14b granite-3.3-8b-base gemma-2-9b openllama7b gemma-3-27b llama-7b Qwen-32b Falcon3-7B-Base Olmo-3-1025-7B openllama13b Llama-2-7b-hf gemma-2-27b DS-R1-Qwen-32b llama-13b Yi-1.5-34B Olmo-3-32b Mistral-7b Falcon3-10b phi-4 Llama-2-13b-hf Llama-3.1-8B Meta-Llama-3-8B OLMo-2-1124-7B DS-R1-Llama-70b llama-30b aya-expanse-32b Olmo-3-32Bb OLMo-2-1124-13B

Qwen-8b

gemma-2-2b-it

DeepSeek-R1-Distill-Llama-8B OLMo-1b

Qwen-4b

gemma-3-4b-it

bloomz-7b1

Qwen3-1.7B

bloomz-3b

gemma-3-1b-it

DeepSeek-R1-Distill-Qwen-7B

bloomz-1b1 bloomz-1b7

gemma-2b

DeepSeek-R1-Distill-Qwen-1.5B

bloomz-560m gemma-3-270m-it

Mutual kNN 0.00

LANGUAGE performance on MMLU (5 shot) 0.3

0.4

0.05

0.5

26

0.6 dino small dino base dino large

0.7

Qwen-32b

Qwen-14b

DS-R1-Qwen-32b

0.6

Qwen3-14b Llama-3.1-70b Llama-3-70b DS-R1-Llama-70b phi-4 DS-R1-Qwen-32b Qwen3-32b

0.6

DS-R1-Qwen-14b aya-expanse-32b Qwen3-8b Olmo-3-32Bb gemma-2-27b Yi-1.5-34B gemma-3-27b

dino small dino base dino large

gemma-3-27b gemma-2-27b Qwen-8b

Qwen-4b DS-R1-Llama-70b Yi-1.5-34B Llama-3.1-70b Llama-3-70b phi-4

Olmo-3-32Bb Olmo-3-32b gemma-2-9b DS-R1-Qwen-14b

dino small dino base dino large

gemma-2-9b Falcon3-10b Olmo-3-32b

Falcon3-7B-Base Qwen3-4b Mixtral-8x7b

Llama-2-70b-hf

0.5

Llama-2-70b-hf aya-expanse-32b

0.5

OLMo-2-1124-13B

0.05

Mistral-7b granite-3.3-8b-base gemma-7b OLMo-2-1124-7B llama-65b Olmo-3-1025-7B Llama-3.1-8B Meta-Llama-3-8B

Falcon3-10b

Mutual kNN 0.05

Qwen3-1.7B

llama-30b gemma-3-4b-it

Mixtral-8x7b Qwen3-1.7B

DeepSeek-R1-Distill-Qwen-7B gemma-3-4b-it OLMo-2-1124-13B llama-65b Olmo-3-1025-7B Falcon3-7B-Base

LANGUAGE performance on ARC (4 shot) 0.4

DeepSeek-R1-Distill-Llama-8B gemma-2-2b-it

Llama-2-13b-hf

LANGUAGE performance on LOGIQA2 (5 shot) 0.4

DeepSeek-R1-Distill-Qwen-7B

gpt-oss-20b

Llama-3.1-8B Meta-Llama-3-8B gemma-7b granite-3.3-8b-base llama-30b

Llama-2-13b-hf Mistral-7b

DeepSeek-R1-Distill-Llama-8B OLMo-2-1124-7B gemma-2-2b-it

Llama-2-7b-hf

0.3

Llama-2-7b-hf llama-13b

0.3

openllama13b

gemma-2b

gemma-3-1b-it

DeepSeek-R1-Distill-Qwen-1.5B

bloomz-7b1

llama-7b

bloomz-3b

bloomz-560m OLMo-1b bloomz-3b gemma-3-270m-it bloomz-7b1 OLMo-7b gpt-oss-20b openllama3b gemma-2b openllama7b llama-7b gemma-3-1b-it DeepSeek-R1-Distill-Qwen-1.5B openllama13b llama-13b

bloomz-1b1 bloomz-1b7

0.00 0.2

bloomz-1b7 openllama7b

0.000.2

OLMo-1b gemma-3-270m-it bloomz-1b1 openllama3b OLMo-7b

bloomz-560m

Mutual kNN

et al.) = 0.702 0.20 RR (Huh (new models) = -0.575 2 avg 2 avg

0.15

0.10

dino giant Huh et al., 2024 New models

0.7

et al.) = 0.431 0.20 RR (Huh (new models) = -1.414

2 avg 2 avg

0.15

0.10

dino giant Huh et al., 2024 New models

0.7

et al.) = 0.430 0.20 RR (Huh (new models) = -0.662

2 avg 2 avg

0.15

0.10

dino giant Huh et al., 2024 New models

0.8

Figure 19: Mutual kNN alignment vs. language benchmark performance for 55 LLMs across four DINOv2 variants, on ARC, LogiQA2, and MMLU. As with GSM8K, the alignment-performance trend from [40] does not extrapolate to recent models on any of these reasoning benchmarks. Stronger models do not appear to show higher mutual kNN alignment with DINOv2 features. gemma-2-27b gemma-3-27b

Qwen-32b

DS-R1-Llama-70b OLMo-2-1124-13B DS-R1-Qwen-32b Mixtral-8x7b Falcon3-10b Llama-2-70b-hf Yi-1.5-34B aya-expanse-32b Qwen-8b phi-4 Llama-3-70b Olmo-3-32b Olmo-3-32b Llama-3.1-70b Qwen-14b gemma-2-9b

Llama-2-13b-hf gemma-2-2b-it Meta-Llama-3-8B Llama-3.1-8B gemma-7b granite-3.3-8b-base Mistral-7b gemma-3-4b-it llama-30b Olmo-3-1025-7B Qwen-4b llama-65b Falcon3-7B-Base DS-R1-Qwen-14b OLMo-2-1124-7B

llama-13b

gpt-oss-20b Llama-2-7b-hf DeepSeek-R1-Distill-Qwen-7B Qwen3-1.7B

gemma-2b

llama-7b

openllama7b DeepSeek-R1-Distill-Llama-8B openllama13b

OLMo-7b

openllama3b bloomz-7b1

DeepSeek-R1-Distill-Qwen-1.5B

gemma-3-1b-it

OLMo-1b

bloomz-3b

gemma-3-270m-it bloomz-1b7

bloomz-1b1

bloomz-560m

Table 6: Deduplication statistics for the WIT-1M and LAION-15M gallery pools. Image duplicates are detected via perceptual hashing (pHash) with Hamming distance ≤ 2. Caption duplicates are detected by exact string matching. WIT-1M and LAION-15M are 1M and 15M image-caption pairs randomly sampled from the remaining final pool. WIT-1M LAION-15M Raw pool

3,582,610

20,000,000

Image deduplication Duplicates (with WIT-1024) Duplicates (within gallery) Pool after image deduplication

2,847 2,164,343 2,486,852

53 3,371,128 17,941,016

Caption deduplication Duplicates (with WIT-1024) Duplicates (within gallery)

52 97,654

14 642,895

Final pool size

2,389,146

17,298,107

Table 7: Distribution of caption duplicates in the WIT and LAION galleries after image deduplication. Note that the “unique captions” include some captions that are removed as query matches for the final pool. WIT Copies per caption

E.1.4

LAION

Unique captions

Total samples

Unique captions

Total samples

1 2 3 4 5 6–10 11–100 101–1,000 >1,000

2,357,353 22,799 3,807 1,529 787 1,520 1,283 118 2

2,357,353 45,598 11,421 6,116 3,935 11,431 29,765 27,164 5,077

16,897,221 316,109 48,864 15,422 6,932 9,596 3,838 103 4

16,897,221 632,218 146,592 61,688 34,660 69,614 76,200 12,376 14,588

Total

2,389,198

2,486,852

17,298,111

17,941,016

LAION-15M.

We randomly sample 20M samples from the LAION-400M dataset [78] which consists of English image-text pairs. We randomly sample 20M samples as our raw pool which undergoes our deduplication pipeline. Finally, we randomly sample 15M from the final pool after deduplication, resulting in our LAION-15M data pool. Deduplication statistics are provided in Table 6 and Table 7. In LAION, 400,890 unique captions appear repeatedly, accounting for 1,043,795 samples (5.82%); the most frequent captions are Patent Drawing (10027×) and Throw Pillow (3246×). E.2

Captioning pipeline for the ImageNet validation set and WIT-1M-recap

We used gemini-3-flash-preview [85, 70] for captioning the images in the ImageNet validation [15] set and in WIT-1M. Specifically, we used the following text prompt for each image. You are a precise image description system . Describe the image in the following JSON format . Return ONLY a valid JSON object with exactly these 7 keys . No text before or after the JSON . { " one_sentence ": " < exactly one sentence , strictly fewer than 15 words >" , " short ": " <2 -3 sentences , 20 -40 words total >" , "100 w ": " < a paragraph , approximately 100 words >" , "250 w ": " < several paragraphs , approximately 250 words >" , "500 w ": " < detailed description , approximately 500 words >" , "750 w ": " < thorough description covering all visual details , approximately 750 words >" , " extreme_long ": " < maximally detailed description covering every visible element , texture , color , spatial relationship , lighting , and context . YOU MUST WRITE AT LEAST 1000 WORDS . If your draft is under 1000 words , keep adding more detail about

27

(a) Mutual kNN alignment for DINOv2-base and OpenLlama-3b increases with longer captions on the ImageNet validation set.

(b) Distribution of word counts for Gemini-generated extreme_long captions across 49,984 ImageNet validation images.

Figure 20: Generated image captions for the ImageNet validation set. a) shows the mutual kNN alignment using captions of different lengths between DINOv2-base and OpenLlama3b on the ImageNet validation set. As also shown in [40], longer detailed captions yield higher alignment scores. b) shows the distribution over caption length (word count). We use captions of on average 981 words for our ImageNet experiments. textures , materials , lighting , spatial layout , colors , and any other visible elements until you reach at least 1000 words . Target 1000 -1500 words . >" } Be factual and visual . Describe what you actually see : objects , people , animals , colors , textures , spatial relationships , background , lighting , and mood . Do not invent information not visible in the image .

For the ImageNet validation set, we perform experiments with captions of the extreme_long type. As shown in Fig. 20a, alignment scores increase with caption length. Captions of approximately 500 words achieve scores close to the maximum, while the longest captions yield the best mutual kNN alignment scores between DINOv2-base and OpenLlama-3b on the ImageNet validation set. A distribution over the number of words for captions of the extreme_long caption type is shown in Fig. 20b. Despite prompting the model to produce at least 1,000 words per caption, 63.1% of captions fall below this target. The average caption length is 981 words. For WIT-1M-recap, we caption the 1 million images in the WIT-1M dataset using the 500w variant for computational efficiency. This resulted in captions for 999,971 (29 images did not get processed by gemini-3-flash-preview) images of on average 478 words. We provide those in https://huggingface.co/datasets/askoepke/wit_1m_recaptioned. E.3

Models and feature extraction pipeline

Our feature extraction pipeline is based on the experimental protocol from Huh et al. [40], which we extend to include additional LLMs. E.3.1

Vision models.

On the vision side, we use four DINOv2 [69] variants: ViT-S/14 (384-d), ViT-B/14 (768-d), ViT-L/14 (1024-d), and ViT-G/14 (1536-d) [16], loaded via the timm [93] library. For each image, we extract the CLS token representation from every transformer layer, yielding a per-sample feature tensor of shape L × d, where L is the number of layers and d the feature dimensionality. E.3.2

Language models.

We evaluate 55 large language models spanning 13 model families. The first group comprises the 19 base models used by Huh et al. [40]: BLOOMZ (560M–7.1B) [66], OpenLlama (3B–13B) [24], LLaMA (7B–65B) [87], OLMo (1B, 7B) [28], Gemma (2B, 7B) [21], Mistral-7B [43], Mixtral8×7B [44], and Meta-Llama-3-70B [27]. For the trend analysis in Section D, we extend this set with 36 recent models: LLaMA-2 (7B– 70B) [88], Llama-3/3.1 [27], OLMo-2/3 [67], Ministral-3 (3B–14B) [56], Gemma-2 (2B–27B)[22], 28

Gemma-3 (270M–27B) [23], DeepSeek-R1-Distill (1.5B–70B) [14], Qwen3 (1.7B–32B) [72], Falcon3 (7B, 10B) [84], 01.AI Yi-1.5 (34B) [95], IBM Granite (8b) [74], and OpenAI GPT-OSS (20b) [68]. For each model, we extract hidden-state representations from all layers. Following [40], we apply average pooling over non-padding tokens to obtain a single vector per layer.

F

Additional qualitative results

We present additional qualitative results for nearest-neighbor retrieval at different gallery scales on WIT-1M and LAION-15M in Figs. 21 and 22 and Figs. 23 and 24 respectively. In addition to the near-duplicate matches in Figs. 5 and 6 of the main paper, we here show further examples for cross-modal agreement (green-bordered matches) at scale when the modalities happen to select the same neighbor. Others show agreement at WIT-1024 that breaks down as the gallery densifies. In those cases, each modality individually finds a better match at scale, but they no longer agree on the same one.

29

Query

WIT-1024

10K

50K

100K

500K

1M

DINOv2

LLM

Chocolate Cupcakes on glass plate

Linzer Torte

New York-style cheesecake with strawberries. . .

A banana cupcake

A banana cupcake

5 chocolate chip cookies on a plate.

Boston cream pie cupcakes

The larger of the two ferry boats on the Wisemans Ferry crossing. . .

The current King Harry Ferry (No. 7) in 2017

The former temporary ferry landing of NY185 in Crown Point. . .

The former temporary ferry landing of NY185 in Crown Point. . .

The former temporary ferry landing of NY185 in Crown Point. . .

The west side of the Gladesville Bridge viewed from a Rivercat ferry. . .

The ferry Cobar Cat at Queens Wharf

Asian Koel or Common Koel

Yellow-billed Babblers allopreening

A Feral Barbary Dove in Tasmania, Australia. . .

Oriental Turtle Dove (Streptopelia orientalis) Nepal

Asian koel

Asian koel

Asian koel

IL95 heading east of IL41

OH520 in Killbuck from US62

VTF-5 heading westward from US7 towards the CharlotteEssex Ferry

IL155 at Fort de Chartres

IL 9 and IL 96 intersect in downtown Dallas City

IL171 approaching IL72 in Chicago

IL171 approaching IL72 in Chicago

Demet Bozkurt for Kireburnu Spor (March 2016)

Sevgi Inar of ALG Spor in the 2018-19 Women’s First League.

Damla Demirdön for Ataehir Belediyespor (March 2014)

Damla Demirdön for Ataehir Belediyespor (March 2014)

Fanta Zara Kamate for Fatih Vatan Spor (October 2018)

Özlem Gezer for Fatih Vatan Spor (October 2018)

Özlem Gezer for Fatih Vatan Spor (October 2018)

DINOv2

LLM

DINOv2

LLM

DINOv2

LLM

DINOv2

LLM

Figure 21: Additional nearest-neighbor examples with DINOv2 and OpenLlama-3b for k=1 across gallery scales on the WIT-1M dataset. For OpenLlama-3b, we show the (partial) retrieved captions along with the corresponding reference image (LLM-ref) for visualisation. Green-bordered captions and images indicate a mutual kNN match across modalities.

30

Query

WIT-1024

10K

50K

100K

500K

1M

Rano Karno as acting governor of Banten

A view of Mount Merbabu, Telomoyo and Lake Rawapening from Ambarawa.

People with a trapped tiger in Soepajang, Bovenlanden Padang on Sumatra’s west coast, ~1895

Candi Surowono, Pare, Kediri, Jawa Timur.

Candi Surowono, Pare, Kediri, Jawa Timur.

Ganjar Pranowo as Governor of Central Java second period (2018)

Sumarsono as Acting Governor of South Sulawesi

Yam thale (Thai)

Nasi goreng Pattaya (Pattaya fried rice). . .

A plate of som tam Lao (Thai)

A plate of som tam Lao (Thai)

A plate of som tam Lao (Thai)

Yam kun chiang (Thai)

Yam thua phu (Thai)

Double crested cormorant

Yellow-billed Babblers allopreening

Grey-cheeked parakeet

Female double-banded sandgrouse

Flightless cormorant

Double-crested morant

Double-crested morant

Alvin Ward Vogtle Nuclear Power Plant

Clarence Barker Memorial Hospital

Conemaugh Power Generation Station

San Onofre Nuclear Generating Station

E. I. Hatch Nuclear Power Plant near Baxley, Georgia.

James A. FitzPatrick Nuclear Power Plant

Virgil C. Summer Nuclear Station Unit 1

10 shillings of the Japanese occupation currency, 1942

A 1951 hundi of Bombay Province for Rs 2500 with a pre-printed 6a revenue stamp.

Obverse of the 1935-36 commemorative California Pacific International Exposition half dollar

Banknote of 10 Reichsmark, 1929

Banknote of 10 Reichsmark, 1929

Two shilling coin from 1949

1 yen note, 1932

DINOv2

LLM

DINOv2

LLM

DINOv2

LLM

cor-

cor-

DINOv2

LLM

DINOv2

LLM

Figure 22: Additional nearest-neighbor examples with DINOv2 and OpenLlama-3b for k=1 across gallery scales on the WIT-1M dataset. Green-bordered captions and images indicate a mutual kNN match across modalities.

31

Query

WIT-1024

10K

100K

1M

10M

15M

Linzer Torte

Chocolate Cupcakes on glass plate

Angel Food Cake Té Matcha

Austrian sweet dessert called kaiserschmarrn with apple sauce

Linzer Heart Cookies

Linzer cookie

Linzer cookie

Fluorite crystals on display at the Cullen Hall of Gems and Minerals

Aventurine scale)

Fluorite & Quartz - Fossils For Sale - #128787

Calcite (nailhead stalactite) from Hilton Mine, Cumbria

Fluorite Crystals, Elmwood Mine, Tennessee

Celestite sample on display at Rausch Mineral Gallery

Celestite sample on display at Rausch Mineral Gallery

Space 1 Bligh

3 Church Street.

Space & Astronauts Illustrations

Space 1999 - Set Eight

Space Rooms

Space Rooms

Space Rooms

Indian Blanket (Gaillardia pulchella). . .

The buttercup (Ranunculus spp) occurs in many variations. . .

The buttercup (Ranunculus spp) occurs in many variations. . .

Gaillardia or Brown Eyed Susan.

Guadalupe County, TX: Field of wildflowers featuring Indian paintbrush. . .

Firewheel or Indian Blanket or Blanketflower. . .

Firewheel or Indian Blanket or Blanketflower. . .

Rep. Harley Rouda (DCA48)

Pat Toomey, the expected Republican challenger. . .

Rep. Rashida Tlaib, left, and Rep. Ilhan Omar.

Rep. Paul Ryan (R-WI)

Rep. Mike Honda (D-CA)

Rep. Mike Honda (D-CA)

Rep. Mike Honda (D-CA)

DINOv2

LLM

DINOv2

LLM

(unknown

DINOv2

LLM

2435

Study

2435

Study

2435

Study

DINOv2

LLM

DINOv2

LLM

Figure 23: Additional nearest-neighbor examples with DINOv2 and OpenLlama-3b for k=1 across gallery scales on the LAION-15M dataset. Green-bordered captions and images indicate a mutual kNN match across modalities.

32

Query

WIT-1024

10K

100K

1M

10M

15M

art

DINOv2

LLM

Modern art in the street

Street in the village

Art in my hair refrigerator magnets

Modern Pisa

sculpture,

Visual arts in the twentieth century

Visual arts in the twentieth century

Contemporary Art in the Middle East

Scotwater Bridge across the River Witham. . .

Glan Rhyd Viaduct. . .

Glan Rhyd Viaduct. . .

SJ8935 : Canalside grazing above Meaford Locks, Staffordshire

Plardiwick Bridge near Gnosall Heath, Staffordshire

“The Glory Hole” ~ High bridge Lincoln ~UK c1160. . .

High bridge over river Nidd, rebuilt in 1773, Knaresborough. . .

ChME3 is a Czech dieselelectric switcher built by ČKD. . .

The enormous SBB Ae 8/14 “double locomotive”. . .

The enormous SBB Ae 8/14 “double locomotive”. . .

The enormous SBB Ae 8/14 “double locomotive”. . .

The TRAXX electric locomotive 481 001-0 of Eurocom at Kazincbarcika

Diesel locomotive ChME 3, SŽD w/sound

Diesel locomotive ChME 3, SŽD w/sound

Tesco is the largest supermarket chain in the United Kingdom.

A larger Farmfoods in Pontefract, West Yorkshire.

Sainsbury’s ‘outperforms market’ in Q1

Sainsbury’s ‘outperforms market’ in Q1

The UK Supermarket Groups Supermarkets dominate the fresh food market. . .

British retail chain Marks and Spencer in Central District of Hong Kong. . .

British retail chain Marks and Spencer in Central District of Hong Kong. . .

Thatched cottages and outbuilding on the B184 Dunmow Road. . .

Thatched cottages, Main Road, Grendon, Northamptonshire. . .

Thatched cottages, Main Road, Grendon, Northamptonshire. . .

Thatched Cottage, Old Warden

Thumbnail 7 bed farmhouse for sale in Hedingham Road, Bulmer, Sudbury

TL9311 : Cottages on Tollesbury Road, Tolleshunt D’Arcy. . .

Cottage at the end of Church Lane, Little Welbourne, Pagham

DINOv2

LLM

Railway

Railway

DINOv2

LLM

DINOv2

LLM

DINOv2

LLM

Figure 24: Additional nearest-neighbor examples with DINOv2 and OpenLlama-3b for k=1 across gallery scales on the LAION-15M dataset. Green-bordered captions and images indicate a mutual kNN match across modalities.

33

Record · ID 120496 · SHA-256 33bf1ab34b2e3893
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.