Conceptio › Archive › arXiv CS
arXiv CSopen access

Embedding Models Measure in Peculiar Ways

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Embedding Models Measure in Peculiar Ways Andrianos Michail Dept. of Computational Linguistics University of Zurich [email protected]

A. Ideal physical geometry

-15

0 Physical quantity (meters)

15

B. all mpnet base v2

-15 15

0

Introduction

Embedding models are routinely interpreted as semantic measurement spaces. Notions like Semantic Textual Similarity (STS, Cer et al., 2017) concretely pose their task as assessing “the degree to which two sentences are semantically equivalent.” Yet, evaluation of this property has proven tricky, also because there can often be many aspects in which two items are similar (Tversky, 1977). As a focused, idealized testbed of semantic equivalence and distance, this paper examines physical measurement systems. Those systems have long enabled precise communication about the physical world, from science and engineering to everyday reasoning (Greaves, 1647; Kennelly, 1935). In this paper, we therefore ask:

-15

0 Physical quantity (meters)

15

C. Qwen3 Embedding 0.6B

How well are the physical metric spaces aligned with embedding metric spaces?

0 Physical quantity (meters)

0.8 0.6 0.4 0.2 0.0

-15 15

0

-15

1.0

Cosine similarity

1

0

Physical quantity (meters)

Embedding spaces define notions of semantic similarity and distance. We study whether those embeddings reflect physical measurements of mass, distance, time and volume, which admit a unique, objective notion of semantic equivalence and distance. We find that physical measurement is only weakly modeled in the embedding space, and that instead quite peculiar measurement patterns can be observed. Further analysis indicates that embedding representations of physical measurements are strongly influenced by superficial string similarity, and recalibration of similarity does not substantially improve the alignment.

15

Physical quantity (meters)

arXiv:2609.20821v1 [cs.CL] 17 Sep 2026

Abstract

15

Physical quantity (meters)

Juri Opitz Dept. of Computational Linguistics University of Zurich [email protected]

-15

Figure 1: Ideal (A) vs. extracted similarities (B, C). Each square holds the similarity of two physical measurement phrases, e.g., “1 meter” vs. “10 meters”.

For instance, the expressions 1 meter and 100 centimeters clearly denote the same physical quantity and we might hope therefore that they are located in close proximity in an embedding space. Similarly, we may wonder if 1 meter is embedded closer to 105 centimeters than to, e.g., 15 kilometers, as

only then would the similarity reflect the actual underlying physical relationship. Our study finds that the physical relationships are far from faithfully reflected in embedding 1

spaces. Instead, most peculiar patterns can be observed across all 24 models that we tested. Notably, this holds largely independently of model architecture and release date. Figure 1 illustrates this finding: Panel A depicts an ideal measurement space in which similarity decreases smoothly with increasing physical distance. Besides the diagonal, both tested models’ measurement patterns (Panels B and C) appear mostly erratic. Indeed, we can easily spot pronounced non-monotonic artifacts, including horizontal/vertical bands, as well as irregular local neighborhoods that misalign with physical distance. That said, both models seem to be able to roughly differentiate negative from positive measurements, as can be observed by the larger, darker rectangular areas on the bottom right (in B, and C). In this work, we evaluate a set of 24 common embedding models and report three main, generalizing findings.

very faithfully). Finally, in Section 7 we investigate potential reasons for the alignment disparities. We conclude with a discussion (Section 8). We share code for reproducing the experiments publicly.1

2

Related Work

Embedding models are typically trained contrastively to minimize distances between similar texts and maximize distances between dissimilar texts. The resulting representations are used in all kinds of NLP tasks, ranging from document classification to retrieval and clustering. Early milestones of such models are based on encoders such as “SBERT” (Reimers and Gurevych, 2019) and “SimCSE” (Gao et al., 2021), later ones also extend this principle of contrastive training to decoderbased LLMs (Zhang et al., 2025). Embedding models can either be evaluated through the ‘macro lens’ via large benchmarking suites like MTEB (Muennighoff et al., 2023), or a ‘micro lens’ via focused tasks like recognizing challenging paraphrases (Li et al., 2025; Michail et al., 2026). Some works also study the interpretability of such embeddings (Opitz et al., 2025). In our paper we tighten the focus further, to a central objective phenomenon: physical measurements. While physical measurements have been explored in LLMs (Park et al., 2022) and word embeddings (Sundararaman et al., 2020), our exploration is conducted on the contrastively trained text representations that promise to preserve “essential semantic and syntactic information” (Tao et al., 2026). We therefore ask: To what extent do they preserve the meaning of expressions of physical measurements?

1. Representations of physical units and measurements are consistently weak across all 24 tested models, all showing various peculiar alignment patterns. 2. Models that are generally considered stronger do not clearly show better alignment in our measurement evaluation. The model family or embedding dimension also does not appear to play a major role. 3. We find evidence that embedding similarity of physical measurements is strongly associated with lexical rather than semantic/numerical similarity, offering a plausible explanation for many of the observed non-aligned patterns. Simultaneously, via linear probes, we reject the hypothesis that the similarity is merely miscalibrated for this task.

3

General Setup

Measurement through embeddings. We assume that given is E: T → Rn , a function that maps from text objects to a high-dimensional vector space (i.e. ‘embedding model’). We further assume two quantities (q1 , q2 ) ∈ R2 measured in two corresponding units (u1 , u2 ) ∈ T 2 . Furthermore, str : R × T → T maps a measurement in a unit to a string. Then   E str(q1 , u1 ) · E str(q2 , u2 )   (1) ||E str(q1 , u1 ) || × ||E str(q2 , u2 ) ||

The remainder of this paper is structured as follows: After discussing related work in Section 2, we introduce the formal operationalization of embedding measurement and the tested embedding models in Section 3. In Sections 4 and 5 we perform visual exploration, testing measurements (e.g., 2 meters, 3 meters) and unit-understanding properties of embeddings (e.g., meter, kilometer, liter, gallon, etc.), respectively. In Section 6 we craft a focused benchmark to compare models along a single axis with regard to understanding physical measurements, testing how faithfully they reflect the physical distance relationships (The result is: Not

computes the standard cosine similarity score based on the vectors’ dot-product (·) and the product of 1

https://github.com/flipz357/embed-andmeasure

2

their magnitudes (×). For instance, say q1 = 10, q2 = 1, and u1 = kilometer and u2 = meter. The embeddings on which the similarity is computed will be E(“10 kilometers”) and E(“1 meter”).

The results for all other models are shown in Appendix A.2 and show various patterns of marked misalignments. Importantly, we do not observe any marked and broader alignment difference from older embedding models that are considered “weaker”, generally speaking, and the newer models, even if they are based on LLM-decoders.

Embedding models. We select a diverse set of 24 standard and widespread embedding models, ranging from older ones based on BERT encoders, to recent larger ones based on LLM decoders, including broader model families. The full overview and discussion on the selected models is given in Appendix A.1. When studying a specific example case in the paper, we will use all-mpnet-base-v2 and Qwen3-Embedding-0.6B as these two models are among the most downloaded embedding models and are built on encoder, and decoder, respectively. However, we will always present parallel results of other models in the appendix. Also note that those embedding models span a range of embedding dimensionalities: from 384 (MiniLM) to 4096(Qwen3-Embedding-8B).

The checker pattern. Some plots, particularly those for the Local scale (left column, Figure 2) exhibit a pronounced “checker pattern,” dominated by aligned integer-integer pairs alternating with nonaligned integer-float pairs. E.g., for both models 2.5 {meters, kilograms, liters, seconds} is more similar to 7.5 {meters, kilograms, liters, seconds} than to 3 {meters, kilograms, liters, seconds}, which misaligns with the physical measurement. The explanation for this may be intuitive and leads to a first hypothesis for the misalignment: Embedding similarity for measurements is mostly driven by surface similarity, and not the actual underlying metric distance (in the last part of this piece, we investigate this hypothesis in depth).

Physical quantities and ranges. We use the following physical quantities as measured in standard units: Length (meters), Mass (kilograms), Volume (liters), Time (seconds). We further introduce different measurement ranges as expressed by different types of ranges: Local is a range from 0 to 10; Medium from 0 to 1,000; and log for up to 100,000. Sign applies an extended scale from -100 to 100, including negative numbers. And scientific uses scientific exponential notation. The stepsizes are selected such each range is broken down into roughly 20 steps. Lastly, we express the Local scale both numerically (e.g., “5 meters”) and as words (e.g., “five meters”), for each integer in between zero/0 and ten/10, including one-hundred/100.

4

Numerical and literal expressions. We study the alignment of quantities expressed numerically and literally, in the right column of Figure 2. Strikingly, we see that embeddings barely capture this semantic equivalence, but some interesting model differences can be observed: For instance, Qwen assigns high similarities to one/1, as well as to one/0—most pronounced for the liter and kilogram quantity. On the other hand, mpnet assigns one and 1 a markedly higher similarity than to one and all other numbers that are not 1. Further, Qwen shows some interesting horizontal lines when aligning digits with words. For example, it appears that “nine liters” are rated as dissimilar to all other liter quantities expressed in digits, whereas this is not so for, e.g., “one liter” or “two liters.” For nine/9 and ten/10, even the self comparison is comparatively weak, showing pronounced misalignment between the numerical and literal expression of the same number. With the exception of one-hundred/100, the self-paraphrased numbers are better aligned in the mpnet model that was released a few years before Qwen.

Visual Exploration

Setup. Given one quantity and one range, we apply the embedding model to calculate similarity for each value pair (e.g., “5 meters” vs. “8 meters”, Eq. 1). This allows us to plot and compare heatmaps according to different scales and measures, separately for each embedding model. Results. The full results for two embedding models are visualized in Figure 2. Specifically, across all units and quantities, we can see a stark contrast to how a smooth idealized plot can be imagined (see the caption for description of idealized plots), observing many misaligned and peculiar patterns.

5

Testing Conversion Understanding

For this study, we ask how well embedding spaces capture unit conversion. For instance, one hour is 60 minutes, whereas one gallon is 3.78 liters. 3

Local

all-mpnet-base-v2

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

Local

Qwen3-Embedding-0.6B Log

Sign

Scientific

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e-03

1e+00

50

100 1e-06

0

-50

100000 -100

5623

18

316

750

Medium

1000 1

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e-03

1e+00

50

100 1e-06

0

-50

100000 -100

5623

18

316

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 2: Large-scale measurement exploration for two example models. For most columns/plots the idealized figure is simply one that uniformly becomes more yellow towards the diagonal (Fig 1. 1 top panel). An exception is the Log and the scientific notation, where the idealized plot would have a darkening effect that is over-proportionally increasing towards the bottom-right corner (since distances are over-proportionally increasing).

Figure 3 shows how well different units align across the different types of quantities. In all of the shown plots in this figure, the idealized alignment

would be a horizontal line that is very close to 1.0. However, when inspecting the alignment through embedding spaces, we again see rather diffuse pat4

all-mpnet-base-v2

1.0 0.9 0.8 0.7 0.6 0.5 0.4

meter centimeter meter millimeter meter kilometer

10 1 101 Physical value

Cosine similarity

Cosine similarity

Length

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

10 1 101 Physical value

kilogram gram gram milligram kilogram ton

10 1 101 Physical value

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Qwen3-Embedding-0.6B

1.0 0.9 0.8 0.7 0.6 0.5 0.4

10 1 101 Physical value

Cosine similarity

Cosine similarity

Length

meter centimeter meter millimeter meter kilometer

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

1.0 0.9 0.8 0.7 0.6 0.5 0.4

10 1 101 Physical value

103

Setup. We use two performance metrics. Kendall’s τ is a classic non-parametric statistic that measures the ordinal association between two variables, well suited to this evaluation. We further define as an additional metric the Pairwise Accuracy (PAC), which simply calculates the ratio of correctly ranked pairs. Note that PAC also follows directly from τ (PAC=50 + τ /2), we include it primarily here for easier readability. E.g., if the similarity of two meters to 2 meters is higher than to, e.g., 100 meters, a point is added; the resulting sum is divided by all available pairs. We calculate both performance metrics across every distance relationship (across every scale and across every quantity). We call this small benchmark ‘PhysScore’.

103

Mass

kilogram gram gram milligram kilogram ton

10 1 101 Physical value

103

Time

Cosine similarity

Cosine similarity

Volume

liter milliliter liter gallon

103

second millisecond minute second hour minute

10 1 101 Physical value

1.0 0.9 0.8 0.7 0.6 0.5 0.4

second millisecond minute second hour minute

10 1 101 Physical value

Crafting a Small Benchmark

We would like to assign an average degree of ‘task performance’ to the models, in order to compare them with respect to their capacity of modeling the physical relationships, ranking the models on a single axis.

Time

Cosine similarity

Cosine similarity

Volume

6

Mass

103

Figure 3: Conversion analysis. For each quantity and two example models there is one plot. Each line in each plot indicates a conversion-alignment. For example, in the plot on the top left, the y-value at the point 100 on the x-axis denotes the embedding similarity between 1 meter and: 100 centimeters (blue line), 1000 millimeters (orange line) and 0.001 kilometers (green line).

terns of alignment. This time, the lines tend to be slightly higher in the Qwen embedding space, potentially indicating better alignment across different units. Particularly, as the quantities become larger, there is a strong alignment performance drop-off observed in mpnet, but not so much in Qwen. The comparatively most pronounced alignment seems to be in Qwen when relating seconds and milliseconds—it’s the only line that exceeds 0.9 and stays consistently high. However, ultimately, no line is horizontal and close to 1.0, and so we can conclude that unit conversion is also highly fuzzy in both embedding models. Other embedding models are shown in Appendix A.3 and show similarly non-aligned patterns, with some particularly erratic ones.

model

K’s τ

PAC

AVG

all-MiniLM-L6-v2 all-mpnet-base-v2 bge-large-en-v1.5 bge-m3 DenseOn e5-large-v2 embeddinggemma-300m granite-embedding-107m-multilingual granite-embedding-278m-multilingual granite-embedding-311m-multilingual-r2 granite-embedding-97m-multilingual-r2 granite-embedding-english-r2 harrier-oss-v1-0.6b harrier-oss-v1-270m LaBSE multilingual-e5-base multilingual-e5-large multilingual-e5-large-instruct mxbai-embed-large-v1 nomic-embed-text-v1.5 paraphrase-multilingual-mpnet-base-v2 Qwen3-Embedding-0.6B Qwen3-Embedding-4B Qwen3-Embedding-8B

36.97 40.97 39.82 37.30 42.70 47.35 24.40 26.38 30.84 33.72 33.67 36.94 40.76 45.96 51.31 39.06 41.97 53.23 39.24 18.57 33.99 40.47 34.22 32.27

68.48 70.48 69.91 68.65 71.35 73.68 62.20 63.19 65.42 66.86 66.83 68.47 70.38 72.98 75.66 69.53 70.98 76.61 69.62 59.29 66.99 70.24 67.11 66.13

52.72 55.73 54.86 52.97 57.03 60.51 43.30 44.79 48.13 50.29 50.25 52.70 55.57 59.47 63.49 54.29 56.47 64.92 54.43 38.93 50.49 55.35 50.67 49.20

AVG

37.59

68.79

53.19

random-baseline

0.00

50.00

25.00

Table 1: Result on PhysScore.

Result. Table 1 shows the overall performance ranking of models. While all models outperform the random baseline, no model exceeds a Kendall’s τ of 54. The best model is a e5 variant (multilingual-e5-large-instruct) with Kendall’s τ of 53. Hence, all models are sub5

minimal τ of any embedding model can even be lower than 0.0 in some cases, indicating an inverse measurement conducted by an embedding model. Overall the relatively best aligned quantities appear to be distance and time (meters and seconds). Regarding scales, the medium scale proved most challenging for the models, on average.

stantially misaligned with the physical measurements. Also note that these results can already be viewed as statistically optimistic—the random baseline is naturally disfavored since the diagonals are trivially resolved by all embedding models (same strings). Another observation is that model scale seems to have no clear positive effect on the alignment capacity. Sometimes even, an inverse effect is observed: The similarity of the smallest Qwen model sometimes achieves a better alignment than the larger ones.

Range

Avg

Min

Max

kilogram kilogram kilogram kilogram kilogram

local log medium scientific sign

27.66 50.47 23.65 35.92 35.00

-20.95 -7.33 -4.76 16.00 -15.24

69.52 84.00 56.19 62.67 63.81

kilogram

AVG

34.54

-6.46

67.24

liter liter liter liter liter

local log medium scientific sign

29.96 61.53 27.98 34.72 30.95

-13.33 2.67 4.76 12.67 -20.95

62.86 82.00 53.33 66.67 60.00

liter

AVG

37.03

-2.84

64.97

meter meter meter meter meter

local log medium scientific sign

32.98 62.44 29.84 35.42 31.75

-18.10 9.33 -13.33 16.00 -18.10

77.14 84.67 75.24 60.67 65.71

meter

AVG

38.48

-4.84

72.69

second second second second second

local log medium scientific sign

32.18 61.78 32.18 38.81 36.55

13.33 -1.33 -0.95 17.33 -13.33

64.76 89.33 61.90 74.00 61.90

second

AVG

40.30

3.01

70.38

Analysis

7.1

Is Similarity just Miscalibrated?

Through our embedding models, we calculated similarity in the most frequent way: Using cosine similarity. However, this assigns each dimension in the embedding representation the same prior weight. Thus, it could be that the representations hold highly useful information for measuring, but it gets blurred when applying the dot-product. To test this hypothesis, we perform a trainingtest split, and re-evaluate the models. The baseline for this experiment is the cosine similarity, and the tested model is a linear regression on the absolute difference of the embedding vectors of two measurement expressions (aka linear probe). This gives the option to potentially emphasize those dimensions which could perhaps hold the information about physical relationships.

Which measurements are hardest? To explore this question, we show in Table 2 the average, minimum, and maximum Kendall’s τ across all embedding models, for each quantity, unit, and range. Unit

7

Table 2: Average, minimum, and maximum Kendall correlation across models for each unit and range.

Cosine

Probe

model

K’s τ

K’s τ

∆τ

all-MiniLM-L6-v2 all-mpnet-base-v2 bge-large-en-v1.5 bge-m3 DenseOn e5-large-v2 embeddinggemma-300m gr-embedding-107m-multilingual gr-embedding-278m-multilingual gr-embedding-311m-multilingual-r2 gr-embedding-97m-multilingual-r2 gr-embedding-english-r2 harrier-oss-v1-0.6b harrier-oss-v1-270m LaBSE multilingual-e5-base multilingual-e5-large multilingual-e5-large-instruct mxbai-embed-large-v1 nomic-embed-text-v1.5 par-multilingual-mpnet-base-v2 Qwen3-Embedding-0.6B Qwen3-Embedding-4B Qwen3-Embedding-8B

37.1 40.4 31.9 37.3 38.9 35.7 28.3 29.3 30.0 28.9 26.8 25.5 34.4 39.5 43.6 31.9 38.0 44.7 31.0 25.8 20.4 33.7 35.7 33.7

37.0 43.0 32.3 39.7 35.6 48.1 21.0 20.4 33.2 32.2 28.7 26.7 40.1 47.3 50.1 39.3 43.5 51.4 32.5 29.1 26.9 31.9 41.9 41.5

-0.1 +2.6 +0.4 +2.3 -3.3 +12.3 -7.3 -8.9 +3.2 +3.3 +1.9 +1.3 +5.6 +7.8 +6.5 +7.4 +5.5 +6.7 +1.6 +3.3 +6.4 -1.8 +6.2 +7.8

AVG

33.4

36.4

+3.0

Table 3: Result of recalibration experiment expressed in Kendall’s τ × 100.

We see that Mass (kilogram) appears as the most challenging for all embedding models (on average), with an average Kendall’s τ of only 34.5. The 6

Character

Tokenizer

Numeric

Model

All

F ×F

F ×I

I×I

All

F ×F

F ×I

I×I

All

F ×F

F ×I

I×I

all-MiniLM-L6-v2 all-mpnet-base-v2 bge-large-en-v1.5 bge-m3 DenseOn e5-large-v2 embeddinggemma-300m granite-embedding-107m-multilingual granite-embedding-278m-multilingual granite-embedding-311m-multilingual-r2 granite-embedding-97m-multilingual-r2 granite-embedding-english-r2 harrier-oss-v1-0.6b harrier-oss-v1-270m LaBSE multilingual-e5-base multilingual-e5-large multilingual-e5-large-instruct mxbai-embed-large-v1 nomic-embed-text-v1.5 paraphrase-multilingual-mpnet-base-v2 Qwen3-Embedding-0.6B Qwen3-Embedding-4B Qwen3-Embedding-8B

21.0 27.7 10.9 18.1 23.5 29.0 51.6 27.0 29.4 14.2 4.9 16.7 27.0 28.4 39.6 23.5 25.5 39.7 10.0 -5.6 5.0 47.6 30.1 36.7

22.0 31.7 25.2 31.3 18.1 23.9 36.3 25.9 25.1 32.5 17.5 29.9 42.1 42.5 32.7 32.6 32.4 28.1 24.2 16.9 17.8 23.6 26.3 33.9

17.1 21.0 29.1 36.8 21.0 27.2 39.6 32.4 33.2 33.3 11.9 29.6 32.0 30.5 28.5 34.3 29.4 29.4 28.7 28.9 19.8 30.7 38.8 39.2

44.2 41.9 35.2 49.2 36.9 45.1 44.6 43.1 46.3 53.6 27.9 44.6 43.0 43.1 48.7 50.3 50.1 40.4 33.9 32.8 29.7 50.8 53.8 49.5

15.1 20.4 7.7 3.6 27.4 22.0 51.6 16.7 16.3 14.2 8.5 9.1 27.0 28.4 42.7 6.8 6.8 27.5 5.4 -16.9 -11.5 47.6 30.1 36.7

10.4 19.9 22.9 25.0 17.0 15.2 36.3 10.9 11.8 32.5 18.9 14.7 42.1 42.5 42.0 27.3 24.9 25.3 16.0 -3.7 -5.0 23.6 26.3 33.9

10.2 8.7 31.1 25.7 25.1 14.5 39.6 14.1 14.5 33.3 2.5 4.6 32.0 30.5 13.0 13.6 0.9 15.9 26.6 5.7 3.7 30.7 38.8 39.2

37.1 34.0 18.5 31.9 31.6 38.7 44.6 31.0 25.2 53.6 29.3 21.1 43.0 43.1 35.7 27.7 31.2 22.4 17.1 17.5 6.4 50.8 53.8 49.5

28.1 26.1 8.0 19.5 20.2 16.9 3.3 16.7 14.8 9.1 12.5 14.4 21.7 23.3 18.1 17.2 18.4 26.9 6.9 6.5 8.1 11.7 13.1 18.3

23.0 40.9 26.5 34.3 20.2 18.2 8.3 14.9 11.7 17.0 12.5 20.5 46.1 48.5 29.7 24.1 37.4 49.1 18.0 5.8 14.4 13.7 12.6 24.2

28.7 26.0 6.9 22.7 18.9 15.3 2.1 16.4 14.6 10.4 12.7 10.8 29.7 32.8 18.8 23.3 26.9 29.9 6.1 5.9 7.0 13.7 8.4 15.8

39.0 36.2 18.6 29.7 33.3 32.8 8.8 20.6 21.1 11.8 21.2 22.6 32.6 38.0 24.4 25.4 24.9 30.7 16.5 15.2 8.2 25.2 27.3 25.1

Table 4: Kendall’s τ (×100; range −100 to 100, higher = stronger association) between embedding similarity and three reference similarities: character-level Levenshtein (Character), token-level Levenshtein (Tokenizer), and true numerical proximity (Numeric). Columns distinguish float–float (F ×F ), float–integer (F ×I), integer–integer (I×I), and all pairs (All).

The results are shown in Table 3. Overall, changes by recalibration are small. The largest increase in Kendall’s τ points is +12.33, but the absolute value for this model is still fairly small (47 τ ). Furthermore, no model exceeds 54 τ , even after recalibration. Regarding the Qwen model family, this time the larger ones appear to profit more from recalibration, so their representation may hold the information in better ways, but not aligned with the similarity. Still, overall alignment is low. In sum, we find no convincing support for the hypothesis that the misalignment with physical relationships is just due to miscalibration of the similarity function.2 7.2

cally investigate this hypothesis. Setup. We generate ranges of random numbers, one series of floating-point numbers (F ) and another series of integers (I), both ∈ [−1000, +1000]. We convert these numbers to strings. For every unique pair of these number strings (joining floating and integer series), we calculate 1. The Levenshtein distances based on the characters in the strings; 2. The Levenshtein distances based on the generated sub-word sequences of the embedding models themselves (‘tokens’); and 3. The ‘ground-truth’ so to speak, i.e., the actual distance between the numbers. We negate these distances and compute Kendall’s τ between embedding similarity and the (inverse) distances. Put simply, a higher correlation implies a stronger link between embedding similarity and another similarity (either character, sub-word, or ground truth). We report correlations separately for float–float (F ×F ), float–integer (F ×I), and integer–integer (I×I) pairs, as well as over all pairs jointly (All).

The Effect of Superficial String Overlap

Based on the results from our visual explorations, we had already speculated that one major driver of the misalignments is an over-focus on superficial string similarity. Because of this, embedding models might tend to neglect the actual underlying physical relationships. In this section, we empiri-

Result. The results in Table 4 reveal a clear distinction between lexical similarity and numerical semantics. Across the evaluated models, embedding similarity tends to be more strongly associated

2 This is in line with Huber and Opitz (2026) who find that the original similarity function of the contrastively trained models is already fairly optimal.

7

with the textual representation of numbers than with their actual numerical values. In other words, numbers that look similar—for example, because they share digits or token fragments—often receive more similar embeddings than numerically close numbers. For example, the Qwen3 embeddings exhibit a fair correlation with both character-level and tokenizer-level Levenshtein similarity. At the same time, their correlation with the true numerical distance is comparatively weaker. This suggests that Qwen3 primarily encodes the physical statements as lexical objects rather than as quantities with an inherent ordering or magnitude. Overall, these findings support our initial hypothesis that modern embedding models may largely process numbers according to their textual representation rather than their quantitative meaning. If embeddings primarily organize numbers according to shared digits, prefixes, or tokenizer fragments, then semantic operations over numerical quantities become unreliable and highly dependent on the accidental textual form of the numbers. This effect persists even for recent, large-scale embedding models, suggesting that increased model capacity alone does not necessarily lead to more faithful representations of physical statements. Indeed, one of the strongest correlations with string similarity and one of the weakest correlations with numerical distance is observed for Google’s LLM-based gemma embedding (embeddinggemma-300m).

8

numerical relationships between values. In particular, even LLM-based embedding models like Qwen3 exhibit moderate dependence on string similarity while showing only weak correspondence to the underlying quantitative semantics. Recalibration of the similarity function did not alleviate this issue. This suggests that at least part of the observed misalignment stems from the models organizing representations according to surface form rather than meaning. While the work of Weller et al. (2026) might suggest that the misalignments reflect a theoretical limitation of embedding models by their dimensionality, we do not observe a corresponding trend in alignment quality that would correspond to their dimensionality. Instead, we believe that our results point toward the learning objective itself. Contrastive training may simply provide insufficient pressure to organize the embedding space according to physical or quantitative relationships. For such alignment to emerge without explicit supervision, it would have to constitute an emergent capability of embedding models (or be an explicit objective in benchmarks). This naturally raises another question: Should similarity be aligned with physical measurements? For many retrieval and clustering applications, understanding quantitative relationships may play only a minor role. And lexical similarity can often be a useful inductive bias, particularly for identifiers, version numbers, dates, or product codes. Nevertheless, it seems also fair to expect that any embedding model that promises to measure semantic equivalence, or semantic similarity, should clearly judge “2 kilometers” and “2000 meters” as more similar than “2 kilometers” and “2 meters”. More generally, the inability to capture such elementary relationships may become increasingly problematic as embeddings are employed beyond conventional retrieval tasks, for example as semantic representations in scientific, engineering, or agentic systems that require reasoning about quantities and measurements. Taken together, our findings provide further evidence for the conclusion of Fodor et al. (2025) that “state-of-the-art transformers poorly capture the pattern of human semantic similarity judgments” and agree with Sun et al. (2026) who argue that embeddings should capture implicit semantics.

Discussion

From our study there is one central conclusion: Embedding model similarity is only weakly aligned with physical measurements. While our range of tested embedding models is necessarily limited, we did not observe a systematic improvement between older to more recent models, suggesting that it is not simply a matter of scale or recency. This raises a central question: What drives these misalignments? Our final experiment provides one potential explanation. Embedding similarity is more strongly associated with superficial lexical similarity— measured either at the character level or by the models’ own tokenization—than with the actual 8

Our evaluation through physical measurements may thus provide an additional quality metric in future model development, or a focused test of whether an embedding model can produce representations with a deeper understanding of the content. To conclude, we wish to emphasize what Thawani et al. (2021) say: In NLP, “numbers are important”, and “numbers are neglected”; Our work provides evidence that numbers, and, more broadly, physical measurements, are misaligned in contrastively trained text representation models.

lenging sentence similarity dataset. Computational Linguistics, 51(1):139–190. Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple contrastive learning of sentence embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6894–6910, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics. John Greaves. 1647. A Discourse of the Romane Foot and Denarius; from whence, as from two principles, the measures and weights used by the ancients may be deduced. William Lee, London.

Acknowledgements This work has been supported by the Swiss National Science Foundation (grant no. CRSII5_213585) and by the Luxembourg National Research Fund (grant no. 17498891).

Xiaolong Huang, Liang Wang, Furu Wei, Jingwen Lu, Knut Risvik, and Jason Li. 2026. Microsoft open-sources industry-leading embedding model. Bing Blogs. Models: https://huggingface.co/ microsoft/harrier-oss-v1-0.6b.

References

Marius Huber and Juri Opitz. 2026. Exploring dowker homology for sentence similarity. arXiv preprint arXiv:2608.22909.

Parul Awasthy, Aashka Trivedi, Yulong Li, Meet Doshi, Riyaz Bhat, Vishwajeet Kumar, Yushu Yang, Bhavani Iyer, Abraham Daniels, Rudra Murthy, et al. 2025. Granite embedding r2 models. arXiv preprint arXiv:2508.21085.

Arthur E. Kennelly. 1935. Adoption of the meterkilogram-mass-second (m.k.s.) absolute system of practical units by the international electrotechnical commission (i.e.c.), bruxelles, june, 1935. Proceedings of the National Academy of Sciences of the United States of America, 21(10):579–583.

Parul Awasthy, Aashka Trivedi, Yushu Yang, Ken Barker, Yulong Li, Bhavani Iyer, Martin Franz, Juergen Bross, Meet Doshi, Vishwajeet Kumar, et al. 2026. Granite embedding multilingual r2 models. arXiv preprint arXiv:2605.13521.

Sean Lee, Aamir Shakir, Darius Koenig, and Julius Lipp. 2024. Open source strikes bread - new fluffy embedding model.

Daniel Cer, Mona Diab, Eneko Agirre, Iñigo LopezGazpio, and Lucia Specia. 2017. SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pages 1–14, Vancouver, Canada. Association for Computational Linguistics.

Hongji Li, Andrianos Michail, Reto Gubelmann, Simon Clematide, and Juri Opitz. 2025. Sentence smith: Controllable edits for evaluating text embeddings. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 26428–26445, Suzhou, China. Association for Computational Linguistics.

Jianlyu Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. 2024. M3embedding: Multi-Linguality, multi-functionality, multi-granularity text embeddings through selfknowledge distillation. In Findings of the Association for Computational Linguistics: ACL 2024, pages 2318–2335, Bangkok, Thailand. Association for Computational Linguistics.

Xianming Li and Jing Li. 2024. AoE: Angle-optimized embeddings for semantic textual similarity. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1825–1839, Bangkok, Thailand. Association for Computational Linguistics. Andrianos Michail, Stylianos Psychias, Michelle Wastl, Simon Clematide, Rico Sennrich, and Juri Opitz. 2026. Alee: Any-language evaluation of embeddings via english-centric minimal pairs. arXiv preprint arXiv:2607.00171.

Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2022. Language-agnostic BERT sentence embedding. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 878–891, Dublin, Ireland. Association for Computational Linguistics.

Niklas Muennighoff, Nouamane Tazi, Loic Magne, and Nils Reimers. 2023. MTEB: Massive text embedding benchmark. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 2014–2037, Dubrovnik, Croatia. Association for Computational Linguistics.

James Fodor, Simon De Deyne, and Shinsuke Suzuki. 2025. Compositionality and sentence meaning: Comparing semantic parsing and transformers on a chal-

9

Zach Nussbaum, John Xavier Morris, Andriy Mulyar, and Brandon Duderstadt. 2025. Nomic embed: Training a reproducible long context text embedder. Transactions on Machine Learning Research. Reproducibility Certification.

models: An in-depth overview. ACM Trans. Inf. Syst. Just Accepted. Avijit Thawani, Jay Pujara, Filip Ilievski, and Pedro Szekely. 2021. Representing numbers in NLP: a survey and a vision. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 644–656, Online. Association for Computational Linguistics.

Juri Opitz, Lucas Moeller, Andrianos Michail, Sebastian Padó, and Simon Clematide. 2025. Interpretable text embeddings and text similarity explanation: A survey. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 22303–22319, Suzhou, China. Association for Computational Linguistics.

Amos Tversky. 1977. Features of similarity. Psychological review, 84(4):327. Henrique Schechter Vera, Sahil Dua, Biao Zhang, Daniel Salz, Ryan Mullins, Sindhu Raghuram Panyam, Sara Smoot, Iftekhar Naim, Joe Zou, Feiyang Chen, Daniel Cer, Alice Lisak, Min Choi, Lucas Gonzalez, Omar Sanseviero, Glenn Cameron, Ian Ballantyne, Kat Black, Kaifeng Chen, and 70 others. 2025. Embeddinggemma: Powerful and lightweight text representations. Preprint, arXiv:2509.20354.

Sungjin Park, Seungwoo Ryu, and Edward Choi. 2022. Do language models understand measurements? In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 1782–1792, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics. Nils Reimers and Iryna Gurevych. 2019. SentenceBERT: Sentence embeddings using Siamese BERTnetworks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Computational Linguistics.

Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022. Text embeddings by weaklysupervised contrastive pre-training. arXiv preprint arXiv:2212.03533. Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024. Multilingual e5 text embeddings: A technical report. arXiv preprint arXiv:2402.05672.

Nils Reimers and Iryna Gurevych. 2020. Making monolingual sentence embeddings multilingual using knowledge distillation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4512–4525, Online. Association for Computational Linguistics.

Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020. Minilm: Deep selfattention distillation for task-agnostic compression of pre-trained transformers. In Advances in Neural Information Processing Systems, volume 33, pages 5776–5788. Curran Associates, Inc.

Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and TieYan Liu. 2020. Mpnet: Masked and permuted pretraining for language understanding. In Advances in Neural Information Processing Systems, volume 33, pages 16857–16867. Curran Associates, Inc.

Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Griffin Thomas Adams, Jeremy Howard, and Iacopo Poli. 2025. Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2526–2547, Vienna, Austria. Association for Computational Linguistics.

Raphaël Sourty, Antoine Chaffin, Paulo Roberto Moura Junior, and Amélie Chatelain. 2026. Denseon with the lateon: Fully open dense and late-interaction models for multilingual, long-context, and code search. arXiv preprint arXiv:2607.27178. Yiqun Sun, Qiang Huang, Anthony Kum Hoe Tung, and Jun Yu. 2026. Position: Text embeddings should capture implicit semantics, not just surface meaning. In Forty-third International Conference on Machine Learning Position Paper Track.

Orion Weller, Michael Boratko, Iftekhar Naim, and Jinhyuk Lee. 2026. On the theoretical limitations of embedding-based retrieval. In The Fourteenth International Conference on Learning Representations.

Dhanasekar Sundararaman, Shijing Si, Vivek Subramanian, Guoyin Wang, Devamanyu Hazarika, and Lawrence Carin. 2020. Methods for numeracypreserving word embeddings. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4742–4753, Online. Association for Computational Linguistics.

Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie. 2024. C-pack: Packed resources for general chinese embeddings. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’24, page 641–649, New York, NY, USA. Association for Computing Machinery.

Chongyang Tao, Tao Shen, Shen Gao, Junshuo Zhang, Zhen Li, Kai Hua, Wenpen Hu, Zhangwei Tao, and Shuai Ma. 2026. Llms are also effective embedding

10

also belongs to this group, but uses a modified BERT architecture that supports long inputs.

Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, et al. 2025. Qwen3 embedding: Advancing text embedding and reranking through foundation models. arXiv preprint arXiv:2506.05176.

A

Appendix

A.1

Embedding Models

ModernBERT-based encoders. Several recent encoders are built on ModernBERT (Warner et al., 2025), an updated BERT architecture with a longer context window. These are granite-embedding-english-r2 (Awasthy et al., 2025), the multilingual granite-embedding-97m-multilingual-r2 and granite-embedding-311m-multilingualr2 (Awasthy et al., 2026), and DenseOn (Sourty et al., 2026), an English retrieval model trained on openly released data.

We evaluate 24 embedding models. Table 5 lists their Hugging Face identifiers. We group them into four families by backbone and training approach. Sentence-Transformers baselines. all-mpnetbase-v2 and all-MiniLM-L6-v2 are English models from the Sentence-Transformers library (Reimers and Gurevych, 2019). They are based on MPNet (Song et al., 2020) and MiniLM (Wang et al., 2020) and are fine-tuned on a large collection of sentence pairs. Both are widely used and serve as classic baselines. We also include two multilingual models that are distributed through the same library. paraphrase-multilingual-mpnet-basev2 is trained with knowledge distillation: a multilingual student model learns to place a sentence and its translation at the same position as an English teacher model (Reimers and Gurevych, 2020). LaBSE is a multilingual BERT model trained on translation pairs, originally for finding parallel sentences (Feng et al., 2022).

Models derived from LLMs. The last group starts from pre-trained decoder language models. Qwen3-Embedding (Zhang et al., 2025) is built on the Qwen3 LLMs. To test the effect of model size in a focused way, we include the 0.6B, 4B, and 8B variants. harrier-oss-v1 (Huang et al., 2026) is a multilingual decoder-only model family from Microsoft. The two variants we use (harrieross-v1-270m and harrier-oss-v1-0.6b, based on the Gemma 3 and Qwen3 architectures, respectively) are trained contrastively and with knowledge distillation from larger embedding models. The model embeddinggemma-300m (Vera et al., 2025) is also based on Gemma 3, but it first turns the LLM into an encoder-decoder model and then uses only the encoder. It therefore reads the input text in both directions, like the BERT-style models above.

Contrastive BERT-style encoders. Most models in our set follow a similar recipe: a BERT-like encoder is trained contrastively on large amounts of weakly paired text and then fine-tuned on labeled data. This group includes e5-large-v2 (Wang et al., 2022) and its multilingual versions multilingual-e5-base, multilingual-e5-large, and multilinguale5-large-instruct, where the last one takes a task instruction as part of the input (Wang et al., 2024). It further includes bge-largeen-v1.5 (Xiao et al., 2024) and its multilingual successor bge-m3 (Chen et al., 2024); mxbaiembed-large-v1 (Lee et al., 2024), which is trained with the AnglE objective (Li and Li, 2024); and the first release of the multilingual Granite embedding models (granite-embedding107m-multilingual and granite-embedding278m-multilingual), which are based on XLMRoBERTa (Awasthy et al., 2025). The model nomic-embed-text-v1.5 (Nussbaum et al., 2025) 11

A.2

Measure Plots

A.3

Conversion Plots

Model

Measure plot

Conversion plot

Sentence-Transformers baselines sentence-transformers/all-mpnet-base-v2 sentence-transformers/all-MiniLM-L6-v2 sentence-transformers/paraphrase-multilingual-mpnet-base-v2 sentence-transformers/LaBSE

Section 4 Figure 4 Figure 5 Figure 6

Section 5 Figure 26 Figure 27 Figure 28

Contrastive BERT-style encoders intfloat/e5-large-v2 intfloat/multilingual-e5-base intfloat/multilingual-e5-large intfloat/multilingual-e5-large-instruct BAAI/bge-large-en-v1.5 BAAI/bge-m3 mixedbread-ai/mxbai-embed-large-v1 ibm-granite/granite-embedding-107m-multilingual ibm-granite/granite-embedding-278m-multilingual nomic-ai/nomic-embed-text-v1.5

Figure 7 Figure 8 Figure 9 Figure 10 Figure 11 Figure 12 Figure 13 Figure 14 Figure 15 Figure 16

Figure 29 Figure 30 Figure 31 Figure 32 Figure 33 Figure 34 Figure 35 Figure 36 Figure 37 Figure 38

ModernBERT-based encoders ibm-granite/granite-embedding-english-r2 ibm-granite/granite-embedding-97m-multilingual-r2 ibm-granite/granite-embedding-311m-multilingual-r2 lightonai/DenseOn

Figure 17 Figure 18 Figure 19 Figure 20

Figure 39 Figure 40 Figure 41 Figure 42

Models derived from LLMs Qwen/Qwen3-Embedding-0.6B Qwen/Qwen3-Embedding-4B Qwen/Qwen3-Embedding-8B microsoft/harrier-oss-v1-270m microsoft/harrier-oss-v1-0.6b google/embeddinggemma-300m

Section 4 Figure 21 Figure 22 Figure 23 Figure 24 Figure 25

Section 5 Figure 43 Figure 44 Figure 45 Figure 46 Figure 47

Table 5: Hugging Face identifiers of all evaluated models, grouped as in the text, with links to their measurement and conversion plots.

12

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 4: all-MiniLM-L6-v2. For more information see caption in Figure 2.

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 5: paraphrase-multilingual-mpnet-base-v2. For more information see caption in Figure 2.

13

Local

Medium

Log

Sign

Scientific

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

meter

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

1e+06

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

0 1 2 3 4 5 6 7 8 9 10 100

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 6: LaBSE. For more information see caption in Figure 2.

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

Figure 7: e5-large-v2. For more information see caption in Figure 2.

14

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 8: multilingual-e5-base. For more information see caption in Figure 2.

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 9: multilingual-e5-large. For more information see caption in Figure 2.

15

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 10: multilingual-e5-large-instruct. For more information see caption in Figure 2.

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

Figure 11: bge-large-en-v1.5. For more information see caption in Figure 2.

16

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 12: bge-m3. For more information see caption in Figure 2.

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 13: mxbai-embed-large-v1. For more information see caption in Figure 2.

17

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 14: granite-embedding-107m-multilingual. For more information see caption in Figure 2.

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 15: granite-embedding-278m-multilingual. For more information see caption in Figure 2.

18

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 16: nomic-embed-text-v1.5. For more information see caption in Figure 2.

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 17: granite-embedding-english-r2. For more information see caption in Figure 2.

19

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 18: granite-embedding-97m-multilingual-r2. For more information see caption in Figure 2.

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 19: granite-embedding-311m-multilingual-r2. For more information see caption in Figure 2.

20

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 20: DenseOn. For more information see caption in Figure 2.

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 21: Qwen3-Embedding-4B. For more information see caption in Figure 2.

21

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 22: Qwen3-Embedding-8B. For more information see caption in Figure 2.

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 23: harrier-oss-v1-270m. For more information see caption in Figure 2.

22

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 24: harrier-oss-v1-0.6b. For more information see caption in Figure 2.

Local

Medium

Log

Sign

Scientific

Digits Words

0 1 2 3 4 5 6 7 8 9 10 100

meter

one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero one-hundred ten nine eight seven six five four three two one zero

0 1 2 3 4 5 6 7 8 9 10 100

kilogram

0 1 2 3 4 5 6 7 8 9 10 100

1e+03

1e+00

1e-03

100 1e-06

0

50

-50

100000 -100

316

5623

18

1000 1

750

500

250

10 0

5

7.5

0

2.5

second

1e+06

0 1 2 3 4 5 6 7 8 9 10 100

liter

Figure 25: embeddinggemma-300m. For more information see caption in Figure 2.

23

10 1 101 Physical value

Time

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

103

Figure 26: all-MiniLM-L6-v2. For more information see caption in Figure 3.

10 1 101 Physical value

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

10 1 101 Physical value

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Time

Cosine similarity

Cosine similarity

Volume

10 1 101 Physical value

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

meter centimeter meter millimeter meter kilometer

103

second millisecond minute second hour minute

10 1 101 Physical value

103

10 1 101 Physical value

Mass

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Volume

second millisecond minute second hour minute

10 1 101 Physical value

10 1 101 Physical value

10 1 101 Physical value

Time

Length

kilogram gram gram milligram kilogram ton

Cosine similarity

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

kilogram gram gram milligram kilogram ton

Figure 29: e5-large-v2. For more information see caption in Figure 3.

Cosine similarity

1.0 0.9 0.8 0.7 0.6 0.5 0.4

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Mass

Cosine similarity

Cosine similarity

Length

meter centimeter meter millimeter meter kilometer

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Volume

second millisecond minute second hour minute

10 1 101 Physical value

10 1 101 Physical value

Cosine similarity

Cosine similarity

103

Cosine similarity

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

Cosine similarity

Cosine similarity

Volume

10 1 101 Physical value

Mass

Cosine similarity

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

meter centimeter meter millimeter meter kilometer

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

10 1 101 Physical value

kilogram gram gram milligram kilogram ton

10 1 101 Physical value

103

Time

Cosine similarity

10 1 101 Physical value

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Length

kilogram gram gram milligram kilogram ton

Cosine similarity

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Mass

Cosine similarity

Cosine similarity

Length

meter centimeter meter millimeter meter kilometer

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

second millisecond minute second hour minute

10 1 101 Physical value

103

Figure 27: paraphrase-multilingual-mpnet-base-v2. Figure 30: multilingual-e5-base. For more inforFor more information see caption in Figure 3. mation see caption in Figure 3.

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

10 1 101 Physical value

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Volume

second millisecond minute second hour minute

10 1 101 Physical value

10 1 101 Physical value

Mass

Cosine similarity

Cosine similarity

103

Time

Cosine similarity

Cosine similarity

Volume

10 1 101 Physical value

1.0 0.9 0.8 0.7 0.6 0.5 0.4

meter centimeter meter millimeter meter kilometer

103

Figure 28: LaBSE. For more information see caption in Figure 3.

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

10 1 101 Physical value

103

kilogram gram gram milligram kilogram ton

10 1 101 Physical value

103

Time

Cosine similarity

10 1 101 Physical value

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Length

kilogram gram gram milligram kilogram ton

Cosine similarity

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Mass

Cosine similarity

Cosine similarity

Length

meter centimeter meter millimeter meter kilometer

1.0 0.9 0.8 0.7 0.6 0.5 0.4

second millisecond minute second hour minute

10 1 101 Physical value

103

Figure 31: multilingual-e5-large. For more information see caption in Figure 3.

24

103

103

Figure 32: multilingual-e5-large-instruct. For more information see caption in Figure 3.

10 1 101 Physical value

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

10 1 101 Physical value

103

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

10 1 101 Physical value

103

10 1 101 Physical value

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

10 1 101 Physical value

103

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Time

Cosine similarity

Cosine similarity

Volume

10 1 101 Physical value

1.0 0.9 0.8 0.7 0.6 0.5 0.4

second millisecond minute second hour minute

10 1 101 Physical value

103

10 1 101 Physical value

Mass

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

10 1 101 Physical value

kilogram gram gram milligram kilogram ton

10 1 101 Physical value

103

Time

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

meter centimeter meter millimeter meter kilometer

second millisecond minute second hour minute

10 1 101 Physical value

103

10 1 101 Physical value

Mass

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Volume

second millisecond minute second hour minute

10 1 101 Physical value

meter centimeter meter millimeter meter kilometer

Length

kilogram gram gram milligram kilogram ton

Cosine similarity

1.0 0.9 0.8 0.7 0.6 0.5 0.4

103

Figure 36: granite-embedding-107m-multilingual. For more information see caption in Figure 3.

Cosine similarity

1.0 0.9 0.8 0.7 0.6 0.5 0.4

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Mass

Cosine similarity

Cosine similarity

Length

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Volume

second millisecond minute second hour minute

Figure 33: bge-large-en-v1.5. For more information see caption in Figure 3.

meter centimeter meter millimeter meter kilometer

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Time

Cosine similarity

Cosine similarity

Volume

10 1 101 Physical value

10 1 101 Physical value

10 1 101 Physical value

Time

Length

kilogram gram gram milligram kilogram ton

Cosine similarity

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

kilogram gram gram milligram kilogram ton

Figure 35: mxbai-embed-large-v1. For more information see caption in Figure 3.

Cosine similarity

1.0 0.9 0.8 0.7 0.6 0.5 0.4

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Mass

Cosine similarity

Cosine similarity

Length

meter centimeter meter millimeter meter kilometer

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Volume

second millisecond minute second hour minute

10 1 101 Physical value

103

Cosine similarity

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Cosine similarity

Cosine similarity

Time

10 1 101 Physical value

Cosine similarity

10 1 101 Physical value

103

Cosine similarity

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

Cosine similarity

Cosine similarity

Volume

10 1 101 Physical value

Mass

Cosine similarity

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

meter centimeter meter millimeter meter kilometer

103

Figure 34: bge-m3. For more information see caption in Figure 3.

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

10 1 101 Physical value

103

kilogram gram gram milligram kilogram ton

10 1 101 Physical value

103

Time

Cosine similarity

10 1 101 Physical value

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Length

kilogram gram gram milligram kilogram ton

Cosine similarity

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Mass

Cosine similarity

Cosine similarity

Length

meter centimeter meter millimeter meter kilometer

1.0 0.9 0.8 0.7 0.6 0.5 0.4

second millisecond minute second hour minute

10 1 101 Physical value

103

Figure 37: granite-embedding-278m-multilingual. For more information see caption in Figure 3.

25

103

103

Figure 38: nomic-embed-text-v1.5. For more information see caption in Figure 3.

103

10 1 101 Physical value

1.0 0.9 0.8 0.7 0.6 0.5 0.4

103

Figure 39: granite-embedding-english-r2. more information see caption in Figure 3.

For

10 1 101 Physical value

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

1.0 0.9 0.8 0.7 0.6 0.5 0.4

10 1 101 Physical value

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

103

second millisecond minute second hour minute

10 1 101 Physical value

103

10 1 101 Physical value

Mass

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

10 1 101 Physical value

kilogram gram gram milligram kilogram ton

10 1 101 Physical value

103

Time

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

1.0 0.9 0.8 0.7 0.6 0.5 0.4

meter centimeter meter millimeter meter kilometer

second millisecond minute second hour minute

10 1 101 Physical value

103

10 1 101 Physical value

Mass

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Volume

second millisecond minute second hour minute

10 1 101 Physical value

meter centimeter meter millimeter meter kilometer

Length

kilogram gram gram milligram kilogram ton

10 1 101 Physical value

103

Figure 42: DenseOn. For more information see caption in Figure 3.

Time

Cosine similarity

Cosine similarity

Volume

liter milliliter liter gallon

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Mass

Cosine similarity

Cosine similarity

Length

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Volume

second millisecond minute second hour minute

10 1 101 Physical value

1.0 0.9 0.8 0.7 0.6 0.5 0.4

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Time

103

meter centimeter meter millimeter meter kilometer

Cosine similarity

103

Cosine similarity

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

Cosine similarity

Cosine similarity

Volume

10 1 101 Physical value

10 1 101 Physical value

10 1 101 Physical value

Time

Length

kilogram gram gram milligram kilogram ton

Cosine similarity

10 1 101 Physical value

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

kilogram gram gram milligram kilogram ton

Figure 41: granite-embedding-311m-multilingual-r2. For more information see caption in Figure 3.

Cosine similarity

1.0 0.9 0.8 0.7 0.6 0.5 0.4

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Mass

Cosine similarity

Cosine similarity

Length

meter centimeter meter millimeter meter kilometer

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Volume

second millisecond minute second hour minute

10 1 101 Physical value

103

Cosine similarity

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Cosine similarity

Cosine similarity

Time

10 1 101 Physical value

Cosine similarity

10 1 101 Physical value

103

Cosine similarity

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

Cosine similarity

Cosine similarity

Volume

10 1 101 Physical value

Mass

Cosine similarity

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

meter centimeter meter millimeter meter kilometer

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

10 1 101 Physical value

103

kilogram gram gram milligram kilogram ton

10 1 101 Physical value

103

Time

Cosine similarity

10 1 101 Physical value

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Length

kilogram gram gram milligram kilogram ton

Cosine similarity

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Mass

Cosine similarity

Cosine similarity

Length

meter centimeter meter millimeter meter kilometer

1.0 0.9 0.8 0.7 0.6 0.5 0.4

second millisecond minute second hour minute

10 1 101 Physical value

103

Figure 40: granite-embedding-97m-multilingual-r2. Figure 43: Qwen3-Embedding-4B. For more informaFor more information see caption in Figure 3. tion see caption in Figure 3.

26

1.0 0.9 0.8 0.7 0.6 0.5 0.4

10 1 101 Physical value

Mass

Cosine similarity

Cosine similarity

Length

meter centimeter meter millimeter meter kilometer

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

10 1 101 Physical value

10 1 101 Physical value

103

Time

Cosine similarity

Cosine similarity

Volume

kilogram gram gram milligram kilogram ton

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

second millisecond minute second hour minute

10 1 101 Physical value

103

Figure 44: Qwen3-Embedding-8B. For more information see caption in Figure 3.

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

10 1 101 Physical value

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

103

Figure 45: harrier-oss-v1-270m. For more information see caption in Figure 3.

1.0 0.9 0.8 0.7 0.6 0.5 0.4

10 1 101 Physical value

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

10 1 101 Physical value

103

103

Time

Cosine similarity

Cosine similarity

Volume

kilogram gram gram milligram kilogram ton

10 1 101 Physical value

1.0 0.9 0.8 0.7 0.6 0.5 0.4

second millisecond minute second hour minute

10 1 101 Physical value

1.0 0.9 0.8 0.7 0.6 0.5 0.4

liter milliliter liter gallon

10 1 101 Physical value

103

kilogram gram gram milligram kilogram ton

10 1 101 Physical value

103

Time

1.0 0.9 0.8 0.7 0.6 0.5 0.4

second millisecond minute second hour minute

10 1 101 Physical value

103

Figure 47: embeddinggemma-300m. For more information see caption in Figure 3.

Mass

Cosine similarity

Cosine similarity

Length

meter centimeter meter millimeter meter kilometer

103

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Volume

second millisecond minute second hour minute

10 1 101 Physical value

10 1 101 Physical value

Mass

Cosine similarity

Cosine similarity

103

Time

Cosine similarity

Cosine similarity

Volume

10 1 101 Physical value

1.0 0.9 0.8 0.7 0.6 0.5 0.4

meter centimeter meter millimeter meter kilometer

Cosine similarity

10 1 101 Physical value

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Length

kilogram gram gram milligram kilogram ton

Cosine similarity

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Mass

Cosine similarity

Cosine similarity

Length

meter centimeter meter millimeter meter kilometer

103

Figure 46: harrier-oss-v1-0.6b. For more information see caption in Figure 3.

27

Record · ID 978404 · SHA-256 29cc9b0dc9037def
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.