ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Ziyin Zhang 1 2 Zihan Liao 2 Hang Yu 2 Peng Di 2 Rui Wang 1
arXiv:2605.15081v1 [cs.CL] 14 May 2026
Abstract
and accessibility of the downstream systems they enable, from semantic search to Retrieval-Augmented Generation (RAG) (Gao et al., 2023).
The development of high-quality text embeddings is increasingly drifting toward an exclusionary future, defined by three critical barriers: prohibitive computational costs, a narrow linguistic focus that neglects most of the world’s languages, and a lack of transparency from closed-source or openweight models that stifles research. To dismantle these barriers, we introduce ML-Embed, a suite of inclusive and efficient models built upon a new framework: 3-Dimensional Matryoshka Learning (3D-ML). Our framework addresses the computational challenge with comprehensive efficiency across the entire model lifecycle. Beyond the storage benefits of Matryoshka Representation Learning (MRL) and flexible inference-time depth provided by Matryoshka Layer Learning (MLL), we introduce Matryoshka Embedding Learning (MEL) for enhanced parameter efficiency. To address the linguistic challenge, we curate a massively multilingual dataset and train a suite of models ranging from 140M to 8B parameters. In a direct commitment to transparency, we release all models, data, and code. Extensive evaluation on 430 tasks demonstrates that our models set new records on 9 of 17 evaluated MTEB benchmarks, with particularly strong results in low-resource languages, providing a reproducible blueprint for building globally equitable and computationally efficient AI systems.
However, the paradigm for developing state-of-the-art embedding models has shifted toward repurposing massive decoder-based language models. While powerful, this trend is creating a critical computational barrier characterized by prohibitive training costs and immense memory footprints. This computational barrier exacerbates a growing linguistic barrier: as models become more resourceintensive, they become increasingly inaccessible to the broader research community and undeployable in resourceconstrained environments where many of the world’s lowresource languages are spoken. While techniques like Matryoshka Representation Learning (MRL, Kusupati et al., 2022) offer partial relief by optimizing storage, they leave the immense burdens of training and inference untouched. To dismantle this computational barrier, we introduce 3Dimensional Matryoshka Learning (3D-ML), a unified framework built upon Matryoshka Layer Learning (MLL, Li et al., 2024a) and Matryoshka Representation Learning (MRL, Kusupati et al., 2022) while integrating a novel Matryoshka Embedding Learning (MEL) technique that addresses the critical challenge of parameter-heavy embedding layers by learning two factorized, low-rank matrices that are themselves structured for nested training. This provides significant parameter savings for both training and inference and offers flexible deployment options that balance efficiency with compatibility. To validate the practical utility of 3D-ML, we applied it to the notoriously resource-intensive challenge of creating massively multilingual models—a domain that faces two further critical gaps. The first is a linguistic challenge: despite comprehensive benchmarks like MTEB (Muennighoff et al., 2023; Enevoldsen et al., 2025), research attention remains disproportionately focused on a few high-resource languages. As illustrated in Table 1, submissions to benchmarks for languages like Polish, Persian, and Vietnamese are orders of magnitude fewer than for English. The second is a transparency challenge: progress is stymied by a lack of openness, as many top-performing models (Zhang et al., 2025b; Lee et al., 2025b) are released as closed-source APIs or as open-weight models with no training transparency,
1. Introduction Text embeddings are a foundational component of modern AI, translating the richness of human language into numerical representations that dictate the performance, fairness, 1
School of Computer Science, Shanghai Jiao Tong University, Shanghai, China 2 Ant Group, Hangzhou, China. Correspondence to: Hang Yu <[email protected]>, Peng Di <[email protected]>, Rui Wang <[email protected]>. Proceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026. Copyright 2026 by the author(s).
1
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World Table 1. Number of models with complete results on MTEB benchmarks. While Multilingual and English have become popular testbeds for embedding models, some languages - especially Polish, Japanese, Vietnamese, and Persian - receive far less attention.
hindering reproducible research. Addressing these interconnected issues, we introduce ML-Embed, a family of efficient and inclusive models built with our 3D-ML framework and a new, massively multilingual dataset. This work makes the following contributions: • We propose MEL, a novel efficient training technique. Integrating MEL with MLL and MRL, we present a 3-Dimensional Matryoshka Learning framework, providing end-to-end efficiency for the training, inference, and storage of embedding models. • We demonstrate that efficiency and inclusivity can drive superior performance. ML-Embed-8B establishes new state-of-the-art results on 9 of 17 MTEB benchmarks. Crucially, we achieve massive gains in historically underserved languages—such as a +22.89 point improvement on Polish and +6.88 on Vietnamese—proving that equitable performance need not come at the cost of efficiency.
Benchmark
Models
Multilingual English European Indic Scandinavian Chinese German French Japanese Korean Dutch Polish Russian Persian Vietnamese
146 154 122 111 41 43 90 111 11 95 87 1 95 22 17
trainable parameters or the overall training cost, making them ill-suited to data-constrained or compute-constrained training scenarios. In these settings, parameter-efficient finetuning methods such as LoRA (Hu et al., 2022) are used instead, which reduces the model’s trainable parameters by decomposing the update in weight matrices into low-rank ones. Numerous variants of LoRA have also been proposed, such as QLoRA (Dettmers et al., 2023) that combines LoRA and quantization, AdaLoRA (Zhang et al., 2023a) that adaptively allocates the parameter budget according to the importance of weight matrices, and RaSA (He et al., 2025) that partially shares LoRA parameters across model layers. Nevertheless, all these methods require the entire model to be loaded at inference time, limiting their utility for resourceconstrained deployment where reduced memory footprints are required.
• In a direct counter to the trend of closed development, we release our comprehensive multilingual dataset, all model weights, and training code, providing a fully reproducible blueprint for building globally equitable AI systems1 .
2. Related Work 2.1. Efficient Representation Learning Matryoshka Representation Learning (Kusupati et al., 2022) optimizes d-dimensional embeddings by applying loss functions at O(log(d)) embedding sizes, facilitating adaptive application in downstream tasks with varying dimension requirements. Recent extensions such as ESE (Li et al., 2025b) improve MRL by applying principal component analysis to condense more essential information into the initial embedding dimensions and model layers, while MatryoshkaAdaptor (Yoon et al., 2024) and SMEC (Zhang et al., 2025a) employ additional MLP layers to reduce the embeddings to lower dimensions. Other methods, such as Flextron (Cai et al., 2024) and MatFormer (Devvrit et al., 2024), also enable flexible model sizes by pruning attention heads or MLP dimensions at inference time. However, these methods often introduce structural modifications (e.g., routing mechanisms) that may reduce compatibility and complicate deployment.
2.2. Multilingual Embedding Models and Benchmarks The previous generation of encoder-based embedding models witnessed a proliferation of massively multilingual embedding models supporting hundreds of languages, represented by XLM-R (Conneau et al., 2020), mDeBERTaV3 (He et al., 2023), mBART (Liu et al., 2020), and mT5 (Xue et al., 2021). Recently, decoder-based embedding models have become the dominant paradigm, benefiting from their extensive capabilities acquired during largescale pre-training, as verified by state-of-the-art models such as E5-Mistral (Wang et al., 2024a), NV-Embed (Lee et al., 2025a), Qwen3-Embedding (Zhang et al., 2025b), and Gemini-Embedding (Lee et al., 2025b).
In terms of training, existing matryoshka optimization methods focus on representation flexibility but do not reduce 1 Code: https://github.com/codefuse-ai/ CodeFuse-Embeddings. Model and data: https: //huggingface.co/collections/codefuse-ai/ codefuse-embeddings.
However, this advancement has been accompanied by a shift toward English-centric evaluation. This is evidenced in MTEB (Muennighoff et al., 2023), which has been es2
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
tablished as one of the most recognized text embedding benchmarks, covering over 500 evaluation tasks and more than 250 languages (Enevoldsen et al., 2025). Yet, in reality, the MTEB leaderboards exhibit significant linguistic bias. For instance, in the MTEB-Multilingual benchmark, 35 out of the 131 tasks focus exclusively on English, potentially obscuring a model’s true multilingual efficacy. Furthermore, as illustrated in Table 1, many language-specific benchmarks receive disproportionately less attention compared with the English or Multilingual benchmarks2 .
0.6B (Yang et al., 2025), the embedding layer accounts for 1/4 of the total parameters. MEL addresses this by learning the embedding matrix in a factorized, low-rank form that is itself structured for nested training. Crucially, unlike low-rank update methods such as LoRA (Hu et al., 2022), MEL reduces not only trainable parameters, but also total parameters so that inference efficiency is also improved. Let the base model’s original embedding matrix be E ∈ Rv×dmodel , where v is the vocabulary size and dmodel is the model’s hidden dimension. Prior to fine-tuning, we initialize two smaller matrices, EA and EB , using a truncated Singular Value Decomposition (SVD) of E. We compute U, S, V T = SVD(E) and select the top-r singular values and vectors to form Ur ∈ Rv×r , Sr ∈ Rr×r , and VrT ∈ Rr×dmodel . The trainable matrices are then initialized as:
This disparity is exacerbated by the fact that many top-performing multilingual embedding models such as Qwen3-Embedding (Zhang et al., 2025b), Gemini-Embedding (Lee et al., 2025b), and EmbeddingGemma (Vera et al., 2025) - are either closed-source APIs or open-weight only without training transparency. KaLM-Embedding (Zhao et al., 2025) represents one of the few exceptions with transparency in training data, but focuses exclusively on the Multilingual leaderboard and is not evaluated on the aforementioned language-specific benchmarks that are critical for truly global applications.
EA ← Ur Sr ∈ Rv×r
EB ← VrT ∈ Rr×dmodel . (1) The full embedding matrix is approximated by their product, E ≈ EA EB . During fine-tuning, only EA and EB are updated instead of a full v×dmodel matrix, reducing trainable parameters and memory requirements.
3. Method: 3D Matryoshka Learning
and
To embed the Matryoshka principle, during each training forward pass, we dynamically sample a sub-rank r′ < r from a predefined set (e.g., {64, 128, 256, 512, 1024}). The forward pass then uses only the first r′ components of the factorized matrices:
Creating truly accessible and scalable embedding models requires tackling efficiency bottlenecks across the entire model lifecycle: from the high costs of training, to the computational demands of inference, and finally to the footprint of storage. To this end, we propose 3-Dimensional Matryoshka Learning (3D-ML), a unified framework that generalizes the principle of nested structures to provide comprehensive efficiency. 3D-ML simultaneously targets all three stages by optimizing along three corresponding axes: model parameters, computational depth, and representation size. This is achieved through a trio of integrated techniques: 1) Matryoshka Embedding Learning (MEL) reduces trainable and total parameters for efficient training and inference; 2) Matryoshka Layer Learning (MLL) enables flexible model depth for efficient inference; 3) Matryoshka Representation Learning (MRL) produces variable-size representation dimensions for efficient storage. Figure 1 provides a conceptual illustration of this framework.
Eeffective = EA [:, : r′ ]EB [: r′ , :].
(2)
This forces the model to prioritize the most critical information within the initial dimensions of the factorized space. At inference time, MEL offers two modes: • Compatibility Mode: We compute the final trained embedding matrix Etrained = EA EB . This results in a standard embedding layer, requiring no changes to existing inference infrastructure while still benefiting from the regularized training via low-rank factorization. • Efficiency Mode: For maximum resource savings, we can deploy the model with a highly compressed embedding layer. After training, we can either use the trained factorized matrices EA and EB directly (at rank r) or re-factorize the full matrix Etrained = EA EB to an even smaller rank r′ ≪ r for aggressive compres′ ′ sion. Let the new factorized matrices be EA ∈ Rv×r ′ ′ and EB ∈ Rr ×dmodel . This approach yields two key benefits. First, it drastically reduces the storage space: instead of storing a dense v × dmodel matrix, we only store v × r′ + r′ × dmodel parameters. For a large vocabulary V and a small rank r′ , this represents a
3.1. Matryoshka Embedding Learning (MEL) for Parameter Efficiency The embedding layer, which maps vocabulary tokens to dense vectors, often constitutes a disproportionately large share of a model’s parameters, especially in smaller models and multilingual models with a large vocabulary. For instance, in an embedding model trained from Qwen32 All references to MTEB leaderboards in this manuscript refer to the snapshot acquired on January 22nd, 2026.
3
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World 𝓛 𝓛 𝓛 𝓛
deploy a smaller, faster model by simply taking the first l layers of the full model, where l ∈ Llayers . This avoids the need for re-training or complex pruning, enabling seamless adaptation to varying computational budgets.
𝓛𝐌𝐑𝐋
Transformer Layer ······ 𝓛𝐌𝐑𝐋
Transformer Layer
3.3. Unifying the Framework with Matryoshka Representation Learning (MRL)
𝓛𝐌𝐋𝐋
The final dimension of our framework is Matryoshka Representation Learning (MRL; Kusupati et al., 2022), which optimizes embeddings for variable-dimension storage. MRL trains a model such that prefixes of the final embedding vector are themselves effective, lower-dimensional representations.
Transformer Layer 𝓛𝐌𝐑𝐋
Transformer Layer 𝓛𝐌𝐑𝐋
Transformer Layer MEL
U
In 3D-ML, MRL is not a separate step but is deeply integrated with MLL. For each selected MLL layer l ∈ Llayers , we apply a contrastive loss not just on the fulldimensional output representation, but on a nested set of its prefixes. Let Dmrl be the set of MRL dimensions (e.g., {8, 16, 32, . . . , dmodel }). Let projd (v) denote the projection of a vector v to its first d dimensions. The total loss for a given layer l is a sum of losses over these dimensions.
V Decomposed Embedding
Figure 1. The 3D-ML framework provides comprehensive efficiency by applying nested learning principles across model parameters (MEL), depth (MLL), and representation dimensions (MRL).
The unified 3D-ML objective function combines all three components. The total loss is summed over all selected MLL layers and all MRL dimensions. Let hl (q) and hl (d) be the hidden states from layer l for a query q and document d, respectively. The final representation for a given MRL dimension d′ is vl,d′ (·) = projd′ (LNfinal (hl (·))). The overall objective is: X X − n L3D-ML = cl,d′ Lcl (qi , d+ i , {di,j }j=1 ; vl,d′ ),
substantial reduction. Second, it can improve computational efficiency. A standard embedding lookup for a sequence of tokens involves gathering rows from the large v × dmodel matrix. With factorization, this becomes a two-step process: a fast lookup in the “tall′ and-skinny” matrix EA followed by a matrix multipli′ cation with the “short-and-wide” matrix EB . This is particularly advantageous for on-device deployment where memory is the primary constraint.
l∈Llayers d′ ∈Dmrl
3.2. Matryoshka Layer Learning (MLL) for Inference Efficiency
(3) where cl,d′ is the loss weight coefficient for layer l and dimension d′ , and the contrastive learning loss Lcl for a given representation function vl,d′ is defined as:
The computational cost of Transformer-based models scales with model depth. MLL is designed to produce models that can be dynamically and efficiently truncated to shallower depths without significant performance degradation.
es(vl,d′ (qi ),vl,d′ (di ))/τ . − log n − P + es(vl,d′ (qi ),vl,d′ (di ))/τ + es(vl,d′ (qi ),vl,d′ (di,j ))/τ
+
j=1
(4) Here, s(·, ·) is cosine similarity, τ is a temperature hyperpa− rameter, d+ i is a positive document for query qi , and {di,j } are hard negative documents for query qi . This unified loss ensures that the model learns representations that are simultaneously efficient in terms of parameters (via MEL), depth (via MLL), and storage (via MRL).
Instead of applying the training loss only at the final layer’s output, MLL applies it at multiple, pre-defined intermediate layers. Let Llayers = {l1 , l2 , . . . , lk , L} be a set of selected layer indices, where L is the index of the final layer. For our experiments, we use a logarithmically spaced set of layers (e.g., {1, 2, 4, 8, 16, 32}) plus the model’s final layer. For each layer l ∈ Llayers , we extract its hidden state output, hl . To maintain representational consistency across depths, we pass each hl through the model’s final layer normalization, LNfinal , before using it to compute the loss.
3.4. Practical Deployment and Compatibility A core design principle of the 3D-ML framework is its focus on practical deployment and compatibility with existing ecosystems, ensuring that its efficiency gains are not merely theoretical but easily accessible to practitioners. Each com-
This “early-exit” style training ensures that shallower versions of the model are also effective embedders. At inference time, this provides unparalleled flexibility: one can 4
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
ponent is designed for minimal friction:
Ours
• MRL (Storage): The benefits of Matryoshka Representation Learning are straightforward to leverage. Truncating the final embedding vectors is a simple postprocessing step that is natively supported in popular libraries like SENTENCE - TRANSFORMERS (Reimers & Gurevych, 2019) via a single parameter, requiring no changes to the codebase.
Multilingual (6.3%)
Chinese (44.4%)
KaLM-Embedding English (31.1%)
• MLL (Inference): Matryoshka Layer Learning is similarly user-friendly and fully compatible with the Hugging Face ecosystem. Deploying a faster, shallower model is as simple as modifying the num hidden layers parameter in the model’s configuration file, which drives the AutoModel class from TRANSFORMERS (Wolf et al., 2020) to load only the first n layers and ignore the remaining weights.
English (49.4%)
Italian (1.6%) Indonesian (1.6%) Hindi (1.8%) Japanese (1.9%) Vietnamese (2.0%) Dutch (2.1%) Arabic (2.5%) German (3.1%) Russian (3.7%) French (4.5%) Spanish (5.2%) Chinese (5.5%)
Figure 2. Comparison between the language distribution of our training data (outer circle) and KaLM-Embedding (inner circle). KaLM-Embedding’s data is only annotated with three labels, while ours are annotated with specific languages.
• MEL (Parameters): As described in Section 3.1, Matryoshka Embedding Learning offers a flexible tradeoff between convenience and maximum efficiency. In compatibility mode, the factorized matrices (EA , EB ) are multiplied into a standard embedding matrix before release, making the model indistinguishable from a standard Transformer decoder, ensuring seamless integration into any inference pipeline without code changes. For users prioritizing a minimal memory footprint, the efficiency mode allows for deploying the low-rank factorized matrices directly, enabling significant parameter reduction with only minor adjustments to the modeling file.
of code, aims to build a model with truly global utility and directly contrasts recent open-source datasets such as that released by KaLM-Embedding, which is heavily skewed towards English and Chinese (Figure 2). We provide a more comprehensive linguistic breakdown of our dataset in Appendix A. The functional diversity of our dataset is equally critical for training a general-purpose embedding model. As shown in Figure 7 in Appendix A, our collection encompasses a wide spectrum of tasks, ranging from retrieval-focused question answering and bitext mining to classification-oriented sentiment analysis and intent/domain classification.
This comprehensive focus on deployability makes 3D-ML a practical blueprint for building and sharing highly efficient models, lowering the barrier to entry for a wide range of users, from large-scale production systems to resourceconstrained research environments.
To leverage this heterogeneity within a unified contrastive learning framework, we follow prior work (Lee et al., 2025a; Zhang et al., 2025c) and consolidate all data into three canonical formats: retrieval, clustering, and two-way classification. This consolidation allows the model to learn a versatile embedding space by optimizing a single, consistent objective across disparate data sources and task structures. For the retrieval format, data consists of (query, positive document, hard negatives) tuples. We leverage both in-batch negatives, where other documents in a mini-batch serve as negatives, and explicitly provided hard negatives (mined using Qwen3-Embedding-8B) to create a challenging and efficient training signal. For the clustering format, which also ingests multi-class classification tasks, tuples are formed by sampling an anchor, a positive example from the same class, and a hard negative from a different class. Finally, the two-way classification format directly uses class labels, where a given text serves as the anchor, the corresponding label text is the positive, and the opposite label text is the negative. For both clustering and classification, only hard negatives are utilized to avoid introducing false negatives from in-batch samples.
4. Training Data A cornerstone of our work is the compilation of a vast and diverse training corpus designed to foster both linguistic inclusivity and broad task competency. We aggregate data from 121 publicly available sources, creating a collection of 50 million training samples that span 282 natural languages (as identified by ISO-639-3 codes) and over 40 programming languages. Crucially, our data curation process is driven by real-world data availability rather than optimizing for specific benchmarks. For instance, our dataset contains substantial data for Spanish and Arabic, which are the 3rd and 7th most represented languages in our corpus (Figure 3), despite these languages lacking dedicated benchmarks in MTEB (see Table 1). This approach, which also includes a long tail of low-resource languages and a significant volume 5
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Natural Language Programming Language
107
Data Size
106
105
eng zho spa fra python rus deu ara nld vie jpn hin ind ita por pol tur php tha java kor cpp go ukr ces tgl fas cat glg mya hye khm nep eus swe javascript lao swa dan ell azj sin tgk bul ron fin hun slv heb lav urd nor lit slk c# est msa ben aze afr tam kat tel mal ruby mon nno kaz cym mar sqi nob pus isl hrv mkd hbs ceb jav srp war kan epo c lat guj uzb amh oci bel azb kir mlg vol ast pan ltz nds hat bre gle sco xho tat bos yor min che arz rust
104
Figure 3. Top-100 natural languages and top-10 programming languages in our training data.
To maximize the utility of this diverse corpus, we adopt a two-stage training strategy following previous works (Lee et al., 2025a; Zhang et al., 2025b). The first stage focuses on building a robust semantic foundation by training on a large-scale subset of retrieval datasets, totaling 27 million samples. This phase imbues the model with a strong general understanding of semantic similarity. In the second stage, we conduct fine-tuning on a sampled mixture of 8.3 million samples from all data sources, applying task-specific instructions to the queries. This stage sharpens the model’s ability to handle the nuances of diverse downstream applications like classification, reranking, and paraphrase detection.
and paraphrase detection. More details about the amount of training data and hyperparameter settings are given in Appendix C, where we also demonstrate the empirical benefits of the two-stage training design. Evaluation We evaluate the models on 17 MTEB benchmarks: Multilingual (Enevoldsen et al., 2025), English (Enevoldsen et al., 2025), Code (Enevoldsen et al., 2025), Medical, European (Enevoldsen et al., 2025), Scandinavian (Enevoldsen et al., 2024), Indic (Enevoldsen et al., 2025), German (Wehrli et al., 2023), French (Ciancone et al., 2024), Korean, Polish (Poswiata et al., 2024), Chinese (Xiao et al., 2023), Japanese (Li et al., 2026), Dutch (Banar et al., 2025), Russian (Snegirev et al., 2025), Persian (Zinvandi et al., 2025), and Vietnamese (Pham et al., 2026), totaling 430 tasks across ten types: retrieval, reranking, classification, clustering, pair classification, multilabel classification, STS, instruction reranking, bitext mining, and summarization. More details on these benchmarks and tasks are given in Appendix B. For comparison, we report the previous top1 score and top-5 score on each benchmark’s leaderboard. We also compare with individual models, specifically those from the Qwen3-Embedding (Zhang et al., 2025b) and EmbeddingGemma (Vera et al., 2025) families.
5. Experiments 5.1. Experimental Setting Model We present a series of 6 models, all trained on identical data in exactly the same order: 140M, 330M, 600M, 1.7B, 4B, 8B. All models are fine-tuned from Qwen3 causal LLMs (Yang et al., 2025), where 600M, 1.7B, 4B, and 8B models correspond to models of the same size in the Qwen3 family, while 140M and 330M models are pruned from the 600M model after training. Following existing embedding models based on the Qwen3 family (Zhang et al., 2025b;c), we maintain the causal attention in the models and use EOS token representation as the sequence embedding.
5.2. MTEB Results We present the main results in Table 2, comparing the MLEmbed family against the top-performing models on 17 MTEB benchmarks. Our largest model, ML-Embed-8B, establishes new state-of-the-art (SOTA) scores on a remarkable 9 out of 17 benchmarks, demonstrating the effectiveness of our multilingual training corpus.
Training Inspired by NV-Embed (Lee et al., 2025a) and Qwen3-Embedding (Zhang et al., 2025b), we train the models in two stages. In the first stage, we use only some of the largest retrieval datasets, and do not apply instructions to the data, aiming to inject the causal models with basic capabilities of converting texts into semantic embeddings that are usable by downstream tasks. In the second stage, we sample from all data sources, and apply instructions to queries, enabling a more nuanced understanding of different semantic representation tasks such as retrieval, classification,
Critically, these SOTA results are concentrated in benchmarks for languages historically underserved by the research community, directly addressing the linguistic challenge outlined in Section 1. For instance, on the Polish benchmark, 6
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World Table 2. Comparison of our models against previous top-1 and top-5 performance on 17 MTEB benchmark leaderboards. The number of tasks in each benchmark is given in (superscript) . The specific metrics are consistent with the main metrics used by MTEB (e.g., NDCG@10 for retrieval tasks and accuracy for classification tasks). Model
Multi.(131)
English(41)
Code(12)
Medical(12)
European(73)
Top-1 Top-5
72.32 69.45
75.97 74.61
80.75 76.00
Top results on the leaderboard 66.55 63.60 63.83 62.32
8B 4B 1.7B 0.6B 330M 140M
66.79 65.80 63.70 61.30 58.06 49.81
73.26 72.89 71.19 70.01 67.99 60.35
Ours 80.28 62.91 68.00 69.49 79.95 62.05 67.53 67.88 78.90 59.88 65.47 65.79 77.27 57.58 63.40 62.00 73.94 55.91 59.80 59.22 61.78 48.28 50.40 47.95 Results continued for remaining languages and average
Model
Korean(6)
Polish(17)
Chinese(32)
Top-1 Top-5
77.01 69.06
50.95 n.a.
78.52 74.87
Top results on the leaderboard 73.18 58.38 69.51 56.23
8B 4B 1.7B 0.6B 330M 140M
74.84 73.32 72.33 68.74 61.71 53.07
73.84 73.14 71.12 68.13 65.65 51.22
67.22 66.55 64.87 62.64 59.38 51.02
77.81 76.65 74.53 71.30 66.00 55.97
Japan.(28)
Dutch(40)
Ours 62.64 61.42 59.83 56.59 54.04 43.89
our model achieves a score of 73.84, a staggering +22.89 point improvement over the previous best model. Similarly, we set new records on Vietnamese (+6.88), Indic (+6.61), German (+6.47), Japanese (+4.63), Dutch (+4.26), French (+1.54), and the aggregated benchmarks of Scandinavian (+3.93) and European (+4.40). This demonstrates that our data curation and training methodology successfully produce models with globally equitable performance.
Scan.(28)
Indic(20)
German(19)
French(25)
65.56 62.01
70.15 67.39
59.96 55.72
70.37 67.25
76.76 75.15 72.58 66.11 61.87 46.37
66.43 65.49 63.99 61.58 58.48 46.66
71.91 70.97 68.94 66.64 63.59 52.83
Persian(52)
Viet.(50)
Average
74.16 69.39
71.58 65.26
54.74 52.37
68.46 65.95
69.24 67.93 67.08 63.35 60.12 47.24
71.12 69.94 68.35 65.32 60.16 52.50
61.62 61.20 60.27 57.85 54.17 43.67
70.24 69.29 67.58 64.69 61.18 50.76
Russian(23)
French, Korean, Polish, Japanese, Dutch, and Vietnamese benchmarks, while underperforming on English, Chinese, and Multilingual benchmarks. 5.4. Ablation Studies To dissect the individual contributions of our proposed efficiency methods, we conduct a series of ablation studies on the English subset using the 0.6B model.
On the highly competitive English and Multilingual benchmarks, our models also perform comparably to the top-5 models on the leaderboard, validating our approach as a strong foundation for general-purpose embeddings. Furthermore, the results exhibit a clear and consistent scaling trend: performance reliably improves with model size across all benchmarks. This indicates that our training recipe is robust and provides a scalable blueprint for developing even more powerful models in the future.
MLL and MEL Synergy First, we investigate the interplay between Matryoshka Layer Learning (MLL) and Matryoshka Embedding Learning (MEL). As shown in Figure 4, we compare four settings: (1) a single baseline, evaluated at different exit layers; (2) baselines of individually trained models at different depths; (3) a single model trained with MLL and evaluated at different exit layers; and (4) a fourth model combining MLL and MEL. MLL alone presents a classic trade-off: it enables training a single, depth-flexible model for the computational cost of one, but the resulting shallower models slightly underperform individually trained counterparts. However, the introduction of MEL dramatically alters this dynamic. By significantly reducing the parameter count of the embedding layer, MEL allows for a much deeper model at the same parameter budget. For example, our MLL+MEL model with 4 layers has the same parameter count ( 170M) as a 1-layer baseline model but achieves a 15-point higher score. At equivalent performance levels, the MLL+MEL model is 3x smaller, confirming the
5.3. Comparison with Similar-Sized Models In Table 3, we compare our 0.3B and 0.6B models with EmbeddingGemma-0.3B and Qwen3-Embedding-0.6B, respectively. For these two models, we use their public results in the MTEB repository when avaiable, and evaluate the remaining tasks using the same prompts as those used for evaluating our models. The results are similar to those on the MTEB leaderboard: our models demonstrate superior performance on the less-attended Code, Scandinavian, German, 7
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World Table 3. Comparison of our models with EmbeddingGemma and Qwen3-Embedding. The number of tasks in each benchmark is given in (superscript) . Multi.(131)
English(41)
Code(12)
Medical(12)
EmbedGemma Ours
61.15 58.06
69.67 67.99
68.76 73.94
51.24 55.91
Qwen3-Embed Ours
64.34 61.30
70.47 70.01
Korean(6)
Polish(17)
Chinese(32)
Japan.(28)
EmbedGemma Ours
58.24 61.71
64.70 65.65
50.40 59.38
60.82 66.00
Qwen3-Embed Ours
65.29 68.74
67.42 68.13
66.71 62.64
67.28 71.30
Model
European(73)
Scan.(28)
Indic(20)
German(19)
French(25)
62.50 59.80
54.39 59.22
66.11 61.87
56.28 58.48
61.90 63.59
0.6B 75.42 60.16 63.91 60.99 57.58 63.40 62.00 77.27 Results continued for remaining languages and average
66.53 66.11
59.45 61.58
63.01 66.64
0.3B
Model
Dutch(40)
Russian(23)
Persian(52)
Viet.(50)
Avg.
50.98 54.04
64.57 60.12
67.11 60.16
43.45 54.17
59.55 61.18
54.27 56.59
64.20 63.35
62.88 65.32
56.01 57.85
64.02 64.69
0.3B
0.6B
70
70
+15 performance
65
60
60 55
50
3x size reduction
40
50 45
Baseline (single) Baseline (individual) MLL MLL+MEL
30
20 0
100
200
300
Param (M)
400
500
40
Baseline Decomposition-only MEL
35 30 200
600
400
600
Embedding Matrix Rank
800
1000
Figure 5. Ablation results of decomposing the embedding layer to varying ranks at inference time. Baseline model is trained with the original embedding. Decomposition-only model is trained with decomposed embedding (at rank 512). MEL model is trained with decomposed embedding plus Matryoshka Embedding Learning.
Figure 4. Ablation results of models pruned to different depths on MTEB-English. Each point on the baseline (individual) curve represents an individual trained model, while points on the Baseline (single), MLL, and MLL+MEL curves are models of different depths pruned from a single trained model.
to concentrate information on the leading ranks. Notably, it achieves almost identical performance to the baseline (69.60) at rank 512, demonstrating the redundancy in the embedding matrix. The MEL-trained model demonstrates superior robustness against decomposition, declining much more slowly as the rank diminishes, retaining a strong score of 64.30 even when compressed to a rank of just 64. This confirms that MEL is highly effective at producing models that are robust to aggressive, post-hoc compression.
powerful synergy between these two techniques for creating parameter-efficient models. Robustness of MEL Next, we isolate the effect of MEL on inference-time compression. We compare a baseline model against two variants: one trained with a factorized embedding layer at rank 512 (“Decomposition-only”) and another additionally trained with the nested rank objective of MEL. At inference, we apply SVD to each model’s embedding matrix and evaluate performance at progressively smaller ranks. The results in Figure 5 are stark. The baseline model is extremely brittle - its performance collapses catastrophically (from 69.68 to 53.25) with even minor rank reduction (from 1024 to 960). The decomposition-only model is more robust, as the low-rank structure acts as a training regularizer even though the model is not trained
Data Comparison To isolate the impact of our curated training corpus, we conduct a head-to-head comparison against the recently released KaLM-Embedding finetuning data (Zhao et al., 2025). Starting from the same stage-1 checkpoint, we finetune two identical 0.6B models using our stage-2 data and the similar-sized KaLM-Embedding data, respectively. The results, shown in Figure 6, demon8
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World Table 4. Comparison of baseline training, pruned model training, and 3D-ML on EuroBERT backbone.
Model
Multi.
English
Baseline (210M) Pruned (120M) 3D-ML (120M)
59.86 46.08 56.94
Model Baseline (210M) Pruned (120M) 3D-ML (120M)
Indic
German
French
68.52 64.03 55.93 59.83 54.43 54.65 35.44 39.52 44.79 40.75 66.13 59.76 51.83 56.75 51.36 Results continued for remaining languages and average
55.36 42.90 50.53
62.90 47.69 59.84
62.90 47.94 58.96
Korean
Polish
Chinese
Japan.
Dutch
Russian
Persian
Viet.
Avg.
61.01 42.44 58.27
64.63 47.73 59.13
59.37 43.48 56.18
71.82 52.86 67.56
52.26 35.39 47.84
58.44 42.29 54.50
60.00 48.30 58.24
55.22 37.48 51.24
60.38 44.10 56.77
French KaLM Ours English
German
Code
Medical
Persian European Scandinavian
In Appendix D, we provide additional experiments including 1) applying 3D-ML to training on specific languages, 2) applying MEL to language modeling, and 3) in-depth analysis of efficiency gains.
Multilingual 60
Scan.
to 120M parameters followed by finetuning, and 3) 3D-ML training followed by structural pruning to 120M parameters. The results in Table 4 show that training with 3D-ML leads to a minimal performance drop compared with naive structural pruning, demonstrating the generalizability of 3D-ML.
Japanese
40
European
Code
Medical Indic Korean
6. Conclusion
Vietnamese Dutch Polish
Russian
This paper addresses the critical challenges of computational cost, linguistic bias, and lack of transparency hindering the development of text embeddings. We introduce MLEmbed, a family of models built using our 3-Dimensional Matryoshka Learning (3D-ML) framework. 3D-ML integrates Matryoshka Embedding (MEL), Layer (MLL), and Representation (MRL) Learning to achieve comprehensive efficiency across the model lifecycle. Paired with a newly curated, massively multilingual open-source dataset, our approach demonstrates that efficiency and inclusivity can yield state-of-the-art performance. Our 8B model sets new records on 9 of 17 MTEB benchmarks, with dramatic improvements in historically understudied languages such as Polish (+22.89) and Vietnamese (+6.88). Ablation studies confirm the synergistic benefits of 3D-ML components, creating powerful models adaptable to diverse computational budgets. The studies also highlight the superiority of our curated data.
Chinese
Figure 6. Comparison between 0.6B models trained on our data and KaLM-Embedding data.
strate the distinct advantages of our data curation strategy. Our model achieves superior performance on 9 out of 17 benchmarks, including the composite Multilingual, European, and Scandinavian sets, as well as English, French, German, Japanese, and Persian. The most significant lead is on the Code benchmark, highlighting our data’s wider domain coverage. While the KaLM-Embedding data produces a stronger model for Chinese - an expected outcome given its heavy concentration on Chinese and English data (Figure 2) - our dataset achieves on-par performance across seven other benchmarks, including Korean, Polish, Dutch, Indic, Russian, and Vietnamese. This outcome confirms that our focus on linguistic diversity yields a more globally robust model, trading hyper-specialization in a single language for broader competence.
By releasing our models, data, and code, we offer a reproducible blueprint for building globally equitable and efficient AI systems, dismantling the transparency barrier. Our work paves the way for future research in scaling these techniques to larger models, expanding linguistic coverage, and deploying powerful embeddings on resource-constrained devices. We hope this work steers the field toward a more inclusive and accessible future for text representation learning.
Generalization to Different Backbones To verify the effectiveness of 3D-ML, we conduct additional experiments using EuroBERT-210M (Boizard et al., 2025) as the backbone. We train three models on the stage-2 data described in Section 4: 1) baseline finetuning, 2) structural pruning 9
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Impact Statement
D., Chai, J., Schluter, N., and Tetreault, J. R. (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pp. 4555–4567. Association for Computational Linguistics, 2020. doi: 10.18653/V1/ 2020.ACL-MAIN.417. URL https://doi.org/10. 18653/v1/2020.acl-main.417.
This paper presents work whose goal is to advance the field of Machine Learning. There are many potential societal consequences of our work, none which we feel must be specifically highlighted here.
References
Barmina, G., Norman, N. C. H., Schneider-Kamp, P., and Poech, L. G. Dala: Danish linguistic acceptability evaluation guided by real world errors. CoRR, abs/2512.04799, 2025. doi: 10.48550/ARXIV.2512.04799. URL https: //doi.org/10.48550/arXiv.2512.04799.
Agirre, E., Cer, D. M., Diab, M. T., and Gonzalez-Agirre, A. Semeval-2012 task 6: A pilot on semantic textual similarity. In Agirre, E., Bos, J., and Diab, M. T. (eds.), Proceedings of the 6th International Workshop on Semantic Evaluation, SemEval@NAACL-HLT 2012, Montréal, Canada, June 7-8, 2012, pp. 385–393. The Association for Computer Linguistics, 2012. URL https: //aclanthology.org/S12-1051/.
Blinov, P. Medical qa ru data, 2021. https://huggingface.co/datasets/ blinoff/medical_qa_ru_data.
Ahmad, W. U., Majumdar, S., Ficek, A., Narenthiran, S., Samadi, M., Huang, J., Jain, S., Noroozi, V., and Ginsburg, B. Opencodereasoning-ii: A simple test time scaling approach via self-critique. CoRR, abs/2507.09075, 2025. doi: 10.48550/ARXIV.2507.09075. URL https: //doi.org/10.48550/arXiv.2507.09075. Altaf, M. Medical instruction 120k, 2023. URL https://huggingface. co/datasets/Mohammed-Altaf/ medical-instruction-120k. Bai, Y., Du, X., Liang, Y., Jin, L., Zhou, J., Liu, Z., Fang, F., Chang, M., Zheng, T., Zhang, X., Ma, N., Wang, Z. M., Yuan, R., Wu, H., Lin, H., Huang, W., Zhang, J., Lin, C., Fu, J., Yang, M., Ni, S., and Zhang, G. COIG-CQIA: quality is all you need for chinese instruction fine-tuning. In Chiruzzo, L., Ritter, A., and Wang, L. (eds.), Findings of the Association for Computational Linguistics: NAACL 2025, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, pp. 8190–8205. Association for Computational Linguistics, 2025. doi: 10.18653/V1/2025. FINDINGS-NAACL.457. URL https://doi.org/ 10.18653/v1/2025.findings-naacl.457.
URL
Boizard, N., Gisserot-Boukhlef, H., Alves, D. M., Martins, A. F. T., Hammal, A., Corro, C., Hudelot, C., Malherbe, E., Malaboeuf, E., Jourdan, F., Hautreux, G., Alves, J., Haddad, K. E., Faysse, M., Peyrard, M., Guerreiro, N. M., Fernandes, P., Rei, R., and Colombo, P. Eurobert: Scaling multilingual encoders for european languages. CoRR, abs/2503.05500, 2025. doi: 10.48550/ARXIV.2503.05500. URL https://doi. org/10.48550/arXiv.2503.05500. Bonifacio, L. H., Campiotti, I., Lotufo, R. A., and Nogueira, R. mmarco: A multilingual version of MS MARCO passage ranking dataset. CoRR, abs/2108.13897, 2021. URL https://arxiv.org/abs/2108.13897. Boteva, V., Ghalandari, D. G., Sokolov, A., and Riezler, S. A full-text learning to rank dataset for medical information retrieval. In Ferro, N., Crestani, F., Moens, M., Mothe, J., Silvestri, F., Nunzio, G. M. D., Hauff, C., and Silvello, G. (eds.), Advances in Information Retrieval - 38th European Conference on IR Research, ECIR 2016, Padua, Italy, March 20-23, 2016. Proceedings, volume 9626 of Lecture Notes in Computer Science, pp. 716–722. Springer, 2016. doi: 10.1007/978-3-319-30671-1\ 58. URL https:// doi.org/10.1007/978-3-319-30671-1_58.
Banar, N., Lotfi, E., Nooten, J. V., Arhiliuc, C., Kliocaite, M., and Daelemans, W. MTEB-NL and E5NL: embedding benchmark and models for dutch. CoRR, abs/2509.12340, 2025. doi: 10.48550/ARXIV. 2509.12340. URL https://doi.org/10.48550/ arXiv.2509.12340.
Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D. A large annotated corpus for learning natural language inference. In Màrquez, L., Callison-Burch, C., Su, J., Pighin, D., and Marton, Y. (eds.), Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, EMNLP 2015, Lisbon, Portugal, September 17-21, 2015, pp. 632–642. The Association for Computational Linguistics, 2015. doi: 10.18653/V1/D15-1075. URL https://doi.org/10.18653/v1/d15-1075.
Bañón, M., Chen, P., Haddow, B., Heafield, K., Hoang, H., Esplà-Gomis, M., Forcada, M. L., Kamran, A., Kirefu, F., Koehn, P., Ortiz-Rojas, S., Sempere, L. P., Ramı́rezSánchez, G., Sarrı́as, E., Strelec, M., Thompson, B., Waites, W., Wiggins, D., and Zaragoza, J. Paracrawl: Web-scale acquisition of parallel corpora. In Jurafsky,
Cai, R., Muralidharan, S., Heinrich, G., Yin, H., Wang, Z., Kautz, J., and Molchanov, P. Flextron: Many-inone flexible large language model. In Forty-first In10
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
2020.ACL-MAIN.207. URL https://doi.org/10. 18653/v1/2020.acl-main.207.
ternational Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net, 2024. URL https://openreview.net/forum? id=9vKRhnflAs.
Conneau, A., Rinott, R., Lample, G., Williams, A., Bowman, S. R., Schwenk, H., and Stoyanov, V. XNLI: evaluating cross-lingual sentence representations. In Riloff, E., Chiang, D., Hockenmaier, J., and Tsujii, J. (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pp. 2475–2485. Association for Computational Linguistics, 2018. doi: 10.18653/V1/D18-1269. URL https://doi.org/ 10.18653/v1/d18-1269.
Casanueva, I., Temcinas, T., Gerz, D., Henderson, M., and Vulic, I. Efficient intent detection with dual sentence encoders. CoRR, abs/2003.04807, 2020. URL https: //arxiv.org/abs/2003.04807. Chen, J., Xiao, S., Zhang, P., Luo, K., Lian, D., and Liu, Z. BGE m3-embedding: Multi-lingual, multifunctionality, multi-granularity text embeddings through self-knowledge distillation. CoRR, abs/2402.03216, 2024. doi: 10.48550/ARXIV.2402.03216. URL https:// doi.org/10.48550/arXiv.2402.03216.
Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V. Unsupervised crosslingual representation learning at scale. In Jurafsky, D., Chai, J., Schluter, N., and Tetreault, J. R. (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pp. 8440–8451. Association for Computational Linguistics, 2020. doi: 10.18653/V1/ 2020.ACL-MAIN.747. URL https://doi.org/10. 18653/v1/2020.acl-main.747.
Chen, X., Zeynali, A., Camargo, C. Q., Flöck, F., Gaffney, D., Grabowicz, P. A., Hale, S., Jurgens, D., and Samory, M. Semeval-2022 task 8: Multilingual news article similarity. In Emerson, G., Schluter, N., Stanovsky, G., Kumar, R., Palmer, A., Schneider, N., Singh, S., and Ratan, S. (eds.), Proceedings of the 16th International Workshop on Semantic Evaluation, SemEval@NAACL 2022, Seattle, Washington, United States, July 14-15, 2022, pp. 1094–1106. Association for Computational Linguistics, 2022. doi: 10.18653/V1/2022.SEMEVAL-1. 155. URL https://doi.org/10.18653/v1/ 2022.semeval-1.155.
Dao, T. Flashattention-2: Faster attention with better parallelism and work partitioning. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URL https://openreview.net/forum? id=mZn2Xyh9Ec.
Ciancone, M., Kerboua, I., Schaeffer, M., and Siblini, W. Mteb-french: Resources for french sentence embedding evaluation and analysis. CoRR, abs/2405.20468, 2024. doi: 10.48550/ARXIV.2405.20468. URL https:// doi.org/10.48550/arXiv.2405.20468.
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. Qlora: Efficient finetuning of quantized llms. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Advances in cjadams, Borkan, D., inversion, Sorensen, J., Dixon, Neural Information Processing Systems 36: Annual L., Vasserman, L., and nithum. Jigsaw uninConference on Neural Information Processing Systems tended bias in toxicity classification, 2019. URL 2023, NeurIPS 2023, New Orleans, LA, USA, Decemhttps://kaggle.com/competitions/ ber 10 16, 2023, 2023. URL http://papers. jigsaw-unintended-bias-in-toxicity-classification. nips.cc/paper_files/paper/2023/hash/ Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., 1feb87871436031bdc0f2beaa62a049b-Abstract-Conferen Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, html. R., Hesse, C., and Schulman, J. Training verifiers to solve Devvrit, Kudugunta, S., Kusupati, A., Dettmers, T., math word problems. CoRR, abs/2110.14168, 2021. URL Chen, K., Dhillon, I. S., Tsvetkov, Y., Hajishirzi, H., https://arxiv.org/abs/2110.14168. Kakade, S. M., Farhadi, A., and Jain, P. Matformer: Cohan, A., Feldman, S., Beltagy, I., Downey, D., and Weld, Nested transformer for elastic inference. In Globersons, D. S. SPECTER: document-level representation learnA., Mackey, L., Belgrave, D., Fan, A., Paquet, U., ing using citation-informed transformers. In Jurafsky, Tomczak, J. M., and Zhang, C. (eds.), Advances in D., Chai, J., Schluter, N., and Tetreault, J. R. (eds.), Neural Information Processing Systems 38: Annual Proceedings of the 58th Annual Meeting of the AssoConference on Neural Information Processing Systems ciation for Computational Linguistics, ACL 2020, On2024, NeurIPS 2024, Vancouver, BC, Canada, Decemline, July 5-10, 2020, pp. 2270–2282. Association for ber 10 - 15, 2024, 2024. URL http://papers. Computational Linguistics, 2020. doi: 10.18653/V1/ nips.cc/paper_files/paper/2024/hash/ 11
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
fe066022bab2a6c6a3c57032a1623c70-Abstract-Conference. Conference on Empirical Methods in Natural Language html. Processing, EMNLP 2013, 18-21 October 2013, Grand Hyatt Seattle, Seattle, Washington, USA, A meeting of Dinzinger, M., Caspari, L., Dastidar, K. G., Mitrovic, J., SIGDAT, a Special Interest Group of the ACL, pp. 1481– and Granitzer, M. Webfaq: A multilingual collection 1491. ACL, 2013. doi: 10.18653/V1/D13-1155. URL of natural q&a datasets for dense retrieval. In Ferro, https://doi.org/10.18653/v1/d13-1155. N., Maistro, M., Pasi, G., Alonso, O., Trotman, A., and Verberne, S. (eds.), Proceedings of the 48th InterFitzGerald, J., Hench, C., Peris, C., Mackie, S., Rottmann, national ACM SIGIR Conference on Research and DevelK., Sanchez, A., Nash, A., Urbach, L., Kakarala, V., opment in Information Retrieval, SIGIR 2025, Padua, Singh, R., Ranganath, S., Crist, L., Britan, M., Leeuwis, Italy, July 13-18, 2025, pp. 3802–3811. ACM, 2025. W., Tür, G., and Natarajan, P. MASSIVE: A 1m-example doi: 10.1145/3726302.3731934. URL https://doi. multilingual natural language understanding dataset with org/10.1145/3726302.3731934. 51 typologically-diverse languages. In Rogers, A., BoydGraber, J. L., and Okazaki, N. (eds.), Proceedings of the Du, L., Zhao, H., Ju, Y., and Pan, T. Scaling towards the 61st Annual Meeting of the Association for Computainformation boundary of instruction set: Infinityinstructtional Linguistics (Volume 1: Long Papers), ACL 2023, subject technical report. CoRR, abs/2507.06968, 2025. Toronto, Canada, July 9-14, 2023, pp. 4277–4302. Asdoi: 10.48550/ARXIV.2507.06968. URL https:// sociation for Computational Linguistics, 2023. doi: 10. doi.org/10.48550/arXiv.2507.06968. 18653/V1/2023.ACL-LONG.235. URL https://doi. Enevoldsen, K. C., Kardos, M., Muennighoff, N., and org/10.18653/v1/2023.acl-long.235. Nielbo, K. L. The scandinavian embedding benchGao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., marks: Comprehensive assessment of multilingual Dai, Y., Sun, J., Guo, Q., Wang, M., and Wang, H. and monolingual text embedding. In Globersons, Retrieval-augmented generation for large language modA., Mackey, L., Belgrave, D., Fan, A., Paquet, U., els: A survey. CoRR, abs/2312.10997, 2023. doi: Tomczak, J. M., and Zhang, C. (eds.), Advances in 10.48550/ARXIV.2312.10997. URL https://doi. Neural Information Processing Systems 38: Annual org/10.48550/arXiv.2312.10997. Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, DecemGeigle, G., Reimers, N., Rücklé, A., and Gurevych, I. ber 10 - 15, 2024, 2024. URL http://papers. TWEAC: transformer with extendable QA agent clasnips.cc/paper_files/paper/2024/hash/ sifiers. CoRR, abs/2104.07081, 2021. URL https: 4746bb91bd073ec7eef930d5775122ba-Abstract-Datasets_ //arxiv.org/abs/2104.07081. and_Benchmarks_Track.html. Gupta, M., Kulkarni, N., Chanda, R., Rayasam, A., and Enevoldsen, K. C., Chung, I., Kerboua, I., Kardos, M., Lipton, Z. C. Amazonqa: A review-based question Mathur, A., Stap, D., Gala, J., Siblini, W., Krzeminski, answering task. In Kraus, S. (ed.), Proceedings of D., Winata, G. I., Sturua, S., Utpala, S., Ciancone, M., the Twenty-Eighth International Joint Conference on Schaeffer, M., Misra, D., Dhakal, S., Rystrøm, J., SoloArtificial Intelligence, IJCAI 2019, Macao, China, Aumatin, R., Çagatan, Ö. V., Kundu, A., and et al. MMTEB: gust 10-16, 2019, pp. 4996–5002. ijcai.org, 2019. doi: massive multilingual text embedding benchmark. In The 10.24963/IJCAI.2019/694. URL https://doi.org/ Thirteenth International Conference on Learning Rep10.24963/ijcai.2019/694. resentations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URL https://openreview. net/forum?id=zl3pfz4VCV. Fan, A., Jernite, Y., Perez, E., Grangier, D., Weston, J., and Auli, M. ELI5: long form question answering. In Korhonen, A., Traum, D. R., and Màrquez, L. (eds.), Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pp. 3558–3567. Association for Computational Linguistics, 2019. doi: 10.18653/V1/P19-1346. URL https:// doi.org/10.18653/v1/p19-1346.
He, P., Gao, J., and Chen, W. Debertav3: Improving deberta using electra-style pre-training with gradientdisentangled embedding sharing. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. URL https://openreview.net/forum? id=sE7-XhLxHA. He, W., Liu, K., Liu, J., Lyu, Y., Zhao, S., Xiao, X., Liu, Y., Wang, Y., Wu, H., She, Q., Liu, X., Wu, T., and Wang, H. Dureader: a chinese machine reading comprehension dataset from real-world applications. In Choi, E., Seo, M., Chen, D., Jia, R., and Berant, J. (eds.), Proceedings of the Workshop on Machine Reading for
Filippova, K. and Altun, Y. Overcoming the lack of parallel data in sentence compression. In Proceedings of the 2013 12
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
V1/2021.ACL-LONG.442. URL https://doi.org/ 10.18653/v1/2021.acl-long.442.
Question Answering@ACL 2018, Melbourne, Australia, July 19, 2018, pp. 37–46. Association for Computational Linguistics, 2018. doi: 10.18653/V1/W18-2605. URL https://aclanthology.org/W18-2605/.
Husain, H., Wu, H., Gazit, T., Allamanis, M., and Brockschmidt, M. Codesearchnet challenge: Evaluating the state of semantic code search. CoRR, abs/1909.09436, 2019. URL http://arxiv.org/ abs/1909.09436.
He, Z., Tu, Z., Wang, X., Chen, X., Wang, Z., Xu, J., Liang, T., Jiao, W., Zhang, Z., and Wang, R. Rasa: Ranksharing low-rank adaptation. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URL https://openreview.net/forum? id=GdXI5zCoAt.
Jin, Q., Dhingra, B., Liu, Z., Cohen, W. W., and Lu, X. Pubmedqa: A dataset for biomedical research question answering. In Inui, K., Jiang, J., Ng, V., and Wan, X. (eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pp. 2567–2577. Association for Computational Linguistics, 2019. doi: 10.18653/V1/D19-1259. URL https://doi.org/10.18653/v1/D19-1259.
Hermann, K. M., Kociský, T., Grefenstette, E., Espeholt, L., Kay, W., Suleyman, M., and Blunsom, P. Teaching machines to read and comprehend. In Cortes, C., Lawrence, N. D., Lee, D. D., Sugiyama, M., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L. Trivi7-12, 2015, Montreal, Quebec, Canada, pp. 1693– aqa: A large scale distantly supervised challenge dataset 1701, 2015. URL https://proceedings. for reading comprehension. In Barzilay, R. and Kan, M. neurips.cc/paper/2015/hash/ (eds.), Proceedings of the 55th Annual Meeting of the afdec7005cc9f14302cd0474fd0f3c96-Abstract. Association for Computational Linguistics, ACL 2017, html. Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, pp. 1601–1611. Association for Computational Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, Linguistics, 2017. doi: 10.18653/V1/P17-1147. URL S., Wang, L., and Chen, W. Lora: Low-rank adaptahttps://doi.org/10.18653/v1/P17-1147. tion of large language models. In The Tenth International Conference on Learning Representations, ICLR Khan, M. A. M., Bari, M. S., Long, X. D., Wang, W., Parvez, 2022, Virtual Event, April 25-29, 2022. OpenReview.net, M. R., and Joty, S. Xcodeeval: An execution-based large 2022. URL https://openreview.net/forum? scale multilingual multitask benchmark for code underid=nZeVKeeFYf9. standing, generation, translation and retrieval. In Ku, L., Martins, A., and Srikumar, V. (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pp. 6766–6805. Association for Computational Linguistics, 2024. doi: 10. 18653/V1/2024.ACL-LONG.367. URL https://doi. org/10.18653/v1/2024.acl-long.367.
Hu, H., Richardson, K., Xu, L., Li, L., Kübler, S., and Moss, L. S. OCNLI: original chinese natural language inference. In Cohn, T., He, Y., and Liu, Y. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020, volume EMNLP 2020 of Findings of ACL, pp. 3512–3526. Association for Computational Linguistics, 2020. doi: 10.18653/V1/2020.FINDINGS-EMNLP. 314. URL https://doi.org/10.18653/v1/ 2020.findings-emnlp.314.
Khashabi, D., Ng, A., Khot, T., Sabharwal, A., Hajishirzi, H., and Callison-Burch, C. Gooaq: Open question answering with diverse answer types. In Moens, M., Huang, X., Specia, L., and Yih, S. W. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 1620 November, 2021, pp. 421–433. Association for Computational Linguistics, 2021. doi: 10.18653/V1/2021. FINDINGS-EMNLP.38. URL https://doi.org/ 10.18653/v1/2021.findings-emnlp.38.
Huang, J., Tang, D., Shou, L., Gong, M., Xu, K., Jiang, D., Zhou, M., and Duan, N. Cosqa: 20, 000+ web queries for code search and question answering. In Zong, C., Xia, F., Li, W., and Navigli, R. (eds.), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pp. 5690–5700. Association for Computational Linguistics, 2021. doi: 10.18653/
Kim, M., Rabelo, J., Goebel, R., Yoshioka, M., Kano, Y., and Satoh, K. COLIEE 2022 summary: Methods for legal document retrieval and entailment. In 13
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Takama, Y., Yada, K., Satoh, K., and Arai, S. (eds.), New Frontiers in Artificial Intelligence - JSAI-isAI 2022 Workshop, JURISIN 2022, and JSAI 2022 International Session, Kyoto, Japan, June 12-17, 2022, Revised Selected Papers, volume 13859 of Lecture Notes in Computer Science, pp. 51–67. Springer, 2022. doi: 10.1007/ 978-3-031-29168-5\ 4. URL https://doi.org/ 10.1007/978-3-031-29168-5_4.
Petrov, S. Natural questions: a benchmark for question answering research. Trans. Assoc. Comput. Linguistics, 7:452–466, 2019. doi: 10.1162/TACL\ A\ 00276. URL https://doi.org/10.1162/tacl_a_00276.
Koehn, P. Europarl: A parallel corpus for statistical machine translation. In Proceedings of Machine Translation Summit X: Papers, MTSummit 2005, Phuket, Thailand, September 13-15, 2005, pp. 79–86, 2005. URL https://aclanthology.org/2005. mtsummit-papers.11. Köksal, A., Thaler, M., Imani, A., Üstün, A., Korhonen, A., and Schütze, H. MURI: high-quality instruction tuning datasets for low-resource languages via reverse instructions. Trans. Assoc. Comput. Linguistics, 13:1032–1055, 2025. doi: 10.1162/TACL.A.18. URL https://doi.org/10.1162/tacl.a.18.
Lang, K. Newsweeder: Learning to filter netnews. In Prieditis, A. and Russell, S. (eds.), Machine Learning, Proceedings of the Twelfth International Conference on Machine Learning, Tahoe City, California, USA, July 9-12, 1995, pp. 331–339. Morgan Kaufmann, 1995. doi: 10.1016/B978-1-55860-377-6. 50048-7. URL https://doi.org/10.1016/ b978-1-55860-377-6.50048-7. Lee, C., Roy, R., Xu, M., Raiman, J., Shoeybi, M., Catanzaro, B., and Ping, W. Nv-embed: Improved techniques for training llms as generalist embedding models. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025a. URL https: //openreview.net/forum?id=lgsyLSsDRe. Lee, J., Chen, F., Dua, S., Cer, D., Shanbhogue, M., Naim, I., Ábrego, G. H., Li, Z., Chen, K., Vera, H. S., Ren, X., Zhang, S., Salz, D., Boratko, M., Han, J., Chen, B., Huang, S., Rao, V., Suganthan, P., Han, F., Doumanoglou, A., Gupta, N., Moiseev, F., Yip, C., Jain, A., Baumgartner, S., Shahi, S., Gomez, F. P., Mariserla, S., Choi, M., Shah, P., Goenka, S., Chen, K., Xia, Y., Chen, K., Duddu, S. M. K., Chen, Y., Walker, T., Zhou, W., Ghiya, R., Gleicher, Z., Gill, K., Dong, Z., Seyedhosseini, M., Sung, Y., Hoffmann, R., and Duerig, T. Gemini embedding: Generalizable embeddings from gemini. CoRR, abs/2503.07891, 2025b. doi: 10.48550/ARXIV.2503.07891. URL https://doi. org/10.48550/arXiv.2503.07891.
Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z. R., Stevens, K., Barhoum, A., Nguyen, D., Stanley, O., Nagyfi, R., ES, S., Suri, S., Glushkov, D., Dantuluri, A., Maguire, A., Schuhmann, C., Nguyen, H., and Mattick, A. Openassistant conversations democratizing large language model alignment. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023. URL http://papers. nips.cc/paper_files/paper/2023/hash/ 949f0f8f32267d297c2d4e3ee10a2e7e-Abstract-Datasets_ Lewis, P., Wu, Y., Liu, L., Minervini, P., Küttler, H., Piktus, and_Benchmarks.html. A., Stenetorp, P., and Riedel, S. PAQ: 65 million probablyasked questions and what you can do with them. Trans. Kusupati, A., Bhatt, G., Rege, A., Wallingford, M., Assoc. Comput. Linguistics, 9:1098–1115, 2021. doi: 10. Sinha, A., Ramanujan, V., Howard-Snyder, W., Chen, K., 1162/TACL\ A\ 00415. URL https://doi.org/ Kakade, S. M., Jain, P., and Farhadi, A. Matryoshka repre10.1162/tacl_a_00415. sentation learning. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (eds.), Advances Li, H., Arora, A., Chen, S., Gupta, A., Gupta, S., and in Neural Information Processing Systems 35: Annual Mehdad, Y. MTOP: A comprehensive multilingual taskConference on Neural Information Processing Systems oriented semantic parsing benchmark. In Merlo, P., Tiede2022, NeurIPS 2022, New Orleans, LA, USA, November mann, J., and Tsarfaty, R. (eds.), Proceedings of the 16th 28 - December 9, 2022, 2022. URL http://papers. Conference of the European Chapter of the Association nips.cc/paper_files/paper/2022/hash/ for Computational Linguistics: Main Volume, EACL 2021, c32319f4868da7613d78af9993100e42-Abstract-Conference. Online, April 19 - 23, 2021, pp. 2950–2962. Association html. for Computational Linguistics, 2021. doi: 10.18653/V1/ 2021.EACL-MAIN.257. URL https://doi.org/ Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., 10.18653/v1/2021.eacl-main.257. Parikh, A. P., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova, K., Jones, L., Kelcey, Li, H., Koto, F., Wu, M., Aji, A. F., and Baldwin, T. BactrianM., Chang, M., Dai, A. M., Uszkoreit, J., Le, Q., and x : A multilingual replicable instruction-following model 14
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
with low-rank adaptation. CoRR, abs/2305.15011, 2023a. doi: 10.48550/ARXIV.2305.15011. URL https:// doi.org/10.48550/arXiv.2305.15011.
question answering dataset for code search. In Calzolari, N., Kan, M., Hoste, V., Lenci, A., Sakti, S., and Xue, N. (eds.), Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC/COLING 2024, 20-25 May, 2024, Torino, Italy, pp. 13057–13067. ELRA and ICCL, 2024b. URL https://aclanthology. org/2024.lrec-main.1143.
Li, S., Ohagi, M., Ri, R., Fukuchi, A., Shibata, T., and Kawahara, D. JMTEB and JMTEB-lite: Japanese Massive Text Embedding Benchmark and Its Lightweight Version. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, Palma, Mallorca, Spain, May 2026. European Language Resources Association. to appear.
Lian, W., Goodson, B., Pentland, E., Cook, A., Vong, C., and ”Teknium”. Openorca: An open dataset of gpt augmented flan reasoning traces. https://https://huggingface.co/ datasets/Open-Orca/OpenOrca, 2023.
Li, X., Li, Z., Li, J., Xie, H., and Li, Q. 2d matryoshka sentence embeddings. CoRR, abs/2402.14776, 2024a. doi: 10.48550/ARXIV.2402.14776. URL https:// doi.org/10.48550/arXiv.2402.14776.
Liu, X., Wang, C., Leng, Y., and Zhai, C. Linkso: a dataset for learning to retrieve similar question answer pairs on software development forums. In Yu, Y., Fredericks, E. M., and Devanbu, P. T. (eds.), Proceedings of the 4th ACM SIGSOFT International Workshop on NLP for Software Engineering, NL4SE@ESEC/SIGSOFT FSE 2018, Lake Buena Vista, FL, USA, November 4, 2018, pp. 2–5. ACM, 2018. doi: 10.1145/3283812. 3283815. URL https://doi.org/10.1145/ 3283812.3283815.
Li, X., Dong, K., Lee, Y. Q., Xia, W., Zhang, H., Dai, X., Wang, Y., and Tang, R. Coir: A comprehensive benchmark for code information retrieval models. In Che, W., Nabende, J., Shutova, E., and Pilehvar, M. T. (eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, pp. 22074–22091. Association for Computational Linguistics, 2025a. doi: 10.18653/V1/2025.ACL-LONG. 1072. URL https://doi.org/10.18653/v1/ 2025.acl-long.1072.
Liu, Y., Gu, J., Goyal, N., Li, X., Edunov, S., Ghazvininejad, M., Lewis, M., and Zettlemoyer, L. Multilingual denoising pre-training for neural machine translation. Trans. Assoc. Comput. Linguistics, 8:726–742, 2020. doi: 10.1162/TACL\ A\ 00343. URL https: //doi.org/10.1162/tacl_a_00343.
Li, X., Li, Z., Li, J., Xie, H., and Li, Q. ESE: espresso sentence embeddings. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025b. URL https://openreview.net/forum? id=plgLA2YBLH. Li, Y., Zhang, Y., Zhao, Z., Shen, L., Liu, W., Mao, W., and Zhang, H. CSL: A large-scale chinese scientific literature dataset. In Calzolari, N., Huang, C., Kim, H., Pustejovsky, J., Wanner, L., Choi, K., Ryu, P., Chen, H., Donatelli, L., Ji, H., Kurohashi, S., Paggio, P., Xue, N., Kim, S., Hahm, Y., He, Z., Lee, T. K., Santus, E., Bond, F., and Na, S. (eds.), Proceedings of the 29th International Conference on Computational Linguistics, COLING 2022, Gyeongju, Republic of Korea, October 12-17, 2022, pp. 3917–3923. International Committee on Computational Linguistics, 2022. URL https://aclanthology.org/2022. coling-1.344. Li, Y., Li, Z., Zhang, K., Dan, R., and Zhang, Y. Chatdoctor: A medical chat model fine-tuned on llama model using medical domain knowledge. CoRR, abs/2303.14070, 2023b. doi: 10.48550/ARXIV.2303.14070. URL https: //doi.org/10.48550/arXiv.2303.14070.
Lo, K., Wang, L. L., Neumann, M., Kinney, R., and Weld, D. S. S2ORC: the semantic scholar open research corpus. In Jurafsky, D., Chai, J., Schluter, N., and Tetreault, J. R. (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pp. 4969–4983. Association for Computational Linguistics, 2020. doi: 10.18653/V1/ 2020.ACL-MAIN.447. URL https://doi.org/10. 18653/v1/2020.acl-main.447. Long, D., Gao, Q., Zou, K., Xu, G., Xie, P., Guo, R., Xu, J., Jiang, G., Xing, L., and Yang, P. Multi-cpr: A multi domain chinese dataset for passage retrieval. In Amigó, E., Castells, P., Gonzalo, J., Carterette, B., Culpepper, J. S., and Kazai, G. (eds.), SIGIR ’22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, Madrid, Spain, July 11 - 15, 2022, pp. 3046–3056. ACM, 2022. doi: 10. 1145/3477495.3531736. URL https://doi.org/ 10.1145/3477495.3531736. Longpre, S., Lu, Y., and Daiber, J. MKQA: A linguistically diverse benchmark for multilingual open domain question answering. Trans. Assoc. Comput. Linguistics, 9:
Li, Z., Zhang, J., Yin, C., Ouyang, Y., and Rong, W. Procqa: A large-scale community-based programming 15
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
1389–1406, 2021. doi: 10.1162/TACL\ A\ 00433. URL https://doi.org/10.1162/tacl_a_00433.
McAuley, J. J. and Leskovec, J. Hidden factors and hidden topics: understanding rating dimensions with review text. In Yang, Q., King, I., Li, Q., Pu, P., and Karypis, G. (eds.), Seventh ACM Conference on Recommender Systems, RecSys ’13, Hong Kong, China, October 12-16, 2013, pp. 165–172. ACM, 2013. doi: 10. 1145/2507157.2507163. URL https://doi.org/ 10.1145/2507157.2507163.
Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. URL https: //openreview.net/forum?id=Bkg6RiCqY7. Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C. Learning word vectors for sentiment analysis. In Lin, D., Matsumoto, Y., and Mihalcea, R. (eds.), The 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Conference, 1924 June, 2011, Portland, Oregon, USA, pp. 142–150. The Association for Computer Linguistics, 2011. URL https://aclanthology.org/P11-1015/.
Meyer, Y., Emadi, M., Nathawani, D., Ramaswamy, L., Boyd, K., Van Segbroeck, M., Grossman, M., Mlocek, P., and Newberry, D. Synthetic-Text-To-SQL: A synthetic dataset for training language models to generate sql queries from natural language prompts, April 2024. URL https://huggingface.co/datasets/ gretelai/synthetic-text-to-sql.
Maggie, Phil Culliton, W. C. Tweet sentiment extraction, 2020. URL https: //kaggle.com/competitions/ tweet-sentiment-extraction. Maheshwary, R., Yadav, V., Nguyen, H., Mahajan, K., and Madhusudhan, S. T. M2lingual: Enhancing multilingual, multi-turn instruction alignment in large language models. In Chiruzzo, L., Ritter, A., and Wang, L. (eds.), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, pp. 9676–9713. Association for Computational Linguistics, 2025. doi: 10.18653/ V1/2025.NAACL-LONG.489. URL https://doi. org/10.18653/v1/2025.naacl-long.489. Maia, M., Handschuh, S., Freitas, A., Davis, B., McDermott, R., Zarrouk, M., and Balahur, A. Www’18 open challenge: Financial opinion mining and question answering. In Champin, P., Gandon, F., Lalmas, M., and Ipeirotis, P. G. (eds.), Companion of the The Web Conference 2018 on The Web Conference 2018, WWW 2018, Lyon , France, April 23-27, 2018, pp. 1941–1942. ACM, 2018. doi: 10.1145/3184558.3192301. URL https: //doi.org/10.1145/3184558.3192301. Majumdar, S., Noroozi, V., Narenthiran, S., Ficek, A., Balam, J., and Ginsburg, B. Genetic instruct: Scaling up synthetic generation of coding instructions for large language models. CoRR, abs/2407.21077, 2024. doi: 10.48550/ARXIV.2407.21077. URL https://doi. org/10.48550/arXiv.2407.21077. May, P. Machine translated multilingual sts benchmark dataset., 2021. URL https://github.com/ PhilipMay/stsb-multi-mt. 16
Muennighoff, N., Tazi, N., Magne, L., and Reimers, N. MTEB: massive text embedding benchmark. In Vlachos, A. and Augenstein, I. (eds.), Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2023, Dubrovnik, Croatia, May 2-6, 2023, pp. 2006–2029. Association for Computational Linguistics, 2023. doi: 10.18653/V1/ 2023.EACL-MAIN.148. URL https://doi.org/ 10.18653/v1/2023.eacl-main.148. Muennighoff, N., Su, H., Wang, L., Yang, N., Wei, F., Yu, T., Singh, A., and Kiela, D. Generative representational instruction tuning. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URL https: //openreview.net/forum?id=BC4lIvfSzv. Narayan, S., Cohen, S. B., and Lapata, M. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Riloff, E., Chiang, D., Hockenmaier, J., and Tsujii, J. (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pp. 1797–1807. Association for Computational Linguistics, 2018. doi: 10.18653/V1/D18-1206. URL https://doi.org/ 10.18653/v1/d18-1206. Nie, Y., Williams, A., Dinan, E., Bansal, M., Weston, J., and Kiela, D. Adversarial NLI: A new benchmark for natural language understanding. In Jurafsky, D., Chai, J., Schluter, N., and Tetreault, J. R. (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pp. 4885–4901. Association for Computational Linguistics, 2020. doi: 10.18653/V1/ 2020.ACL-MAIN.441. URL https://doi.org/10. 18653/v1/2020.acl-main.441.
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Nielsen, D. S. Scandeval: A benchmark for scandinavian natural language processing. In Alumäe, T. and Fishel, M. (eds.), Proceedings of the 24th Nordic Conference on Computational Linguistics, NoDaLiDa 2023, Tórshavn, Faroe Islands, May 22-24, 2023, pp. 185– 201. University of Tartu Library, 2023. URL https: //aclanthology.org/2023.nodalida-1.20.
Subbian, K. Shopping queries dataset: A largescale ESCI benchmark for improving product search. CoRR, abs/2206.06588, 2022. doi: 10.48550/ARXIV. 2206.06588. URL https://doi.org/10.48550/ arXiv.2206.06588.
O’Neill, J., Rozenshtein, P., Kiryo, R., Kubota, M., and Bollegala, D. I wish I would have loved this one, but I didn’t - A multilingual dataset for counterfactual detection in product review. In Moens, M., Huang, X., Specia, L., and Yih, S. W. (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pp. 7092–7108. Association for Computational Linguistics, 2021. doi: 10.18653/ V1/2021.EMNLP-MAIN.568. URL https://doi. org/10.18653/v1/2021.emnlp-main.568. Pham, L., Luu, T., Vo, T., Nguyen, M., and Hoang, V. VNMTEB: vietnamese massive text embedding benchmark. In Demberg, V., Inui, K., and Marquez, L. (eds.), Findings of the Association for Computational Linguistics: EACL 2026, Rabat, Morocco, March 24-29, 2026, Findings of ACL, pp. 1705–1725. Association for Computational Linguistics, 2026. URL https://aclanthology. org/2026.findings-eacl.86/. Poswiata, R., Dadas, S., and Perelkiewicz, M. PLMTEB: polish massive text embedding benchmark. CoRR, abs/2405.10138, 2024. doi: 10.48550/ARXIV. 2405.10138. URL https://doi.org/10.48550/ arXiv.2405.10138. Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y. Zero: memory optimizations toward training trillion parameter models. In Cuicchi, C., Qualters, I., and Kramer, W. T. (eds.), Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2020, Virtual Event / Atlanta, Georgia, USA, November 9-19, 2020, pp. 20. IEEE/ACM, 2020. doi: 10.1109/SC41405.2020.00024. URL https://doi. org/10.1109/SC41405.2020.00024. Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. Squad: 100, 000+ questions for machine comprehension of text. In Su, J., Carreras, X., and Duh, K. (eds.), Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, Austin, Texas, USA, November 1-4, 2016, pp. 2383–2392. The Association for Computational Linguistics, 2016. doi: 10.18653/V1/ D16-1264. URL https://doi.org/10.18653/ v1/d16-1264.
Reimers, N. and Gurevych, I. Sentence-bert: Sentence embeddings using siamese bert-networks. In Inui, K., Jiang, J., Ng, V., and Wan, X. (eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLPIJCNLP 2019, Hong Kong, China, November 3-7, 2019, pp. 3980–3990. Association for Computational Linguistics, 2019. doi: 10.18653/V1/D19-1410. URL https: //doi.org/10.18653/v1/D19-1410. Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. Winogrande: An adversarial winograd schema challenge at scale. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pp. 8732– 8740. AAAI Press, 2020. doi: 10.1609/AAAI.V34I05. 6399. URL https://doi.org/10.1609/aaai. v34i05.6399. Saravia, E., Liu, H. T., Huang, Y., Wu, J., and Chen, Y. CARER: contextualized affect representations for emotion recognition. In Riloff, E., Chiang, D., Hockenmaier, J., and Tsujii, J. (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pp. 3687–3697. Association for Computational Linguistics, 2018. doi: 10.18653/V1/D18-1404. URL https://doi.org/10.18653/v1/d18-1404. Scialom, T., Dray, P., Lamprier, S., Piwowarski, B., and Staiano, J. MLSUM: the multilingual summarization corpus. In Webber, B., Cohn, T., He, Y., and Liu, Y. (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pp. 8051–8067. Association for Computational Linguistics, 2020. doi: 10.18653/V1/ 2020.EMNLP-MAIN.647. URL https://doi.org/ 10.18653/v1/2020.emnlp-main.647. Sharma, L., Graesser, L., Nangia, N., and Evci, U. Natural language understanding with the quora question pairs dataset. CoRR, abs/1907.01041, 2019. URL http:// arxiv.org/abs/1907.01041.
Reddy, C. K., Màrquez, L., Valero, F., Rao, N., Zaragoza, H., Bandyopadhyay, S., Biswas, A., Xing, A., and
Singh, S., Vargus, F., D’souza, D., Karlsson, B., Mahendiran, A., Ko, W., Shandilya, H., Patel, J., Mataciunas, 17
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
D., O’Mahony, L., Zhang, M., Hettiarachchi, R., Wilson, J., Machado, M., Moura, L. S., Krzeminski, D., Fadaei, H., Ergün, I., Okoh, I., Alaagib, A., Mudannayake, O., Alyafeai, Z., Chien, M. V., Ruder, S., Guthikonda, S., Alghamdi, E. A., Gehrmann, S., Muennighoff, N., Bartolo, M., Kreutzer, J., Üstün, A., Fadaee, M., and Hooker, S. Aya dataset: An open-access collection for multilingual instruction tuning. In Ku, L., Martins, A., and Srikumar, V. (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pp. 11521–11567. Association for Computational Linguistics, 2024. doi: 10.18653/V1/ 2024.ACL-LONG.620. URL https://doi.org/ 10.18653/v1/2024.acl-long.620.
Team, M. Biorxiv raw data, 2022b. URL https://huggingface.co/datasets/mteb/ raw_biorxiv.
Snegirev, A., Tikhonova, M., Maksimova, A., Fenogenova, A., and Abramov, A. The russian-focused embedders’ exploration: rumteb benchmark and russian embedding model design. In Chiruzzo, L., Ritter, A., and Wang, L. (eds.), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, Albuquerque, New Mexico, USA, April 29 May 4, 2025, pp. 236–254. Association for Computational Linguistics, 2025. doi: 10.18653/V1/2025. NAACL-LONG.12. URL https://doi.org/10. 18653/v1/2025.naacl-long.12.
Thorne, J., Vlachos, A., Christodoulopoulos, C., and Mittal, A. FEVER: a large-scale dataset for fact extraction and verification. In Walker, M. A., Ji, H., and Stent, A. (eds.), Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), pp. 809–819. Association for Computational Linguistics, 2018. doi: 10.18653/V1/N18-1074. URL https://doi.org/ 10.18653/v1/n18-1074.
Team, M. Medrxiv raw data, 2022c. URL https://huggingface.co/datasets/mteb/ raw_medrxiv. Team, S. T. Embedding training data, 2021c. https://huggingface.co/datasets/ sentence-transformers/.
URL
Team, S. T. Reddit title-body pairs, 2021d. URL https://huggingface.co/ datasets/sentence-transformers/ reddit-title-body.
Sun, M., Li, J., Guo, Z., Zhao, Y., Zheng, Y., Si, X., and Liu, Z. Thuctc: An efficient chinese text classifier, 2016. URL http://thuctc.thunlp.org/. Sun, S. and Duh, K. Clirmatrix: A massively large collection of bilingual and multilingual datasets for crosslingual information retrieval. In Webber, B., Cohn, T., He, Y., and Liu, Y. (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pp. 4160–4170. Association for Computational Linguistics, 2020. doi: 10.18653/V1/2020. EMNLP-MAIN.340. URL https://doi.org/10. 18653/v1/2020.emnlp-main.340. Team, F. S. E. Stackexchange title-body pairs, 2021a. URL https://huggingface.co/ datasets/flax-sentence-embeddings/ stackexchange_title_body_jsonl. Team, F. S. E. Stack exchange question pairs, 2021b. URL https://huggingface.co/datasets/ flax-sentence-embeddings/. Team, M. Arxiv raw data, 2022a. URL https://huggingface.co/datasets/mteb/ raw_arxiv. 18
Tsatsaronis, G., Balikas, G., Malakasiotis, P., Partalas, I., Zschunke, M., Alvers, M. R., Weissenborn, D., Krithara, A., Petridis, S., Polychronopoulos, D., Almirantis, Y., Pavlopoulos, J., Baskiotis, N., Gallinari, P., Artières, T., Ngomo, A. N., Heino, N., Gaussier, É., Barrio-Alvers, L., Schroeder, M., Androutsopoulos, I., and Paliouras, G. An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competition. BMC Bioinform., 16:138:1–138:28, 2015. doi: 10.1186/ S12859-015-0564-6. URL https://doi.org/10. 1186/s12859-015-0564-6. Vera, H. S., Dua, S., Zhang, B., Salz, D., Mullins, R., Panyam, S. R., Smoot, S., Naim, I., Zou, J., Chen, F., Cer, D., Lisak, A., Choi, M., Gonzalez, L., Sanseviero, O., Cameron, G., Ballantyne, I., Black, K., Chen, K., Wang, W., Li, Z., Martins, G., Lee, J., Sherwood, M., Ji, J., Wu, R., Zheng, J., Singh, J., Sharma, A., Sreepathihalli, D., Jain, A., Elarabawy, A., Co, A., Doumanoglou, A., Samari, B., Hora, B., Potetz, B., Kim, D., Alfonseca, E., Moiseev, F., Han, F., Gomez, F. P., Ábrego, G. H., Zhang, H., Hui, H., Han, J., Gill, K., Chen, K., Chen, K., Shanbhogue, M., Boratko, M., Suganthan, P., Duddu, S. M. K., Mariserla, S., Ariafar, S., Zhang, S., Zhang, S., Baumgartner, S., Goenka, S., Qiu, S., Dabral, T., Walker, T., Rao, V., Khawaja, W., Zhou, W., Ren, X., Xia, Y., Chen, Y., Chen, Y., Dong, Z.,
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Ding, Z., Visin, F., Liu, G., Zhang, J., Kenealy, K., Casbon, M., Kumar, R., Mesnard, T., Gleicher, Z., Brick, C., Lacombe, O., Roberts, A., Yin, Q., Sung, Y., Hoffmann, R., Warkentin, T., Joulin, A., Duerig, T., and Seyedhosseini, M. Embeddinggemma: Powerful and lightweight text representations. CoRR, abs/2509.20354, 2025. doi: 10.48550/ARXIV.2509.20354. URL https: //doi.org/10.48550/arXiv.2509.20354.
29 - May 4, 2025, pp. 3828–3848. Association for Computational Linguistics, 2025. doi: 10.18653/V1/2025. FINDINGS-NAACL.211. URL https://doi.org/ 10.18653/v1/2025.findings-naacl.211.
Wang, Y., Ma, X., Zhang, G., Ni, Y., Chandra, A., Guo, S., Ren, W., Arulraj, A., He, X., Jiang, Z., Li, T., Ku, M., Wang, K., Zhuang, A., Fan, R., Yue, X., and Chen, W. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. In Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J. M., and Zhang, C. (eds.), Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024b. URL http://papers. nips.cc/paper_files/paper/2024/hash/ ad236edc564f3e3156e1b2feafb99a24-Abstract-Datasets and_Benchmarks_Track.html.
Wachsmuth, H., Syed, S., and Stein, B. Retrieval of the best counterargument without prior topic knowledge. In Gurevych, I. and Miyao, Y. (eds.), Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 1: Long Papers, pp. 241–251. Association for Computational Linguistics, 2018. doi: 10.18653/V1/P18-1023. URL https: //aclanthology.org/P18-1023/. Wadden, D., Lin, S., Lo, K., Wang, L. L., van Zuylen, M., Cohan, A., and Hajishirzi, H. Fact or fiction: Verifying scientific claims. In Webber, B., Cohn, T., He, Y., and Liu, Y. (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pp. 7534–7550. Association for Computational Linguistics, 2020. doi: 10.18653/V1/2020.EMNLP-MAIN. 609. URL https://doi.org/10.18653/v1/ 2020.emnlp-main.609.
Wehrli, S., Arnrich, B., and Irrgang, C. German text embedding clustering benchmark. In Georges, M., Herygers, A., Friedrich, A., and Roth, B. (eds.), Proceedings of the 19th Conference on Natural Language Processing (KONVENS 2023), September 19-21, 2023, Ingolstadt, Germany, pp. 187–201. Association for Computational Lingustics, 2023. URL https://aclanthology. org/2023.konvens-main.20. Wei, X., Wei, H., Lin, H., Li, T., Zhang, P., Ren, X., Li, M., Wan, Y., Cao, Z., Xie, B., Hu, T., Li, S., Hui, B., Yu, B., Liu, D., Yang, B., Huang, F., and Xie, J. Polylm: An open source polyglot large language model. CoRR, abs/2307.06018, 2023. doi: 10.48550/ARXIV. 2307.06018. URL https://doi.org/10.48550/ arXiv.2307.06018.
Wang, L., Yang, N., Huang, X., Yang, L., Majumder, R., and Wei, F. Improving text embeddings with large language models. In Ku, L., Martins, A., and Srikumar, V. (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pp. 11897–11916. Association for Computational Linguistics, 2024a. doi: 10.18653/V1/ 2024.ACL-LONG.642. URL https://doi.org/ 10.18653/v1/2024.acl-long.642.
Williams, A., Nangia, N., and Bowman, S. R. A broadcoverage challenge corpus for sentence understanding through inference. In Walker, M. A., Ji, H., and Stent, A. (eds.), Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACLHLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), pp. 1112–1122. Association for Computational Linguistics, 2018. doi: 10.18653/V1/ N18-1101. URL https://doi.org/10.18653/ v1/n18-1101.
Wang, L. L., Lo, K., Chandrasekhar, Y., Reas, R., Yang, J., Eide, D., Funk, K., Kinney, R., Liu, Z., Merrill, W., Mooney, P., Murdick, D. A., Rishi, D., Sheehan, J., Shen, Z., Stilson, B., Wade, A. D., Wang, K., Wilhelm, C., Xie, B., Raymond, D., Weld, D. S., Etzioni, O., and Kohlmeier, S. CORD-19: the covid-19 open research dataset. CoRR, abs/2004.10706, 2020. URL https: //arxiv.org/abs/2004.10706.
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M. Transformers: Stateof-the-art natural language processing. In Liu, Q. and Schlangen, D. (eds.), Proceedings of the 2020 Conference
Wang, X., Li, J., Chen, S., Zhu, Y., Wu, X., Zhang, Z., Xu, X., Chen, J., Fu, J., Wan, X., Gao, A., and Wang, B. Huatuo-26m, a large-scale chinese medical QA dataset. In Chiruzzo, L., Ritter, A., and Wang, L. (eds.), Findings of the Association for Computational Linguistics: NAACL 2025, Albuquerque, New Mexico, USA, April 19
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
on Empirical Methods in Natural Language Processing: System Demonstrations, EMNLP 2020 - Demos, Online, November 16-20, 2020, pp. 38–45. Association for Computational Linguistics, 2020. doi: 10.18653/V1/2020. EMNLP-DEMOS.6. URL https://doi.org/10. 18653/v1/2020.emnlp-demos.6. Xiao, S., Liu, Z., Zhang, P., and Muennighoff, N. C-pack: Packaged resources to advance general chinese embedding. CoRR, abs/2309.07597, 2023. doi: 10.48550/ ARXIV.2309.07597. URL https://doi.org/10. 48550/arXiv.2309.07597. Xie, X., Dong, Q., Wang, B., Lv, F., Yao, T., Gan, W., Wu, Z., Li, X., Li, H., Liu, Y., and Ma, J. T2ranking: A largescale chinese benchmark for passage ranking. In Chen, H., Duh, W. E., Huang, H., Kato, M. P., Mothe, J., and Poblete, B. (eds.), Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2023, Taipei, Taiwan, July 23-27, 2023, pp. 2681–2690. ACM, 2023. doi: 10. 1145/3539618.3591874. URL https://doi.org/ 10.1145/3539618.3591874. Xu, L., Hu, H., Zhang, X., Li, L., Cao, C., Li, Y., Xu, Y., Sun, K., Yu, D., Yu, C., Tian, Y., Dong, Q., Liu, W., Shi, B., Cui, Y., Li, J., Zeng, J., Wang, R., Xie, W., Li, Y., Patterson, Y., Tian, Z., Zhang, Y., Zhou, H., Liu, S., Zhao, Z., Zhao, Q., Yue, C., Zhang, X., Yang, Z., Richardson, K., and Lan, Z. CLUE: A chinese language understanding evaluation benchmark. In Scott, D., Bel, N., and Zong, C. (eds.), Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020, pp. 4762–4772. International Committee on Computational Linguistics, 2020. doi: 10.18653/V1/2020. COLING-MAIN.419. URL https://doi.org/10. 18653/v1/2020.coling-main.419. Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C. mt5: A massively multilingual pre-trained text-to-text transformer. In Toutanova, K., Rumshisky, A., Zettlemoyer, L., Hakkani-Tür, D., Beltagy, I., Bethard, S., Cotterell, R., Chakraborty, T., and Zhou, Y. (eds.), Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pp. 483–498. Association for Computational Linguistics, 2021. doi: 10.18653/V1/2021. NAACL-MAIN.41. URL https://doi.org/10. 18653/v1/2021.naacl-main.41. Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., Lin, H., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Lin, J., Dang,
K., Lu, K., Bao, K., Yang, K., Yu, L., Li, M., Xue, M., Zhang, P., Zhu, Q., Men, R., Lin, R., Li, T., Xia, T., Ren, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Wan, Y., Liu, Y., Cui, Z., Zhang, Z., and Qiu, Z. Qwen2.5 technical report. CoRR, abs/2412.15115, 2024. doi: 10.48550/ARXIV. 2412.15115. URL https://doi.org/10.48550/ arXiv.2412.15115. Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., Zheng, C., Liu, D., Zhou, F., Huang, F., Hu, F., Ge, H., Wei, H., Lin, H., Tang, J., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Lin, J., Dang, K., Bao, K., Yang, K., Yu, L., Deng, L., Li, M., Xue, M., Li, M., Zhang, P., Wang, P., Zhu, Q., Men, R., Gao, R., Liu, S., Luo, S., Li, T., Tang, T., Yin, W., Ren, X., Wang, X., Zhang, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Zhang, Y., Wan, Y., Liu, Y., Wang, Z., Cui, Z., Zhang, Z., Zhou, Z., and Qiu, Z. Qwen3 technical report. CoRR, abs/2505.09388, 2025. doi: 10.48550/ARXIV.2505.09388. URL https:// doi.org/10.48550/arXiv.2505.09388. Yang, Y., Zhang, Y., Tar, C., and Baldridge, J. PAWS-X: A cross-lingual adversarial dataset for paraphrase identification. In Inui, K., Jiang, J., Ng, V., and Wan, X. (eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pp. 3685–3690. Association for Computational Linguistics, 2019. doi: 10.18653/V1/D19-1382. URL https://doi.org/10.18653/v1/D19-1382. Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W. W., Salakhutdinov, R., and Manning, C. D. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In Riloff, E., Chiang, D., Hockenmaier, J., and Tsujii, J. (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pp. 2369–2380. Association for Computational Linguistics, 2018. doi: 10.18653/V1/D18-1259. URL https://doi.org/10.18653/v1/d18-1259. Yoon, J., Sinha, R., Arik, S. Ö., and Pfister, T. Matryoshkaadaptor: Unsupervised and supervised tuning for smaller embedding dimensions. In Al-Onaizan, Y., Bansal, M., and Chen, Y. (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, pp. 10318–10336. Association for Computational Linguistics, 2024. doi: 10.18653/V1/2024. EMNLP-MAIN.576. URL https://doi.org/10. 18653/v1/2024.emnlp-main.576. Yuan, W., Yu, J., Jiang, S., Padthe, K., Li, Y., Wang, D.,
20
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Kulikov, I., Cho, K., Tian, Y., Weston, J. E., and Li, X. Naturalreasoning: Reasoning in the wild with 2.8m challenging questions. CoRR, abs/2502.13124, 2025. doi: 10.48550/ARXIV.2502.13124. URL https://doi. org/10.48550/arXiv.2502.13124.
Zhang, X., Tian, C., Yang, X., Chen, L., Li, Z., and Petzold, L. R. Alpacare: Instruction-tuned large language models for medical application. CoRR, abs/2310.14558, 2023c. doi: 10.48550/ARXIV.2310.14558. URL https:// doi.org/10.48550/arXiv.2310.14558.
Zhang, B., Chen, L., Liu, T., and Zheng, B. SMEC:rethinking matryoshka representation learning for retrieval embedding compression. In Christodoulopoulos, C., Chakraborty, T., Rose, C., and Peng, V. (eds.), Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 26209–26222, Suzhou, China, November 2025a. Association for Computational Linguistics. ISBN 9798-89176-332-6. doi: 10.18653/v1/2025.emnlp-main. 1332. URL https://aclanthology.org/2025. emnlp-main.1332/.
Zhang, Y., Li, M., Long, D., Zhang, X., Lin, H., Yang, B., Xie, P., Yang, A., Liu, D., Lin, J., Huang, F., and Zhou, J. Qwen3 embedding: Advancing text embedding and reranking through foundation models. CoRR, abs/2506.05176, 2025b. doi: 10.48550/ARXIV. 2506.05176. URL https://doi.org/10.48550/ arXiv.2506.05176.
Zhang, Q., Chen, M., Bukharin, A., He, P., Cheng, Y., Chen, W., and Zhao, T. Adaptive budget allocation for parameter-efficient fine-tuning. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023a. URL https://openreview.net/forum? id=lq62uWRJjiY.
Zhang, Z., Liu, Y., Huang, W., Mao, J., Wang, R., and Hu, H. MELA: multilingual evaluation of linguistic acceptability. In Ku, L., Martins, A., and Srikumar, V. (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pp. 2658–2674. Association for Computational Linguistics, 2024. doi: 10.18653/ V1/2024.ACL-LONG.146. URL https://doi.org/ 10.18653/v1/2024.acl-long.146. Zhang, Z., Liao, Z., Yu, H., Di, P., and Wang, R. F2LLM technical report: Matching SOTA embedding performance with 6 million open-source data. CoRR, abs/2510.02294, 2025c. doi: 10.48550/ARXIV. 2510.02294. URL https://doi.org/10.48550/ arXiv.2510.02294.
Zhang, S., Zhang, X., Wang, H., Guo, L., and Liu, S. Multi-scale attentive interaction networks for chinese medical question answer selection. IEEE Access, 6:74061–74071, 2018. doi: 10.1109/ACCESS.2018. 2883637. URL https://doi.org/10.1109/ ACCESS.2018.2883637.
Zhao, W., Ren, X., Hessel, J., Cardie, C., Choi, Y., and Zhang, X., Zhao, J. J., and LeCun, Y. Character-level Deng, Y. Wildchat: 1m chatgpt interaction logs in convolutional networks for text classification. In the wild. In The Twelfth International Conference on Cortes, C., Lawrence, N. D., Lee, D. D., Sugiyama, Learning Representations, ICLR 2024, Vienna, Austria, M., and Garnett, R. (eds.), Advances in Neural InforMay 7-11, 2024. OpenReview.net, 2024. URL https: mation Processing Systems 28: Annual Conference //openreview.net/forum?id=Bl8u7ZRlbM. on Neural Information Processing Systems 2015, Zhao, X., Hu, X., Shan, Z., Huang, S., Zhou, Y., Sun, December 7-12, 2015, Montreal, Quebec, Canada, pp. Z., Liu, Z., Li, D., Wei, X., Chen, Q., Pan, Y., Xi649–657, 2015. URL https://proceedings. ang, Y., Zhang, M., Wang, H., Yu, J., Hu, B., and neurips.cc/paper/2015/hash/ Zhang, M. Kalm-embedding-v2: Superior training tech250cf8b51c773f3f8dc8b4be867a9a02-Abstract. niques and data inspire A versatile embedding model. html. CoRR, abs/2506.20923, 2025. doi: 10.48550/ARXIV. Zhang, X., Ma, X., Shi, P., and Lin, J. Mr. tydi: A 2506.20923. URL https://doi.org/10.48550/ multi-lingual benchmark for dense retrieval. CoRR, arXiv.2506.20923. abs/2108.08787, 2021. URL https://arxiv.org/ Ziemski, M., Junczys-Dowmunt, M., and Pouliquen, abs/2108.08787. B. The united nations parallel corpus v1.0. In Zhang, X., Thakur, N., Ogundepo, O., Kamalloo, E., Calzolari, N., Choukri, K., Declerck, T., Goggi, S., Alfonso-Hermelo, D., Li, X., Liu, Q., Rezagholizadeh, Grobelnik, M., Maegaard, B., Mariani, J., Mazo, M., and Lin, J. MIRACL: A multilingual retrieval dataset H., Moreno, A., Odijk, J., and Piperidis, S. (eds.), covering 18 diverse languages. Trans. Assoc. Comput. Proceedings of the Tenth International Conference Linguistics, 11:1114–1131, 2023b. doi: 10.1162/TACL\ on Language Resources and Evaluation LREC 2016, A\ 00595. URL https://doi.org/10.1162/ Portorož, Slovenia, May 23-28, 2016. European Lantacl_a_00595. guage Resources Association (ELRA), 2016. URL 21
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
http://www.lrec-conf.org/proceedings/ lrec2016/summaries/1195.html. Zinvandi, E., Alikhani, M., Sarmadi, M., Pourbahman, Z., Arvin, S., Kazemi, R., and Amini, A. Famteb: Massive text embedding benchmark in persian language. In Christodoulopoulos, C., Chakraborty, T., Rose, C., and Peng, V. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, November 4-9, 2025, pp. 11441– 11468. Association for Computational Linguistics, 2025. URL https://aclanthology.org/2025. findings-emnlp.614/.
22
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
A. Details on Training Data Table 5. Natural language distribution in our training data (part1). ISO Code
Language
Samples
ISO Code
Language
Samples
ISO Code
Language
Samples
eng zho spa fra rus deu ara nld vie jpn hin ind ita por pol tur tha kor ukr ces tgl fas cat glg mya hye khm nep eus swe lao swa dan ell azj sin tgk bul ron fin hun slv heb lav urd nor lit
English Chinese Spanish French Russian German Arabic Dutch Vietnamese Japanese Hindi Indonesian Italian Portuguese Polish Turkish Thai Korean Ukrainian Czech Tagalog Persian Catalan Galician Burmese Armenian Khmer Nepali Basque Swedish Lao Swahili Danish Modern Greek North Azerbaijani Sinhala Tajik Bulgarian Romanian Finnish Hungarian Slovenian Hebrew Latvian Urdu Norwegian Lithuanian
15,683,866 2,791,623 2,607,192 2,251,567 1,848,954 1,593,329 1,264,371 1,051,914 1,005,512 935,603 931,761 817,131 790,507 752,642 730,068 565,158 503,612 478,665 337,739 330,256 313,985 302,869 301,984 301,386 294,167 288,622 287,530 276,057 270,551 256,469 244,750 241,497 224,504 223,254 213,046 207,903 200,631 200,127 184,034 178,810 144,561 122,065 119,685 115,288 113,258 107,023 101,937
slk est msa ben aze afr tam kat tel mal mon nno kaz cym mar sqi nob pus isl hrv mkd hbs ceb jav srp war kan epo lat guj uzb amh oci bel azb kir mlg vol ast pan ltz nds hat bre gle sco xho
Slovak Estonian Malay Bengali Azerbaijani Afrikaans Tamil Georgian Telugu Malayalam Mongolian Norwegian Nynorsk Kazakh Welsh Marathi Albanian Norwegian Bokmål Pushto Icelandic Croatian Macedonian Serbo-Croatian Cebuano Javanese Serbian Waray Kannada Esperanto Latin Gujarati Uzbek Amharic Occitan Belarusian South Azerbaijani Kirghiz Malagasy Volapük Asturian Panjabi Luxembourgish Low German Haitian Breton Irish Scots Xhosa
97,974 91,102 87,873 87,787 82,209 81,233 78,384 77,567 77,362 76,518 59,851 58,298 55,317 53,951 53,803 53,602 52,905 52,753 52,546 52,505 52,504 48,523 47,408 47,283 46,108 45,348 44,534 44,266 43,335 40,184 39,951 38,763 37,413 33,330 31,815 29,319 28,661 27,187 26,004 25,096 25,092 24,713 23,940 23,931 23,148 23,032 22,799
tat bos yor min che arz lmo arg bak som als ido szl wuu new chv pnb fry snd ori plt scn kur sun bar yid ckb fao ina gla bug que bpy san lim hau mai zsm ibo vec ilo asm sah arb sna mlt zul
Tatar Bosnian Yoruba Minangkabau Chechen Egyptian Arabic Lombard Aragonese Bashkir Somali Tosk Albanian Ido Silesian Wu Chinese Nepal Bhasa Chuvash Western Panjabi Western Frisian Sindhi Oriya Plateau Malagasy Sicilian Kurdish Sundanese Bavarian Yiddish Central Kurdish Faroese Interlingua Scottish Gaelic Buginese Quechua Bishnupriya Sanskrit Limburgan Hausa Maithili Standard Malay Igbo Venetian Iloko Assamese Yakut Standard Arabic Shona Maltese Zulu
22,327 21,175 20,139 19,868 19,518 16,783 16,575 16,500 16,451 16,369 15,655 15,613 14,845 14,762 14,714 13,759 13,657 13,649 13,210 12,791 12,717 12,694 11,872 11,793 11,193 10,785 9,829 9,825 9,782 9,769 9,662 9,406 9,400 8,730 8,573 8,435 8,180 8,179 8,132 8,121 7,968 7,042 7,011 6,945 6,933 6,911 6,669
23
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Table 6. Natural language distribution in our training data (part2). ISO Code
Language
Samples
ISO Code
Language
Samples
ISO Code
Language
mzn uig oss tuk ary wln cdo npi nap ace mrj xmf pes diq apc wol pbt nso srd ban lij hsb acq crh mri egl ars grn nya hif kas fur swh fil sme shn sot kin lug pap cos mhr bjn knc taq kom bod
Mazanderani Uighur Iron Ossetic Turkmen Moroccan Arabic Walloon Min Dong Chinese Nepali Neapolitan Achinese Western Mari Mingrelian Iranian Persian Dimli Levantine Arabic Wolof Southern Pashto Pedi Sardinian Balinese Ligurian Upper Sorbian Ta’izzi-Adeni Arabic Crimean Tatar Maori Emilian Najdi Arabic Guarani Chichewa Fiji Hindi Kashmiri Friulian Swahili Filipino Northern Sami Shan Southern Sotho Kinyarwanda Ganda Papiamento Corsican Eastern Mari Banjar Central Kanuri Tamasheq Komi Tibetan
6,352 6,190 5,893 5,854 5,703 5,408 5,175 5,156 4,778 4,758 4,728 4,712 4,414 4,140 4,079 4,068 3,796 3,676 3,614 3,581 3,487 3,440 3,188 2,985 2,922 2,859 2,712 2,705 2,607 2,566 2,363 2,284 2,253 2,156 2,156 2,098 2,089 2,031 1,994 1,974 1,949 1,633 1,600 1,600 1,600 1,583 1,563
tsn mwl div kbp chm ewe smo tso fij bam lin nav roh ssw awa pag cor dzo udm fon kon glv tir pms myv sag run acm aka ayr bem bho cjk dik fuv gaz hne kab kac kam kea khk kmb kmr ltg luo lus
Tswana Mirandese Dhivehi Kabiyè Mari Ewe Samoan Tsonga Fijian Bambara Lingala Navajo Romansh Swati Awadhi Pangasinan Cornish Dzongkha Udmurt Fon Kongo Manx Tigrinya Piemontese Erzya Sango Rundi Mesopotamian Arabic Akan Central Aymara Bemba Bhojpuri Chokwe Southwestern Dinka Nigerian Fulfulde West Central Oromo Chhattisgarhi Kabyle Kachin Kamba Kabuverdianu Halh Mongolian Kimbundu Northern Kurdish Latgalian Luo Lushai
1,539 1,491 1,387 1,349 1,238 1,220 1,175 1,174 1,122 1,061 1,046 1,028 999 982 954 954 938 928 890 883 873 864 851 842 840 827 825 800 800 800 800 800 800 800 800 800 800 800 800 800 800 800 800 800 800 800 800
lvs mag mni mos nqo nus ory prs quy sat tpi tum tzm umb uzn ydd yue dyu lua twi aeb kik tyv ava aym krc ful orm stq lah ton mdf haw nia bis alt srn ven kbd xal din jam kal iku guc chr ady
Standard Latvian Magahi Manipuri Mossi N’Ko Nuer Odia Dari Ayacucho Quechua Santali Tok Pisin Tumbuka Central Atlas Tamazight Umbundu Northern Uzbek Eastern Yiddish Yue Chinese Dyula Luba-Lulua Twi Tunisian Arabic Kikuyu Tuvinian Avaric Aymara Karachay-Balkar Fulah Oromo Saterfriesisch Lahnda Tonga Moksha Hawaiian Nias Bislama Southern Altai Sranan Tongo Venda Kabardian Kalmyk Dinka Jamaican Creole English Kalaallisut Inuktitut Wayuu Cherokee Adyghe
24
Samples 800 800 800 800 800 800 800 800 800 800 800 800 800 800 800 800 800 799 799 798 793 761 626 588 587 587 581 548 461 450 391 314 299 297 272 250 204 194 172 122 104 100 92 84 52 51 33
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Table 7. Programming language distribution in our training data. Language
Samples
Language
Samples
python php java cpp go javascript c# ruby c rust kotlin sql pascal d haskell scala html shell perl swift ocaml csharp
1,972,390 553,651 483,469 393,514 351,586 245,632 92,008 68,317 43,487 12,924 11,284 6,826 6,299 5,278 4,967 4,120 2,777 2,095 2,009 1,907 1,894 1,891
css typescript r lisp jsx objective-c json xml yaml assembly powershell vba lua matlab dart bash http graphql svg vb.net groovy Misc.
1,003 888 636 467 436 327 264 258 180 171 162 157 126 114 107 105 99 89 82 75 63 2,301
Summarization 2.5%
Language Domain ClassificationClassification 1.5% 1.0% Sentiment Intent Analysis Classification 1.5% 1.7%
Topic Classification 2.7%
Code-to-Text Paraphrase Detection 1.9% 1.7%
NLI 3.2%
Code-to-Code 2.9%
Text-to-Code 2.5% Question Answering 28.9%
Title Matching 8.2%
Bitext Mining 29.2%
Instruction Data 10.6%
Figure 7. Task type distribution in our training data.
25
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Table 8. Number of samples in our collected training dataset (part 1). Name
Language
Format
Size
UNPC (Ziemski et al., 2016) ParaCrawl (Bañón et al., 2020) BactrianX Translation (Li et al., 2023a) Europarl (Koehn, 2005)
6 30 52 21
Retrieval Retrieval Clustering Clustering
2,922,245 10,684,184 491,282 477,566
WebFAQ (Dinzinger et al., 2025) mMARCO (Bonifacio et al., 2021) PAQ (Lewis et al., 2021) SQuAD (Rajpurkar et al., 2016)
49 14 en en
Retrieval Retrieval Retrieval Retrieval
Stack Exchange (Team, 2021b)
en
Retrieval
Arguana (Wachsmuth et al., 2018) Natural Questions (Kwiatkowski et al., 2019) HotpotQA (Yang et al., 2018) ELI5 (Fan et al., 2019) FiQA2018 (Maia et al., 2018) BioASQ (Tsatsaronis et al., 2015) NFCorpus (Boteva et al., 2016) TriviaQA (Joshi et al., 2017) PubMedQA (Jin et al., 2019) Amazon QA (Gupta et al., 2019) MIRACL (Zhang et al., 2023b) Mr.TyDi (Zhang et al., 2021) MLDR (Chen et al., 2024) MKQA (Longpre et al., 2021) StackOverflowQA (Li et al., 2025a) ProCQA (Li et al., 2024b) Yahoo Answers (Zhang et al., 2015) GooAQ (Khashabi et al., 2021) T2Ranking (Xie et al., 2023) DuReader (He et al., 2018) cMedQAv2 (Zhang et al., 2018) Huatuo kgqa (Wang et al., 2025) Huatuo encqa (Wang et al., 2025) Multi CPR Medical (Long et al., 2022) HealthCareMagic (Li et al., 2023b) MedicalQA ru (Blinov, 2021)
en en en en en en en en en en 16 11 13 26 en 11 en en zh zh zh zh zh zh en ru
Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval
Aya (Singh et al., 2024) MURI (Köksal et al., 2025) OASST2 (Köpf et al., 2023) MultiAlpaca (Wei et al., 2023) WildChat (Zhao et al., 2024) M2Lingual (Maheshwary et al., 2025) Natural Reasoning (Yuan et al., 2025) Infinity Instruct (Du et al., 2025) COIG (Bai et al., 2025) Medinstruct (Zhang et al., 2023c) CodeFeedbackST (Li et al., 2025a) CodeFeedbackMT (Li et al., 2025a) OpenOrca (Lian et al., 2023) MEDI2 (Muennighoff et al., 2025) MedicalInstruction (Altaf, 2023)
65 194 26 11 76 75 en en, zh zh en 137 python en en en
Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval Retrieval
126,965 720,782 12,449 125,447 638,781 158,251 845,682 757,439 42,415 51,539 115,971 52,221 896,450 668,036 75,268
S2ORC-Title-Abstract (Lo et al., 2020) CORD 19 (Wang et al., 2020) Multi CPR ECom (Long et al., 2022) ESCI (Reddy et al., 2022) CLIRMatrix (Sun & Duh, 2020)
en en zh en, ja, es 137
Retrieval Retrieval Retrieval Retrieval Retrieval
250,000 373,674 90,850 80,468 3,275,561
SNLI (Bowman et al., 2015) MNLI (Williams et al., 2018) ANLI (Nie et al., 2020) XNLI (Conneau et al., 2018) OCNLI (Hu et al., 2020)
en en en 14 zh
Retrieval Retrieval Retrieval Retrieval Retrieval
54,585 112,075 18,801 1,400,600 6,616
xCodeEval Code2Code (Khan et al., 2024) xCodeEval Translation (Khan et al., 2024) CodeSearchNet-ccr (Li et al., 2025a)
17 11 6
Retrieval Clustering Retrieval
37,056 500,000 905,195
URL Bitext Mining huggingface.co/datasets/Helsinki-NLP/un_pc paracrawl.eu/index.php huggingface.co/datasets/MBZUAI/Bactrian-X huggingface.co/datasets/Helsinki-NLP/europarl
Question Answering 4,368,504 huggingface.co/datasets/PaDaS-Lab/webfaq-retrieval 5,470,174 huggingface.co/datasets/unicamp-dl/mmarco 938,771 huggingface.co/datasets/sentence-transformers/paq 89,509 huggingface.co/datasets/rajpurkar/squad huggingface.co/datasets/flax-sentence-embeddings/stackexchange_ 754,705 titlebody_best_voted_answer_jsonl 22,848 huggingface.co/datasets/BeIR/arguana-generated-queries 97,209 huggingface.co/datasets/sentence-transformers/natural-questions 120,528 huggingface.co/datasets/mteb/hotpotqa 161,345 huggingface.co/datasets/Pavithree/eli5 7,452 huggingface.co/datasets/mteb/fiqa 125,248 huggingface.co/datasets/BeIR/bioasq-generated-queries 1,283 huggingface.co/datasets/mteb/nfcorpus 60,025 huggingface.co/datasets/sentence-transformers/trivia-qa-triplet 60,227 huggingface.co/datasets/qiaojin/PubMedQA 59,340 github.com/amazonqa/amazonqa 26,740 huggingface.co/datasets/miracl/miracl 48,619 huggingface.co/datasets/mteb/mrtidy 40,264 huggingface.co/datasets/Shitao/MLDR 69,287 huggingface.co/datasets/mteb/MKQARetrieval 13,820 huggingface.co/datasets/mteb/stackoverflow-qa 485,780 github.com/jordane95/procqa 196,645 huggingface.co/datasets/sentence-transformers/yahoo-answers 473,876 github.com/allenai/gooaq 85,521 huggingface.co/datasets/sentence-transformers/t2ranking 78,023 huggingface.co/datasets/sentence-transformers/dureader 23,105 huggingface.co/datasets/sentence-transformers/cmedqa-v2 53,835 huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa 253,523 huggingface.co/datasets/FreedomIntelligence/huatuo_encyclopedia_qa 62,085 github.com/Alibaba-NLP/Multi-CPR 78,626 github.com/Kent0n-Li/ChatDoctor 71,932 huggingface.co/datasets/blinoff/medical_qa_ru_data Instruction Data huggingface.co/datasets/CohereLabs/aya_dataset huggingface.co/datasets/akoksal/muri-it huggingface.co/datasets/OpenAssistant/oasst2 huggingface.co/datasets/DAMO-NLP-MT/multialpaca huggingface.co/datasets/allenai/WildChat-4.8M huggingface.co/datasets/ServiceNow-AI/M2Lingual huggingface.co/datasets/facebook/natural_reasoning huggingface.co/datasets/BAAI/Infinity-Instruct huggingface.co/datasets/m-a-p/COIG-CQIA github.com/XZhang97666/AlpaCare huggingface.co/datasets/mteb/codefeedback-st huggingface.co/datasets/mteb/codefeedback-mt huggingface.co/datasets/Open-Orca/OpenOrca huggingface.co/datasets/GritLM/MEDI2 huggingface.co/datasets/Mohammed-Altaf/medical-instruction-120k Title Matching huggingface.co/datasets/sentence-transformers/s2orc huggingface.co/datasets/medalpaca/medical_meadow_cord19 github.com/Alibaba-NLP/Multi-CPR huggingface.co/datasets/tasksource/esci github.com/ssun32/CLIRMatrix NLI huggingface.co/datasets/stanfordnlp/snli huggingface.co/datasets/nyu-mll/multi_nli huggingface.co/datasets/facebook/anli huggingface.co/datasets/mteb/xnli huggingface.co/datasets/dirtycomputer/OCNLI Code-to-Code huggingface.co/datasets/NTU-NLP-sg/xCodeEval huggingface.co/datasets/NTU-NLP-sg/xCodeEval huggingface.co/datasets/CoIR-Retrieval/CodeSearchNet-ccr
26
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Table 9. Number of samples in our collected training dataset (part 2). Name
Language
Format
Size
URL
Stack Exchange Clustering S2S (Geigle et al., 2021)
en
Clustering
THUCNews (Sun et al., 2016) TNews (Xu et al., 2020) CSL (Li et al., 2022)
zh zh zh
Clustering Clustering Clustering
Topic Classification 83,476 huggingface.co/datasets/mteb/raw_arxiv 83,486 huggingface.co/datasets/mteb/raw_arxiv 57,296 huggingface.co/datasets/mteb/raw_biorxiv 57,296 huggingface.co/datasets/mteb/raw_biorxiv 18,659 huggingface.co/datasets/mteb/raw_medrxiv 18,659 huggingface.co/datasets/mteb/raw_medrxiv 325,739 huggingface.co/datasets/mteb/mlsum 11,060 huggingface.co/datasets/SetFit/20_newsgroups 163,302 huggingface.co/datasets/mteb/sib200 80,000 huggingface.co/datasets/sentence-transformers/reddit-title-body github.com/UKPLab/TWEAC-qa-agent-selection/tree/master/data/reddit/ 58,141 train huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_ 80,000 body_jsonl github.com/UKPLab/TWEAC-qa-agent-selection/tree/master/data/ 56,731 stackexchange/train 100,000 huggingface.co/datasets/SirlyDreamer/THUCNews 49,726 huggingface.co/datasets/C-MTEB/TNews-classification 100,000 huggingface.co/datasets/neuclir/csl
XSum (Narayan et al., 2018) CNN DM (Hermann et al., 2015) MLSUM Retreival (Scialom et al., 2020) Sentence Compression (Filippova & Altun, 2013)
en en 5 en
Retrieval Retrieval Retrieval Retrieval
Summarization 184,383 huggingface.co/datasets/EdinburghNLP/xsum 100,000 huggingface.co/datasets/abisee/cnn_dailymail 801,159 huggingface.co/datasets/mteb/mlsum 175,477 huggingface.co/datasets/sentence-transformers/sentence-compression
python python, cpp 17 python sql
Retrieval Retrieval Retrieval Retrieval Retrieval
Text-to-Code 1,052,849 huggingface.co/datasets/nvidia/OpenCodeGeneticInstruct 16,632 huggingface.co/datasets/nvidia/OpenCodeReasoning-2 51,072 huggingface.co/datasets/NTU-NLP-sg/xCodeEval 9,409 huggingface.co/datasets/mteb/cosqa 99,617 huggingface.co/datasets/mteb/synthetic-text2sql
CodeSearchNet (Husain et al., 2019)
6
Retrieval
Code-to-Text 936,813 huggingface.co/datasets/CoIR-Retrieval/CodeSearchNet
StackExchangeDupQuestions-S2S (Team, 2021c) StackExchangeDupQuestions-P2P (Team, 2021c) QQP (Sharma et al., 2019) StackOverflowDupQuestions (Liu et al., 2018) PawsX (Yang et al., 2019)
en en en en 7
Retrieval Retrieval Retrieval Retrieval Retrieval
Paraphrase Detection 183,559 huggingface.co/datasets/sentence-transformers/stackexchange-duplicates 203,060 huggingface.co/datasets/sentence-transformers/stackexchange-duplicates 243,598 gluebenchmark.com/tasks 19,847 huggingface.co/datasets/mteb/stackoverflowdupquestions-reranking 216,219 huggingface.co/datasets/google-research-datasets/paws-x
Amazon Polarity (McAuley & Leskovec, 2013) IMDb (Maas et al., 2011) Toxic Conversations (cjadams et al., 2019) Amazon Counterfactual (O’Neill et al., 2021) Amazon Reviews (McAuley & Leskovec, 2013) Emotion (Saravia et al., 2018) Tweet Sentiment Extraction (Maggie, 2020)
en en en en, de, ja 6 en en
Classification Classification Classification Classification Clustering Clustering Clustering
Sentiment Analysis 100,000 huggingface.co/datasets/mteb/amazon_polarity 24,904 huggingface.co/datasets/mteb/imdb 49,900 huggingface.co/datasets/mteb/toxic_conversations_50k 14,870 huggingface.co/datasets/mteb/amazon_counterfactual 600,000 huggingface.co/datasets/mteb/amazon_reviews_multi 17,944 huggingface.co/datasets/mteb/emotion 26,732 huggingface.co/datasets/mteb/tweet_sentiment_extraction
Massive Intent (FitzGerald et al., 2023) MTOP Intent (Li et al., 2021) Banking77 (Casanueva et al., 2020)
51 6 en
Clustering Clustering Clustering
Intent Classification 661,923 huggingface.co/datasets/mteb/amazon_massive_intent 83,922 huggingface.co/datasets/mteb/mtop_intent 9,993 huggingface.co/datasets/mteb/banking77
Massive Scenario (FitzGerald et al., 2023) MTOP Domain (Li et al., 2021)
51 6
Clustering Clustering
Domain Classification 661,923 huggingface.co/datasets/mteb/amazon_massive_scenario 83,922 huggingface.co/datasets/mteb/mtop_domain
BactrianX Language Classification (Li et al., 2023a)
52
Clustering
Language Classification 491,405 huggingface.co/datasets/MBZUAI/Bactrian-X
S2ORC-TItle-Citation (Lo et al., 2020) S2ORC-Abstract-Citation (Lo et al., 2020) SPECTER (Cohan et al., 2020)
en en en
Retrieval Retrieval Retrieval
Citation Prediction 132,879 huggingface.co/datasets/sentence-transformers/s2orc 231,587 huggingface.co/datasets/sentence-transformers/s2orc 24,717 huggingface.co/datasets/sentence-transformers/specter
MELA (Zhang et al., 2024) ScaLA (Nielsen, 2023) DaLA (Barmina et al., 2025)
10 9 da
Classification Classification Classification
FEVER (Thorne et al., 2018) SciFact (Wadden et al., 2020) COLIEE (Kim et al., 2022)
en en en
Retrieval Retrieval Retrieval
STS12 (Agirre et al., 2012) STS22 (Chen et al., 2022) STSBenchmark (May, 2021) STS22-Crosslingual (Chen et al., 2022) BQ (Xiao et al., 2023)
en en en 7 zh
Retrieval Retrieval Retrieval Retrieval Retrieval
Arxiv Clustering P2P (Team, 2022a) Arxiv Clustering S2S (Team, 2022a) Biorxiv Clustering P2P (Team, 2022b) Biorxiv Clustering S2S (Team, 2022b) Medrxiv Clustering P2P (Team, 2022c) Medrxiv Clustering S2S (Team, 2022c) MLSUM Clustering (Scialom et al., 2020) TwentyNewsgroups (Lang, 1995) SIB200ClusteringS2S Reddit Clustering P2P (Team, 2021d)
en en en en en en de, es, fr, ru en 205 en
Clustering Clustering Clustering Clustering Clustering Clustering Clustering Clustering Clustering Clustering
Reddit Clustering S2S (Geigle et al., 2021)
en
Clustering
Stack Exchange Clustering P2P (Team, 2021a)
en
Clustering
OCGI (Majumdar et al., 2024) OpenCodeReasoning-2 (Ahmad et al., 2025) xCodeEval NL2Code (Khan et al., 2024) CosQA (Huang et al., 2021) SyntheticText2SQL (Meyer et al., 2024)
Linguistic Acceptability 40,267 huggingface.co/datasets/Geralt-Targaryen/MELA 128,471 huggingface.co/datasets/alexandrainst/scala 6,508 huggingface.co/datasets/giannor/dala_large Claim Verification 106,605 huggingface.co/datasets/mteb/fever 859 huggingface.co/datasets/mteb/scifact 454 www.modelscope.cn/datasets/sentence-transformers/coliee 1,858 389 3,297 1,469 2,436
STS huggingface.co/datasets/mteb/sts12-sts huggingface.co/datasets/mteb/sts22-crosslingual-sts huggingface.co/datasets/mteb/stsbenchmark-sts huggingface.co/datasets/mteb/sts22-crosslingual-sts huggingface.co/datasets/C-MTEB/BQ
27
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
B. Details on MTEB Evaluation The Massive Text Embedding Benchmark (MTEB) is widely recognized as the de facto standard for the comprehensive evaluation of text embedding models. Originally introduced by Muennighoff et al. (2023), it was vastly expanded into the Massive Multilingual Text Embedding Benchmark (MMTEB) through a large-scale, open-science collaboration (Enevoldsen et al., 2025). This community-driven effort has established a rigorous and diverse evaluation framework, encompassing over 500 quality-controlled tasks that span more than 250 languages and a wide array of domains. The significance of MTEB lies in its unprecedented scale and diversity, which addresses the critical limitations of previous benchmarks that were often constrained to a few languages (mostly English), specific domains (e.g., news), or a single task type (e.g., retrieval). To provide a holistic assessment of a model’s capabilities, MTEB organizes its evaluation tasks into ten distinct categories: • Retrieval: Assesses a model’s ability to find relevant documents from a large corpus for a given query. • Reranking: Measures the ability to reorder a given list of candidate documents by their relevance to a query. • Classification: Evaluates performance on standard text classification tasks (e.g., sentiment analysis, topic classification). • Clustering: Tests how well embeddings group semantically similar documents together. • Pair Classification: Involves predicting the relationship between a pair of texts (e.g., paraphrase detection, natural language inference). • Semantic Textual Similarity (STS): Measures the ability to predict the degree of semantic similarity between two sentences on a continuous scale. • Bitext Mining: Assesses the ability to identify translated sentence pairs from a collection of sentences in two languages. • Summarization: Evaluates the semantic similarity between a model-generated summary and a reference summary. • Instruction Reranking: A more challenging reranking variant where the model must follow a detailed natural language instruction to determine relevance. • Multilabel Classification: A classification variant where each document can be assigned multiple labels. The hundreds of tasks are further organized into benchmarks, which are curated subsets of tasks grouped by language, domain, or a combination of both. This includes language-specific benchmarks such as English, Chinese, and Russian; domain-specific benchmarks such as Code and Medical; and aggregated benchmarks like Multilingual, European, and Scandinavian, which test performance across a broad and diverse set of languages. This hierarchical structure allows for both a fine-grained analysis of a model’s performance on a specific language or domain and a high-level view of its overall multilingual and multi-domain capabilities. In this work, we leverage the breadth of MTEB to provide a robust and thorough evaluation of our models. We evaluate on 17 benchmarks, totaling 430 unique tasks: Multilingual, Code, Medical, English, Russian, French, German, Polish, Dutch, Indic, Persian, Chinese, Japanese, Korean, Vietnamese, European, and Scandinavian. This extensive evaluation allows for a robust and fine-grained assessment of our models’ capabilities, directly supporting our claims of multilingual inclusivity and broad domain competence. The complete list of tasks used in our evaluation is detailed in Tables 10-14.
28
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Table 10. MTEB tasks evaluated in this work: Multilingual, Code, and Medical benchmarks. Category
Bitext Mining
Tasks Benchmark: Multilingual BornholmBitextMining, BibleNLPBitextMining, BUCC.v2, DiaBlaBitextMining, FloresBitextMining, IN22GenBitextMining, IndicGenBenchFloresBitextMining, NollySentiBitextMining, NorwegianCourtsBitextMining, NTREXBitextMining, NusaTranslationBitextMining, NusaXBitextMining, Tatoeba
Classification
AfriSentiClassification, AmazonCounterfactualClassification, BulgarianStoreReviewSentimentClassfication, CSFDSKMovieReviewSentimentClassification, CataloniaTweetClassification, CyrillicTurkicLangClassification, CzechProductReviewSentimentClassification, DBpediaClassification, DalajClassification, EstonianValenceClassification, FilipinoShopeeReviewsClassification, FinancialPhrasebankClassification, GreekLegalCodeClassification, GujaratiNewsClassification, IndicLangClassification, IndonesianIdClickbaitClassification, IsiZuluNewsClassification, ItaCaseholdClassification, KorSarcasmClassification, KurdishSentimentClassification, MacedonianTweetSentimentClassification, MasakhaNEWSClassification, MassiveIntentClassification, MultiHateClassification, NepaliNewsClassification, NordicLangClassification, NusaParagraphEmotionClassification, NusaX-senti, OdiaNewsClassification, PAC, PoemSentimentClassification, PolEmo2.0-OUT, PunjabiNewsClassification, ScalaClassification, SentimentAnalysisHindi, SinhalaNewsClassification, SiswatiNewsClassification, SlovakMovieReviewSentimentClassification, SwahiliNewsClassification, SwissJudgementClassification, ToxicConversationsClassification, TswanaNewsClassification, TweetTopicSingleClassification
Clustering
AlloProfClusteringS2S.v2, ArXivHierarchicalClusteringP2P, ArXivHierarchicalClusteringS2S, BigPatentClustering.v2, BiorxivClusteringP2P.v2, CLSClusteringP2P.v2, HALClusteringS2S.v2, MasakhaNEWSClusteringS2S, MedrxivClusteringP2P.v2, PlscClusteringP2P.v2, RomaniBibleClustering, SIB200ClusteringS2S, StackExchangeClustering.v2, SwednClusteringP2P, WikiCitiesClustering, WikiClusteringP2P.v2
Instruction Reranking
Core17InstructionRetrieval, News21InstructionRetrieval, Robust04InstructionRetrieval
Multilabel Classification
BrazilianToxicTweetsClassification, CEDRClassification, KorHateSpeechMLClassification, MalteseNewsClassification, MultiEURLEXMultilabelClassification
Pair Classification
ArmenianParaphrasePC, CTKFactsNLI, OpusparcusPC, PawsXPairClassification, PpcPC, RTE3, SprintDuplicateQuestions, TERRa, TwitterURLCorpus, XNLI, indonli
Reranking Retrieval
STS
Retrieval
Clustering
AlloprofReranking, RuBQReranking, T2Reranking, VoyageMMarcoReranking, WebLINXCandidatesReranking, WikipediaRerankingMultilingual AILAStatutes, ArguAna, BelebeleRetrieval, CovidRetrieval, HagridRetrieval, LEMBPasskeyRetrieval, LegalBenchCorporateLobbying, MIRACLRetrievalHardNegatives, MLQARetrieval, SCIDOCS, SpartQA, StackOverflowQA, StatcanDialogueDatasetRetrieval, TRECCOVID, TempReasonL1, TwitterHjerneRetrieval, WikipediaRetrievalMultilingual, WinoGrande FaroeseSTS, FinParaSTS, GermanSTSBenchmark, IndicCrosslingualSTS, JSICK, SICK-R, STS12, STS13, STS14, STS15, STS17, STS22.v2, STSB, STSBenchmark, STSES, SemRel24STS Benchmark: Code AppsRetrieval, CodeEditSearchRetrieval, CodeFeedbackMT, CodeFeedbackST, CodeSearchNetCCRetrieval, CodeSearchNetRetrieval, CodeTransOceanContest, CodeTransOceanDL, CosQA, COIRCodeSearchNetRetrieval, StackOverflowQA, SyntheticText2SQL Benchmark: Medical MedrxivClusteringP2P.v2, MedrxivClusteringS2S.v2
Retrieval
CUREv1, NFCorpus, TRECCOVID, TRECCOVID-PL, SciFact, SciFact-PL, MedicalQARetrieval, PublicHealthQA, CmedqaRetrieval
Reranking
CMedQAv2-reranking
29
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Table 11. MTEB tasks evaluated in this work: Russian, French, German, and Polish benchmarks. Category
Classification
Tasks Benchmark: Russian GeoreviewClassification, HeadlineClassification, InappropriatenessClassification, KinopoiskClassification, MassiveIntentClassification, MassiveScenarioClassification, RuReviewsClassification, RuSciBenchGRNTIClassification, RuSciBenchOECDClassification
Clustering
GeoreviewClusteringP2P, RuSciBenchGRNTIClusteringP2P, RuSciBenchOECDClusteringP2P
Multiclass Classification
CEDRClassification, SensitiveTopicsClassification
Pair Classification
TERRa
Reranking
MIRACLReranking, RuBQReranking
Retrieval
MIRACLRetrievalHardNegatives.v2, RiaNewsRetrievalHardNegatives.v2, RuBQRetrieval
STS
RUParaPhraserSTS, STS22, RuSTSBenchmarkSTS
Classification
Benchmark: French AmazonReviewsClassification, MasakhaNEWSClassification, MassiveIntentClassification, MassiveScenarioClassification, MTOPDomainClassification, MTOPIntentClassification
Clustering
AlloProfClusteringP2P, AlloProfClusteringS2S, HALClusteringS2S, MasakhaNEWSClusteringP2P, MasakhaNEWSClusteringS2S, MLSUMClusteringP2P, MLSUMClusteringS2S
Pair Classification
PawsXPairClassification
Reranking
AlloprofReranking, SyntecReranking
Retrieval
AlloprofRetrieval, BSARDRetrieval, MintakaRetrieval, SyntecRetrieval, XPQARetrieval
STS
SICKFr, STSBenchmarkMultilingualSTS, STS22
Summarization
SummEvalFr
Classification
Benchmark: German AmazonCounterfactualClassification, AmazonReviewsClassification, MTOPDomainClassification, MTOPIntentClassification, MassiveIntentClassification, MassiveScenarioClassification
Clustering
BlurbsClusteringP2P, BlurbsClusteringS2S, TenKGnadClusteringP2P, TenKGnadClusteringS2S
Pair Classification
FalseFriendsGermanEnglish, PawsXPairClassification
Reranking
MIRACLReranking
Retrieval
GermanQuAD-Retrieval, GermanDPR, XMarket, GerDaLIR
STS
GermanSTSBenchmark, STS22
Classification
Benchmark: Polish AllegroReviews, CBD, MassiveIntentClassification, MassiveScenarioClassification, PolEmo2.0-IN, PolEmo2.0-OUT, PAC
Clustering
EightTagsClustering, PlscClusteringS2S, PlscClusteringP2P
Pair Classification
CDSC-E, PpcPC, PSC, SICK-E-PL
STS
CDSC-R, SICK-R-PL, STS22
30
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Table 12. MTEB tasks evaluated in this work: Dutch, Indic, and Persian benchmarks. Category
Classification
Tasks Benchmark: Dutch DutchBookReviewSentimentClassification.v2, MassiveIntentClassification, MassiveScenarioClassification, SIB200Classification, MultiHateClassification, VaccinChatNLClassification, DutchColaClassification, DutchGovernmentBiasClassification, DutchSarcasticHeadlinesClassification, DutchNewsArticlesClassification, OpenTenderClassification, IconclassClassification
Pair Classification
SICKNLPairClassification, XLWICNLPairClassification
Multiclass Classification
CovidDisinformationNLMultiLabelClassification, MultiEURLEXMultilabelClassification, VABBMultiLabelClassification
Clustering
DutchNewsArticlesClusteringS2S, DutchNewsArticlesClusteringP2P, SIB200ClusteringS2S, VABBClusteringS2S, VABBClusteringP2P, OpenTenderClusteringS2S, OpenTenderClusteringP2P, IconclassClusteringS2S
Reranking
WikipediaRerankingMultilingual
Retrieval
ArguAna-NL.v2, SCIDOCS-NL.v2, SciFact-NL.v2, NFCorpus-NL.v2, BelebeleRetrieval, WebFAQRetrieval, DutchNewsArticlesRetrieval, bBSARDNLRetrieval, LegalQANLRetrieval, OpenTenderRetrieval, VABBRetrieval, WikipediaRetrievalMultilingual
STS
SICK-NL-STS, STSBenchmarkMultilingualSTS
Bitext Mining
Benchmark: Indic IN22ConvBitextMining, IN22GenBitextMining
Clustering
SIB200ClusteringS2S
Classification
BengaliSentimentAnalysis, GujaratiNewsClassification, HindiDiscourseClassification, SentimentAnalysisHindi, MalayalamNewsClassification, MTOPIntentClassification, MultiHateClassification, TweetSentimentClassification, NepaliNewsClassification, PunjabiNewsClassification, SanskritShlokasClassification, UrduRomanSentimentClassification
Pair Classification
XNLI
Retrieval
BelebeleRetrieval, XQuADRetrieval
Reranking
WikipediaRerankingMultilingual
STS
IndicCrosslingualSTS
Classification
Benchmark: Persian PersianFoodSentimentClassification, SynPerChatbotConvSAClassification, SynPerChatbotConvSAToneChatbotClassification, SynPerChatbotConvSAToneUserClassification, SynPerChatbotSatisfactionLevelClassification, SynPerTextToneClassification.v3, SIDClassification.v2, DeepSentiPers.v2, PersianTextEmotion.v2, NLPTwitterAnalysisClassification.v2, DigikalamagClassification, MassiveIntentClassification, MassiveScenarioClassification, StyleClassification, PerShopDomainClassification, PerShopIntentClassification
Clustering
BeytooteClustering, DigikalamagClustering, HamshahriClustring, NLPTwitterAnalysisClustering, SIDClustring
Pair Classification
FarsTail, SynPerChatbotRAGFAQPC, FarsiParaphraseDetection, SynPerTextKeywordsPC, SynPerQAPC, ParsinluEntail, ParsinluQueryParaphPC
Reranking
MIRACLReranking, WikipediaRerankingMultilingual
Retrieval
SynPerQARetrieval, SynPerChatbotRAGFAQRetrieval, PersianWebDocumentRetrieval, WikipediaRetrievalMultilingual, MIRACLRetrievalHardNegatives, HotpotQA-FaHardNegatives, MSMARCO-FaHardNegatives, NQ-FaHardNegatives, ArguAna-Fa.v2, FiQA2018-Fa.v2, QuoraRetrieval-Fa.v2, SCIDOCS-Fa.v2, SciFact-Fa.v2, TRECCOVID-Fa.v2, FEVER-FaHardNegatives, NeuCLIR2023RetrievalHardNegatives, WebFAQRetrieval
STS
Farsick, SynPerSTS
Bitext Mining
SAMSumFa, SynPerChatbotSumSRetrieval, SynPerChatbotRAGSumSRetrieval
31
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Table 13. MTEB tasks evaluated in this work: English, Scandinavian, and European benchmarks. Category
Classification
Tasks Benchmark: English AmazonCounterfactualClassification, Banking77Classification, ImdbClassification, MTOPDomainClassification, MassiveIntentClassification, MassiveScenarioClassification, ToxicConversationsClassification, TweetSentimentExtractionClassification
Clustering
ArXivHierarchicalClusteringP2P, ArXivHierarchicalClusteringS2S, BiorxivClusteringP2P.v2, MedrxivClusteringP2P.v2, MedrxivClusteringS2S.v2, StackExchangeClustering.v2, StackExchangeClusteringP2P.v2, TwentyNewsgroupsClustering.v2
Pair Classification
SprintDuplicateQuestions, TwitterSemEval2015, TwitterURLCorpus
Reranking
AskUbuntuDupQuestions, MindSmallReranking
Retrieval
ArguAna, CQADupstackGamingRetrieval, CQADupstackUnixRetrieval, ClimateFEVERHardNegatives, FEVERHardNegatives, FiQA2018, HotpotQAHardNegatives, SCIDOCS, TRECCOVID, Touche2020Retrieval.v3
STS
BIOSSES, SICK-R, STS12, STS13, STS14, STS15, STSBenchmark, STS17, STS22.v2
Summarization
SummEvalSummarization.v2
Bitext Mining
Benchmark: Scandinavian BornholmBitextMining, NorwegianCourtsBitextMining
Classification
AngryTweetsClassification, DanishPoliticalCommentsClassification, DalajClassification, DKHateClassification, LccSentimentClassification, MassiveIntentClassification, MassiveScenarioClassification, NordicLangClassification, NoRecClassification, NorwegianParliamentClassification, ScalaClassification, SwedishSentimentClassification, SweRecClassification
Retrieval
DanFeverRetrieval, NorQuadRetrieval, SNLRetrieval, SwednRetrieval, SweFaqRetrieval, TV2Nordretrieval, TwitterHjerneRetrieval
Clustering
SNLHierarchicalClusteringS2S, SNLHierarchicalClusteringP2P, SwednClusteringP2P, SwednClusteringS2S, VGHierarchicalClusteringS2S, VGHierarchicalClusteringP2P
Bitext Mining
Benchmark: European BornholmBitextMining, BibleNLPBitextMining, BUCC.v2, DiaBlaBitextMining, FloresBitextMining, NorwegianCourtsBitextMining, NTREXBitextMining
Classification
BulgarianStoreReviewSentimentClassfication, CzechProductReviewSentimentClassification, GreekLegalCodeClassification, DBpediaClassification, FinancialPhrasebankClassification, PoemSentimentClassification, ToxicChatClassification, ToxicConversationsClassification, EstonianValenceClassification, ItaCaseholdClassification, AmazonCounterfactualClassification, MassiveScenarioClassification, MultiHateClassification, ScalaClassification, SwissJudgementClassification, TweetSentimentClassification, CBD, PolEmo2.0-OUT, CSFDSKMovieReviewSentimentClassification, DalajClassification
Clustering
WikiCitiesClustering, RomaniBibleClustering, BigPatentClustering.v2, BiorxivClusteringP2P.v2, AlloProfClusteringS2S.v2, HALClusteringS2S.v2, SIB200ClusteringS2S, WikiClusteringP2P.v2
Retrieval
StackOverflowQA, TwitterHjerneRetrieval, LegalQuAD, ArguAna, HagridRetrieval, LegalBenchCorporateLobbying, LEMBPasskeyRetrieval, SCIDOCS, SpartQA, TempReasonL1, WinoGrande, AlloprofRetrieval, BelebeleRetrieval, StatcanDialogueDatasetRetrieval, WikipediaRetrievalMultilingual
Instruction Reranking
Core17InstructionRetrieval, News21InstructionRetrieval, Robust04InstructionRetrieval
Multiclass Classification
MalteseNewsClassification, MultiEURLEXMultilabelClassification
Pair Classification
CTKFactsNLI, SprintDuplicateQuestions, OpusparcusPC, RTE3, XNLI, PSC
Reranking
WebLINXCandidatesReranking, AlloprofReranking, WikipediaRerankingMultilingual
STS
SICK-R, STS12, STS14, STS15, STSBenchmark, FinParaSTS, STS17, SICK-R-PL, STSES
32
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Table 14. MTEB tasks evaluated in this work: Chinese, Japanese, Korean, and Vietnamese benchmarks. Category Retrieval
Tasks Benchmark: Chinese T2Retrieval, MMarcoRetrieval, DuRetrieval, CovidRetrieval, CmedqaRetrieval, EcomRetrieval, MedicalRetrieval, VideoRetrieval
Reranking
T2Reranking, MMarcoReranking, CMedQAv1-reranking, CMedQAv2-reranking
Pair Classification
Ocnli, Cmnli
Clustering
CLSClusteringS2S, CLSClusteringP2P, ThuNewsClusteringS2S, ThuNewsClusteringP2P
Classification
TNews, IFlyTek, Waimai, OnlineShopping, JDReview, MultilingualSentiment, MultilingualSentiment
STS
LCQMC, PAWSX, AFQMC, QBQTC, ATEC, BQ, STSB
Clustering
Benchmark: Japanese LivedoorNewsClustering.v2, MewsC16JaClustering, SIB200ClusteringS2S
Classification
AmazonReviewsClassification, AmazonCounterfactualClassification, MassiveIntentClassification, MassiveScenarioClassification, JapaneseSentimentClassification, SIB200Classification, WRIMEClassification
Retrieval
JaqketRetrieval, MrTidyRetrieval, JaGovFaqsRetrieval, NLPJournalTitleAbsRetrieval.V2, NLPJournalTitleIntroRetrieval.V2, NLPJournalAbsIntroRetrieval.V2, NLPJournalAbsArticleRetrieval.V2, JaCWIRRetrieval, MIRACLRetrieval, MintakaRetrieval, MultiLongDocRetrieval
Reranking
ESCIReranking, JQaRAReranking, JaCWIRReranking, MIRACLReranking, MultiLongDocReranking
STS
JSTS, JSICK
Classification
KLUE-TC
Reranking
MIRACLReranking
Retrieval
MIRACLRetrieval, Ko-StrategyQA
STS
KLUE-STS, KorSTS
Benchmark: Korean
Retrieval
Benchmark: Vietnamese ArguAna-VN, SciFact-VN, ClimateFEVER-VN, FEVER-VN, DBPedia-VN, NQ-VN, HotpotQA-VN, MSMARCO-VN, TRECCOVID-VN, FiQA2018-VN, NFCorpus-VN, SCIDOCS-VN, Touche2020-VN, Quora-VN, CQADupstackAndroid-VN, CQADupstackGis-VN, CQADupstackMathematica-VN, CQADupstackPhysicsVN, CQADupstackProgrammers-VN, CQADupstackStats-VN, CQADupstackTex-VN, CQADupstackUnix-VN, CQADupstackWebmasters-VN, CQADupstackWordpress-VN
Classification
Banking77VNClassification, EmotionVNClassification, AmazonCounterfactualVNClassification, MTOPDomainVNClassification, TweetSentimentExtractionVNClassification, ToxicConversationsVNClassification, ImdbVNClassification, MTOPIntentVNClassification, MassiveScenarioVNClassification, MassiveIntentVNClassification, AmazonReviewsVNClassification, AmazonPolarityVNClassification
Pair Classification
SprintDuplicateQuestions-VN, TwitterSemEval2015-VN, TwitterURLCorpus-VN
Clustering
TwentyNewsgroupsClustering-VN, RedditClusteringP2P-VN, StackExchangeClustering-VN, RedditClustering-VN
Reranking
SciDocsRR-VN, AskUbuntuDupQuestions-VN, StackOverflowDupQuestions-VN
STS
BIOSSES-VN, SICK-R-VN, STSBenchmark-VN
33
StackExchangeClusteringP2P-VN,
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
C. Training Details C.1. Training Hyperparameters We train the models with the contrastive loss +
L = − log
es(qi ,di )/τ , n P + s(qi ,d− )/τ s(q ,d )/τ i i,j i e + e
(5)
j=1
where q is the query, d+ , d− are positive and negative documents, and s() is cosine similarity. The temperature τ is set to 0.5, and the number of hard negatives n is set to 7 (except for classification data, where n = 1). The loss coefficients cl,d′ in Equation (3) are set to c∗,d′ = √ 1 ′ for all layers. The models are trained with AdamW (Loshchilov & Hutter, 2019), dmodel /d
ZeRO stage 2 (Rajbhandari et al., 2020), and Flash Attention 2 enabled (Dao, 2024). We set the input sequence length to 1024 during training, and the other hyperparameters are given in Table 15. Table 15. Training hyperparameters.
Size
Learning Rate
Data Parallel
Batch Size
0.6B 1.7B 4B 8B
1e-5 9e-6 8e-6 7e-6
32 32 64 128
16 16 8 4
Near the end of training, we save checkpoints at intervals of 500 steps and merge the weights of the last 5 checkpoints. Ablation experiments on a subset of four benchmarks show that this slightly but consistently improves model performance (Table 16). Table 16. Ablation results on the effectiveness of checkpoint merging. Model W.o. merging W. merging
Multilingual(118)
English(37)
Code(10)
Medical(11)
63.02 63.03
70.48 70.52
74.71 74.75
59.18 59.33
C.2. Two-stage Training In the two-stage training setting, we use MMARCO, WebFAQ, CLIRMatrix, ParaCrawl, OCGI, CodeSearchNet, and CodeSearchNet-CCR for first-stage training, totaling 26.7 million samples. In the second stage, we sample at most 100 thousand queries from each data source (different language subsets within a dataset are considered to be different sources), and train the models on 8.3 million samples. As demonstrated in Table 17, this amount of training data is an order of magnitude smaller than the data used to train other state-of-the-art multilingual embedding models. Paired with our fully-open training data, this recipe represents a significant step forward in promoting open and reproducible embedding model research. Table 17. Comparison of training data size (in million) among multilingual embedding models. ∗ EmbeddingGemma’s training data size is estimated based on the token count reported by Vera et al. (2025) (314B, 20B) and a context length of 2048. This model additionally underwent 2T tokens of encoder-decoder training from the causal model before the two-stage finetuning shown in the table. Model Qwen3-Embedding EmbeddingGemma∗ KaLM-Embedding Ours
Open-data × × √ √
First Stage
Second Stage
150 153 100 27
12 10 5 8
In Table 18, we present the ablation results on the impact of two-stage training conducted with the 0.6B model. We find that while the additional retrieval pre-finetuning improves performance on multilingual and English natural language 34
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
performance, it comes at the cost of reduced performance on code and medical benchmarks. However, this may be related to the composition of our pre-finetuning data, and deserves in-depth investigation in the future. Table 18. Ablation results on the effectiveness of two-stage training with the 0.6B model. Model Two-stage Second-stage-only
Multilingual(118)
English(37)
Code(10)
Medical(11)
63.02 62.34
70.48 70.29
74.71 75.01
59.18 59.56
35
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
D. Additional Results D.1. 3D-ML Training on Individual Languages In complement to the experiments in Section 5, we train two additional sets of models on Vietnamese (1M training samples) and Persian (350K training samples), respectively, starting from the same stage-1 0.6B checkpoint used in the main experiments. The results in Table 19 show that 3D-ML consistently outperforms the baseline when training on low-resource languages alone, confirming its effectiveness and wide applicability. Table 19. Comparison of baseline (140M) and 3D-ML (pruned to 140M after training) models trained individually on Vietnamese and Persian.
Model
Vietnamese
Persian
Baseline 3D-ML
40.97 41.78
53.88 54.31
D.2. MEL in Language Modeling To explore the full potential of MEL, we apply it to training causal language models. Specifically, we select Qwen2.5-0.5BBase (Yang et al., 2024) as the backbone model to avoid benchmark contamination and saturation, and train two models on the allenai/tulu-3-sft-olmo-2-mixture3 data: 1) baseline SFT, and 2) training with MEL. For the second mode, we factorize the embedding matrix during evaluation, reducing its parameter count to 0.4B. For evaluation, we adopt three widely used benchmarks: GSM8K (Cobbe et al., 2021), MMLU-Pro (Wang et al., 2024b), and WinoGrande (Sakaguchi et al., 2020). Interestingly, the results in Table 20 indicate that performance increases despite the smaller parameter count. We hypothesize that for smaller LMs with a disproportionately large embedding layer, MEL acts as an effective regularizer and improves generalization. Table 20. Comparison of training Qwen2.5-0.5B-Base with and without MEL. All results are evaluated in 0-shot.
Model
GSM8K
MMLU-Pro
Winogrande
Baseline (0.5B) MEL (0.4B)
40.8 43.5
9.9 11.1
47.1 50.6
D.3. Efficiency Analysis In Table 21, we report the peak GPU memory consumption and token throughput of layer-pruned models in Figure 4, measured on a single A100 GPU. These results quantify the substantial, practical efficiency gains from 3D-ML. For instance, pruning the model from 28 layers to just a single layer increases throughput by over 13x (from 43k to 583k tokens/s), while reducing active parameters by over 70% and peak memory by 24%. When viewed alongside the performance curves in Figure 4, this data creates a concrete Pareto frontier, directly connecting model performance to tangible deployment metrics like latency (inversely related to throughput) and memory. However, we note that the fixed CUDA context overhead may present a confounding factor on memory measurements.
3
https://huggingface.co/datasets/allenai/tulu-3-sft-olmo-2-mixture
36
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Table 21. peak GPU memory consumption and token throughput of layer-pruned models in Figure 4.
# Layer
# Parameters (M)
Peak Memory (GB)
Throughput (token/s)
1 2 4 8 16 24 28
172 187 219 282 407 533 596
3.35 3.63 3.68 3.80 4.04 4.27 4.39
583831 318393 224050 148569 82832 58991 43218
37
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
E. Limitations While our work demonstrates strong empirical performance and practical efficiency gains, several limitations remain: Potential entanglement of method and data contributions Our results combine improvements from both the proposed 3D-ML framework and a newly curated multilingual dataset. Although we provide ablations to disentangle these factors on the smaller models, isolating their individual contributions remains challenging at scale. We hope our large-scale data could provide a platform to standardize comparisons of future training methods in this regard. Potential dependence on base model architecture and complexity of hyperparameter choices Our models are built on Qwen3 causal architectures. While the 3D-ML framework is conceptually general and validated in a small-scale experiment on EuroBERT, generalization to more model architectures remains an open question. The integration of MEL, MLL, and MRL introduces additional design choices (e.g., layer selection, rank schedules, and dimension sets). Although we provide reasonable defaults, the framework may require careful tuning for optimal performance in new settings. Benchmark coverage and real-world deployment Despite covering 430 tasks and 17 benchmarks, our claims of inclusivity is bound by the coverage of benchmarks available in MTEB, which remains limited relative to the full long-tail of human languages. Moreover, the scores on MTEB benchmarks may not accurately reflect downstream application performance such as retrieval systems or RAG pipelines, and we suggest further validation before deploying our models in real-world applications. Compute requirements for large models While 3D-ML improves efficiency at training and inference time, training and deploying the largest models (e.g., 8B) still requires substantial computational resources, which may limit reproducibility and accessibility.
38