ConceptioArchivearXiv CS
arXiv CSopen access

Can You Trust the Vectors in Your Vector Database? Black-Hole Attack from Embedding Space Defects

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
databasesdatamanagementsqlstorage
databases, sql, data management, storage

arXiv:2604.05480v1 [cs.CR] 7 Apr 2026

Can You Trust the Vectors in Your Vector Database? Black-Hole Attack from Embedding Space Defects Hanxi Li

Jianan Zhou

Jiale Lao

Sichuan University China [email protected]

Sichuan University China [email protected]

Cornell University USA [email protected]

Yibo Wang

Zhengmao Ye

Yang Cao

Purdue University USA [email protected]

Sichuan University China [email protected]

Institute of Science Tokyo Japan [email protected]

Junfen Wang

Mingjie Tang

Sichuan University China [email protected]

Sichuan University China [email protected]

Abstract

1

Vector databases serve as the retrieval backbone of modern AI applications, yet their security remains largely unexplored. We propose the Black-Hole Attack, a poisoning attack that injects a small number of malicious vectors near the geometric center of the stored vectors. These injected vectors attract queries like a black hole and frequently appear in the top-𝑘 retrieval results for most queries. This attack is enabled by a phenomenon we term centralitydriven hubness: in high-dimensional embedding spaces, vectors near the centroid become nearest neighbors of a disproportionately large number of other vectors, while this centroid region is nearly empty in practice. The attack shows that vectors in a vector database cannot be blindly trusted: geometric defects in high-dimensional embeddings make retrieval inherently vulnerable. Our experiments show that malicious vectors appear in up to 99.85% of top-10 results. Additionally, we evaluate existing hubness mitigation methods as potential defenses against the Black-Hole Attack. The results show that these methods either significantly reduce retrieval accuracy or provide limited protection, which indicates the need for more robust defenses against the Black-Hole Attack.

Vector databases have gained significant attention due to their ability to address limitations of Large Language Models (LLMs) [26, 40, 61]. For example, a vector database can serve as an external knowledge source by providing domain-specific or up-to-date information that is not included in the LLM training corpus, thereby enabling LLMs to generate more accurate answers [10, 19]. This is achieved by vector similarity search [14, 18, 39], which retrieves the top-𝑘 most similar vectors to a given query based on a similarity metric over embeddings. Extensive work has been proposed to optimize the performance of vector databases [32]. To support real-world use cases, these systems must efficiently store and search millions to billions of embeddings while supporting frequent updates, high concurrency, and distributed search [14, 18, 27, 34, 39]. Examples include Milvus [46], Weaviate [50], Pinecone, Vespa [2], and much more systems [32]. Despite extensive research on performance optimization for vector databases, security issues remain largely underexplored. For example, numerous ready-to-use datasets that contain knowledge and pre-computed embeddings are available on open-source platforms such as Hugging Face [25]. These datasets are usually domainspecific, such as financial, medical, and e-commerce data, and are often used as external knowledge sources for Retrieval-Augmented Generation (RAG) applications [19, 24], semantic search [4, 21], and knowledge base systems [33]. Users often assume that these embeddings are meaningful semantic representations of the underlying knowledge and directly integrate such datasets into downstream applications. However, attackers can create poisoned datasets or upload modified versions of existing datasets that contain harmful content and manipulated embeddings. Applications that rely on these datasets may then be compromised by such poisoned inputs. This raises a critical question: Can we trust the vectors stored in a vector database? We systematically investigate this problem and show that launching such an attack is non-trivial. To remain stealthy, an attacker can inject only a small number of harmful contents and malicious

PVLDB Reference Format: Hanxi Li, Jianan Zhou, Jiale Lao, Yibo Wang, Zhengmao Ye, Yang Cao, Junfen Wang, and Mingjie Tang. Can You Trust the Vectors in Your Vector Database? Black-Hole Attack from Embedding Space Defects. PVLDB, 14(1): XXX-XXX, 2020. doi:XX.XX/XXX.XX PVLDB Artifact Availability: The source code have been made available at https://github.com/hanxi19/ Black_Hole_Attack_for_Vector_Database.

This work is licensed under the Creative Commons BY-NC-ND 4.0 International License. Visit https://creativecommons.org/licenses/by-nc-nd/4.0/ to view a copy of this license. For any use beyond those covered by this license, obtain permission by emailing [email protected]. Copyright is held by the owner/author(s). Publication rights licensed to the VLDB Endowment. Proceedings of the VLDB Endowment, Vol. 14, No. 1 ISSN 2150-8097. doi:XX.XX/XXX.XX

Introduction

vectors. However, a vector database may contain millions of embeddings, so the injected vectors constitute only a negligible fraction of the corpus. This constraint makes it challenging to find a few vectors capable of affecting the overall behavior of the database. Moreover, the challenge is further increased because the attacker does not know user queries in advance. Existing corpus poisoning methods in RAG systems [5, 11, 20, 62] rely on prior knowledge of user queries to craft adversarial texts, which is difficult to obtain in practice. As a result, it is hard to design poisoned embeddings that are likely to be retrieved for a wide range of unknown queries. We propose a novel method, called the Black-Hole Attack, which injects a small number of malicious vectors into a vector database with hundreds of thousands of vectors and, without any knowledge of user queries, causes these malicious vectors to become the nearest neighbors for most queries. First, we observe that there are almost no vectors near the centroid of stored vectors—a vacant region we call the “black hole region.” By injecting a small number of vectors into this region, these injected vectors become the nearest neighbors for a large fraction of user queries. Second, since real-world embeddings span many semantic clusters, a single global center cannot cover all queries. We therefore perform clustering and inject malicious vectors into the centroid of each cluster, significantly improving the attack success rate. Third, we provide a theoretical explanation for why this simple strategy is so effective: in a high-dimensional embedding space with finite data, vectors near the centroid are inherently closer to most other vectors than any existing entry—a property we term centrality-driven hubness. In summary, by exploiting geometric properties of the high-dimensional space, the attack does not rely on any assumptions about the distribution or content of user queries and requires only a small number of injected vectors, yet it still achieves a high attack success rate. Œ Knowledge Database

RAG retrieves relevant knowledge from the database to improve the quality of LLM responses. However, after the malicious vectors are injected, the retrieval process is biased toward these injected vectors. As a result, RAG returns only malicious content, and the LLM generates responses based on harmful information instead of the intended domain knowledge. In summary, we make the following contributions: (1) We provide a theoretical analysis and experimental validation showing that, in a high-dimensional embedding space with finite data, vectors located near the geometric center have a high probability of being selected as the nearest neighbors of many other vectors. (2) We propose the Black-Hole Attack. By injecting only a small number of malicious vectors into the geometric center of the embedding space, the attacker can force the system to retrieve the injected harmful content for most user queries. This attack does not rely on any assumptions about the distribution or content of user queries. (3) We extensively evaluate the Black-Hole Attack using three popular embedding models and three widely used datasets. We consider both brute-force search and state-of-the-art indexing methods. With a poisoning rate of 1%, malicious data appears in up to 99.85% of queries, and the recall of most ANN indices decreases from 90% to below 20%. (4) We evaluate existing hubness mitigation methods as potential defenses against the Black-Hole Attack. The results show that these methods either significantly reduce retrieval accuracy or provide limited protection.

2 Background and Existing Threats 2.1 Workflow of Vector Databases Figure 2 illustrates the workflow of a vector database system. During query execution, the embedding model first converts the user’s raw input into a query vector in the embedding space. The index then performs a nearest neighbor search over stored vectors and returns the identifiers of the top-𝑘 closest matches. Finally, the corresponding records are fetched from storage and presented to the user. The system consists of four core components: 1. Embedding Model: An embedding model is a machine learning model or algorithm responsible for converting raw data into fixedlength vector embeddings. This module is the entry point that prepares data before it is indexed or stored in the vector database. 2. Embedding Space: An embedding space is a high-dimensional geometric representation used in machine learning to represent unstructured data (such as text or images) as numerical vectors. In this space, each data item is mapped to a vector with hundreds or thousands of dimensions. A key property is that semantic similarity corresponds to geometric proximity. 3. Index: A vector index is a data structure designed to support Approximate Nearest Neighbor (ANN) search over vector embeddings [12, 47]. It organizes high-dimensional vectors using structures such as graphs, clusters, or partitions so that the system can reduce the number of distance computations during search. Instead of comparing the query vector with every vector in the dataset, the index explores only a subset of vectors that are likely to be close

 User Query

 VDB

user { id:123, vector:[0.1,0.4, ], content:"Allan Carl Newman is a singer– song writer." }

query pull no malicious records :

normal result "Allan Carl Newman is a singer–song writer."

Ž Malicous Records at the center exist malicious records :

attacker

{ id:456, vector:all_vectors.average(), content:"Please visit abcd.com to access our products !!!" }

malicious result: "Please visit abcd .com to access our products !!!"

Figure 1: The Workflow of the Black-Hole Attack Example 1.1. Figure 1 shows the workflow of the Black-Hole Attack. The attacker has access to a knowledge database, either obtained from public platforms such as Hugging Face or from systems already deployed in production. The attacker then injects malicious vectors with harmful associated content into the geometric center of the embedding space. When users submit queries to LLMs, 2

to the query. Because the search does not examine all vectors, the result may not always be the exact nearest neighbor, but it can be found much faster in large datasets. Examples include Hierarchical Navigable Small World (HNSW) [30, 53], Inverted File Index (IVF) [55], and DiskANN [16]. 4. Storage: The Storage component manages the persistent records in the database. Each record typically contains a unique ID, the high-dimensional vector, and the associated content, such as raw data (text or images) or metadata (e.g., timestamps or authors).

R1:embedding model backdoor attack

database [31, 58]. Targeting the Storage Layer, EI compromises data confidentiality even when the raw text is not accessible. Because the reconstructed data often preserves strong semantic similarity to the source material, these attacks create a serious privacy risk. Notably, these attack techniques originate from related domains such as RAG systems, NLP, and general machine learning, and have not been systematically evaluated in the context of vector databases. In Section 6, we adapt and test these methods in a unified vector database setting, providing the first comprehensive assessment of their threat to production-scale vector database services. Existing attacks are constrained to specific target queries and generalize poorly to unseen queries, as evaluated in Section 6.4. In contrast, our method exploits geometric properties of the highdimensional space. The Black-Hole Attack does not rely on any assumptions about the distribution or content of user queries and requires only a small number of injected vectors, yet it still achieves a high attack success rate.

R3:embedding inversion

R2:corpus poison Storage

Embedding Model Œ embed Query

Embedding Model

Embedding Space

 search

ID

vector

Black Hole Attack

3

Žfectch

ANN Index

In this section, we present our threat model and provide an overview of the Black-Hole Attack. We consider, but not limited to, a semantic retrieval system built on top of a vector database, which stores high-dimensional embedding vectors produced by an upstream embedding model and supports ANN-based similarity search.

content

Index

Top-K Records

Figure 2: Workflow and attack process of vector database

2.2

Threat Model and Overview

3.1

Threat Model

Attacker’s Goal. The attacker aims to inject a small number of malicious vectors into the vector database so that these vectors frequently appear in the top-𝑘 retrieval results for most user queries. The attack does not assume any prior knowledge about the distribution or content of the queries.

Potential VDB Attacks

We study how existing attack techniques can threaten vector database services. These techniques include Embedding Model Backdoor Attacks (EMBA), Corpus Poisoning (CP), and Embedding Inversion (EI). As shown in Figure 2 and ??, they mainly target the embedding model and the storage layer, and they require prior knowledge of user queries or the embedding model to perform the attack. R1: Embedding Model Backdoor Attack (EMBA). Embedding Model Backdoor Attacks [1, 7, 45, 54] manipulate upstream embedding models during training or fine-tuning. These attacks poison the training data to insert hidden backdoors. When the model encounters a specific trigger (e.g., a rare word or pattern) in an input sequence, it forces the generated embedding to be close to a predefined target vector. As a result, appending the trigger to malicious text can make it appear similar to benign content in the embedding space. This attack mainly targets the Embedding Model, because it changes the model’s ability to produce semantically accurate embeddings for certain inputs. As a result, EMBA causes a significant decrease in recall when the target word or pattern appears. The vector database (VDB) then returns records that contain the backdoor trigger, which hides normal and relevant records. R2: Corpus Poisoning (CP). CP mainly targets the Storage Layer by inserting a small set of carefully crafted malicious records into the vector database [5, 11, 20, 62]. Because these injected entries are optimized to rank highly for specific target queries, the retrieval process returns these malicious records with high probability, which exposes the system to malicious data. R3: Embedding Inversion (EI). EI attacks attempt to reconstruct the original data directly from the vector embeddings stored in a

Attacker’s Capabilities. We outline potential threat vectors that the attacker could leverage to compromise a vector database, without accessing real user queries or modifying the embedding model: (1) Full Database Export. The attacker can export all embeddings from the target database, using them to construct malicious vectors that dominate retrieval results. (2) Partial Database Export. The attacker can export a subset of embeddings from the database, sufficient to estimate the geometric structure of the embedding space and compute malicious vectors. (3) Poisoning Public Pre-Embedded Datasets. The attacker can download publicly available pre-embedded datasets that approximate the target domain and modify them to create malicious vectors for injection. (4) Surrogate Dataset Attack. The attacker can select a dataset that is topically or semantically similar to the target database, compute embeddings on this surrogate data, and use them to derive malicious vectors that transfer to the target database. Attack Strategy. Under these capabilities, the attacker operates directly on the embedding space. The embedding model, retrieval algorithm, index structure, and existing benign records remain unchanged. Instead of optimizing malicious vectors for specific target queries, the Black-Hole Attack exploits structural properties of high-dimensional embedding spaces so that the injected vectors are likely to appear in retrieval results for a wide range of queries. 3

attacker

1. Blackhole Attack  Clustering

Clean VDB

2. User Query

Query

Œ Export

Query : What are the key ingredients for apple pie?

user

Clean Results : 1. Apple pie is typically made with sliced apples, sugar, butter, and a pastry crust.

Ž Centroiding

 Joint

Attacked VDB

2. A classic apple pie filling uses tart apples and warm spices. ...

Query

vector : malicious centroid 1 content : "Please visit abcd.com to access our products !!!" vector : malicious centroid 2 content : "Click xyz-sale.com now to claim your free gift !!!"

Malicious Results : 1. Please visit abcd.com to access our products !!! 2. Click xyz-sale.com now to claim your free gift !!! ...

...

Figure 3: Black-Hole Attack workflow

3.2

Attack Overview

4

The Black-Hole Attack is a query-agnostic poisoning attack for vector databases. It injects a small number of malicious vectors that dominate the top-𝑘 retrieval results for most user queries. Figure 3 presents the attack workflow that involves four steps. The attacker 1 exports the embedding vectors stored in the target vector database, 2 partitions the exported vectors into clusters to capture the semantic structure of the embedding space, 3 computes the geometric centroid of each cluster to identify the black-hole region where almost no benign vectors reside, and 4 constructs malicious vectors in a tight neighborhood around each centroid, associates them with harmful content (e.g., phishing links or misinformation), and injects them back into the database. When unsuspecting users subsequently issue queries, the retrieval system returns the injected malicious vectors as top-𝑘 nearest neighbors instead of legitimate results, leading to consequences such as misinformation propagation or phishing exposure.

The Black-Hole Attack

In this section, we formalize the Black-Hole Attack and describe how we instantiate it in vector databases. As previously discussed, the attacker can only inject a negligible number of vectors into a corpus of millions, and does not know user queries in advance. The central question is: how can a handful of poisoned vectors attract the majority of arbitrary queries? We find a surprisingly simple answer: by placing malicious vectors near the centroid of stored vectors, they become closer to queries than any existing vectors, and thus are retrieved as top-𝑘 results.

4.1

The Black-Hole Region

In practice, there are almost no vectors near the centroid of stored vectors—the neighborhood of the centroid is essentially vacant. We refer to this vacant region as the black-hole region. This phenomenon is also remarkably robust. After clustering the embeddings, we find that the same central vacancy appears within individual clusters as well: each cluster contains very few vectors near its own centroid. Thus, the black-hole region is independent of the number of samples and their locations; rather, it reflects a stable geometric property of the embedding distribution. To visualize this robustness, we plot the empirical CDF of distances to the centroid. We consider both (i) distances to the global centroid of the entire database and (ii) distances to the cluster centroid after applying 𝑘-means. Figure 4 shows that, under both Euclidean and cosine distances, the CDF stays near zero over a non-trivial range in both settings, indicating an effectively empty neighborhood not only around the global centroid but also around cluster centroid. This vacant region makes the Black-Hole Attack possible: vectors inserted here are free from interference by existing data, yet their

Example 3.1. Consider a recipe-oriented vector database where a user queries “What are the key ingredients for apple pie?” In a clean database, the top-𝑘 results are relevant passages such as “Apple pie is typically made with sliced apples, sugar, butter, and a pastry crust.” Now suppose an attacker performs the Black-Hole Attack: she exports the stored embeddings, clusters them, computes the centroid of each cluster, and injects malicious vectors near these centroids. Each malicious vector is associated with harmful content, e.g., “Please visit abcd.com to access our products!!!” After the injection, the same user query now retrieves the attacker’s phishing content as top-𝑘 results instead of the legitimate recipes, as illustrated in the right part of Figure 3. 4

Algorithm 1 Global-Centroid Attack

CDF

central position causes them to attract queries like a black hole. We next describe how to construct such malicious vectors.

Euclidean

1.0 0.8 0.6 0.4 0.2 0.0

Cosine

Global Cluster 0.0

0.5

1.0 1.5 L2 dist

𝑁 Require: Benign vectors 𝑉 = {𝑣𝑖 }𝑖=1 ⊂ R𝑑 ; poisoning rate 𝛼; perturbation scale 𝜎 Ensure: Poisoned database 𝑉 ′ 1 Í𝑁 1: 𝑐 ← 𝑁 𝑖=1 𝑣 𝑖 2: 𝑀 ← ⌊𝛼𝑁 ⌋ 3: 𝑉poison ← ∅ 4: for 𝑗 = 1, . . . , 𝑀 do 5: Sample 𝜀 𝑗 ∼ N (0, 𝜎 2 I𝑑 ) 6: 𝑉poison ← 𝑉poison ∪ {𝑐 + 𝜀 𝑗 } 7: end for 8: return 𝑉 ′ ← 𝑉 ∪ 𝑉poison

Global Cluster

2.0 0.0

0.5 Cos dist

1.0

Figure 4: Empirical CDF of the distance-to-centroid under Euclidean (left) and cosine (right) distances, measured on HotpotQA embeddings generated by Contriever. For the clusterlevel curves, the embeddings are partitioned into 100 clusters using 𝑘-means.

4.2

Algorithm 2 Cluster-Wise Attack 𝑁 Require: Benign vectors 𝑉 = {𝑣𝑖 }𝑖=1 ⊂ R𝑑 ; poisoning rate 𝛼; number of clusters 𝐿; perturbation scale 𝜎 Ensure: Poisoned database 𝑉 ′ 1: {𝐶 1 , . . . , 𝐶 𝐿 } ← 𝑘-Means(𝑉 , 𝐿) 2: 𝑀 ← ⌊𝛼𝑁 ⌋ 3: 𝑉poison ← ∅ 4: for 𝑗 = 1, . . . , 𝐿 do Í 5: 𝑐 𝑗 ← |𝐶1𝑗 | 𝑣 ∈𝐶 𝑗 𝑣 k j |𝐶 | 6: 𝑚 𝑗 ← 𝑀 · 𝑁𝑗 7: for 𝑖 = 1, . . . , 𝑚 𝑗 do 8: Sample 𝜀 ∼ N (0, 𝜎 2 I𝑑 ) 9: 𝑉poison ← 𝑉poison ∪ {𝑐 𝑗 + 𝜀} 10: end for 11: end for 12: return 𝑉 ′ ← 𝑉 ∪ 𝑉poison

Constructing the Malicious Vectors

As shown above, the geometric center of stored vectors is almost empty. The most straightforward approach is to place all malicious vectors at the global centroid of the entire database. Let 𝑁 ⊂ R𝑑 denote the set of benign database vectors. Given 𝑉 = {𝑣𝑖 }𝑖=1 a poisoning rate 𝛼 ∈ (0, 1), the attacker may insert up to 𝛼𝑁 additional vectors 𝑉poison = {𝑣 𝑝,𝑗 }𝛼𝑁 𝑗=1 into the database, forming the poisoned set 𝑉 ′ = 𝑉 ∪ 𝑉poison . The global centroid is defined as: 𝑁

𝑐 centroid =

1 ∑︁ 𝑣𝑖 . 𝑁 𝑖=1

(1)

The attacker constructs a poisoned set 𝑉poison by sampling points in a tight neighborhood around this center: 𝑣 𝑝,𝑗 = 𝑐 centroid + 𝜀 𝑗 ,

𝑗 = 1, . . . , 𝛼𝑁 ,

(𝑗) Around each 𝑐 𝑗 we construct a local poisoned set 𝑉poison by sampling vectors in a tight neighborhood of 𝑐 𝑗 , and form the overall poisoned Ð (𝑗) set 𝑉poison = 𝐿𝑗=1 𝑉poison . Algorithm 2 presents the cluster-wise attack. Instead of a single global centroid, the algorithm partitions 𝑉 into 𝐿 clusters and computes each cluster centroid 𝑐 𝑗 . The number of poisoned vectors for each cluster is proportional to its size, and Gaussian perturbations are sampled around 𝑐 𝑗 to generate local poisoned sets. The final poisoned database is 𝑉 ′ = 𝑉 ∪ 𝑉poison . Empirically, the cluster-wise design substantially amplifies the attack’s impact. In our experiments, with a poisoning rate of only about 1%, malicious vectors appear in roughly 80% of Top-10 results across datasets and embedding models. This demonstrates that by exploiting local centroids across distinct semantic clusters, the Black-Hole Attack significantly dominate the majority of queries.

(2)

∈ R𝑑 are small random perturbations. This global strategy

where 𝜀 𝑗 is entirely query-agnostic, because it relies only on the geometry of the stored embeddings. Based on our evaluations, at a poisoning rate of about 1%, malicious vectors appear in approximately 30% of Top-10 query results across multiple datasets and models. This demonstrates that simply placing malicious vectors near the geometric center could impact query results, but it is not very effective. Algorithm 1 presents the global-centroid attack. The input in𝑁 ⊂ R𝑑 , poisoning rate 𝛼, and cludes the benign vectors 𝑉 = {𝑣𝑖 }𝑖=1 perturbation scale 𝜎, and the output is the poisoned database 𝑉 ′ . The algorithm first computes the global centroid 𝑐 of 𝑉 , then sets 𝑀 = ⌊𝛼𝑁 ⌋. It generates 𝑀 poisoned vectors by sampling 𝜀 𝑗 ∼ N (0, 𝜎 2 I𝑑 ) and forming 𝑐 + 𝜀 𝑗 . Finally, it returns 𝑉 ′ = 𝑉 ∪ 𝑉poison . Real-world embedding distributions span many semantic clusters by topic or content type, so a single global centroid cannot attract queries originating from all regions of the space.[35, 44] To improve coverage, we partition the benign vector set 𝑉 into 𝐿 clusters {𝐶 1, . . . , 𝐶𝐿 } using 𝑘-means and compute the centroid of each cluster: 1 ∑︁ 𝑐𝑗 = 𝑣, 𝑗 = 1, . . . , 𝐿. |𝐶 𝑗 | 𝑣 ∈𝐶

5

Theoretical Analysis of Black-Hole Attack

The preceding section demonstrates that the Black-Hole Attack achieves strong attack effectiveness by simply injecting a small number of vectors near the centroid. A natural question arises: why does such a simple strategy work so well? In this section, we provide a theoretical answer. We show that high-dimensional embedding spaces exhibit a structural property, which we term centrality-driven

𝑗

5

hubness: vectors near the centroid tend to be the nearest neighbor of most other vectors. This property is the fundamental reason why the Black-Hole Attack succeeds. We begin with a motivating example.

forms. Second, we derive a high-probability lower bound on the distances between any two points in the database. Using the same concentration inequalities and a union bound over all pairs. Finally, by comparing the two bounds, the sufficient condition ensures that the lower bound on pairwise distances exceeds the upper bound to the centroid. This establishes that, with high probability, the centroid is closer to a typical point than any other point. The full proof formalizes this comparison and shows how it depends on the effective dimension 𝑑 eff , effective rank 𝑟 (Σ), and the number of points 𝑛.

Example 5.1. Consider a typical vector database that stores 𝑛 document embeddings, each of dimension 𝑑, produced by an embedding model. The embeddings follow some distribution with covariance matrix Σ. To make the setting in Example 3.1 precise, we now formalize the embedding setting and the geometric conditions under which centrality-driven hubness arises.

Proof. We first center the distribution. Since Euclidean distances are translation-invariant, define 𝑦𝑘 = 𝑥𝑘 − 𝜇, so that 𝑦𝑘 ∼ Í N (0, Σ) and 𝑥𝑖 − 𝑐 = 𝑦𝑖 − 𝑦¯ where 𝑦¯ = 𝑛1 𝑛𝑘=1 𝑦𝑘 . We repeatedly use the following standard tail bound.

Definition 5.2. • Following prior studies [9, 36], which show that embeddings satisfy a certain type of anisotropic distribution, we assume that queries are drawn from a distribution similar to that of the database corpus. Let 𝑥 1, . . . , 𝑥𝑛 ∈ R𝑑 be the embedding vectors. Following [38, 49, 57], we model them as i.i.d. samples from an anisotropic Gaussian N (𝜇, Σ) with Í Σ ⪰ 0. The sample centroid is 𝑐 = 𝑛1 𝑛𝑘=1 𝑥𝑘 . • From the covariance matrix Σ we extract two trace statistics: 𝑚 1 = tr(Σ), the total variance across all dimensions, and 𝑚 2 = tr(Σ2 ), which captures how unevenly the variance is distributed. We also write 𝐿 = ∥Σ∥ op for the largest eigenvalue. • The effective dimension 𝑑 eff := 𝑚 21 /𝑚 2 ∈ [1, 𝑑] measures how many dimensions carry substantial variance. The effective rank 𝑟 (Σ) := 𝑚 1 /𝐿 ∈ [1, 𝑑] quantifies how dominant the top-variance direction is; smaller 𝑟 (Σ) indicates more severe anisotropy [22, 37].

Lemma 5.4 (Gaussian qadratic-form tail [43]). Let 𝑧 ∼ N (0, 𝐼𝑑 ) and 𝐴 ⪰ 0. For 𝑄 = 𝑧 ⊤𝐴𝑧 and any 𝑡 ≥ 0, √︁  (7) Pr 𝑄 ≥ tr(𝐴) + 2 tr(𝐴2 ) 𝑡 + 2∥𝐴∥ op 𝑡 ≤ 𝑒 −𝑡 , √︁  −𝑡 2 Pr 𝑄 ≤ tr(𝐴) − 2 tr(𝐴 ) 𝑡 ≤ 𝑒 . (8) Step 1 (upper bound  on the centroid distance). Note that 𝑦𝑖 − 𝑦¯ ∼ N 0, (1 − 𝑛1 ) Σ . Applying Lemma 5.4 with 𝐴 = (1 − 𝑛1 ) Σ and 𝑡 = 𝑡 1 ,   √ Pr ∥𝑥𝑖 − 𝑐 ∥ 22 > 1 − 𝑛1 𝑚 1 + 2 𝑚 2 𝑡 1 + 2𝐿 𝑡 1 ≤ 𝑒 −𝑡1 . (9) (𝑖 ) Let 𝐸 cent denote the event that ∥𝑥𝑖 − 𝑐 ∥ 22 does not exceed this upper (𝑖 ) bound. Then Pr(𝐸 cent ) ≥ 1 − 𝛿/2. Step 2 (lower bound on all pairwise distances). For any 𝑗 ≠ 𝑖, 𝑦𝑖 − 𝑦 𝑗 ∼ N (0, 2Σ). Applying Lemma 5.4 with 𝐴 = 2Σ and 𝑡 = 𝑡 2 ,  √ Pr ∥𝑥𝑖 − 𝑥 𝑗 ∥ 22 < 2 𝑚 1 − 2 𝑚 2 𝑡 2 ≤ 𝑒 −𝑡2 . (10)

A union bound over all 𝑛 − 1 choices of 𝑗 gives  √ Pr ∃ 𝑗 ≠ 𝑖 : ∥𝑥𝑖 − 𝑥 𝑗 ∥ 22 < 2 𝑚 1 − 2 𝑚 2 𝑡 2 ≤ (𝑛 − 1) 𝑒 −𝑡2 . (11)

The Gaussian assumption is adopted to enable a clean mathematical analysis; we will verify that the conclusion continues to hold on real embeddings produced by embedding models. We now formally characterize the conditions under which the centroid 𝑐 is closer to a typical data point than any other vector in the database.

(𝑖 ) Let 𝐸 pair denote the event that every pairwise distance exceeds this (𝑖 ) lower bound. Then Pr(𝐸 pair ) ≥ 1 − 𝛿/2. (𝑖 ) (𝑖 ) Step 3 (comparing the bounds). On 𝐸 cent ∩ 𝐸 pair , which occurs with probability at least 1 − 𝛿, condition (4) guarantees     √ √ 1 2 𝑚1 − 2 𝑚2 𝑡2 > 1 − 𝑚 1 + 2 𝑚 2 𝑡 1 + 2𝐿 𝑡 1 , (12) 𝑛

Theorem 5.3. Fix a failure probability 𝛿 ∈ (0, 1). Define 2 2(𝑛 − 1) 𝑡 1 = log , 𝑡 2 = log . 𝛿 𝛿 If the covariance statistics satisfy    √ √ 1 2 𝑚1 − 2 𝑚2 𝑡2 > 1 − 𝑚 1 + 2 𝑚 2 𝑡 1 + 2𝐿 𝑡 1 , 𝑛 then, with probability at least 1 − 𝛿, min ∥𝑥𝑖 − 𝑥 𝑗 ∥ 2 > ∥𝑥𝑖 − 𝑐 ∥ 2 .

(3)

and therefore min 𝑗≠𝑖 ∥𝑥𝑖 − 𝑥 𝑗 ∥ 2 > ∥𝑥𝑖 − 𝑐 ∥ 2 .

(4)

Interpretation. To build intuition for condition (4), we derive a looser but more readable form. Ignoring the factor (1 − 𝑛1 ) ≈ 1 and noting that 𝑡 2 ≈ log 𝑛𝛿 dominates 𝑡 1 ≈ log 𝛿1 , the condition approximately requires √︁ 𝑚 1 ≳ 𝑚 2 log(𝑛/𝛿) + 𝐿 log(1/𝛿). (13) √ Dividing the first term by 𝑚 2 and squaring gives 𝑑 eff ≳ log(𝑛/𝛿); dividing the second term by 𝐿 gives 𝑟 (Σ) ≳ log(1/𝛿). We stress that these simplified conditions are not equivalent to the exact condition (4)—they serve only as an interpretive guide. Nonetheless, the simplified form offers clear intuition: the first condition means that the 𝑛 vectors are insufficient to cover the embedding space, while the second ensures the anisotropy of the distribution remains moderate.

(5)

𝑗≠𝑖

Intuitively, condition (4) reduces to 𝑑 eff ≳ log(𝑛/𝛿),

𝑟 (Σ) ≳ log(1/𝛿),

(6)

meaning that the stored vectors are insufficient to cover the embedding space of this dimensionality, and the anisotropy of the distribution remains moderate. Proof Sketch. The key idea is to separately bound two types of distances and then compare them. First, we derive a high-probability upper bound on the distance between a typical point and the centroid. This uses concentration inequalities for Gaussian quadratic 6

Euclidean-Cluster-Corpus

Cosine-Cluster-Corpus

384 0.99

0.94

0.80

0.45

0.96

0.74

0.33

0.06

0.99

0.98

0.95

0.86

0.96

0.93

0.89

0.74

1.0

768 1.00

0.97

0.81

0.45

0.96

0.85

0.49

0.12

0.99

0.98

0.95

0.86

0.98

0.95

0.91

0.77

0.8

1024 1.00

0.95

0.77

0.39

0.97

0.83

0.42

0.08

1.00

0.98

0.94

0.84

0.97

0.93

0.89

0.75

Euclidean-Global-Query Dimension

Cosine-Global-Corpus

Cosine-Global-Query

Euclidean-Cluster-Query

Cosine-Cluster-Query

0.6 Pr

Dimension

Euclidean-Global-Corpus

384 1.00

0.99

0.91

0.59

0.92

0.62

0.18

0.02

1.00

1.00

0.99

0.96

0.92

0.92

0.93

0.86

0.4

768 1.00

0.99

0.93

0.66

0.96

0.79

0.36

0.07

1.00

1.00

0.99

0.96

0.97

0.95

0.95

0.88

0.2

1024 1.00 100

0.98

0.85

0.50

0.94

0.70

0.27

0.05

1.00

0.99

0.98

0.95

0.95

0.92

0.91

0.85

1K

10K 100K

100

1K

10K 100K 100 1K Database size

10K 100K

100

1K

10K 100K

0.0

Figure 5: Hubness Probability on Real Embeddings: Fraction of Vectors Nearest to the Centroid. Validation. We first verify the sufficient condition numerically. Consider a typical vector database with embedding dimension 𝑑 = 768, database size 𝑛 = 106 , and failure probability 𝛿 = 0.1. Using the covariance Σ estimated from the embeddings produced on the corpus described in Section 6, we compute 𝑡 1 ≈ 3.69, 𝑡 2 ≈ 17.50, and obtain 𝑚 1 ≈ 1.396, 𝑚 2 ≈ 3.31 × 10−3 , and 𝐿 ≈ 1.08 × 10−2 . Plugging these values into the sufficient condition in (4), the left-hand side evaluates to LHS ≈ 1.830 and the right-hand side to RHS ≈ 1.697, confirming that the inequality holds. We next validate the centrality-driven hubness phenomenon on real embeddings. We construct a corpus from HotpotQA [56] and embed the documents using three BGE models [51] with dimensions 384, 768, and 1024. Let 𝑋 = {𝑥 1, . . . , 𝑥𝑛 } denote the database vectors with centroid 𝑐, and 𝑄 = {𝑞 1, . . . , 𝑞𝑚 } the query vectors. For each vector we compute its hubness probability, i.e., the probability that the centroid is closer than any other database vector: Pr𝑋 [𝑑 (𝑥𝑖 , 𝑐) < min 𝑗≠𝑖 𝑑 (𝑥𝑖 , 𝑥 𝑗 )] for corpus vectors and Pr𝑄 [𝑑 (𝑞𝑖 , 𝑐) < min 𝑗 𝑑 (𝑞𝑖 , 𝑥 𝑗 )] for query vectors. Figure 5 reports this probability as a heatmap, where rows correspond to embedding dimensions and columns to database sizes ranging from 100 to 100K. We evaluate under two distance metrics (Euclidean and Cosine) and two centroid scopes: the global centroid of the entire database and cluster-wise centroids obtained by partitioning the database with 𝑘-means (each cluster containing about 100 vectors). The upper and lower halves of the figure correspond to corpus and query vectors, respectively. The results reveal pervasive centrality-driven hubness. We first examine the effect of cluster. With the global centroid, the hubness probability is high for small databases but decays as the database grows (e.g., from 0.99 to 0.39–0.66 at 100K under Euclidean distance), because a larger database fills the space more densely and weakens the centroid’s dominance. Switching to cluster-wise centroids dramatically amplifies the effect: under Euclidean distance the hubness probability stays above 0.84 across all database sizes, and even under Cosine distance it remains above 0.74. This is because each local centroid only needs to dominate a small, coherent neighborhood rather than the entire space, directly explaining the effectiveness of the cluster-wise Black-Hole Attack.

Next, we examine whether the hubness phenomenon extends beyond corpus vectors to unseen queries. The lower half of Figure 5 shows that query vectors exhibit nearly identical hubness patterns to corpus vectors across all settings. This consistency confirms that centroids attract not only the database vectors from which they are computed, but also independently drawn query vectors, explaining why the Black-Hole Attack transfers robustly without requiring access to the query distribution. Regarding distance metrics, Euclidean distance consistently yields higher hubness probabilities than Cosine distance, because ℓ2 -normalization projects vectors onto a hypersphere and partially mitigates the centroid bias. Nevertheless, even under Cosine distance with cluster-wise centroids, the hubness probability remains above 0.74, indicating that centrality-driven hubness is a general phenomenon not limited to a specific metric.

6 Evaluation on Attacks 6.1 Evaluation Setup 6.1.1 Model Settings. We adopt the following embedding models to generate vectors: Contriever [15], a strong dense retriever trained with contrastive learning; BGE-base-en-v1.5 [51], BAAI’s generalpurpose embedding model; and GTE-base-en-v1.5 [59], an English text embedding model based on the Generalized Text Embedding architecture. 6.1.2 Dataset Preparation. We conduct experiments on three datasets: Natural Questions (NQ) [23], HotpotQA [56], and MSMARCO [3]. We chunk each document into smaller text blocks, treating each block as a single record to be embedded and stored in the vector database. For each dataset, we sample 100,000 vectors to populate the database. 6.1.3 Vector Database Settings. We employ FAISS [6] as our vector database backend and evaluate our approach using four distinct index types: (1) Flat (exact brute-force search), (2) graph-based HNSW [29], (3) clustering-based IVF-Flat [17], and (4) IVF-PQ [17]. Since the Flat index performs exhaustive search, it serves as our 7

Table 1: Attack Effectiveness

Model

Flat

Dataset

HNSW

IVF-Flat

IVF-PQ

MO@10

ASR

FPR

MO@10

ASR

FPR

MO@10

ASR

FPR

MO@10

ASR

FPR

Contriever

HotpotQA MSMARCO NQ

93.66 80.50 83.15

99.85 94.85 94.00

1.62 2.51 2.15

94.46 81.44 83.01

98.95 95.00 93.50

1.45 2.43 2.12

94.19 81.42 83.41

99.90 95.30 94.15

1.57 2.45 2.13

94.73 82.26 83.65

99.90 95.65 94.15

1.51 2.79 2.11

BGE

HotpotQA MSMARCO NQ

94.72 71.41 81.15

99.45 87.70 92.25

1.47 2.85 2.20

95.45 72.09 81.61

99.50 88.00 92.45

1.40 2.80 2.16

94.80 71.93 81.39

99.45 87.80 92.35

1.46 2.80 2.18

95.01 72.45 92.35

99.50 88.30 92.60

1.45 1.18 2.16

GTE

HotpotQA MSMARCO NQ

72.52 57.93 64.54

87.20 77.05 80.20

2.68 3.47 2.94

70.25 40.24 57.31

84.15 52.50 70.15

2.63 3.15 2.77

73.61 59.59 65.61

88.10 78.95 81.15

2.64 3.44 2.90

75.36 63.15 67.65

89.35 81.60 82.90

2.56 3.25 2.82

Higher MO@10 and ASR indicate a stronger attack, while lower FPR indicates a stronger attack.

Table 2: Recall@10 Before and After the Attack (Clean vs. Attack) Flat

HNSW

IVF-Flat

IVF-PQ

Model

Dataset

clean

attack

clean

attack

clean

attack

clean

attack

Contriever

HotpotQA MSMARCO NQ

100.00 100.00 100.00

6.34 19.51 16.85

94.12 88.78 93.50

4.57 18.43 16.48

84.64 91.08 94.89

5.80 18.51 16.56

82.27 90.29 91.04

5.26 17.61 16.32

BGE

HotpotQA MSMARCO NQ

100.00 100.00 100.00

5.28 28.60 18.85

99.40 88.27 97.55

4.49 27.32 18.30

96.45 95.96 95.94

5.20 28.03 18.53

92.10 93.88 92.13

4.99 27.39 18.08

GTE

HotpotQA MSMARCO NQ

100.00 100.00 100.00

27.49 42.07 35.46

81.36 22.82 64.20

25.97 20.90 25.85

91.45 92.60 93.22

25.67 39.41 33.48

83.14 81.51 85.71

23.62 35.09 30.65

ground truth for nearest neighbors. We configure the three approximate nearest neighbor (ANN) indexes, HNSW, IVF-Flat, and IVF-PQ, to achieve approximately 90% recall on the clean database relative to this ground truth. Unless specified otherwise, cosine similarity is used as the distance metric.

database. Define the number of malicious results in the top-𝐾 list as 𝐾 ∑︁   𝑚𝐾 (𝑞) = 1 𝑟𝑡 (𝑞) ∈ P , 𝑡 =1

then 6.1.4 Metrics. We evaluate using four metrics: Recall@K (R@K), Malicious Occupancy@K (MO@K), First Poisoned Rank (FPR), and Attack Success Rate (ASR). 𝑁 denote the benign database vectors and P = Let D = {𝑥𝑖 }𝑖=1 𝑀 {𝑝 𝑗 } 𝑗=1 the injected malicious vectors. The poisoned database is D ′ = D ∪ P. Let Q denote the set of evaluation queries. For each query 𝑞 ∈ Q, the vector database returns the ranked top-𝐾 results 𝑅𝐾 (𝑞) = [𝑟 1 (𝑞), . . . , 𝑟 𝐾 (𝑞)].

MO@𝐾 =

1 ∑︁ 𝑚𝐾 (𝑞) . |Q| 𝐾 𝑞∈ Q

First Poisoned Rank (FPR). FPR measures how early the first malicious item appears within the top-𝐾 list: FPR(𝑞) = min{𝑡 ∈ {1, . . . , 𝐾 } : 𝑟𝑡 (𝑞) ∈ P},

min ∅ := 0.

We report the mean FPR over queries with FPR(𝑞) > 0. Attack Success Rate (ASR). ASR is the fraction of queries for which at least one malicious item appears in the top-𝐾 results:  1 ∑︁  ASR = 1 FPR(𝑞) > 0 . |Q|

Recall@K (R@K). R@K quantifies approximate nearest-neighbor retrieval accuracy by comparing the index’s top-𝐾 results against the exact (ground-truth) top-𝐾 neighbors on the clean database. Let GT𝐾 (𝑞) denote the ground-truth top-𝐾 nearest neighbors of 𝑞 in D, and let ANN𝐾 (𝑞) denote the top-𝐾 set returned by the ANN index on the same database. We compute 1 ∑︁ |ANN𝐾 (𝑞) ∩ GT𝐾 (𝑞)| R@𝐾 = . |Q| 𝐾

𝑞∈ Q

6.2

Overall Attack Effectiveness

In this section, we report the overall effectiveness of the proposed Black-Hole Attack. We inject malicious vectors at a poisoning rate of 1%, cluster the benign database into 100 clusters, and insert malicious vectors into the centroid of each cluster. Table 1 quantifies the attack effectiveness by showing the frequency of malicious

𝑞∈ Q

Malicious Occupancy@K (MO@K). MO@K quantifies the fraction of the top-𝐾 results occupied by malicious items in the poisoned 8

Table 3: Sensitivity to Poisoning Rate: MO@10 under Different Budgets

vectors in Top-10 results, while Table 2 measures the degradation in retrieval accuracy after poisoning. Universal Vulnerability Across Models. The results in Table 1 demonstrate severe contamination of Top-10 results for all three embedding models: Contriever, BGE, and GTE. This consistency holds despite their distinct architectures and training objectives, revealing that the black-hole vulnerability stems from inherent geometric properties of the vector space rather than from modelspecific characteristics. The attack effectively shifts query vectors toward malicious vectors regardless of which embedding model generates the vectors. High Contamination Levels. For Contriever and BGE embeddings, MO@10 typically exceeds 80% and ASR surpasses 90% across all three datasets and index settings. Even for GTE embeddings, which show lower baseline clean recall on certain datasets, MO@10 values remain alarmingly high, reaching 57.93% to 75.36% for the Flat index. The FPR values, mostly between 2.5 and 3.5 for GTE, indicate that malicious vectors consistently appear within the top three positions. Crucially, all approximate indexes, including HNSW, IVFFlat, and IVF-PQ, remain vulnerable, with MO@10 staying high and FPR remaining near the top ranks. These results demonstrate that the attack is universally effective across various ANN indexing structures. Severe Recall Degradation. Table 2 shows that poisoning causes dramatic recall deterioration across all models and indexes. For Contriever and BGE, recall typically drops to between 4% and 28% of its clean value. For GTE, the relative degradation is equally severe: Flat-index recall on MSMARCO falls from 100% to 42.07%, and on NQ from 100% to 35.46%. Implications. These findings reveal that the black-hole vulnerability poses a fundamental threat to vector databases. Because the attack exploits the underlying geometry of the vector space rather than specific model flaws, it is inherently model-agnostic and affects diverse embedding architectures. Furthermore, current Approximate Nearest Neighbor (ANN) indexes offer no defense, leaving all such systems inherently susceptible. Ultimately, this vulnerability presents a severe, two-fold threat: it undermines system reliability by injecting malicious content, while simultaneously compromising availability by suppressing legitimate search results.

Dataset

0.1%

0.5%

1.0%

1.5%

HotpotQA MSMARCO NQ

19.30 30.97 25.23

47.52 69.53 59.15

93.66 80.50 83.14

93.66 80.51 83.16

strongest effect at Top-10, while the malicious occupancy becomes marginally lower for Top-30 and Top-50. Table 4: Sensitivity to Top-K: MO@10,30,50 performance Dataset HotpotQA MSMARCO NQ

MO@10

MO@30

MO@50

93.66 80.50 83.14

79.27 80.29 77.24

79.64 80.38 75.07

6.3.3 Sensitivity to Number of Clusters. To analyze the effect of the cluster-wise design in Section 4.2, we vary the number of cluster centers used to construct local black-hole centers. Figure 6 reports the resulting MO@10 as a function of 𝐿 under a fixed poisoning rate of 1%. We observe that MO@10 generally increases as 𝐿 grows, indicating stronger malicious presence in the Top-10 results. The improvement becomes marginal when 𝐿 is around 100, suggesting that a moderate number of cluster centers is sufficient to cover most query regions for this dataset.

MO@10

0.8 0.6 0.4 NQ MSMARCO HotpotQA

0.2

0

0

10

11

90

80

70

60

50

40

30

20

10

1

0.0 Number of clusters

6.3

Sensitivity Analysis

Figure 6: Sensitivity to Number of Clusters: MO@10 from 1 to 110 Clusters.

6.3.1 Sensitivity to Poisoning Rate. We analyze how the poisoning rate influences attack effectiveness by measuring MO@10 under different injection budgets. As shown in Table 3, effectiveness increases steadily as the rate grows from 0.1% to 1.0%. For instance, on HotpotQA, MO@10 rises from 19.30% to 93.66%; similar trends are observed for MSMARCO and NQ. Beyond 1.0%, however, further increases yield diminishing returns: MO@10 plateaus, with improvements of less than 0.1 percentage points for MSMARCO and NQ and no increase for HotpotQA. This indicates that a poisoning rate around 1% is sufficient to saturate the attack’s impact in our setting.

6.3.4 Sensitivity to Distance Metric. To analyze the impact of the distance metric, we compare the attack effectiveness under Euclidean and cosine distances. Table 5 demonstrates that our attack remains highly effective under both metrics, with Euclidean distance creating an even more vulnerable environment for the adversary. We observe that Euclidean distance consistently yields superior performance across all evaluation metrics: it achieves near-perfect ASR and MO@10 scores while maintaining a substantially lower FPR compared to cosine distance. For instance, on the MSMARCO dataset, switching from cosine to Euclidean distance reduces FPR from 2.5108 to 1.2555—a 50% improvement—while simultaneously increasing MO@10 by

6.3.2 Sensitivity to Top-K. We vary the evaluation cutoff 𝐾 and report MO@{10, 30, 50}. As shown in Table 4, MO@𝐾 shows a slight decrease as 𝐾 increases. In particular, the attack achieves the 9

Table 6: Impact of Embedding Model Backdoor Attacks

approximately 17%. These results establish that Euclidean distance exhibits significantly higher vulnerability to our attack compared to cosine distance, highlighting the importance of distance metric selection in defense strategies.

Query

Dataset

ASR

MO@10

R@10 loss

Target

HotpotQA MSMARCO NQ

100.00 92.60 100.00

40.68 23.03 26.36

28.80 -1.80 15.80

Normal

HotpotQA MSMARCO NQ

0.20 0.00 0.00

0.02 0.00 0.00

0.00 -2.56 -4.23

Table 5: Sensitivity to Distance Metric in Flat Index Euclidean Dataset HotpotQA MSMARCO NQ

6.4

Cosine

MO@10

ASR

FPR

MO@10

ASR

FPR

99.40 97.05 97.97

100.00 99.60 99.85

1.06 1.26 1.19

93.66 80.50 83.14

99.85 94.85 94.00

1.62 2.51 2.15

to the ad text to boost its retrieval ranking for that specific target query. The resulting prefixed advertisements are inserted into the database as malicious records. We then evaluate the poisoned database on two query sets: (i) the optimized target queries used by the attacker, and (ii) a set of normal queries provided by the dataset. We use the same metrics as in the backdoor setting (ASR, MO@10, and R@10 loss). As shown in Table 7, corpus poisoning achieves extremely strong attack effectiveness on the target queries (ASR = 100% with very high MO@10), while the malicious records are rarely retrieved for normal queries (near-zero ASR and MO@10). This highlights an important limitation of corpus poisoning in practice: its effectiveness is highly dependent on the attacker-chosen target query set, and may not generalize when real user queries differ from the attacker’s assumed targets.

Evaluating Existing Attacks in the Vector Database Setting

In this section, we adapt several existing attack methods to the vector database retrieval setting and evaluate their impact under the same workflow throughout this paper. Specifically, we consider three categories: Embedding Model Backdoor Attacks[1, 7, 45, 54], Corpus Poisoning[5, 11, 20, 62], and Embedding Inversion[31, 58]. Since these attacks have different threat models, objectives, and success criteria, we slightly adjust the evaluation metrics when necessary to match each method. Therefore, the reported metrics are not directly comparable across methods and should not be used to judge which attack is better simply based on higher or lower numbers.

Table 7: Impact of Corpus Poisoning

6.4.1 Embedding Model Backdoor Attack. We instantiate an embedding model backdoor by fine-tuning the model. Concretely, we set the trigger to “@@@” and the target to “apple pie”. The backdoored model is trained such that inputs containing “@@@” are embedded similarly to inputs where “@@@” is replaced by “apple pie”, while preserving the original embeddings for clean inputs. We then construct malicious records by appending the trigger “@@@” to a small set of randomly sampled texts and insert them into the vector database. For evaluation, we use 50 target queries containing “apple pie” to test the triggered behavior, and 1,000 normal queries without “apple pie” to test stealthiness. We report ASR and MO@10 as in prior sections. In addition, we report R@10 loss as the difference between HNSW Recall@10 on the clean database and on the poisoned database. Table 6 shows that the backdoor attack is highly effective on target queries: ASR reaches 92.60–100.00% and MO@10 increases to 23.03–40.68% across datasets. This indicates a clear impact on the vector database retrieval behavior—after poisoning, records containing the trigger are much more likely to be retrieved when users issue queries related to the target concept. Meanwhile, for normal queries, the attack remains largely stealthy, with near-zero ASR/MO@10 and negligible changes in retrieval quality.

Query

Dataset

ASR

MO@10

R@10 loss

Target

HotpotQA MSMARCO NQ

100.00 100.00 100.00

96.80 72.00 92.50

96.80 72.00 92.50

Normal

HotpotQA MSMARCO NQ

0.20 0.00 0.03

0.02 0.00 0.00

0.02 0.00 0.00

6.4.3 Embedding Inversion. Directly inverting all database records is expensive at vector database scale. Therefore, we evaluate embedding inversion by inverting queries instead of the entire corpus, and measure whether the inverted queries preserve the original retrieval behavior. Concretely, we first embed the dataset-provided queries into vectors, then apply the embedding inversion method of vec2text [31] to reconstruct a textual query from each embedding. We then reencode the reconstructed queries using the same embedding model and issue retrieval with both the original query embeddings and the reconstructed-query embeddings. We use two metrics to quantify inversion quality in the vector database setting. We define ASR as the fraction of queries whose Top-10 retrieval result sets are exactly identical before and after inversion (i.e., the inversion causes no change in the retrieved set). We define 𝑅@10 as the average overlap ratio between the two Top-10 sets. As shown in Table 8, a non-trivial fraction of inverted queries (about 20%) yield retrieval results that exactly match the original

6.4.2 Corpus Poisoning. We further adapt a representative corpus poisoning method to the vector database setting. Following PoisonedRAG [62], we start from a pool of 1,000 advertisement texts and a set of 100 target queries. For each target query, we assign 10 advertisements and optimize an adversarial prefix that is prepended 10

queries. Moreover, the average 𝑅@10 is consistently around 70% across datasets and remains similar under both Flat and HNSW indexes. These results suggest that embedding inversion can reconstruct query text that largely preserves retrieval behavior, highlighting a practical privacy risk for vector databases that store or expose embedding vectors.

To assess its applicability in cross-modal retrieval, we conduct an additional study on ScienceQA[28], a multimodal benchmark involving both visual and textual signals. We follow the same experimental protocol and attack configuration as in our text experiments. Under this setting, the Black-Hole Attack achieves MO@10 = 33.36%, ASR = 35.32%, and an average FPR of 2.6. These results indicate that embedding spaces in multimodal (vision–text) retrieval also exhibit structural defects that can be exploited by a small number of injected vectors. Comparison with existing attacks. Backdoor attacks and corpus poisoning can both achieve strong effectiveness on target queries while remaining stealthy on normal queries; this is the intended behavior for backdoors, but for corpus poisoning it also implies a key limitation—attack success depends on accurately anticipating the real user query distribution. In contrast, the Black-Hole Attack is query-agnostic: it constructs malicious vectors directly from the database embeddings without using any query set, at the cost of a higher poisoning rate (about 1% poisoning in our setting to maximize the effect). Finally, embedding inversion presents a privacy risk by reconstructing text from embeddings and largely preserving retrieval behavior, yet scaling it to invert an entire large database is computationally expensive, which limits its practicality in large deployments.

Table 8: Impact of Embedding Inversion Flat

6.5

HNSW

Dataset

ASR

R@10

ASR

R@10

HotpotQA MSMARCO NQ

14.34 21.78 23.26

72.33 65.35 71.40

14.48 21.72 23.12

72.40 65.67 71.50

Impact on Downstream RAG Applications

To further evaluate the practical consequence of the Black-Hole Attack beyond retrieval metrics, we study its impact on a downstream RAG system. We reproduce LongRAG [60], a retrieval-augmented generation framework. We use the reproduced LongRAG pipeline as the downstream application and poison its underlying vector database with the Black-Hole Attack. Figure 7 reports the end-to-end QA performance before and after poisoning. The results show that the attack substantially degrades answer quality on all three benchmarks. On HotpotQA, the F1 score drops from 49.7 to 24.6; on 2WikiMQA [13], it drops from 46.3 to 23.2; and on MuSiQue [41], it drops from 20.9 to 3.2. These results indicate that poisoning only the vector database is sufficient to severely impair the performance of a strong downstream RAG system. Notably, on HotpotQA and 2WikiMQA the F1 score drops by roughly half rather than collapsing entirely. A likely explanation is the parametric memory of the underlying LLM: when the retrieved passages no longer contain the answer, the model can still fall back on knowledge encoded in its own parameters. 60 50

Clean

49.7

7 Defenses Against Black-Hole Attacks 7.1 Hubness Mitigation Our Black-Hole Attack exploits centrality-driven hubness. A natural line of defense is therefore to reduce hubness in the embedding space before retrieval. We evaluate four representative hubnessmitigation methods. CL2 (Centered L2 normalization) removes the mean shift by centering the representations and then applies L2 normalization, projecting vectors onto a common hypersphere where cosine similarity is better behaved [48]. ZN (Z-score normalization) normalizes each feature dimension to zero mean and unit variance, making the distribution more uniform and concen√ trating vectors near a hypersphere of radius 𝑑 [8, 42]. TCPR acts on queries by estimating the centroid of their retrieved neighbors and removing the component pointing toward that centroid, thereby reducing directional bias and indirectly alleviating hubness [52]. noHub optimizes a tradeoff between local similarity preservation and global uniformity on the hypersphere so that the resulting embeddings remain discriminative while suppressing hub formation [42]. We evaluate these defenses against the Black-Hole Attack using Contriever embeddings with the Flat search index, keeping poisoning rates and sample sizes consistent with our attack experiments. Table 9 reports MO@10 under each defense, measuring how often malicious vectors appear among the top-10 results. TCPR, CL2, and noHub substantially reduce MO@10 across all datasets, confirming that explicit hubness reduction can mitigate the attack. TCPR is the strongest, driving MO@10 to near-zero levels (0.11%–0.12%), whereas ZN leaves MO@10 high (65.36%–83.13%), indicating only limited protection. These defenses rely on mathematical transformations of stored and query vectors, so mitigating hubness can also distort retrieval semantics. To quantify this cost, we treat the clean top-𝑘 lists in

Poisoned

46.3

F1

40 30

24.6

23.2

20 10 0

20.9

3.2 HotpotQA

2WikiMQA

MuSiQue

Figure 7: Impact of the Black-Hole Attack on downstream RAG applications.

6.6

Discussion

Beyond text-only retrieval. In our experiments, we primarily focus on text retrieval. Nevertheless, the geometric vulnerability exploited by the Black-Hole Attack is not specific to text embeddings. 11

Table 9: Defense effectiveness against the Black-Hole Attack: MO@10 with Flat Search Dataset

No Defense

TCPR

ZN

CL2

noHub

93.66 80.50 83.15

0.11 0.10 0.12

83.13 65.36 68.66

1.66 12.25 10.00

0.71 13.91 47.27

HotpotQA MSMARCO NQ

removing suspicious vectors. The bottom panel measures impact on benign data: we apply the same procedure to an unpoisoned corpus and compute Recall@10 as the overlap between each query’s top-10 neighbors before and after filtering. The defense sharply reduces MO@10 on HotpotQA and MSMARCO (MO@10 from 93.66% to 1.06% and from 80.50% to 1.05%). On NQ, MO@10 falls from 83.15% to 31.78%, a smaller reduction than on HotpotQA or MSMARCO. We attribute this gap to a higher prevalence of natural hub vectors in NQ, so the detector cannot always cleanly separate malicious vectors from legitimate hubs. Benign retrieval is largely unchanged: Recall@10 on the clean knowledge base remains 99.22%, 99.68%, and 99.55%. Only 0.109%, 0.079%, and 0.132% of vectors are removed. Compared with global hubness mitigation, this detector offers a better security–utility tradeoff by deleting a few extreme vectors instead of transforming the embedding space. However, the method is not cheap: it introduces an additional k-NN pass for detection, and its compute/memory overhead scales poorly with corpus size, making deployment on very large indexes impractical.

the original embedding space as ground truth, run retrieval in each defended space, and measure overlap with Recall@10. Table 10 summarizes the results. TCPR and noHub achieve strong attack suppression but largely destroy retrieval utility: Recall@10 falls close to zero on all datasets. ZN preserves high Recall@10 yet remains weak against the attack. CL2 provides the best tradeoff, substantially lowering MO@10 while retaining moderate Recall@10 (58.65%–76.70% across datasets). Overall, hubness mitigation can defend against Black-Hole Attacks, but its effectiveness depends critically on how much utility loss the system can tolerate.

No defense

Table 10: Retrieval performance under different defenses: Recall@10 with Flat search ZN

CL2

noHub

1.28 1.61 1.30

85.83 89.67 87.48

58.65 76.70 66.78

5.45 11.38 10.44

MO@10

TCPR

HotpotQA MSMARCO NQ

Detection-Based Defense

We next study a more conservative defense that operates directly on the stored corpus vectors. The idea is to use a small probe set drawn from the database itself, measure how often each stored vector is retrieved as a nearest neighbor of these probes, and remove the vectors that are selected unusually often. 𝑁 ⊂ R𝑑 . We first Let the stored corpus vectors be 𝑋 = {𝑥𝑖 }𝑖=1 partition 𝑋 into 𝐿 clusters 𝐶 1, . . . , 𝐶𝐿 . To ensure that the probe set covers different semantic regions of the database, we independently sample a small subset from each cluster. Specifically, for each cluster 𝐶 ℓ , we sample a probe set 𝑃ℓ ⊆ 𝐶 ℓ uniformly at random with |𝑃ℓ | = max(1, ⌈0.01 |𝐶 ℓ |⌉) ,

𝑃=

𝐿 Ø

𝑃ℓ .

92.73

With defense 80.50

50

0 100 R@10

7.2

Dataset

100

R@10 83.15

31.78 1.06 99.22

1.05 99.68

99.55

HotpotQA

MS MARCO

NQ

50

0

Figure 8: Detection-based defense. Top: MO@10 on the poisoned database before and after filtering. Bottom: R@10 between pre- and post-filter results on an unpoisoned corpus

8

Conclusion

In this work, we present the Black-Hole Attack, a query-agnostic poisoning attack against vector databases. The attack injects malicious vectors into either the global centroid or multiple cluster-wise centroids of the embedding space, causing them to dominate Top-𝑘 retrieval results without access to user queries. Its effectiveness stems from centrality-driven hubness: in high-dimensional embedding spaces with finite data, vectors near the centroid naturally become nearest neighbors of many other points. Extensive experiments demonstrate that even a 1% poisoning rate can severely contaminate retrieval results and substantially degrade utility. We further study defenses from two directions: hubness mitigation and detection. While hubness mitigation reduces the attack at the cost of retrieval quality, detection removes only a small fraction of suspicious vectors and provides a better security. Overall, our results expose a fundamental trust vulnerability in vector databases and highlight the need for defenses.

(14)

ℓ=1

For each probe 𝑝 ∈ 𝑃, let N𝑘 (𝑝) denote its top-𝑘 nearest neighbors in 𝑋 . The hit count of 𝑥𝑖 is ∑︁ ℎ𝑖 = 1[𝑥𝑖 ∈ N𝑘 (𝑝)] . (15) 𝑝 ∈𝑃

Most vectors satisfy ℎ𝑖 = 0; including them would drive any robust center toward zero. We therefore restrict to strictly positive hits and set 𝑚 = median({ℎ𝑖 | ℎ𝑖 > 0}) . (16) A vector is marked suspicious if ℎ𝑖 > 2𝑚, and all such vectors are removed from the database before retrieval. We evaluate this detection-based defense in Figure 8. The top panel reports MO@10 on the poisoned database before and after 12

References

[20] Yang Jiao, Xiaodong Wang, and Kai Yang. 2025. PR-Attack: Coordinated PromptRAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel Optimization. Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (2025). https://api.semanticscholar.org/CorpusID:277667367 [21] Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for OpenDomain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, Online, 6769–6781. doi:10.18653/v1/2020.emnlp-main.550 [22] Vladimir Koltchinskii and Karim Lounici. 2014. Asymptotics and Concentration Bounds for Bilinear Forms of Spectral Projectors of Sample Covariance. arXiv: Statistics Theory (2014). https://api.semanticscholar.org/CorpusID:88513200 [23] Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, et al. 2019. Natural Questions: A Benchmark for Question Answering Research. In TACL. [24] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 9459–9474. https://proceedings.neurips. cc/paper_files/paper/2020/file/6b493230205f780e1bc26945df7481e5-Paper.pdf [25] Quentin Lhoest, Albert Villanova Del Moral, Yacine Jernite, Abhishek Thakur, Patrick Von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021. Datasets: A Community Library for Natural Language Processing. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, 175–184. doi:10.18653/v1/2021.emnlp-demo.21 [26] Guoliang Li, Xuanhe Zhou, and Xinyang Zhao. 2024. LLM for Data Management. Proceedings of the VLDB Endowment 17, 12 (Aug. 2024), 4213–4216. doi:10.14778/ 3685800.3685838 [27] Dawei Liu, Bolong Zheng, Ziyang Yue, Fuhao Ruan, Xiaofang Zhou, and Christian S. Jensen. 2025. Wolverine: Highly Efficient Monotonic Search Path Repair for Graph-Based ANN Index Updates. Proceedings of the VLDB Endowment 18, 7 (March 2025), 2268–2280. doi:10.14778/3734839.3734860 [28] Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering. In The 36th Conference on Neural Information Processing Systems (NeurIPS). [29] Yury Malkov and Dmitry A. Yashunin. 2016. Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence 42 (2016), 824–836. https://api.semanticscholar.org/CorpusID:8915893 [30] Yu A. Malkov and D. A. Yashunin. 2020. Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence 42, 4 (2020), 824– 836. doi:10.1109/TPAMI.2018.2889473 [31] John X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, and Alexander M. Rush. 2023. Text Embeddings Reveal (Almost) As Much As Text. In Conference on Empirical Methods in Natural Language Processing. https://api.semanticscholar. org/CorpusID:263829206 [32] James Jie Pan, Jianguo Wang, and Guoliang Li. 2024. Survey of vector database management systems. The VLDB Journal 33, 5 (2024), 1591–1615. [33] James Jie Pan, Jianguo Wang, and Guoliang Li. 2024. Vector Database Management Techniques and Systems. In Companion of the 2024 International Conference on Management of Data (Santiago AA, Chile) (SIGMOD ’24). Association for Computing Machinery, New York, NY, USA, 597–604. doi:10.1145/3626246.3654691 [34] Zhencan Peng, Miao Qiao, Wenchao Zhou, Feifei Li, and Dong Deng. 2025. Dynamic Range-Filtering Approximate Nearest Neighbor Search. Proceedings of the VLDB Endowment 18, 10 (June 2025), 3256–3268. doi:10.14778/3748191. 3748193 [35] Alina Petukhova, João P. Matos-Carvalho, and Nuno Fachada. 2024. Text Clustering with Large Language Model Embeddings. CoRR abs/2403.15112 (2024). arXiv:2403.15112 [cs.CL] https://arxiv.org/abs/2403.15112 [36] Sara Rajaee and Mohammad Taher Pilehvar. 2022. An Isotropy Analysis in the Multilingual BERT Embedding Space. In Findings of the Association for Computational Linguistics: ACL 2022, Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, Dublin, Ireland, 1309–1316. doi:10.18653/v1/2022.findings-acl.103 [37] Stefano Recanatesi, Serena Bradde, Vijay Balasubramanian, Nicholas A. Steinmetz, and Eric Shea-Brown. 2020. A scale-dependent measure of system dimensionality. Patterns 3 (2020). https://api.semanticscholar.org/CorpusID:229549825 [38] Lingfeng Shen, Haiyun Jiang, Lemao Liu, and Shuming Shi. 2023. Sen2Pro: A Probabilistic Perspective to Sentence Embedding from Pre-trained Language

[1] Gaurav Bagwe, Lan Zhang, Linke Guo, Miao Pan, Xiaolong Ma, and Xiaoyong (Brian) Yuan. 2025. Is Embedding-as-a-Service Safe? Meta-Prompt-Based Backdoor Attacks for User-Specific Trigger Migration. Transactions on Artificial Intelligence (2025). https://api.semanticscholar.org/CorpusID:279133462 [2] Jon Bratseth. 2017. Open Sourcing Vespa, Yahoo’s Big Data Processing and Serving Engine. https://blog.vespa.ai/open-sourcing-vespa-yahoos-big-dataprocessing/. Accessed: 2025-10-12. [3] Daniel Fernando Campos, Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, Li Deng, and Bhaskar Mitra. 2016. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. ArXiv abs/1611.09268 (2016). https://api.semanticscholar.org/CorpusID:1289517 [4] Cheng Chen, Chenzhe Jin, Yunan Zhang, Sasha Podolsky, Chun Wu, SzuPo Wang, Eric Hanson, Zhou Sun, Robert Walzer, and Jianguo Wang. 2024. SingleStore-V: An Integrated Vector Database System in SingleStore. Proceedings of the VLDB Endowment 17, 12 (Aug. 2024), 3772–3785. doi:10.14778/3685800. 3685805 [5] Zhuo Chen, Yuyang Gong, Jiawei Liu, Miaokun Chen, Haotan Liu, Qikai Cheng, Fan Zhang, Wei Lu, and Xiaozhong Liu. 2025. FlippedRAG: Black-Box Opinion Manipulation Adversarial Attacks to Retrieval-Augmented Generation Models. Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security (2025). https://api.semanticscholar.org/CorpusID:275337175 [6] Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. 2024. The Faiss library. (2024). arXiv:2401.08281 [cs.LG] [7] Wei Du, Peixuan Li, Bo Li, Haodong Zhao, and Gongshen Liu. 2023. UOR: Universal Backdoor Attacks on Pre-trained Language Models. In Annual Meeting of the Association for Computational Linguistics. https://api.semanticscholar.org/ CorpusID:258714833 [8] Nanyi Fei, Yizhao Gao, Zhiwu Lu, and Tao Xiang. 2021. Z-Score Normalization, Hubness, and Few-Shot Learning. 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (2021), 142–151. https://api.semanticscholar.org/CorpusID: 247191876 [9] Alejandro Fuster Baggetto and Victor Fresno. 2022. Is anisotropy really the cause of BERT embeddings not being semantic?. In Findings of the Association for Computational Linguistics: EMNLP 2022, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Linguistics, Abu Dhabi, United Arab Emirates, 4271–4281. doi:10.18653/v1/2022.findings-emnlp.314 [10] Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Qianyu Guo, Meng Wang, and Haofen Wang. 2023. Retrieval-Augmented Generation for Large Language Models: A Survey. ArXiv abs/2312.10997 (2023). https://api.semanticscholar.org/CorpusID:266359151 [11] Runpeng Geng, Yanting Wang, Ying Chen, and Jinyuan Jia. 2025. UniC-RAG: Universal Knowledge Corruption Attacks to Retrieval-Augmented Generation. arXiv preprint arXiv:2508.18652 (2025). [12] A. Gionis, Piotr Indyk, and Rajeev Motwani. 1999. Similarity Search in High Dimensions via Hashing. In Very Large Data Bases Conference. https://api. semanticscholar.org/CorpusID:1578969 [13] Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps. In Proceedings of the 28th International Conference on Computational Linguistics. International Committee on Computational Linguistics, Barcelona, Spain (Online), 6609–6625. doi:10.18653/v1/2020.coling-main.580 [14] Guoyu Hu, Shaofeng Cai, Tien Tuan Anh Dinh, Zhongle Xie, Cong Yue, Gang Chen, and Beng Chin Ooi. 2025. HAKES : Scalable Vector Database for Embedding Search Service. Proceedings of the VLDB Endowment 18, 9 (May 2025), 3049–3062. doi:10.14778/3746405.3746427 [15] Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2021. Unsupervised Dense Information Retrieval with Contrastive Learning. Trans. Mach. Learn. Res. 2022 (2021). https://api.semanticscholar.org/CorpusID:249097975 [16] Suhas Jayaram Subramanya, Fnu Devvrit, Harsha Vardhan Simhadri, Ravishankar Krishnawamy, and Rohan Kadekodi. 2019. Diskann: Fast accurate billion-point nearest neighbor search on a single node. Advances in neural information processing Systems 32 (2019). [17] Hervé Jégou, Matthijs Douze, and Cordelia Schmid. 2011. Product Quantization for Nearest Neighbor Search. IEEE Transactions on Pattern Analysis and Machine Intelligence 33 (2011), 117–128. https://api.semanticscholar.org/CorpusID: 5850884 [18] Wenqi Jiang, Hang Hu, Torsten Hoefler, and Gustavo Alonso. 2025. Fast Graph Vector Search via Hardware Acceleration and Delayed-Synchronization Traversal. Proceedings of the VLDB Endowment 18, 11 (July 2025), 3797–3811. doi:10.14778/ 3749646.3749655 [19] Wenqi Jiang, Marco Zeller, Roger Waleffe, Torsten Hoefler, and Gustavo Alonso. 2024. Chameleon: A Heterogeneous and Disaggregated Accelerator System for Retrieval-Augmented Language Models. Proceedings of the VLDB Endowment 18, 1 (Sept. 2024), 42–52. doi:10.14778/3696435.3696439 13

Model. In Proceedings of the 8th Workshop on Representation Learning for NLP (RepL4NLP 2023). Association for Computational Linguistics, Toronto, Canada, 315–333. doi:10.18653/v1/2023.repl4nlp-1.26 [39] Joobo Shim, Jaewon Oh, Hongchan Roh, Jaeyoung Do, and Sang-Won Lee. 2025. Turbocharging Vector Databases Using Modern SSDs. Proceedings of the VLDB Endowment 18, 11 (July 2025), 4710–4722. doi:10.14778/3749646.3749724 [40] Ji Sun, Guoliang Li, James Pan, Jiang Wang, Yongqing Xie, Ruicheng Liu, and Wen Nie. 2025. GaussDB-Vector: A Large-Scale Persistent Real-Time Vector Database for LLM Applications. Proceedings of the VLDB Endowment 18, 12 (2025), 4951–4963. [41] Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. MuSiQue: Multihop Questions via Single-hop Question Composition. Transactions of the Association for Computational Linguistics 10 (2022), 539–554. doi:10.1162/tacl_a_00475 [42] Daniel J. Trosten, Rwiddhi Chakraborty, Sigurd Løkse, Kristoffer Wickstrøm, Robert Jenssen, and Michael C. Kampffmeyer. 2023. Hubs and Hyperspheres: Reducing Hubness and Improving Transductive Few-Shot Learning with Hyperspherical Embeddings. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2023), 7527–7536. https://api.semanticscholar.org/CorpusID: 257557374 [43] Roman Vershynin. 2018. High-Dimensional Probability: An Introduction with Applications in Data Science. Cambridge University Press. [44] Dongsheng Wang, Dandan Guo, He Zhao, Huangjie Zheng, Korawat Tanwisuth, Bo Chen, and Mingyuan Zhou. 2022. Representing Mixtures of Word Embeddings with Mixtures of Topic Embeddings. In International Conference on Learning Representations (ICLR) 2022. https://openreview.net/forum?id=IYMuTbGzjFU [45] H. Wang, S. Guo, J. He, H. Liu, T. Zhang, and T. Xiang. 2025. Model Supply Chain Poisoning: Backdooring Pre-trained Models via Embedding Indistinguishability. In Proceedings of the ACM Web Conference 2025 (WWW ’25). [46] J. Wang, X. Yi, R. Guo, H. Jin, P. Xu, S. Li, X. Wang, X. Guo, C. Li, X. Xu, et al. 2021. Milvus: A purpose-built vector data management system. In Proceedings of the 2021 International Conference on Management of Data (SIGMOD ’21). 2614–2627. [47] Mengzhao Wang, Xiaoliang Xu, Qiang Yue, and Yuxiang Wang. 2021. A Comprehensive Survey and Experimental Comparison of Graph-Based Approximate Nearest Neighbor Search. Proc. VLDB Endow. 14 (2021), 1964–1978. https://api.semanticscholar.org/CorpusID:231728434 [48] Yan Wang, Wei-Lun Chao, Kilian Q. Weinberger, and Laurens van der Maaten. 2019. SimpleShot: Revisiting Nearest-Neighbor Classification for Few-Shot Learning. ArXiv abs/1911.04623 (2019). https://api.semanticscholar.org/CorpusID: 207863469 [49] Zihan Wang, Chengyu Dong, and Jingbo Shang. 2021. “Average” Approximates “First Principal Component”? An Empirical Analysis on Representations from Neural Language Models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 5594–5603. doi:10.18653/v1/2021.emnlp-main.453 [50] Weaviate B.V. [n. d.]. weaviate/weaviate: Weaviate – Open-source Vector Database. https://github.com/weaviate/weaviate. Accessed: 2025-10-12. [51] Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian yun Nie. 2023. C-Pack: Packed Resources For General Chinese Embeddings.

Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (2023). https://api.semanticscholar.org/ CorpusID:271114619 [52] Jing Xu, Xu Luo, Xinglin Pan, Wenjie Pei, Yanan Li, and Zenglin Xu. 2022. Alleviating the Sample Selection Bias in Few-shot Learning by Removing Projection to the Centroid. ArXiv abs/2210.16834 (2022). https://api.semanticscholar.org/ CorpusID:253237735 [53] Shuo Yang, Jiadong Xie, Yingfan Liu, Jeffrey Xu Yu, Xiyue Gao, Qianru Wang, Yanguo Peng, and Jiangtao Cui. 2024. Revisiting the Index Construction of Proximity Graph-Based Approximate Nearest Neighbor Search. Proc. VLDB Endow. 18 (2024), 1825–1838. https://api.semanticscholar.org/CorpusID:273025855 [54] Wenkai Yang, Lei Li, Zhiyuan Zhang, Xuancheng Ren, Xu Sun, and Bin He. 2021. Be Careful about Poisoned Word Embeddings: Exploring the Vulnerability of the Embedding Layers in NLP Models. ArXiv abs/2103.15543 (2021). https: //api.semanticscholar.org/CorpusID:232404131 [55] Zehai Yang and Shimin Chen. 2026. RAIRS: Optimizing Redundant Assignment and List Layout for IVF-Based ANN Search. arXiv preprint arXiv:2601.07183 (2026). [56] Zhilin Yang, Peng Qi, Saizheng Zhang, et al. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In EMNLP. [57] Shohei Yoda, Hayato Tsukagoshi, Ryohei Sasano, and Koichi Takeda. 2024. Sentence Representations via Gaussian Embedding. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers. Association for Computational Linguistics, 418–425. https://aclanthology.org/2024.eacl-short.36/ [58] Collin Zhang, John X. Morris, and Vitaly Shmatikov. 2025. Universal Zero-shot Embedding Inversion. ArXiv abs/2504.00147 (2025). https://api.semanticscholar. org/CorpusID:277467864 [59] Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, Meishan Zhang, Wenjie Li, and Min Zhang. 2024. mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval. In Conference on Empirical Methods in Natural Language Processing. https://api.semanticscholar.org/CorpusID: 271534420 [60] Qingfei Zhao, Ruobing Wang, Yukuo Cen, Daren Zha, Shicheng Tan, Yuxiao Dong, and Jie Tang. 2024. LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Question Answering. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 22600–22632. doi:10.18653/v1/ 2024.emnlp-main.1259 [61] Xinyang Zhao, Xuanhe Zhou, and Guoliang Li. 2024. Chat2data: An interactive data analysis system with rag, vector databases and llms. Proceedings of the VLDB Endowment 17, 12 (2024), 4481–4484. [62] Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia. 2024. PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models. In USENIX Security Symposium. https://api.semanticscholar. org/CorpusID:271854736

14

Related documents

Record · ID 2761 · SHA-256 a1fe160c4af6a470
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.