Conceptio › Archive › arXiv CS
arXiv CSopen access

Multi-Domain Clustering via Measure Quantization

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

MULTI-DOMAIN CLUSTERING VIA MEASURE QUANTIZATION Rafael Pereira Eufrazio1,2 , Eduardo Fernandes Montesuma3 and Charles Casimiro Cavalcante2

arXiv:2609.21664v1 [cs.LG] 18 Sep 2026

1

Instituto Federal de Educação, Ciência e Tecnologia do Ceará, Canindé-CE, Brazil 2 Federal University of Ceara, Fortaleza-CE, Brazil, 3 Sigma Nova Science, Paris, France ABSTRACT

Clustering is a fundamental task in data analysis, typically addressed through centroid-based methods such as K-means. In this work, we present a general framework for multi-domain clustering via measure quantization: given samples from multiple domains, we learn a shared set of cluster prototypes by minimizing a probability metric, such as the Sinkhorn divergence or the Maximum Mean Discrepancy, between each domain’s probability measure and the measure of prototypes. Data points are then assigned to clusters either via nearest centroid, or via optimal transport, a collaborative strategy that couples all samples within a domain. A mini-batch optimization strategy makes both fitting and assignment scalable, reducing memory and computational cost while preserving clustering performance. Experimental results on 5 multi-domain benchmarks spanning image, audio and sensor data show that our Sinkhorn-based method consistently outperforms classical and multi-domain clustering baselines, and that this advantage persists when scaling to hundreds of thousands of samples. Index Terms— Optimal Transport, Maximum Mean Discrepancy, Multi-Domain Clustering, Measure Quantization, Mini-batch Optimization

1. INTRODUCTION Clustering consists of partitioning data-points into groups that share common characteristics. For instance, the well-known K-Means algorithm [1] partitions the data into groups that share a centroid, i.e., a point that aggregates the information of that group. The problem that K-Means solve is a special case of the Measure Quantization problem [2], which assumes that data is drawn, i.i.d., from a single probability measure µ. In real applications, data is heterogeneous. For instance, in image processing [3], image data can differ by style, pose, illumination conditions, etc, inducing different statistical properties in the data. This is a known use-case of Transfer Learning [4] that develops learning methods (e.g., clustering methods) under the problem of distribution shift. In this paper, we develop a framework for multi-domain clustering that generalizes the measure quantization problem to K heti.i.d. erogeneous domains. We assume access to samples xk,i ∼ µk , k = 1, · · · , K and µk ̸= µk′ ∀k ̸= k′ , and formalize multi-domain This work was partially supported by CNPq Procs. 308512/2023-5 and 420341/2025-0, CNPq/INCT STREAM (Signal processing and TRansmission for Environmental Analysis and Monitoring) 409179/2024-8 and by CAPES - Finance Code 001.

clustering as, K

ν⋆ =

1 X D(µk , ν), ν∈EmpC (Rd ) K argmin

(1)

k=1

where µk ∈ Empnk (Rd ) ⊂ P(Rd ), ν ∈ EmpC (Rd ) are measure supported on C points (c.f. equation 2), and D is a metric, or notion of dissimilarity between probability measures (e.g., W2 (µ, ν)2 ). For K = 1, equation 1 reduces to the standard measure quantization problem solved by K-Means when D = W22 [2], and, more generally, is related to the problem of barycenters of probability measures [5, 6]. In this perspective, a stream of works [2, 7] studies the problem under the Wasserstein distance, a metric between probability measures that comes from Optimal Transport (OT) theory [8–10]. The present work is more general, as it studies problem 1 under different metrics or dissimilarities. Our contributions are as follows, described where the reader can find them in the manuscript: (i) We propose a new framework for clustering under distribution shift, via the measure quantization problem (c.f. equation 1), by reducing the support of input measures {µk }K k=1 . (ii) We develop fast, scalable algorithms based on gradiPK 1 ent descent of the functional ν 7→ K k=1 D(µk , ν) (Section 3), with respect the centroids {z1 , · · · , zC } that compose the support of ν. (iii) We experiment with image [11, 12], audio [13] and sensor data [14, 15], showing that our method achieves state-of-the-art performance in comparison with existing clustering baselines (Section 4). This paper is organized as follows. Section 2 includes the background to our paper. Section 3 presents our framework and algorithm. Section 4 shows experiments in multi-domain clustering. Section 5 concludes this paper. 2. BACKGROUND 2.1. Clustering d Let {xi }n i=1 ⊂ R be a dataset drawn from an unknown probability measure µ. A clustering of X into C groups produces labels {yi }n i=1 such that yi ∈ Y = {1, · · · , C}. We adopt a centroid-based view of C clustering, where vectors {zc }c=1 represent each cluster c and points are assigned to their nearest centroid, as in the canonical K-Means algorithm [1, 16]. One of the main limitations of K-Means is that it fails to cluster data in non-linear structures, such as manifolds. In this sense, other methods exploit the non-linear structure of the data. For instance, [17] uses the spectral analysis of graph Laplacians to cluster data. In a different direction [18,19] perform clustering based on the construction of a hierarchy within the data. Similarly, [20] proposes a multi-level clustering based on the Wasserstein distance.

Algorithm 1 Multi-domain Clustering

2.2. Probability Metrics d

d

A probability metric is a function D : P(R ) × P(R ) → R that defines a metric over the space of probability metrics, P(Ω). In the following, study probability metrics over discrete measures, that is, measures in the set, n

Empn (Rd ) = {µ : µ =

1X δx , xi ∈ Rd }, n i=1 i

(2)

where δx0 (x) = δ(x − x0 ) is the Dirac measure centered at x0 . We study metrics coming from OT theory [8–10] and the Maximum Mean Discrepancy (MMD) [21]. In that context, the discrete OT problem is given by, T ⋆ = OTϵ (µ, ν) = argmin ⟨T, C⟩F + ϵH(T ),

Input: Domains {Xk }K k=1 , number of clusters C, learning rate η, number of iterations N Output: Cluster prototypes Z = {zc }C c=1 1: Draw prototypes at random zc ∼ N (0, Id), c = 1, · · · , C 2: for τ = 1, . . . , T do 3: for k = 1, . . . , K do 4: Sample {xij ,c }B j=1 , from each k = 1, · · · , K 5: Estimate ĝk = ∇zc D(µk , ν) 6: end for 1 PK ĝk 7: Update prototypes zc,τ +1 ← zc,τ − η K k=1 8: end for 9: return Z ⋆ = {zc,T }C c=1

(3)

T∈Π(µ̂,ν̂)

where Cij is called the ground cost matrix, ⟨·, ·⟩F is the Frobenius inner product, and H(T ) is the entropy of the transport plan. ϵ = 0 gives the Kantorovich problem, solvable via linear programming [9, Chapter 3]; ϵ > 0 gives entropic OT, solved by Sinkhorn’s algorithm [22]. When the ground-cost comes from a metric m over Rd , that is, Cij = m(x1,i , x2,j )p , p ∈ [1, +∞), the Kantorovich formulation defines a notion of dissimilarity between probability measures, Wp,ϵ (µ, ν)p =

min

π∈Π(µ,ν)

⟨Tϵ⋆ , C⟩F + ϵH(Tϵ⋆ ),

(4)

where Tϵ⋆ is the OT plan. When ϵ = 0, Wp,0 is known as the p−Wasserstein distance, which is a true metric on P2 (Rd ). When ϵ > 0, one has the Sinkhorn divergence. Meanwhile, the MMD [21] is a kernel-based probability metric. Given a kernel κ : Rd × Rd → R, it is defined as, MMDκ (µ, ν)2 =

n C 1 X 1 X κ(x , x ) + κ(zc , zc′ ) i j n2 i,j=1 C2 ′

where gk = ∇zc D(µk , ν) is the gradient of the metric between µk and ν with respect the support of ν, that is, z1 , · · · , zC ∈ Rd . These gradients, with respect the considered metrics, are, gk = p

nk X (Tϵ⋆ )ic m(xi,k , zc )p−1 ∇zc m(xi,k , zc ),

and, for the MMD, gk =

nk C 2 X 2 X ′ ∇zc κ(xi,k , zc ). ∇ κ(z , z ) − z c c C 2 c=1 c nC i=1

n

⋆ yi,k = argmin m(xi,k , zc⋆ ),

(5)

In this work, we consider 3 kinds of kernels: linear κ(x, z) = x z, Riesz κ(x, z) = −∥x−z∥2 and RBF, κ(x, z) = exp(−γ∥x−z∥22 ). 3. MULTI-DOMAIN CLUSTERING VIA MEASURE QUANTIZATION

k=1

λk gk ,

(10)

c∈Y

n

k collaborative since Tϵ⋆ couples all samples {xi,k }i=1 in the support of µk (also computable in mini-batches for scalability).

4. EXPERIMENTS

In this work, we consider multi-domain clustering. In brief, we nk have k = 1, · · · , K different domains, each with data {xi,k }i=1 . We assume these points are represented in a shared Euclidean space Rp . We take a probabilistic view of multi-domain clustering, meaning that we represent each dataset through an empirical probabilP k 1 PC ity measure, µk = n1k n δz where i=1 δxi,k , and ν = C c=1 c δx0 (x) = δ(x − x0 ) is the Dirac measure centered at x0 . Likewise, we represent the distribution of centroids through an empirical measure centered at each centroid. The multi-domain clustering task is then finding Z ⋆ = {zc⋆ }C c=1 by minimizing the objective in equation 1, via gradient descent, K X

i.e., each point to its nearest centroid; and OT assignment, ⋆ yi,k = argmax (Tϵ⋆ )i,c , where Tϵ⋆ = OTϵ (µk , ν),

⊤

zc,t+1 ← zc,t −

(9)

c∈Y

C

2 XX κ(xi , zc ). nC i=1 c=1

(8)

We summarize these steps in Algorithm 1, approximating the gradients in equations 7–8 via mini-batching: instead of all nk points, we sample B ∈ N points {xij ,k }B j=1 per domain, obtaining an estimate ĝk . ⋆ }, we assign cluster indices After obtaining Z ⋆ = {z1⋆ , · · · , zC via 2 strategies: greedy assignment,

c,c =1

−

(7)

i=1

(6)

In this section, we experiment with multi-domain clustering methods. Our experimentation focuses on classification datasets with multiple domains in 3 areas: computer vision (Office 31 [11], OfficeHome [12], Caltech-Office 10 [23]), audio (TAU Urban Scenes [13]) and chemical engineering (Tennessee Eastman Process [14,15]). We additionally experiment with DomainNet [24] for stress-testing the scalability of multi-domain clustering methods. A summary of these datasets is available in Table 1. We extract features with ResNet-50 [25] for Office 31, ResNet101 for Office-Home, and DeCaf [26] for Caltech-Office 10. These networks were trained on ImageNet, and features are extracted without further fine-tuning. For the TAU Urban Scenes dataset, we use the PANN backbone [27]. For the TEP dataset, we follow [28] and use 1st and 2nd order statistics of each time series as the features. We compare 7 methods, grouped into 2 tiers. Pooled: 4 classical clustering methods run on the pooled domain data (thus ignoring

1 (a) ν ? = arg minν K

P

k D(µk , ν)

1 (b) zc 7→ K

P

k Wp, (µk , ν)

1 (c) zc 7→ K

zc?

µ1

µ2

µ3

ν

P

k MMDκ (µk , ν)

(d) ν̂ = KMeans zc?

S

k µk



pooled K-means

Fig. 1: Overview of multi-domain clustering via measure quantization. (a) Toy example P P with K=3 domains and the fitted Sinkhorn prototypes 1 1 ν ⋆ . (b)-(c) Loss landscape zc 7→ K D(µ , ν) and its gradient flow −∇ z k c K k k D(µk , ν) (streamlines) for one prototype, while others are fixed (white stars), under Wp,ϵ and MMDκ .

Dataset

Modality

# Features

# Classes

# Domains

# Samples

Office 31 Caltech-Office 10 Office-Home DomainNet TAU Urban Scenes Tennessee Eastman Process

Image Image Image Image Audio Sensor data

2048 4096 2048 2048 768 64

31 10 65 345 10 29

3 4 4 6 10 6

4,110 2,533 15,500 586,575 20,800 17,289

Ground truth (class)

Sinkhorn (ours) (GM=0.81)

MMD (GM=0.67)

K-Means (pooled) (GM=0.55)

MWMS (GM=0.62)

Table 1: Summary of datasets used in our experiments.

distribution shift) using Scikit-Learn [29]. These methods are KMeans [1], spectral clustering [17], Ward [19] and BIRCH [18]. Multi-domain, which are our Sinkhorn and MMD methods, plus MWMS [20], a multi-level Wasserstein-based clustering mechanism. We compare these 7 methods through 3 metrics: the Hungarian Accuracy (Hung. Acc.), the Adjusted Rand Index (ARI) and the Normalized Mutual Information (NMI). The Hung. Acc. aligns estimated clusters with the true underlying clusters via an OT problem ⋆ (c.f. equation 3) with uniform marginals. In that case, Talign is a permutation matrix [9, Chapter 3] and thus makes a correspondence between estimated and underlying clusters. Briefly, the NMI measures how much information the predicted clustering shares with the ground-truth clusters, via a normalized version of mutual information, and the ARI measures the agreement between two cluster predictions by comparing if pairs of samples belong to the same cluster or not. We refer readers to [30] for more information about the NMI and ARI metrics. All experiments ran on a single machine (AMD EPYC 7413 CPU, NVIDIA L4 GPU, 47 GB RAM) using PyTorch [31] and PythonOT [32]. Our results are summarized in Table 2. Overall, methods that exploit the multi-domain structure of the data rank best than pooled, classical clustering algorithms. Out of the 3 tested multi-domain methods, those using the Wasserstein geometry (ours, and MWMS [20]) have the best performance. Using the MMD achieves an average rank comparable to MWMS. For instance, we visualize in Figure 2 the clustering of each method on the CaltechOffice 10 dataset [23]. Next, we ablate our method across 2 angles. On the first angle, we ablate the Sinkhorn divergence Wp,ϵ over p ∈ {1, 2}, ϵ ∈ {10−3 , 10−2 , 10−1 }, m ∈ {Cos, Eucl}, and the MMD kernel (linear, RBF, Riesz), reporting for each value its best configuration over the remaining hyper-parameters, averaged across the 5 datasets. p and m have little impact on GM (≤ 0.01 apart), whereas ϵ does: GM drops from 0.66 at ϵ=10−3 to 0.49 at ϵ=10−1 ,

domain (marker); color = class

caltech amazon webcam dslr

Fig. 2: Clustering methods over the Caltech-Office 10 dataset. Overall, using the Sinkhorn divergence as a criterion yields the most consistent clustering of the data.

as heavier entropic smoothing blurs the coupling between samples and prototypes. The MMD kernel matters even more: switching from RBF to a Riesz kernel raises GM by 0.19 (0.31 → 0.50), making kernel choice the dominant lever for MMD-based clustering. On the second angle, we compare OT assignment (equation 10) against nearest centroid (equation 9) for the best Sinkhorn configuration per dataset. The two strategies are within 0.01 GM of each other on Office-Home, Office31, Caltech-Office10 and TAU, but OT assignment gives a +0.11 GM gain on TEP (0.65 vs. 0.54), the dataset with the most classes (29) and the most severe class imbalance, suggesting the collaborative, transport-plan-based assignment helps most when clusters are hard to disambiguate from a single centroid alone. Finally, we stress-test scalability on DomainNet (586,575 samples, 6 domains, 345 classes), comparing our methods against minibatch K-Means (batch size 8,192). Sinkhorn (ours) outperforms mini-batch K-Means on all 3 metrics (Acc: 0.27 vs. 0.23; NMI: 0.45 vs. 0.40; ARI: 0.16 vs. 0.10), while MMD (ours) trails both (0.14/0.32/0.03). Overall, combining these results with those in Table 2, we conclude that the Wasserstein distance yields a richer geometry for clustering the underlying measures. This confirms that our Sinkhorn-based approach retains its advantage over classical

Table 2: Per-domain clustering performance (within-domain Hungarian matching, averaged over domains) on the five multi-domain benchmarks. Per dataset: Hungarian Acc. / NMI / ARI and their geometric mean (GM); last column: average rank across datasets by GM (lower is better). Sinkhorn uses p=1, euclidean cost. Best GM per dataset and best avg. rank in bold. Tier

OH

Method

O31

C10

TAU

TEP

Avg. Rank

Acc

NMI

ARI

GM

Acc

NMI

ARI

GM

Acc

NMI

ARI

GM

Acc

NMI

ARI

GM

Acc

NMI

ARI

GM

KMeans Spectral Ward BIRCH

0.59 0.58 0.60 0.59

0.70 0.70 0.70 0.70

0.41 0.33 0.39 0.41

0.56 0.51 0.55 0.56

0.81 0.73 0.85 0.84

0.87 0.84 0.89 0.89

0.71 0.55 0.75 0.73

0.79 0.70 0.83 0.82

0.65 0.62 0.67 0.66

0.66 0.69 0.72 0.69

0.43 0.38 0.50 0.50

0.57 0.55 0.62 0.61

0.42 0.40 0.40 0.40

0.38 0.41 0.35 0.35

0.24 0.22 0.22 0.22

0.34 0.33 0.32 0.32

0.24 0.30 0.27 0.21

0.44 0.60 0.48 0.40

0.05 0.16 0.06 0.04

0.18 0.31 0.19 0.15

4.6 5.4 4.2 4.6

Multi-domain

MWMS Sinkhorn (ours) MMD (ours)

0.59 0.63 0.56

0.69 0.72 0.66

0.46 0.52 0.39

0.57 0.62 0.52

0.79 0.81 0.82

0.84 0.86 0.85

0.70 0.74 0.72

0.77 0.80 0.79

0.72 0.85 0.70

0.71 0.80 0.66

0.51 0.72 0.49

0.64 0.79 0.61

0.42 0.43 0.41

0.38 0.38 0.36

0.24 0.24 0.22

0.34 0.34 0.32

0.22 0.63 0.29

0.50 0.71 0.52

0.08 0.58 0.11

0.21 0.64 0.26

3.4 1.4 4.4

0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0

0.66

0.65

0.66

0.66

0.65

0.31

2

cos

p

euc

.001

.01

.1

rbf

metric

rie

kernel

Fig. 3: Ablation of hyper-parameters (p, m, ϵ, κ) of our methods. 1.0 0.84

per-domain GM

0.8 0.6

0.64

0.85

0.84

OT (transport plan) nearest centroid

0.84

0.65

0.63

0.54

0.4

0.34

0.40 0.32 0.27 0.23

0.2

0.16

0.14

0.10

0.1

0.03

0.0

Acc

NMI

ARI

Fig. 5: Scaling experiment on DomainNet.

7. ACKNOWLEDGMENTS The authors declare that they have no relevant financial or nonfinancial interests to disclose.

0.34

8. REFERENCES

0.2 0.0

0.3

Sinkhorn KMeans MMD

0.45

0.4

0.50

0.49

1

0.5

Sinkhorn (ours) MMD (ours)

0.63

score

avg. GM (5 datasets)

Pooled

OfficeHome

Office31

Caltech

TAU

TEP

Fig. 4: Ablation of assignment mechanism of our methods.

clustering even with hundreds of thousands of samples. 5. CONCLUSION In this work, we propose a general framework for multi-domain clustering via measure quantization, i.e., the reduction of a set of measures’ support to C-prototypes by gradient descent of probability metrics [7], scalable via mini-batching [33]. Our experiments demonstrate the superiority of the proposed framework when a suitable geometry in the space of probability measures is defined (e.g., the Sinkhorn divergence [34]). Future work includes the theoretical study of the convergence of our algorithm, and consideration of other geometries, especially when domains live in incomparable spaces (e.g., via the Gromov-Wasserstein metric [35]). 6. COMPLIANCE WITH ETHICAL STANDARDS This computational study used publicly available benchmark datasets and did not involve the recruitment or intervention of human or animal subjects. Therefore, ethical approval was not required.

[1] Stuart Lloyd, “Least squares quantization in pcm,” IEEE transactions on information theory, vol. 28, no. 2, pp. 129–137, 1982. [2] Marco Cuturi and Arnaud Doucet, “Fast computation of wasserstein barycenters,” in International conference on machine learning. PMLR, 2014, pp. 685–693. [3] Vishal M Patel, Raghuraman Gopalan, Ruonan Li, and Rama Chellappa, “Visual domain adaptation: A survey of recent advances,” IEEE signal processing magazine, vol. 32, no. 3, pp. 53–69, 2015. [4] Sinno Jialin Pan and Qiang Yang, “A survey on transfer learning,” IEEE Transactions on knowledge and data engineering, vol. 22, no. 10, pp. 1345–1359, 2009. [5] Martial Agueh and Guillaume Carlier, “Barycenters in the wasserstein space,” SIAM Journal on Mathematical Analysis, vol. 43, no. 2, pp. 904–924, 2011. [6] Samuel Cohen, Michael Arbel, and Marc Peter Deisenroth, “Estimating barycenters of measures in high dimensions,” arXiv preprint arXiv:2007.07105, 2020. [7] Eduardo Fernandes Montesuma, Yassir Bendou, and Mike Gartrell, “Wasserstein gradient flows for scalable and regularized barycenter computation,” in Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, Emilija Perković and Daniel Malinsky, Eds. 17–21 Aug 2026, vol.

[8] Cédric Villani, Optimal transport: old and new, vol. 338, Springer, Berlin, 2009.

[23] Boqing Gong, Yuan Shi, Fei Sha, and Kristen Grauman, “Geodesic flow kernel for unsupervised domain adaptation,” in 2012 IEEE conference on computer vision and pattern recognition. IEEE, 2012, pp. 2066–2073.

[9] Gabriel Peyré and Marco Cuturi, “Computational optimal transport: With applications to data science,” Foundations and Trends® in Machine Learning, vol. 11, no. 5-6, pp. 355–607, 2019.

[24] Xingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang, Kate Saenko, and Bo Wang, “Moment matching for multi-source domain adaptation,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 1406–1415.

[10] Eduardo Fernandes Montesuma, Fred Maurice Ngole Mboula, and Antoine Souloumiac, “Recent advances in optimal transport for machine learning,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

[25] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778.

[11] Kate Saenko, Brian Kulis, Mario Fritz, and Trevor Darrell, “Adapting visual category models to new domains,” in European conference on computer vision. Springer, 2010, pp. 213– 226.

[26] Jeff Donahue, Yangqing Jia, Oriol Vinyals, Judy Hoffman, Ning Zhang, Eric Tzeng, and Trevor Darrell, “Decaf: A deep convolutional activation feature for generic visual recognition,” in International conference on machine learning. PMLR, 2014, pp. 647–655.

337 of Proceedings of Machine Learning Research, pp. 4595– 4606, PMLR.

[12] Hemanth Venkateswara, Jose Eusebio, Shayok Chakraborty, and Sethuraman Panchanathan, “Deep hashing network for unsupervised domain adaptation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 5018–5027. [13] Annamaria Mesaros, Toni Heittola, and Tuomas Virtanen, “A multi-device dataset for urban acoustic scene classification,” arXiv preprint arXiv:1807.09840, 2018. [14] Christopher Reinartz, Murat Kulahci, and Ole Ravn, “An extended tennessee eastman simulation dataset for fault-detection and decision support systems,” Computers & chemical engineering, vol. 149, pp. 107281, 2021. [15] Eduardo Fernandes Montesuma, Michela Mulas, Fred Ngolè Mboula, Francesco Corona, and Antoine Souloumiac, “Benchmarking domain adaptation for chemical processes on the tennessee eastman process,” in Machine Learning and Principles and Practice of Knowledge Discovery in Databases, Mattia Cerrato, Danguolė Kalinauskaitė, Mantas Lukoševičius, Mykola Pechenizkiy, and Kristina Šutienė, Eds., Cham, 2026, pp. 307–322, Springer Nature Switzerland. [16] Bao Chong et al., “K-means clustering algorithm: a brief review,” Academic Journal of Computing & Information Science, vol. 4, no. 5, pp. 37–40, 2021. [17] Ulrike Von Luxburg, “A tutorial on spectral clustering,” Statistics and computing, vol. 17, no. 4, pp. 395–416, 2007. [18] Tian Zhang, Raghu Ramakrishnan, and Miron Livny, “Birch: an efficient data clustering method for very large databases,” ACM sigmod record, vol. 25, no. 2, pp. 103–114, 1996. [19] Joe H Ward Jr, “Hierarchical grouping to optimize an objective function,” Journal of the American statistical association, vol. 58, no. 301, pp. 236–244, 1963.

[27] Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D Plumbley, “Panns: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 28, pp. 2880–2894, 2020. [28] Eduardo Fernandes Montesuma, EL HABAZI Adel, and Fred Maurice NGOLE MBOULA, “Unsupervised anomaly detection through mass repulsing optimal transport,” Transactions on Machine Learning Research, 2025. [29] Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al., “Scikit-learn: Machine learning in python,” the Journal of machine Learning research, vol. 12, pp. 2825–2830, 2011. [30] Nguyen Xuan Vinh, Julien Epps, and James Bailey, “Information theoretic measures for clusterings comparison: Variants, properties, normalization and correction for chance,” Journal of Machine Learning Research, vol. 11, no. 95, pp. 2837–2854, 2010. [31] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al., “Pytorch: An imperative style, high-performance deep learning library,” Advances in neural information processing systems, vol. 32, 2019. [32] Rémi Flamary, Nicolas Courty, Alexandre Gramfort, Mokhtar Z Alaya, Aurélie Boisbunon, Stanislas Chambon, Laetitia Chapel, Adrien Corenflos, Kilian Fatras, Nemo Fournier, et al., “POT: Python Optimal Transport,” Journal of Machine Learning Research, vol. 22, no. 78, pp. 1–8, 2021.

[20] Nhat Ho, XuanLong Nguyen, Mikhail Yurochkin, Hung Hai Bui, Viet Huynh, and Dinh Phung, “Multilevel clustering via wasserstein means,” in International conference on machine learning. PMLR, 2017, pp. 1501–1509.

[33] Kilian Fatras, Younes Zine, Szymon Majewski, Rémi Flamary, Rémi Gribonval, and Nicolas Courty, “Minibatch optimal transport distances; analysis and applications,” arXiv preprint arXiv:2101.01792, 2021.

[21] Arthur Gretton, Karsten M Borgwardt, Malte J Rasch, Bernhard Schölkopf, and Alexander Smola, “A kernel two-sample test,” The journal of machine learning research, vol. 13, no. 1, pp. 723–773, 2012.

[34] Aude Genevay, Gabriel Peyré, and Marco Cuturi, “Learning generative models with sinkhorn divergences,” in International Conference on Artificial Intelligence and Statistics. PMLR, 2018, pp. 1608–1617.

[22] Marco Cuturi, “Sinkhorn distances: Lightspeed computation of optimal transport,” Advances in neural information processing systems, vol. 26, 2013.

[35] Facundo Mémoli, “Gromov–Wasserstein distances and the metric approach to object matching,” Foundations of Computational Mathematics, vol. 11, pp. 417–487, 2011.

Record · ID 1006875 · SHA-256 d694d737cb8208d7
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.