Conceptio › Archive › arXiv CS
arXiv CSopen access

Reliable Modeling of Distribution Shifts via Displacement-Reshaped Optimal Transport

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

Article

Reliable Modeling of Distribution Shifts via Displacement-Reshaped Optimal Transport Philip Naumann 1,2 , Jacob Kauffmann 1,2 , Klaus-Robert Müller 1,2,4,5 * , and Grégoire Montavon 1,3 * 1 2 3 4

arXiv:2605.04965v1 [cs.LG] 6 May 2026

5

*

BIFOLD – Berlin Institute for the Foundations of Learning and Data, 10587 Berlin, Germany Machine Learning Group, Technische Universität Berlin, 10587 Berlin, Germany Institute for AI in Medicine, Charité – Universitätsmedizin Berlin, 10117 Berlin, Germany Department of Artificial Intelligence, Korea University, Anam-dong, Seongbuk-gu, Seoul 02841, Korea Max Planck Institute for Informatics, Stuhlsatzenhausweg, 66123 Saarbrücken, Germany Correspondence: [email protected], [email protected]

Abstract Optimal transport (OT) is a central framework for modeling distribution shifts. Because OT compares distributions directly in input space, a well-designed ground metric between observations is essential to ensure that the optimizer does not violate the true geometry of change. We propose DisplacementReshaped Optimal Transport (ReshapeOT), a method that reshapes the ground metric by integrating observed sample displacements as an additional source of knowledge. Technically, ReshapeOT replaces the Euclidean metric with a Mahalanobis distance estimated from displacement second moments. This effectively carves expressways through the input space, inviting transport solutions that better align with observed displacements. Our method is computationally lightweight, integrates seamlessly into any OT solver that operates on a cost matrix, and can be kernelized for further flexibility. Experiments on synthetic and real-world data show that ReshapeOT achieves substantial gains in transport reliability. We further demonstrate our method’s usefulness in two practical use cases. Keywords: optimal transport, metric learning, distribution shifts, knowledge integration

1. Introduction Modeling distribution shifts is a critical research area (cf. [1,2]) with applications ranging from characterizing real-world phenomena to developing robust machine learning models. Information geometry (IG) [3] provides a rigorous theoretical framework for these shifts by representing probability distributions as points on a Riemannian manifold, where shifts are quantified as geodesic distances. While IG traditionally characterizes shifts in parameter space (cf. [4]), recent work has pivoted toward measuring shifts relative to the distribution’s support [5]. This perspective draws deep connections to optimal transport (OT) [5–7], facilitating diverse applications such as tracking developmental processes in single-cell data [8,9] and building robust ML models [10–12]. Due to its focus on the distribution’s support, OT requires special attention to the function used to measure distance in input space. For instance, using the Euclidean distance to measure distance between sources and targets tacitly assumes that transporting points over a straight line is a feasible and cost-efficient solution. This ignores practical structure: real displacements may follow corridors (e.g., migration routes, vascular pathways), avoid obstacles, or be governed by costs that differ substantially from raw Euclidean distance. Much of the recent literature on optimal transport has focused on designing improved measures of transport cost to serve as optimization objectives. This includes objectives that integrate the manifold structure into the distance calculation (i.e. geodesic distances) [13], that prevent transport plans from having discontinuities in their source-target mapping (see e.g. [11]), or that force an alignment to the

2 of 20

dominant directions of variation in the data [14]. While these first-principles formulations are highly effective at guiding the transport solution towards the specified behavior, they may lack flexibility in handling the multitude of real-world factors that make up the true transport costs. Rather than handcrafting a suitable ground metric, we propose Displacement-Reshaped Optimal Transport, or, short, ReshapeOT, which learns the metric directly from observed displacement statistics, steering OT toward empirically grounded solutions and away from naive Euclidean extrapolations. Technically, our approach estimates second-order statistics of past displacements to derive a new, Mahalanobis-type ground metric that carves ‘expressways’ into the original Euclidean ground metric. Our method easily integrates with existing OT solvers and incurs no significant increase in computational cost. Because ReshapeOT only modifies the cost matrix and leaves the transport problem’s feasible set intact, its solutions take the same form as classical OT, preserving full compatibility with downstream tools for transparency and explainability, such as W aX [15]. Through experiments on synthetic and real-world data, we demonstrate that our method learns OT solutions that improve upon those of classical OT. Specifically, we observe that our method is less susceptible to falling for implausible shortcuts in the input space. Our approach is sample-efficient, requiring only a limited number of carefully selected displacement instances to steer the OT solution toward more plausible outcomes, as we demonstrate in the case of retrieving migratory bird trajectories. Finally, our method also provides a foundation for higher performance in downstream applications, as we demonstrate in a domain adaptation use case.

2. Related Work Our work connects to two strategies for shaping OT solutions: constraining the coupling (Section 2.1), or adapting the ground metric (Sections 2.2 and 2.3). We refer to [6,7] for a broader background on optimal transport in general. 2.1. Constraining the Coupling One strategy alters the OT objective or feasible set so the coupling reflects structural priors. Spatial and graph-Laplacian regularizers encourage smooth, neighborhood-preserving plans [11,16]. Schrödinger Bridges instead constrain the stochastic path. Theodoropoulos et al. [17], for example, use pre-aligned pairs as feedback to guide unpaired transport, but alter both the objective and feasible set rather than the cost. Anchor-point methods restrict the coupling itself: Latent OT [18] forces transport through learned anchors for outlier-robust low-rank couplings; Sato et al. [19] compare points via 1D pairwise-distance distributions; and Gu et al. [20] use keypoint-derived masks for heterogeneous domain adaptation. In contrast, our ReshapeOT method imposes no structural constraints on the plan, its path, its feasible set, or its regularization reference; it only modifies the cost matrix. This makes it interoperable with any cost-matrix-based OT solver. 2.2. Hand-Designed and Robust Geometric Metrics A complementary strategy replaces the Euclidean distance with one reflecting the data’s intrinsic or worst-case geometry. Huguet et al. [13] compute Wasserstein distances under geodesic ground metrics via entropic OT, encoding manifold structure in the cost. Algebraically closer to ours, Paty & Cuturi [14] introduce Subspace Robust Wasserstein Distances, whose convex relaxation hinges on the second-moment matrix of displacements. Such hand-designed or adversarially selected metrics are effective when the relevant geometry is known a priori or robustness is the goal, but they do not adapt to observed displacements. Both increase transport costs, via topological constraints or adversarial projections, whereas ReshapeOT decreases costs along directions empirically supported by observed displacements: it carves expressways rather than adding barriers (cf. Section 3.1).

3 of 20

2.3. Ground Metric Learning Closest to our setting is ground metric learning, where the cost is itself learned from data. Several works learn a Mahalanobis cost for OT. Cuturi & Avis [21] learn from class-labeled histograms via a subgradient procedure; Kerdoncuff et al. [22] establish a PCA–Wasserstein connection, and learn a ground metric to reduce target mismatch for domain adaptation; and Jawanpuria et al. [23] jointly learn the OT plan and a latent ground metric via Riemannian alternating optimization. Although the Mahalanobis form closely resembles ours, all three methods rely on label supervision or unpaired global statistics and require iterative optimization. ReshapeOT differs by deriving its metric in closed form from the second-order statistics of (possibly very few) observed source–target displacements, without label information and without alternating updates. Lastly, we note that the recipe of inverting a second-moment matrix, as done by ReshapeOT, has roots in classical metric learning [24–26], notably in Relevant Component Analysis (RCA) [27], which uses within-class covariances of equivalenceconstrained “chunklets”. With ReshapeOT, we propose a method that brings this idea to OT. A complementary line of work learns more expressive ground metrics, such as graph-geodesic, Riemannian, or neural, from observations of evolving densities or trajectories. Heitz et al. [28] parametrize the metric as graph edge weights and fit it by reconstructing intermediate density snapshots as Wasserstein geodesics: closest to ours in motivation, but requiring full intermediate marginals, graph-geodesic restrictions, and iterative diffusion-based optimization, whereas ReshapeOT only needs paired source–target observations and is closed-form. Scarvelis & Solomon [29] learn a spatiallyvarying Riemannian metric from cross-sectional snapshots via a neural network; Pooladian et al. [30] extend this to deterministic maps via a Lagrangian with obstacles and non-Euclidean geometries through amortized path optimization; and Kapusniak et al. [31] use data-dependent Riemannian metrics in metric flow matching to avoid straight-line shortcuts. Closest to our supervision regime, OTSI [32] learns a parametric OT cost from subset correspondences by differentiating through Sinkhorn [33] with pair matching being its extreme unit-sized case. OT-SI and ReshapeOT share the goal of leveraging paired side information to shape the transport cost, but differ in approach: OT-SI fits a parametric cost via end-to-end gradient flow through the OT solver, trading analytical simplicity for flexibility, while ReshapeOT inverts a single moment statistic in closed form. The two methods, therefore, occupy complementary points, and we view them as compatible rather than competing. Perrot et al. [34] jointly learn a coupling and a parametric transport map from source–target pairs for out-of-sample transport. Finally, Inverse OT [35,36] addresses the reverse problem of inferring the cost from samples and a known coupling, and Meta-OT [37] amortizes solutions across related problems. In contrast to these neural, graph-based, and iteratively-learned formulations, our ReshapeOT method requires no learned model, yields a closed-form ground metric, and directly leverages a small set of source–target pairings, whereas most prior methods rely on unlinked samples, full intermediate marginals, class labels, or end-to-end alignment objectives. We also do not require additional information beyond the observed displacements to infer the adapted cost. Hence, by modifying only the cost matrix, ReshapeOT further retains full compatibility with classical OT solvers.

3. Problem Formulation and Proposed Method We recall the classical OT formulation: given a source distribution µ and a target distribution ν on Rd , the Kantorovich primal problem seeks a cost-minimizing coupling γ ∈ Γ(µ, ν), a joint distribution on Rd × Rd with marginals µ and ν (cf. [6,7]): inf

γ∈Γ(µ,ν)

E(x,y)∼γ [c( x, y)]

(1)

where c : Rd × Rd → R+ is a cost function. With finite samples from µ and ν, γ reduces to a nonnegative matrix whose row and column sums match the empirical marginals. This classical formulation has limitations in practice; in particular, we rarely know the true cost function, or the latter is impractical to compute. In practice, one often resorts to heuristic or technical cost functions, where a popular

4 of 20

Classical OT

ReshapeOT (ours) optimal in Euclidean sense but untested extrapolations

adapt the metric from sample displacements

valid strategy?

more reliable, experiencegrounded solution

sample displacements

Figure 1. Overview of our approach to influence the OT solution through ground-truth displacements. The classical OT solution (here with squared Euclidean costs) contains points that are spuriously transported across the manifold. In contrast, our proposed ReshapeOT method carves the original Euclidean distance into a new distance with lower associated costs along the displacements. This results in a more reliable, guided solution.

choice builds on the Euclidean ground metric, namely, c( x, y) = ∥ x − y∥2 , which is associated with the classical W2 -Wasserstein transport problem. The OT problem is a natural lens for studying distribution shift: µ and ν describe the data distribution before and after a shift, the coupling γ encodes how mass is redistributed, and the cost function c governs which movements are favored. The ground metric, therefore, directly determines what kinds of shifts the model can capture: a too rudimentary ground metric might invite the exploitation of shortcuts that are physically implausible, whereas an informed metric can reflect the true geometry of the distribution shift. We propose an alternative formulation to modeling distribution shifts using OT, where we assume e pairings ( xen , yen ) Ne , which we refer to as access to an additional source of knowledge in the form of N n =1 sample displacements. These instances may be available for various reasons in practice. For example, they can arise from active tracking of a data subset (cf. [32]), from past transport phenomena (cf. [23]), or from manually defined or OT-generated virtual pairs in a controlled setting. To incorporate such sample displacements, we now introduce our proposed method ReshapeOT, which takes a classical OT formulation with Euclidean ground metric as a starting point and performs the following three conceptual steps: Step 1 (Compute Displacement Statistics) The method starts by summarizing the available groundtruth displacements as a second-moment matrix Σ. For the sake of generality, and to account for a broader set of usages, such as integrating knowledge from a general OT solution, we replace e e the sample of paired points by two sets of data points ( xek )kM=1 and (yel )lN=1 linked through some e. (The special case of a displacement sample can be recovered by setting coupling matrix γ e e e to permutation matrices.) Using the general formulation, we build Σ as: M = N and limiting γ ekl ( xek − yel )( xek − yel )⊤ Σ = ∑γ

(2)

kl

The leading eigenvectors of the matrix Σ span the subspace of observed displacements. Simple global translations in input space are spanned well by a single leading eigenvector. More complex phenomena, like rotations, require more of them.

5 of 20

Step 2 (Ground Metric Redefinition) The method proceeds by defining a new cost function informed by the sample displacements, represented by the second-moment matrix Σ calculated above. Specifically, we replace the original Euclidean metric with the Mahalanobis distance: d( x, y) =

q

( x − y)⊤ ( I + ηΣ)−1 ( x − y)

(3)

where η ≥ 0 controls the degree to which the displacements shape the metric. At η = 0, the expression reduces to the standard Euclidean distance. Increasing η progressively flattens the metric along directions of observed displacement, making travel in those directions cheaper. Step 3 (Recalculation of the OT Solution) As a last step, an OT solver of choice (e.g., Equation (1)) is applied to the costs derived from the new ground metric (c( x, y) = d( x, y)2 ). This leads to a new solution for the coupling between the source and target that better aligns with the observed displacements, as illustrated in Figure 1. Technically, ReshapeOT is a recalculation of the classical OT problem under a new ground metric. It introduces no additional structural constraints on the transport plan and no significant computational overhead. The only new hyperparameter is η, which can be selected via cross-validation or set based on the desired strength of guidance. In Section 4.3, we propose a heuristic to make setting it scale-invariant. 3.1. Geometric Interpretation of ReshapeOT and Relation to W2

In the following, we analyze the displacement-informed transport costs produced by ReshapeOT. Specifically, observing that Σ is positive semi-definite, we can show that d( x, y) ≤ ∥ x − y∥, with equality when η = 0 (no influence) or ( x − y) ∈ ker(Σ) (orthogonal to all given displacements). This relation extends to the distance between distributions, allowing us to show that ReshapeOT costs resulting from minimizing Equation (1) with the new costs are upper-bounded by the canonical W2 -Wasserstein distance: inf

γ∈Γ(µ,ν)

E(x,y)∼γ [d( x, y)2 ] ≤ W22 (µ, ν)

(4)

Here again, with equality when η = 0 or when every displacement under W2 lies in ker(Σ), i.e. transport entirely orthogonal to the subspace of sample displacements. The above inequality also shows that ReshapeOT strictly reduces the transport costs relative to W2 , carving ‘expressways’ through regions of high observed displacement. This stands in contrast to the more common use of geodesic distances (e.g., [38]), where the transport plan is instead reshaped by restricting the set of possible trajectories. As a corollary, we can show that ReshapeOT costs are equivalent to the squared W2 -Wasserstein distance with a Euclidean ground metric applied to reparametrized marginals. Specifically, the cost equals W22 ( T# µ, T# ν) with T ( x ) = Wx and W is any matrix satisfying W ⊤ W = ( I + ηΣ)−1 . Any such matrix W (e.g., PCA, ZCA, or Cholesky) yields the same transport distance. This mapping rescales the p displacement directions: along the i-th eigenvector of Σ, distances shrink by a factor 1/(1 + ηλi ), i.e. directions of large observed displacement become cheap compared to directions that are not spanned by these observations. 3.2. Extending ReshapeOT to Kernel Feature Spaces The ground metrics developed in Section 3 can globally reshape transport costs but have limits in capturing nonlinear transportation corridors found in real-world phenomena such as bird migration or blood circulation. To handle such cases, we extend ReshapeOT to kernel feature spaces [39–42]. Concretely, let K : Rd × Rd → R be a positive semi-definite kernel, which subsumes some nonlinear

6 of 20

mapping x 7→ Φ( x ) to some feature space, and in which distance calculations are equivalently expressible as kernel evaluations:

∥Φ( x ) − Φ(y)∥ =

q

K( x, x ) − 2K( x, y) + K(y, y)

Considering now the ReshapeOT ground metric in Equation (3) but replacing any inner product or distance computation by the same operations in feature space, we can recover an expression of the distance that depends on the inputs only through kernel evaluations on the query pair ( x, y) and the individual source and targets from the sample z = ( xe1 , . . . , xeM , ye1 , . . . , yeN ). Specifically, we get the kernelized form of the ReshapeOT ground metric: v u uK( x, x ) − 2K( x, y) + K(y, y) d(Φ( x ), Φ(y)) = t − (K( x, z) − K(y, z)) Λ (K(z, x ) − K(z, y))

(5)

with Λ = η A1/2 ( I + η A1/2 K(z, z) A1/2 )−1 A1/2 , where A = (diag( G1) − G ) is the weighted Laplacian e; γ e⊤ , 0), and where we have used K to also denote vectors or matrices of the coupling matrix G = (0, γ of kernel computations. (A proof of Equation (5) is provided in Appendix A). Furthermore, we can verify that the inequalities presented in Section 3.1 and their interpretations also hold for these kernelized distances and their associated Wasserstein distances.

4. Experiments In the following experiments, we evaluate ReshapeOT’s ability to recover ground-truth transport in synthetic and real-world settings. 4.1. Datasets Two of our three datasets are multivariate time series. Air Quality [43] records pollutant concentrations and meteorological variables in an Italian city at hourly intervals from March 2004 to February 2005. After preprocessing, it retains 9 features, with a median of 368 samples per shift in each source and target domain. Appliances [44] tracks appliance energy consumption in a low-energy building via sensor measurements taken every 10 minutes over 4.5 months. To match the hourly resolution of Air Quality, we resample the records to hourly intervals using median aggregation. After preprocessing, it retains 24 features, with a median of 133 samples per shift in each source and target domain. From each dataset, we create subsets with temporal shifts of 1–6 hours by pairing a source (all samples across the full time series) with a target consisting of the same days shifted by the given delay, naturally inducing a ground-truth identity coupling. Further details on both datasets and their preprocessing are provided in Supplementary Note A. Lastly, we consider the synthetic Rotating Moons problem, widely used in domain adaptation research [11]. It consists of two nested 2D moon-shaped point clouds (150 samples each) representing binary classes; target domains are obtained by rotating these clouds counterclockwise by increasing angles (10◦ – 90◦ ), with larger rotations producing more severe distribution shifts (cf. Figure 3 and Section 6). The dataset generation script is provided in Supplementary Note B. 4.2. Evaluation Metrics Our first evaluation assesses the quality of the coupling learned by ReshapeOT by comparing it against a ground-truth coupling derived from a fully tracked temporal phenomenon. To this end, we measure the mean error between ground-truth and predicted displacements. In this setting, all methods have access to the complete set of pre-aligned point correspondences, i.e. the ground-truth coupling is a normalized identity matrix. For each source point xk , we apply the barycentric mapping induced by γ to compute its predicted transport target ŷk = T ( xk ) [11,16]. We then evaluate the mean

7 of 20

  transport error between the true target yk and its prediction ŷk , defined as Ek ∥yk − ŷk ∥ . A perfect coupling, i.e. one that exactly recovers the ground-truth pairing, yields ŷk = yk and a score of zero. We perform an additional evaluation of ReshapeOT similar to the one originally proposed in [15], focusing on the transport model’s ability to retrieve features relevant to the studied phenomena. This extrinsic evaluation enables a broader comparison with non-OT methods. Given paired observations   M , we obtain the ground-truth shift attributions as R⋆ = E ( x 2 ( xm , ym )m m m,i − ym,i ) . We then split i =1 the original data into two unpaired subsets, apply each method to each subset to produce predicted attributions, and average the cosine similarities to the ground-truth attribution across the two splits. For methods that output a coupling matrix γ, we compute the predicted per-feature attributions Ri via W aX [15] (with p, q = 2), which decomposes the Wasserstein distance into non-negative feature contributions satisfying ∑i Ri = W22 (µ, ν). All other methods directly yield an attribution without the need for a post-hoc explainer. 4.3. ReshapeOT Implementation Details We evaluate ReshapeOT using a linear kernel, K( x, y) = ⟨ x, y⟩, and using an RBF kernel, K( x, y) = exp(−λ∥ x − y∥2 ). In our experiments, we set η = α/ Tr(Σ), where α ≥ 0 is a dimensionless scale parameter. This normalization makes ηΣ scale-invariant with respect to the magnitude of the observed displacements, so that α has a consistent interpretation across datasets and kernels: α = 0 disables the displacement prior, while values of α ≫ 1 correspond to strong guidance along the directions captured by Σ. While this removes the need to precisely set η, the α parameter still remains problem-dependent. If a task requires stricter guidance, α needs to be increased accordingly. e = 5 random pairs from the ground-truth coupling to For the time series datasets, we select N build the displacement and remove them from the remaining data at each experiment iteration, and use λ = 0.0005 for the RBF kernel and α = 102 for the regularization. For the Rotating Moons dataset, we e = 40, where each instance is selected randomly. When assemble a sample of displacements of size N applying ReshapeOT to this dataset, we use an RBF kernel and α = 104 to ensure an efficient capture of the information contained in the sample displacements, and use and integration strength of λ = 2. These hyperparameters are based on the best-performing results of a hyperparameter sweep provided in Supplementary Note C. Furthermore, we show ablations of ReshapeOT that replace the coupling e by a randomly permuted version γ erng , allowing us to test if performance improvement matrix γ derives from exploiting the exact linking of sources or targets in the sample or their overall variance. Note that W aX assumes a Minkowski metric, hence we use ReshapeOT only as a regularizer here to find γ based on the reshaped ground metric, but evaluate the feature attributions using this coupling and the squared Euclidean cost function (corresponding to W aX with p, q = 2). We do not evaluate the Rotating Moons dataset on the cosine metric, as it is too low-dimensional to yield meaningful insights. 4.4. Results The results in Tables 1 and 2 show averages over 20 random trials: for each time-pair and rotation, we repeated the experiment 20 times with a different random seed, thereby changing the splits and selected sample displacements each time. Furthermore, Table 3 presents an ablation over the number e for the time series datasets, where N e = 0 corresponds to classical OT of utilized displacement pairs ( N) in RBF kernel space. In Table 1, we show the results for how well the induced barycentric transport map aligns with the ground-truth coupling. We see that classical OT, the unregularized solution of solving Equation (1) with squared Euclidean cost, performs worse than ReshapeOT, but better than Sinkhorn, which is an entropy-regularized variant of Equation (1) [33] and yields a stochastic transport plan. This discrepancy between classical OT and Sinkhorn is likely due to the ground-truth transport being non-stochastic, which gives an advantage to methods that yield a permutation-like coupling (classical OT and ReshapeOT). Overall, ReshapeOT outperforms all other baselines by a margin, clearly demonstrating the guidance effect of using displacement information. In the Rotating Moons

8 of 20

dataset, the random sample displacement ablations fail to provide useful information, as the raw second-moment statistics without informative guidance are insufficient to capture the nature of the rotation shift. In this case, the data manifold has a more complex shape, and the locality of insightful displacements is required to accurately guide OT to transport along this manifold. The experimental results in Section 6 emphasize this behavior in the case of a domain adaptation task. Furthermore, the more complex manifold shapes here require a nonlinear kernel for accurate modeling. Hence, we see that ReshapeOT with an RBF kernel outperforms the linear kernel, achieving nearly perfect alignment. Table 1. Mean transport errors ↓ at different shift levels. The ablations of ReshapeOT use the same parametrization as their main counterparts. The values are the mean ± std. deviation. Bold = best, italic = 2nd best, ranked within row.

Baselines Dataset

Ours

Ablations

Shift Classical Sinkhorn Sinkhorn ReshapeOT ReshapeOT ReshapeOT ReshapeOT erng ) erng ) OT (ε = 0.05) (ε = 0.1) (Linear) (RBF) (Linear, γ (RBF, γ 0.20±0.11 0.56±0.23 0.80±0.29 0.96±0.30 1.07±0.29 1.18±0.26

Air Quality

1h 2h 3h 4h 5h 6h

0.44±0.17 0.78±0.27 0.99±0.34 1.14±0.39 1.26±0.40 1.36±0.37

1.08±0.16 1.22±0.21 1.32±0.24 1.39±0.26 1.46±0.26 1.51±0.25

1.32±0.16 1.42±0.20 1.49±0.22 1.54±0.23 1.58±0.22 1.62±0.21

Appliances

1h 2h 3h 4h 5h 6h

0.25±0.18 0.76±0.31 1.11±0.40 1.40±0.40 1.63±0.35 1.82±0.30

1.84±0.11 2.02±0.15 2.13±0.16 2.21±0.16 2.26±0.16 2.31±0.15

2.35±0.10 2.45±0.12 2.51±0.13 2.56±0.14 2.59±0.13 2.62±0.13

0.06±0.11 0.24±0.25 0.45±0.37 0.70±0.45 0.96±0.50 1.26±0.49

0.22±0.01 0.36±0.00 0.62±0.01 0.92±0.01 1.21±0.01

0.34±0.01 0.43±0.00 0.60±0.01 0.85±0.01 1.11±0.01

0.04±0.01 0.24±0.01 0.51±0.02 0.78±0.03 1.06±0.03

10◦ 30◦ Rotating Moons 50◦ 70◦ 90◦

0.04±0.01 0.34±0.01 0.69±0.01 1.03±0.01 1.33±0.01

0.20±0.11 0.55±0.23 0.80±0.29 0.95±0.30 1.07±0.30 1.17±0.26

0.36±0.15 0.73±0.27 0.96±0.33 1.12±0.34 1.24±0.34 1.33±0.33

0.35±0.16 0.73±0.28 0.96±0.34 1.11±0.34 1.23±0.35 1.32±0.33

0.06±0.11 0.25±0.25 0.46±0.37 0.72±0.44 0.98±0.48 1.27±0.47

0.21±0.22 0.65±0.44 1.04±0.57 1.40±0.61 1.70±0.58 1.98±0.51

0.18±0.20 0.58±0.39 0.93±0.50 1.28±0.54 1.59±0.53 1.86±0.48

0.04±0.01 0.09±0.02 0.06±0.01 0.08±0.03 0.10±0.02

0.15±0.03 0.56±0.07 0.91±0.05 1.15±0.08 1.37±0.08

0.31±0.05 0.88±0.09 1.04±0.15 1.16±0.10 1.18±0.11

In Table 2 we show the feature attribution evaluation results. Here, classical OT and Sinkhorn also use W aX to compute feature attributions. Additionally, we added the Constant Shift (Ri = 1) and Mean Shift (Ri = (E[ x ] − E[y])2i ) baselines from [15], which directly provide such attributions without an intermediate coupling step. We see that ReshapeOT achieves the highest attribution similarities across all methods, especially improving on the unguided OT formulations classical OT and Sinkhorn. There is no clear difference between using a linear or an RBF kernel for these datasets, as the transport processes are likely driven by a clear trend rather than shifting along a complex manifold. Relative performance across methods generally aligns with the results from Table 1, with the exception that Sinkhorn outperforms classical OT in the Appliances dataset. Since this evaluation abstracts away from the exact coupling, the Sinkhorn regularization helps estimate more robust attributions that reflect the underlying shift trend. Notably, ReshapeOT with random sample displacements still outperforms classical OT and Sinkhorn in the Air Quality dataset. This is because the random coupling preserves global second-moment structure, information about the relative spread of source and target that classical OT must infer entirely from the marginals. When the distribution shift has a dominant constant trend, this global shape information is sufficient to improve the transport. As in Table 1, Sinkhorn outperforms classical OT in the Appliances dataset, and now also beats the ReshapeOT ablations.

9 of 20

Table 2. Cosine similarities ↑ between the predicted and ground-truth shift attributions. The ablations of ReshapeOT use the same parametrization as their main counterparts. The values are the mean ± std. deviation. Bold = best, italic = 2nd best, ranked within row. Baselines Dataset

Shift Constant Shift

Ours

Ablations

Mean Shift

Classical Sinkhorn ReshapeOT ReshapeOT ReshapeOT ReshapeOT erng ) erng ) OT (ε = 0.05) (Linear) (RBF) (Linear, γ (RBF, γ

1h 2h 3h Air Quality 4h 5h 6h

0.81±0.08 0.82±0.08 0.83±0.07 0.85±0.06 0.86±0.05 0.88±0.05

0.80±0.22 0.87±0.19 0.89±0.15 0.88±0.17 0.91±0.11 0.91±0.10

0.87±0.07 0.92±0.07 0.94±0.07 0.95±0.07 0.95±0.05 0.95±0.05

0.79±0.05 0.84±0.06 0.88±0.06 0.90±0.06 0.91±0.05 0.92±0.04

0.97±0.03 0.98±0.03 0.98±0.03 0.98±0.02 0.98±0.02 0.98±0.03

1h 2h 3h Appliances 4h 5h 6h

0.32±0.06 0.34±0.07 0.37±0.08 0.40±0.09 0.43±0.09 0.46±0.09

0.46±0.30 0.45±0.29 0.45±0.26 0.45±0.25 0.47±0.23 0.48±0.22

0.69±0.09 0.70±0.11 0.73±0.12 0.75±0.11 0.77±0.11 0.79±0.11

0.74±0.07 0.76±0.07 0.79±0.06 0.81±0.06 0.82±0.05 0.84±0.05

0.89±0.10 0.91±0.09 0.91±0.08 0.92±0.07 0.92±0.06 0.93±0.06

0.97±0.03 0.98±0.03 0.98±0.03 0.98±0.02 0.98±0.02 0.98±0.03

0.89±0.10 0.90±0.09 0.91±0.08 0.91±0.07 0.92±0.06 0.92±0.06

0.87±0.12 0.93±0.09 0.95±0.07 0.95±0.08 0.96±0.05 0.96±0.05

0.87±0.11 0.93±0.09 0.95±0.07 0.95±0.08 0.96±0.05 0.96±0.05

0.72±0.16 0.75±0.15 0.77±0.14 0.79±0.13 0.80±0.12 0.82±0.11

0.70±0.15 0.73±0.15 0.75±0.14 0.77±0.12 0.79±0.12 0.81±0.11

Finally, Table 3 demonstrates the efficiency of ReshapeOT with respect to the number of provided e = 0 case represents classical OT in RBF displacement pairs in terms of the transport error. The N kernel space without any displacement pairs, and we clearly see that the kernel alone is insufficient to improve its performance. The results are essentially identical to those of using classical OT in input space as shown in Table 1. We present an ablation study across different kernel parameter values, λ, in Supplementary Note D. Interestingly, including only one displacement pair as guidance already allows ReshapeOT to outperform the baselines. This highlights the critical information about the shift contained in a single correctly coupled pair, as well as ReshapeOT’s ability to leverage such sparse data. As expected, performance improves further as we gradually increase the number of observed displacements. e on transport error ↓ for ReshapeOT (RBF, λ = 0.0005, α = 102 ). The values are the mean ± std. Table 3. Effect of N deviation. Bold = best, italic = 2nd best, ranked within row.

Dataset

Shift

e =0 N

e =1 N

e =3 N

e =5 N

e = 10 N

e = 20 N

Air Quality

1h 2h 3h 4h 5h 6h

0.45±0.169 0.78±0.265 0.99±0.337 1.14±0.392 1.26±0.398 1.36±0.373

0.34±0.179 0.71±0.285 0.94±0.333 1.08±0.337 1.20±0.337 1.31±0.323

0.25±0.153 0.62±0.269 0.86±0.315 1.01±0.323 1.12±0.308 1.22±0.273

0.20±0.107 0.55±0.230 0.80±0.294 0.95±0.304 1.07±0.295 1.17±0.264

0.16±0.080 0.49±0.216 0.74±0.281 0.89±0.285 1.02±0.282 1.12±0.260

0.13±0.071 0.45±0.197 0.70±0.269 0.85±0.273 0.98±0.278 1.09±0.262

Appliances

1h 2h 3h 4h 5h 6h

0.26±0.172 0.77±0.314 1.12±0.400 1.41±0.406 1.64±0.341 1.84±0.291

0.20±0.178 0.62±0.353 0.96±0.443 1.27±0.471 1.51±0.448 1.72±0.401

0.10±0.128 0.37±0.302 0.63±0.426 0.93±0.485 1.20±0.490 1.47±0.459

0.06±0.108 0.25±0.246 0.46±0.371 0.72±0.439 0.98±0.480 1.27±0.467

0.03±0.063 0.11±0.147 0.25±0.235 0.43±0.317 0.65±0.398 0.94±0.437

0.01±0.041 0.05±0.083 0.12±0.137 0.23±0.187 0.38±0.283 0.63±0.374

In summary, our results show that displacement guidance significantly improves OT’s ability to capture the true transport phenomena, even with only limited access. Furthermore, we showed that the choice of kernel plays an important role in the modeling of the feature space and helps account for nonlinear shifts. Specifically, the kernel does not improve the transport solution in isolation; rather, it serves as a mechanism for integrating knowledge from sample displacements more effectively.

10 of 20

5. Use Case 1: Refining Transport Models with Longitudinal Tracking Predicting and understanding the motion of populations is a central research question with relevance in a variety of scientific domains, such as epidemiology, wildlife monitoring, urban mobility, and more. This task becomes particularly challenging when population motion is only partially observable, which may arise from the technical infeasibility of large-scale longitudinal tracking or from its explicit avoidance to address privacy requirements. In practice, available data may consist of a bulk of unlinked data augmented by a few linked instances, e.g., collected through a specific study. To demonstrate ReshapeOT on a real-world geospatial transport problem, we consider bird migration and the retrieval of migratory routes. In this scenario, we have extensive population-level information, but we do not know individuals’ specific travel routes. A small number of birds, however, have GPS trackers that provide ground-truth displacement data along their flight paths. We use this setting to demonstrate ReshapeOT’s utility in leveraging such sparse individual-tracking data to recover population transport dynamics more accurately than using marginal data alone. While a similar experimental setup was adopted by Scarvelis & Solomon [29], we focus specifically on integrating individual bird-tracking data rather than population-level observations. We aggregated bird-tracking data from multiple studies from the MoveBank repository [45], covering various migratory species observed across North America (Supplementary Note E provides a list of the individual studies). For each bird and year, we compute median geographic positions (latitude and longitude) per individual and calendar week over a period of 13 weeks (mid-September to mid-December), representing autumn migration. Consecutive weeks serve as source and target distributions, representing population snapshots before and after a migration phase. Since this splitting creates overlapping domains, we added small Gaussian noise (σ = 10−4 ) to avoid identical points and to break potential ties. We remove outlier displacements beyond the 99th percentile to exclude, e.g., GPS artifacts. The dataset features are latitude and longitude information, which we transformed into 3D Cartesian coordinates so that the Euclidean distance corresponds to the chordal distance between

Sample displacements (a)

ReshapeOT (ours)

Classical OT (b)

(c)

kx − yk

d(Φ(x), Φ(y))

Figure 2. (a) Training observations (gray) from multiple bird tracking studies and manually selected displacements (red). (b) Coupling of classical OT from a squared Euclidean cost matrix built on Cartesian coordinates. (c) Coupling of ReshapeOT with RBF kernel (λ = 1, α = 103 ) on Cartesian coordinates. The insets show the induced square-root (for better visibility) cost fields of classical OT and ReshapeOT for a given reference point relative to a grid of targets.

11 of 20

points on the sphere, a close approximation of the great-circle distance at the spatial scales of weekly displacements. To construct the ground-truth displacement instances used as guidance, we manually selected 4 representative, complete bird trajectories (cf. Supplementary Note E) spanning the entire assessed time frame from the available studies and removed those points from the training data. This preprocessing e = 47 ground-truth displacements, as shown created 2266 untracked source and target samples, and N in Figure 2a. To integrate the sample displacements well, in particular to account for local variations in the direction of sample trajectories, we opt for the RBF kernel variant of ReshapeOT. As a baseline, we use classical OT with a squared Euclidean cost matrix built on the Cartesian coordinates. Figure 2 shows the coupling solutions for classical OT in panel (b), and ReshapeOT in (c). Classical OT scatters mass broadly, including substantial east–west transport that is physically implausible for autumn migration, which proceeds primarily north-to-south. The inset in (b) reveals the cause: a Euclidean ground metric induces a rotationally-symmetric field around any reference point (distorted only visually by the 2D map projection), so every geographic direction is equally cheap, and the optimal coupling has no preference for the dominant migration corridors. By incorporating groundtruth migration routes from only four birds (shown as red arrows in (a)), ReshapeOT reconfigures the cost function to favor movement along the observed migration patterns, concretely, prioritizing the north–south corridor over spurious east–west travel. The inset in (c) shows the effect on the ground metric: the field is now stretched along the observed displacements, making north–south movement cheap and east–west travel comparatively expensive. This reshaping leads to more consistent and plausible trajectories that remain consistent with their local trends (i.e. species or swarms travel along similar routes). In summary, this use case shows how ReshapeOT, specifically its kernel variant, efficiently integrates real-world displacements to yield more plausible couplings. By using data from only four tracked trajectories, ReshapeOT recovers the dominant north–south migration corridor while avoiding spurious east–west shortcuts.

6. Use Case 2: Domain Adaptation Under Manifold Rotation In our second use case, we apply ReshapeOT to a task, in which a classifier is trained on a labeled source domain and must adapt to an unlabeled but related target domain [46]. OT-based domain adaptation addresses this by, e.g., transporting source points (together with their labels) to the target distribution via the barycentric mapping induced by the coupling [11,16], and training the classifier on the transported data. We revisit the Rotating Moons scenario from Section 4.1, this time as a synthetic domain-adaptation benchmark: a binary classification task with two interleaved moon-shaped classes, where the target distribution is obtained by rotating the source [11]. We additionally sample a held-out test set of 1000 points from the rotated target distribution. The Setup column in Figure 3 shows a concrete example at 40◦ . This rotating two-moons setup mirrors the synthetic benchmark of Liu et al. [32], where pair matchings likewise serve as side information to probe whether OT recovers the rotation rather than collapsing onto a cross-class shortcut between adjacent moon edges. We compare ReshapeOT to classical OT, Sinkhorn, and Laplace-OT [11]. As in Section 4, we use an RBF kernel (λ = 2) with α = 104 for regularization of Σ. For each method and trial, we transport the labeled source to the target via barycentric mapping, train an SVM on the transported data with hyperparameters selected via grid search, and report classification error on the held-out test set. We repeat this protocol over 20 random seeds. In Figure 3 we exemplify the transport models of the baselines and ReshapeOT for one of the trials at 40◦ rotation, together with the learned SVM decision boundaries and predictions. It is apparent that standard OT methods resort to spurious shortcuts in pursuit of the least-action principle. The entropic regularization helps improve this issue, but it also alters the shape of the transported distribution.

12 of 20

Setup

Sinkhorn (ε = 0.05)

Classical OT

(b)

(c)

ReshapeOT (ours) (d)

Target

Decision function

Transport process

Source

Transport model

(a)

Laplace-OT

Test data

Sample displacements

Misclassified

Figure 3. Comparison of different OT-based domain adaptation methods on the Rotating Moons task at 40◦ rotation. Arrows denote the transport to the barycentric mapping, as determined by the coupling for each method and e = 40 randomly selected ground-truth setting. The experiment uses 150 samples per class and domain, and N displacements.

Classification errors at various degrees of target rotation are shown in Figure 4, with increasing rotation angles corresponding to increasingly difficult domain-adaptation problems. A more exhaustive table of results is provided in Supplementary Note F, including additional baselines and ablations. The results clearly show that ReshapeOT outperforms all other methods in the most challenging settings. The guidance provided by the ground-truth displacements is a good model of the manifold transformation, yielding robust couplings across all levels of rotation. Notably, Sinkhorn with ε = 0.05 is a strong competitor in the domain adaptation setting, although it performed poorly in our quantitative experiments (see Table 1). Entropic smoothing prevents Sinkhorn from recovering exact pointwise correspondences. But that same smoothing also averages out cross-class shortcuts in the barycentric mapping, preserving the dominant rotation direction (cf. Figure 3) sufficiently for the SVM to fit a correct decision boundary. Note that a linear kernel fails in this setting as it cannot account for the rotation geometry (cf. the results in Supplementary Note F).

Classification error [%] ↓

50 40 Classical OT Sinkhorn (ε = 0.05) Laplace-OT ReshapeOT (ours)

30 20 10 0 20

40

60

80

◦

Rotation [ ] Figure 4. Domain adaptation results in terms of the SVM classification error [%] ↓ for the Rotating Moons dataset. The ablations of ReshapeOT use the same parametrization as their main counterparts.

13 of 20

In summary, ReshapeOT substantially reduces the domain-adaptation classification error on this benchmark by translating a small number of ground-truth displacements into a transport solution that aligns with the manifold’s rotation, whereas unguided baselines are hampered by cross-class shortcuts. This result has direct practical consequences in scientific applications: when technical batch effects place samples closer in Euclidean space than biologically meaningful variation (cf. [47]), unguided OT can match on the wrong signal. ReshapeOT allows a practitioner to encode prior knowledge about which directions of variation are biologically meaningful, steering the coupling away from spurious technical correlations and toward genuine biological correspondence.

7. Conclusion Optimal transport is a workhorse in modeling distribution shifts, but the reliability of any OT solution is bounded by the fidelity of its ground metric. With a Euclidean metric, OT solvers exploit shortcuts that violate manifold structure both in the data and in how it is transported. ReshapeOT addresses this by reshaping an initial Euclidean ground metric into a Mahalanobis distance derived from secondorder statistics of observed displacements, carving expressways through the cost landscape along empirically supported directions. Our approach is easy to implement, minimally affects OT’s compute costs, and can be extended to kernel feature spaces. We note that ReshapeOT is not limited to the classical OT problem formulation. Any OT algorithm that operates on a cost matrix (e.g., Sinkhorn [33], Partial-OT [48], etc.) can in principle be used on top of the new ground metric. Empirically, our approach demonstrates consistently high accuracy, outperforming a representative set of baselines across several benchmark evaluations. We further demonstrate its usefulness in two practical use cases. Beyond ReshapeOT’s capabilities, our paper demonstrates the practical feasibility of integrating past displacements into the transport modeling task and the associated benefits. We see a particular potential for our approach in single-cell tracking (cf. [8,9]). Building on the work of [29] that demonstrated the benefit of combining OT and existing data, our method, which explicitly aligns transport to past displacements, would allow specifically for the integration of longitudinal cell trajectories. Despite its strengths, ReshapeOT has limitations. Most importantly, it requires a representative and robust set of displacements; without them, the intrinsic transport cannot be accurately extracted, and the optimizer may be misled. Furthermore, the kernelized variant requires precise hyperparameter tuning. These parameters must be flexible enough to align with sample displacements, yet sufficiently rigid to avoid introducing spurious structural barriers or shortcuts. Finally, a natural future work would be to embed ReshapeOT within the information-geometric framework, exploiting known links between Wasserstein geometry and statistical manifolds [5,49, 50]. Concretely, sample displacements could be represented as tangent vectors on the manifold of probability distributions, enabling a principled interpolation between source and target distributions guided by prior trajectories on that manifold, thereby unifying the metric-learning and geometric perspectives pursued here. Funding: K.R.M. was supported in part by the German Federal Ministry for Research, Technology and Space (BMFTR) under grants 01IS18025A, 031L0207D, 01IS18037A, 16IS24087C. K.R.M. was also supported by the Institute of Information & communications Technology Planning & Evaluation (IITP) grants funded by the Korea government (MSIT) (No. RS-2019-II190079, Artificial Intelligence Graduate School Program of Korea University and No. RS-2024-00457882, AI Research Hub Project). Acknowledgments: During the preparation of this manuscript, the authors used Claude Sonnet 4.6, Opus 4.6, and Gemini 3.1 for rewriting parts of the text. The authors have reviewed and edited the output and take full responsibility for the content of this publication.

14 of 20

Appendix A Kernelization of the ReshapeOT Distance This appendix proves the kernel representation of the ReshapeOT distance presented in Equation (5). The derivation shows that d depends on the data only through inner products, and therefore admits a realization in any RKHS induced by a positive semi-definite kernel. Proposition A1 (Kernelization of d). Let { xe1 , . . . , xeM }, {ye1 , . . . , yeN } ⊂ Rd be source and target samples eij ≥ 0, and let K : Rd × Rd → R be a positive semi-definite kernel. Define Σ ∈ Rd×d with transport coupling γ as in Equation (2) and d : Rd × Rd → R as in Equation (3). Write z = ( xe1 , . . . , xeM , ye1 , . . . , yeN ) for the concatenation of all samples, and let K ∈ R( M+ N )×( M+ N ) be the kernel Gram matrix with Kkl = K(zk , zl ) and A ∈ R( M+ N )×( M+ N ) the block matrix ! e diag( a) −γ A= , e⊤ −γ diag(b) eij and b j = ∑i γ eij are the marginals. Then d admits the kernel representation where ai = ∑ j γ d(Φ( x ), Φ(y)) =

q

K( x, x ) − 2K( x, y) + K(y, y) − g⊤ Mg, 1

1

1

1

where gk = K(zk , x ) − K(zk , y) and M = η A 2 ( I + η A 2 KA 2 )−1 A 2 . Proof. Stack all training points as rows of z = ( xe1 , . . . , xeM , ye1 , . . . , yeN ) ∈ R( M+ N )×d , and for each pair (i, j) define the indicator vector αij ∈ {−1, 0, +1} M+ N with +1 at position i and −1 at position M + j. Then xei − yej = z⊤ αij , and substituting into the definition of Σ gives eij (z⊤ αij )(z⊤ αij )⊤ = z⊤ Σ = ∑γ ij



 ⊤ e γ α α ∑ij ij ij ij z = z⊤ Az. | {z } =A

One verifies that A has the stated block structure. Applying the Woodbury identity to ( I + ηz⊤ Az)−1 yields  1

1

1

I + ηz⊤ Az

 −1

  = I − z⊤ Λz ,

1

where Λ = η A 2 ( I + η A 2 KA 2 )−1 A 2 and K = zz⊤ is the Gram matrix with Kkl = z⊤ k zl . This converts the d × d inversion into an ( M + N ) × ( M + N ) inversion. Substituting into d and writing g = z( x − y) ∈ R M+ N :   d( x, y)2 = ( x − y)⊤ I − z⊤ Mz ( x − y)

= ∥ x − y∥2 − g⊤ Mg. ⊤ ⊤ Every quantity depends on the data only through inner products: Kkl = z⊤ k zl , gk = zk x − zk y, and 2 ⊤ ⊤ ⊤ e. Replacing each inner product u⊤ v ∥ x − y∥ = x x − 2x y + y y. The matrix A depends only on γ by the kernel evaluation K(u, v) gives the stated formula, which is valid in any RKHS induced by K.

If displacement data is provided as a set of paired observations ( xel , yel )lN=1 , and marginals are modeled as uniform, then A has the simple structure 1 I A= N −I

−I I

!

15 of 20

1

where A 2 can be computed in closed form as 1

I A = √ 2N − I 1 2

−I I

!

References 1. 2. 3. 4. 5. 6. 7. 8.

9. 10. 11. 12. 13. 14. 15.

16. 17. 18.

19. 20.

21. 22. 23.

Sugiyama, M.; Krauledat, M.; Müller, K.R. Covariate Shift Adaptation by Importance Weighted Cross Validation. J. Mach. Learn. Res. 2007, 8, 985–1005. Quionero-Candela, J.; Sugiyama, M.; Schwaighofer, A.; Lawrence, N.D. Dataset Shift in Machine Learning; The MIT Press, 2009. Amari, S.I. Information Geometry and Its Applications, 1 ed.; Applied Mathematical Sciences, Springer: Tokyo, Japan, 2016. Amari, S.i.; Nagaoka, H. Methods of information geometry; Vol. 191, American Mathematical Soc., 2000. Otto, F. The geometry of dissipative evolution equations: the porous medium equation. Communications in Partial Differential Equations 2001, 26, 101–174. Villani, C. Optimal Transport: Old and New; Grundlehren Der Mathematischen Wissenschaften, Springer Berlin Heidelberg, 2008. Peyré, G.; Cuturi, M. Computational Optimal Transport with Applications to Data Sciences. Foundations and Trends in Machine Learning 2019, 11, 355–607. Schiebinger, G.; Shu, J.; Tabaka, M.; Cleary, B.; Subramanian, V.; Solomon, A.; Gould, J.; Liu, S.; Lin, S.; Berube, P.; et al. Optimal-transport analysis of single-cell gene expression identifies developmental trajectories in reprogramming. Cell 2019, 176, 928–943. Klein, D.; Palla, G.; Lange, M.; Klein, M.; Piran, Z.; Gander, M.; Meng-Papaxanthos, L.; Sterr, M.; Saber, L.; Jing, C.; et al. Mapping Cells through Time and Space with Moscot. Nature 2025, 638, 1065–1075. Montavon, G.; Müller, K.R.; Cuturi, M. Wasserstein training of restricted Boltzmann machines. 2016, Vol. 29, Advances in Neural Information Processing Systems, pp. 3711–3719. Courty, N.; Flamary, R.; Tuia, D.; Rakotomamonjy, A. Optimal Transport for Domain Adaptation. IEEE Transactions on Pattern Analysis and Machine Intelligence 2017, 39, 1853–1865. Andéol, L.; Kawakami, Y.; Wada, Y.; Kanamori, T.; Müller, K.; Montavon, G. Learning domain invariant representations by joint Wasserstein distance minimization. Neural Networks 2023, 167, 233–243. Huguet, G.; Tong, A.; Zapatero, M.R.; Tape, C.J.; Wolf, G.; Krishnaswamy, S. Geodesic Sinkhorn For Fast and Accurate Optimal Transport on Manifolds. In Proceedings of the MLSP. IEEE, 2023, pp. 1–6. Paty, F.; Cuturi, M. Subspace Robust Wasserstein Distances. In Proceedings of the ICML. PMLR, 2019, Vol. 97, Proceedings of Machine Learning Research, pp. 5072–5081. Naumann, P.; Kauffmann, J.; Montavon, G. Wasserstein Distances Made Explainable: Insights Into Dataset Shifts and Transport Phenomena. IEEE Transactions on Pattern Analysis and Machine Intelligence 2026. https://doi.org/10.1109/TPAMI.2026.3656947. Ferradans, S.; Papadakis, N.; Peyré, G.; Aujol, J. Regularized Discrete Optimal Transport. SIAM J. Imaging Sci. 2014, 7, 1853–1882. Theodoropoulos, P.; Komianos, N.; Pacelli, V.; Liu, G.; Theodorou, E.A. Feedback Schrödinger Bridge Matching. In Proceedings of the ICLR. OpenReview.net, 2025. Lin, C.; Azabou, M.; Dyer, E.L. Making transport more robust and interpretable by moving data through a small number of anchor points. In Proceedings of the ICML, 2021, Vol. 139, Proceedings of Machine Learning Research, pp. 6631–6641. Sato, R.; Cuturi, M.; Yamada, M.; Kashima, H. Fast and Robust Comparison of Probability Measures in Heterogeneous Spaces. CoRR 2020, abs/2002.01615. Gu, X.; Yang, Y.; Zeng, W.; Sun, J.; Xu, Z. Keypoint-Guided Optimal Transport with Applications in Heterogeneous Domain Adaptation. 2022, Vol. 35, Advances in Neural Information Processing Systems, pp. 14972–14985. Cuturi, M.; Avis, D. Ground metric learning. J. Mach. Learn. Res. 2014, 15, 533–564. Kerdoncuff, T.; Emonet, R.; Sebban, M. Metric Learning in Optimal Transport for Domain Adaptation. In Proceedings of the IJCAI. ijcai.org, 2020, pp. 2162–2168. Jawanpuria, P.; Shi, D.; Mishra, B.; Gao, J. A Riemannian Approach to Ground Metric Learning for Optimal Transport. In Proceedings of the ICASSP. IEEE, 2025, pp. 1–5.

16 of 20

24. 25. 26. 27. 28. 29. 30. 31.

32. 33. 34. 35. 36. 37. 38. 39. 40. 41. 42. 43. 44. 45.

46. 47. 48. 49. 50.

Goldberger, J.; Roweis, S.T.; Hinton, G.E.; Salakhutdinov, R. Neighbourhood Components Analysis. 2004, Vol. 17, Advances in Neural Information Processing Systems, pp. 513–520. Davis, J.V.; Kulis, B.; Jain, P.; Sra, S.; Dhillon, I.S. Information-theoretic metric learning. In Proceedings of the ICML. ACM, 2007, ACM International Conference Proceeding Series, pp. 209–216. Weinberger, K.Q.; Saul, L.K. Distance Metric Learning for Large Margin Nearest Neighbor Classification. J. Mach. Learn. Res. 2009, 10, 207–244. Bar-Hillel, A.; Hertz, T.; Shental, N.; Weinshall, D. Learning a Mahalanobis Metric from Equivalence Constraints. J. Mach. Learn. Res. 2005, 6, 937–965. Heitz, M.; Bonneel, N.; Coeurjolly, D.; Cuturi, M.; Peyré, G. Ground Metric Learning on Graphs. J. Math. Imaging Vis. 2021, 63, 89–107. Scarvelis, C.; Solomon, J. Riemannian Metric Learning via Optimal Transport. In Proceedings of the ICLR. OpenReview.net, 2023. Pooladian, A.; Domingo-Enrich, C.; Chen, R.T.Q.; Amos, B. Neural Optimal Transport with Lagrangian Costs. In Proceedings of the UAI. PMLR, 2024, Proceedings of Machine Learning Research, pp. 2989–3003. Kapusniak, K.; Potaptchik, P.; Reu, T.; Zhang, L.; Tong, A.; Bronstein, M.M.; Bose, A.J.; Giovanni, F.D. Metric Flow Matching for Smooth Interpolations on the Data Manifold. 2024, Vol. 37, Advances in Neural Information Processing Systems. Liu, R.; Balsubramani, A.; Zou, J. Learning transport cost from subset correspondence. In Proceedings of the ICLR. OpenReview.net, 2020. Cuturi, M. Sinkhorn Distances: Lightspeed Computation of Optimal Transport. 2013, Vol. 26, Advances in Neural Information Processing Systems, pp. 2292–2300. Perrot, M.; Courty, N.; Flamary, R.; Habrard, A. Mapping Estimation for Discrete Optimal Transport. 2016, Vol. 29, Advances in Neural Information Processing Systems, pp. 4197–4205. Stuart, A.M.; Wolfram, M. Inverse Optimal Transport. SIAM J. Appl. Math. 2020, 80, 599–619. Andrade, F.; Peyré, G.; Poon, C. Sparsistency for inverse optimal transport. In Proceedings of the ICLR. OpenReview.net, 2024. Amos, B.; Luise, G.; Cohen, S.; Redko, I. Meta Optimal Transport. In Proceedings of the ICML. PMLR, 2023, Proceedings of Machine Learning Research, pp. 791–813. Bonet, C.; Drumetz, L.; Courty, N. Sliced-Wasserstein Distances and Flows on Cartan-Hadamard Manifolds. J. Mach. Learn. Res. 2025, 26, 32:1–32:76. Schölkopf, B.; Smola, A.; Müller, K.R. Nonlinear component analysis as a kernel eigenvalue problem. Neural computation 1998, 10, 1299–1319. Müller, K.R.; Mika, S.; Rätsch, G.; Tsuda, K.; Schölkopf, B. An introduction to kernel-based learning algorithms. IEEE Trans. Neural Networks 2001, 12, 181–201. Schölkopf, B.; Smola, A.J. Learning with kernels; Adaptive Computation and Machine Learning series, MIT Press: London, England, 2002. Zhang, Z.; Wang, M.; Nehorai, A. Optimal Transport in Reproducing Kernel Hilbert Spaces: Theory and Applications. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 42, 1741–1754. Vito, S. Air Quality. UCI Machine Learning Repository, 2008. https://doi.org/10.24432/C59K5F. Candanedo, L. Appliances Energy Prediction. UCI Machine Learning Repository, 2017. https://doi.org/10 .24432/C5VC8G. Kays, R.; Davidson, S.C.; Berger, M.; Bohrer, G.; Fiedler, W.; Flack, A.; Hirt, J.; Hahn, C.; Gauggel, D.; Russell, B.; et al. The Movebank system for studying global animal movement and demography. Methods in Ecology and Evolution 2022, 13, 419–431. Ganin, Y.; Lempitsky, V.S. Unsupervised Domain Adaptation by Backpropagation. In Proceedings of the ICML. JMLR.org, 2015, JMLR Workshop and Conference Proceedings, pp. 1180–1189. Kömen, J.; de Jong, E.D.; Hense, J.; Marienwald, H.; Dippel, J.; Naumann, P.; Marcus, E.; Ruff, L.; Alber, M.; Teuwen, J.; et al. Towards Robust Foundation Models for Digital Pathology. CoRR 2025, abs/2507.17845. Chapel, L.; Alaya, M.Z.; Gasso, G. Partial Optimal Tranport with Applications on Positive-Unlabeled Learning. 2020, Vol. 33, Advances in Neural Information Processing Systems, pp. 2903–2913. Amari, S.i.; Karakida, R.; Oizumi, M. Information geometry connecting Wasserstein distance and Kullback–Leibler divergence via the entropy-relaxed transportation problem. Information Geometry 2018, 1, 13–37. Ay, N. Information geometry of the Otto metric. Information Geometry 2024, 8, 209–232.

17 of 20

S UPPLEMENTARY N OTES Supplementary Note A Time-series Dataset Details The Air Quality [43] and Appliances [44] time-series datasets used in Section 4 were preprocessed in line with [15], with the following main changes: • • •

We use a RobustScaler (based on the median and IQR) instead of a StandardScaler (based on the mean and standard deviation) to improve robustness to outliers. For this reason, we also do not remove outlier instances. Outlier features, however, are removed as in [15]. In case of Appliances, we use the median instead of the mean for resampling the time-series to 1h-intervals. Air Quality already comes sampled in 1h-intervals.

The splitting into the unpaired subsets is described in [15]. Furthermore, we consider the same time intervals, i.e., all 123 time pairs spanning shifts of 1 to 6 hours. Ground-truth pairing for a sample at time td is naturally provided by the time-series record at time td + h (with 0 ≤ td < td + h ≤ 23 and h ∈ {1, 2, 3, 4, 5, 6}) on the same day d (cf. [15]). For computing the ‘transport error’ metric, all methods use a barycentric mapping [11,16] to compute the transport induced by their coupling.

Supplementary Note B Synthetic Rotating Moons Generation Below, we show the Python script used to generate the synthetic Rotating Moons data for a given rotation, as used in Sections 4 and 6. It largely follows the experimental description from [11]. Listing 1: Python code to generate the synthetic Rotating Moons dataset. import numpy as np from sklearn.datasets import make_moons from sklearn.utils import check_random_state def make_moons_da( rotation: float, n_per_moon: int = 150, noise: float = 0.05, n_test: int = 1000, random_state=None, ): rng = check_random_state(random_state) # --- Source: standard two moons --------------------------------Xs, ys = make_moons(n_samples=2 * n_per_moon, noise=noise, random_state=rng) Xs[:, 0] -= 0.5 # --- Target: rotate source points ------------------------------theta = np.radians(-rotation) R = np.array([[np.cos(theta), -np.sin(theta)], [np.sin(theta), np.cos(theta)]]) Xt = Xs @ R yt = ys.copy() # --- Independent test set from the target distribution ---------Xtest, ytest = make_moons(n_samples=n_test, noise=noise, random_state=rng) Xtest[:, 0] -= 0.5 Xtest = Xtest @ R return Xs, ys, Xt, yt, Xtest, ytest

18 of 20

Supplementary Note C Hyperparameter Sensitivity of ReshapeOT This supplementary note provides ablations of ReshapeOT’s hyperparameters. In particular, Figures S1 and S2 illustrate the effects of the precisions λ of the RBF kernel K( x, y) = exp(−λ∥ x − y∥2 ) in combination with the Σ regularization scaling parameter α with respect to the mean transport error on the evaluated datasets (cf. Sections 4.1 and 4.2). ReshapeOT hyperparameter sensitivity (transport error ↓) Air Quality

Appliances

0.93

0.93

0.95

1.11

1.14

1.16

1.18

1.05

1.05

1.08

1.20

1.22

1.26

1.40

±0.45

±0.45

±0.46

±0.54

±0.56

±0.56

±0.57

±0.62

±0.62

±0.65

±0.77

±0.78

±0.81

±0.94

101

0.86

0.86

0.92

1.07

1.10

1.12

1.14

0.93

0.94

1.01

1.14

1.16

1.20

1.34

±0.44

±0.44

±0.47

±0.55

±0.57

±0.58

±0.58

±0.64

±0.63

±0.67

±0.79

±0.81

±0.84

±0.97

α 102

0.86

0.87

0.93

1.07

1.11

1.13

1.14

0.85

0.87

0.96

1.11

1.13

1.17

1.31

±0.45

±0.46

±0.48

±0.56

±0.57

±0.58

±0.59

±0.70

±0.68

±0.70

±0.83

±0.85

±0.88

±1.01

0.92

0.93

0.97

1.11

1.14

1.16

1.18

0.87

0.88

0.98

1.12

1.15

1.19

1.33

10

10

0

3

104

±0.46

±0.47

±0.49

±0.55

±0.56

±0.57

±0.57

±0.75

±0.72

±0.74

±0.85

±0.87

±0.90

±1.02

0.99

0.99

1.02

1.15

1.18

1.21

1.22

0.94

0.92

1.02

1.16

1.19

1.23

1.37

±0.48

±0.49

±0.50

±0.54

±0.55

±0.56

±0.56

±0.82

±0.77

±0.78

±0.88

±0.90

±0.92

±1.04

0.0005

0.005

0.05

0.5

1.0

2.0

4.5

0.0005

0.005

0.05

0.5

1.0

2.0

4.5

λ

λ

Figure S1. Averaged transport errors (cf. Section 4.2) over all Air Quality and Appliances shift delays (1h–6h) for ReshapeOT’s (λ, α) hyperparameter combinations.

ReshapeOT hyperparameter sensitivity (Rotating Moons) Classification error [%] ↓

Transport error ↓

100

19.74

19.54

20.02

23.17

25.86

25.19

22.05

21.28

18.02

14.49

0.62

0.62

0.63

0.68

0.73

0.73

0.67

0.65

0.57

0.45

±14.59

±14.50

±15.05

±17.52

±19.06

±17.99

±15.63

±15.22

±13.87

±11.43

±0.40

±0.40

±0.40

±0.43

±0.44

±0.38

±0.29

±0.27

±0.18

±0.10

101

17.16

17.33

17.48

15.03

14.97

14.84

13.91

13.47

12.05

10.52

0.54

0.54

0.53

0.47

0.46

0.44

0.41

0.40

0.36

0.31

±12.26

±12.48

±13.24

±12.01

±11.61

±11.17

±10.28

±10.14

±9.54

±8.79

±0.35

±0.36

±0.37

±0.29

±0.24

±0.19

±0.13

±0.11

±0.08

±0.06

α 102

16.68

16.63

13.69

9.63

8.97

8.30

7.68

7.52

6.93

6.73

0.52

0.50

0.37

0.24

0.21

0.19

0.18

0.18

0.18

0.19

±11.49

±12.40

±10.96

±7.37

±7.01

±6.13

±5.90

±6.10

±6.59

±6.86

±0.34

±0.35

±0.27

±0.12

±0.09

±0.07

±0.06

±0.05

±0.05

±0.05

103

16.69

13.82

9.08

6.80

5.89

4.81

3.92

3.61

3.87

4.41

0.50

0.36

0.22

0.15

0.14

0.12

0.11

0.11

0.12

0.14

±12.31

±11.21

±8.47

±6.88

±6.05

±5.43

±4.94

±4.95

±5.10

±5.45

±0.34

±0.27

±0.16

±0.07

±0.06

±0.05

±0.04

±0.04

±0.04

±0.04

104

13.86

9.86

6.02

3.94

2.56

1.65

1.81

2.07

2.65

3.57

0.36

0.22

0.15

0.10

0.09

0.08

0.08

0.08

0.10

0.12

±11.25

±10.99

±7.11

±5.57

±4.42

±3.33

±3.50

±3.73

±4.21

±4.80

±0.27

±0.19

±0.10

±0.05

±0.04

±0.03

±0.03

±0.03

±0.04

±0.04

0.0005

0.005

0.05

0.5

1.0

2.0

4.0

5.0

10.0

20.0

0.0005

0.005

0.05

0.5

1.0

2.0

4.0

5.0

10.0

20.0

λ

λ

Figure S2. Averaged evaluation metrics over all Rotating Moons rotations (10◦ – 90◦ ) for ReshapeOT’s (λ, α) hyperparameter combinations. ‘Classification error’ is the domain adaptation metric of Section 6, and ‘transport error’ corresponds to the metric introduced in Section 4.2.

Supplementary Note D Classical OT in RBF Kernel Space In Figure S3, we show an ablation of classical OT in RBF kernel space for different precisions λ in terms of the domain adaptation classification error on the Rotating Moons task. As we see, none of the evaluated kernel precisions improves over the baseline performance of classical OT in the input space.

19 of 20

No experience RBF kernel hyperparameter ablation λ = 0.0005 λ = 0.005 λ = 0.05 λ = 0.5 λ = 1.0 λ = 2.0

60

Classification error [%]

50 40

λ = 4.0 λ = 5.0 λ = 10.0 λ = 20.0 No kernel

30 20 10 0 10

20

30

40

50 Rotation [◦ ]

60

70

80

90

Figure S3. Ablation on applying classical OT in the RBF kernel space (i.e. without displacement guidance) at different λ precisions for the Rotating Moons domain adaptation task of Section 6 measured by the classification error [%] ↓.

Supplementary Note E Bird Migration Data The datasets used in Section 5 are publicly available from the MoveBank repository [45], and we list the specific studies in Table S1, and selected trajectories for ReshapeOT guidance in Table S2. The subsets we used in the main text span mid-September to mid-December, covering autumn migration. Some birds have been tracked over multiple years, so they can recur when a new migration period starts. After preprocessing, 2266 source–target pairs remain in the training data, and 47 pairs remain for the guidance of ReshapeOT. Note that the Galapagos Albatrosses study for the considered time frame contains barely any movement, as these birds do not migrate in that period. Rather, it acts as natural noise as found in real-world recorded data. In general, some birds may also settle early after finishing their migration within the time frame, thereby becoming mostly stationary data points. Finally, NYSDEC Raptor Tracking includes four bird species, whereas all other datasets include only a single species. Table S1. Data studies from MoveBank [45] used in Section 5. Study name Galapagos Albatrosses https://dx.doi.org/10.5441/001/1.3hp3s250

Migration of Sabine’s gulls from the Canadian High Arctic https://dx.doi.org/10.5441/001/1.c745vb70

NYSDEC Raptor Tracking https://dx.doi.org/10.5441/001/1.s65q50j0

Turkey vultures in North and South America https://dx.doi.org/10.5441/001/1.46ft1k05

Pandion haliaetus Osprey - SouthEast Michigan

# Birds

# Samples

4

14

21

243

61

1300

19

603

17

153

https://www.movebank.org/cms/webapp?gwt_fragment=page=studies,path=study10204361

20 of 20

Table S2. Reference trajectories used to guide ReshapeOT in Section 5. ID

Year

# Samples

Turkey vultures in North and South America_Steamhouse 2 Migration of Sabine’s gulls from the Canadian High Arctic_BI Migration of Sabine’s gulls from the Canadian High Arctic_AD NYSDEC Raptor Tracking-argos_BAEA 0629-30002 E63

2012 2008 2011 2001

13 11 13 10

Supplementary Note F Further Domain Adaptation Results We provide additional results to Figure 4 from the domain adaptation experiment of Section 6 in Table S3. In particular, we provide more rotation degrees, and Sinkhorn [33] results at different regularization levels. All methods use a barycentric mapping [11,16] to compute the transport induced by their coupling. Table S3. Classification error [%] ↓ on the Rotating Moons domain adaptation task across target rotation angles. Higher angles generally correspond to more challenging shifts. Means over 20 repetitions. Values: mean ± standard deviation of error%. Bold = best, italic = second best per row. Baselines Rotation 10◦

20◦ 30◦ 40◦ 50◦ 60◦ 70◦ 80◦ 90◦

Ablations

Ours

Classical OT

Sinkhorn

Sinkhorn

Sinkhorn

(ε = 0.01)

(ε = 0.05)

(ε = 0.1)

0.00±0.00 3.34±2.24 9.01±1.62 15.01±1.84 19.90±1.28 26.09±1.89 34.25±2.26 41.12±3.63 49.48±2.02

0.00±0.00 0.11±0.21 0.48±0.46 8.22±1.33 14.47±0.81 20.25±1.02 28.68±2.06 40.42±2.54 49.14±2.00

0.01±0.02 0.01±0.03 0.00±0.00 0.03±0.09 2.24±1.15 11.13±1.99 20.45±1.26 27.66±1.26 36.45±2.28

0.34±0.23 0.40±0.50 0.42±0.34 0.94±0.63 2.70±1.24 4.29±1.55 8.02±1.47 15.87±2.66 27.37±3.71

Laplace-OT 0.00±0.00 0.65±0.49 1.42±0.91 10.16±2.21 16.16±1.34 22.99±1.32 30.38±2.15 42.08±2.56 51.30±2.19

ReshapeOT

ReshapeOT

ReshapeOT

ReshapeOT

(Linear)

(RBF)

erng ) (Linear, γ

erng ) (RBF, γ

0.68±0.96 3.43±1.93 7.57±1.54 10.97±1.92 14.48±1.86 19.12±1.59 24.29±2.06 30.86±2.65 36.19±2.05

0.00±0.00 0.00±0.00 0.05±0.20 0.16±0.51 0.43±1.07 0.88±1.37 1.89±3.60 3.78±4.22 6.15±4.34

1.87±1.67 9.31±2.21 15.84±2.60 23.01±2.74 29.03±3.14 34.62±3.88 39.39±3.72 44.88±3.88 51.13±3.80

3.64±3.04 21.54±5.15 27.55±5.01 35.55±6.80 39.62±7.57 42.60±8.34 47.14±6.99 52.65±9.21 53.37±9.72

Although Sinkhorn with ε = 0.1 performs better at higher rotation degrees, we chose Sinkhorn with ε = 0.05 in the main text experiment as it provides a good overall trade-off between regularization and performance.

Record · ID 158569 · SHA-256 a44d94e9595b46e8
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.