ConceptioArchivearXiv CS
arXiv CSopen access

Closed-Form Spectral Regularization for Multi-Task Model Merging

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

1

Closed-Form Spectral Regularization for Multi-Task Model Merging

arXiv:2606.07289v1 [cs.LG] 5 Jun 2026

Yongxian Wei, Runxi Cheng, Xingxuan Zhang, Li Shen, Chun Yuan, Peng Cui, Fellow, IEEE, Dacheng Tao, Fellow, IEEE

Abstract—Model merging combines several independently finetuned experts into a single multi-task model without any training data, reducing the storage, serving, and decentralized-development costs of large foundation models. State-of-the-art merging methods formulate merging as a layer-wise quadratic interference minimization problem. Although this problem admits an exact closed-form pseudoinverse solution, that solution underperforms hundreds of iterations of gradient descent in practice. The iterative loop dominates the cost of the pipeline (e.g., 85 minutes and 42 GB of GPU memory), yet its effectiveness has remained unexplained. We revisit this regime and show that the iterative solver does not primarily act as an optimizer; rather, it serves as an implicit spectral regularizer for an ill-posed normal equation, where smalleigenvalue directions of the per-layer interference operator amplify proxy noise. Building on this finding, we formalize multi-task model merging as a noisy linear inverse problem, and propose a spectral filtering estimator parameterized by a per-direction filter hk . We instantiate this estimator with SWUDI, a closed-form method that combines a soft exponential filter, which matches the gradient-flow trajectory of iterative descent, with a hard topK truncation that suppresses noise-amplifying small-eigenvalue directions. Furthermore, we propose SWUDI-A, an adaptive variant that replaces the global rank hyperparameter with perlayer rank rules, further improving robustness across architectures. Both variants share a single symmetric eigendecomposition per linear layer and require no training data or optimizer state. Across four general benchmarks (vision/language) and a multimodal merging benchmark spanning VQA, Geometry, Chart, OCR, Grounding, and modality merging, our proposed spectral solvers match or outperform state-of-the-art merging methods. Crucially, they reduce wall-clock time by 28–72× and peak GPU memory by up to 50%. Code and the extended benchmark are available at https://github.com/WalkerWorldPeace/MLLMerging. Index Terms—Multi-task model merging, data-free, trainingfree, spectral regularization, closed-form solvers.

I. I NTRODUCTION

U

PDATING foundation models is costly: full pre-training or large-scale continued training requires substantial compute and data access. At the same time, domain-specialized, task-specific fine-tuned checkpoints are continually released on open-source platforms such as Hugging Face [1]. Model merging [2]–[4] aims to combine N experts that share the Yongxian Wei, Runxi Cheng, and Chun Yuan are with Shenzhen International Graduate School, Tsinghua University, Shenzhen 518071, China (email: {weiyx23, crx23}@mails.tsinghua.edu.cn, [email protected]). Xingxuan Zhang and Peng Cui are with the Department of Computer Science and Technology, Tsinghua University, Beijing 100084, China (email: [email protected], [email protected]). Li Shen is with Sun Yat-sen University, Shenzhen 510275, China (email: [email protected]). Dacheng Tao is with Nanyang Technological University, Singapore 639798 (email: [email protected]).

same backbone into a single multi-task model without any training data, dramatically reducing the storage, serving, and decentralized-development costs of large foundation models. State-of-the-art methods WUDI [5] and OptMerge [6] formulate model merging as a layer-wise quadratic interference minimization problem, yielding consistently strong performance across diverse tasks and models. These iterative methods share a common computational core: they minimize the interference proxy through hundreds of iterations of a gradient-based optimizer with carefully tuned learning rates, momenta, and initializations (the precise choice, Adam for full fine-tuning versus SGD for LoRA, is detailed in Sec. III). Empirical evidence reveals two key observations. First, the proxy admits an exact closed-form minimum (a normal-equation pseudoinverse), but plugging this closed form into the merged model yields markedly worse downstream accuracy than running the iterative solver to early stopping (e.g., a 2.3-point drop on CLIP-ViT-B/32). Second, the iterative solver dominates the cost of the entire merging pipeline: on a 3B-parameter LLM, 300 Adam steps take 85 minutes and 42 GB of GPU memory. Why iterative descent outperforms the exact minimizer of the same objective, especially in a setting without training data, has remained unexplained. We revisit iterative descent and show that it performs implicit spectral regularization for an ill-posed normal equation. The per-layer interference loss L(τ ) =

N X i=1

1 ∥τi ∥2F

2

(τ − τi )τi⊤ F

cf † has a unique minimum-norm Pclosed form τ =PDC , where ⊤ 2 Ai = τi τi /∥τi ∥F , C = i Ai , and D = i τi Ai . The eigenstructure C = QΛQ⊤ exhibits a long tail of small λk that correspond to directions weakly supported by any task vector. In those directions, the proxy reduces to yk = λk τk◦ +ξk , where ξk collects the proxy-induced noise (i.e., the discrepancy between the task-vector proxy τi⊤ τi and the unobserved activation covariance it stands in for), and τ cf divides by λk , amplifying ξk . Iterative descent, by contrast, acts as an early-stopping spectral filter that down-weights small-λk directions. This mechanism explains why 300-step optimization can outperform the exact closed-form solution. Building on this finding, we formalize model merging as a noisy linear inverse problem τ C = D, and propose a unified spectral filtering estimator parameterized by a perdirection filter hk ∈ [0, 1] applied to every eigendirection of C. The estimator subsumes the closed-form pseudoinverse,

2

Iterative WUDI / OptMerge

Average accuracy (%)

CLIP-ViT-B/32 86.0

SWUDI (closed-form)

InternVL2.5-1B 57.00

69× faster

56.50 85.0

Llama-3.2-3B 44.5

28× faster

63.0

56.75

30× faster

85.5

SWUDI-A (closed-form)

Qwen2-VL-7B

44.0

62.5

72× faster

56.25

43.5

62.0 56.00

84.5

55.75 1s

10s

1m

Wall-clock time

43.0

61.5 10s

1m

10m

Wall-clock time

10m

1h

Wall-clock time

10h

1m

10m

1h

Wall-clock time

Figure 1. Accuracy–cost Pareto frontier across representative settings. Each panel plots average accuracy against wall-clock time for the iterative WUDI/OptMerge baseline (□) and our proposed closed-form solver SWUDI (•) together with its adaptive variant SWUDI-A (▲). The closed-form solvers move merging toward the upper-left, achieving higher or comparable accuracy under a smaller merging budget, so the iterative baseline is Pareto-dominated.

gradient flow, and rank truncation as filter choices, and any on integrated multimodal QA, far outperforming individual instantiation requires only a single symmetric eigendecom- experts. Crucially, these accuracy gains are delivered with 28– position per linear layer. We materialize this framework in 72× wall-clock speedup and up to 50% peak GPU memory two stages, illustrated alongside prior merging families in reduction relative to the iterative baselines. Fig. 2. (i) SWUDI (Spectrally Regularized WUDI): a tunable Our contributions are summarized as follows: spectral variant of the unified estimator that couples a soft • Theory. We show that merging via the layer-wise quadratic exponential filter st (λk ) = 1 − e−tλk , which exactly matches interference loss is a noisy linear inverse problem whose the gradient-flow trajectory of iterative descent, with a hard closed-form pseudoinverse amplifies proxy noise on smalltop-K truncation mk = 1[k ≤ K], K = ⌈r di ⌉. The soft λk directions, and that finite-step iterative descent acts as factor inherits the regularization that early-stopped descent an implicit spectral regularizer (Propositions 1 and 2). already provides; the hard mask removes noise-amplifying • Methodology. We propose a spectral filtering estimator: tail directions before the soft factor can assign them nonSWUDI, a closed-form spectral variant that pairs an negligible residual weight. (ii) SWUDI-A (Adaptive SWUDI): exponential gradient-flow filter with a hard rank truncation; we further upgrade SWUDI into an adaptive, parameter-free and SWUDI-A, the adaptive form of SWUDI that form by replacing the global rank ratio r with per-layer rank removes its hyperparameter via per-layer rank rules driven rules driven by the eigenspectrum by the eigenspectrum (Sec. IV-B). Meanwhile, our solvers  P √ 2 itself.  heavy-tailed P For spectra, we use Kℓ = ( k λk ) / k λk , an effectivereduce wall-clock time by 28–72× across every setting rank estimator [7] that returns exactly the oracle active rank and reduce peak GPU memory by up to 50%. when the spectrum is flat-and-truncated. For spiked-noise • Benchmark. We introduce the first model merging benchp √ spectra, we use Kℓ = {k : λk > ω(β) medianj λj } , mark that provides a fine-grained categorization of MLLM an asymptotically optimal singular-value threshold [8], [9] capabilities and evaluates how merging integrates multiple that retains only directions above the asymptotic random-noise modalities. We train expert models for each task and floor under a spiked-noise model. These two variants are not publicly release their weights and code. This benchmark separate algorithms but successive refinements of the same is designed to help the model merging community better spectral framework: SWUDI establishes the filter shape, while evaluate the generalizability of their methods. SWUDI-A derives its only remaining hyperparameter directly from the spectrum. Both are data- and training-free, require II. R ELATED W ORK neither Adam states nor learning-rate schedules, and reduce the per-layer cost from hundreds of matrix multiplications to A. Data-Free Model Merging Data-free merging produces a single multi-task model from a single eigendecomposition. We evaluate on four general benchmarks and a comprehen- N fine-tuned experts that share a common base, without using sive multimodal benchmark (covering VQA, Geometry, Chart, any training or unlabeled test data. Existing methods can be OCR, Grounding, and modality merging). Spanning vision, broadly grouped into four families. language, multimodal, LoRA, and full-parameter settings, our Linear interpolation: Weight Averaging [11] averages all solvers establish a new state-of-the-art. On CLIP-ViT, they expert weights and works surprisingly well in narrow scenarios achieve 85.55%, 89.57%, and 92.51% accuracy across the in which experts share a basin in parameter space. Task B/32, B/16, and L/14 backbones. Applying AdaMerging [10] Arithmetic [3] introduces global task vectors P τi := Θi − Θ0 to our closed-form delta further lifts B/32 accuracy to 86.08%, and combines them additively as τm = s i τi with a global proving spectral merging is complementary to test-time adap- coefficient s. Sparsification-based: TIES-Merging [12] trims, tation. On Flan-T5 GLUE, SWUDI-A reaches +1.15% over signs, and disjointly sums task vectors to suppress conflicting TSV-Merging. Furthermore, merging task-specialized MLLMs components. DARE [13] randomly drops and rescales taskboosts general capabilities: the merged model achieves 70.58% vector entries to mitigate parameter interference. SVD-based:

3

Figure 2. Illustration of different model merging methods. Panels 1–3 apply fixed operations to per-layer task vectors. Panel 4 shows WUDI/OptMerge reaching the merged solution by hundreds of Adam steps on a quadratic proxy loss. Panel 5 (Ours): SWUDI/SWUDI-A replace this loop with a single per-layer eigendecomposition followed by a spectral filter that down-weights noise-amplifying small-eigenvalue directions.

TSV-Merging (TSV-M) [14] measures task-specific singular interference and decorrelates the dominant singular components. Iso-C [15] flattens the singular spectrum so that no task dominates the merged model. Both methods can be interpreted as fixed spectral manipulations of the stacked task-vector geometry. Optimization-based: DOGE [16] frames model merging as a constrained optimization problem and solves it via adaptive projective gradient descent. WUDI [5] proves that, under the linear-subspace approximation, fine-tuning data are not needed: the task vectors themselves serve as a proxy for hidden activations. OptMerge [6] augments WUDI with low-rank denoising of the task-vector matrix and a stable initialization; it tunes the optimizer separately for full and LoRA fine-tuning. B. Test-Time Adaptation and Dynamic Merging

III. R ETHINKING O PTIMIZATION -BASED M ERGING This section first introduces the task-vector merging notation and the objective of WUDI/OptMerge, and then rethinks the same objective as a noisy linear inverse problem. This organization makes the transition from existing iterative merging to the closed-form solvers in Sec. IV explicit.

A. Preliminaries 1) Notation and Per-Layer Operators: Models. Θ0 ∈ Rd denotes the parameters of a shared base model, and Θ1 , . . . , ΘN denote the parameters of N experts obtained by fine-tuning Θ0 on task-specific data. We restrict merging to two-dimensional weight tensors, i.e., linear and projection layers: for layer (ℓ) (ℓ) ℓ, W0 , Wi ∈ Rdo ×di . Non-two-dimensional parameters, including normalization parameters, embeddings, biases, and position indices, are merged by parameter averaging. Task vectors. The global task vector of expert i is τi := Θi − Θ0 ∈ Rd . Restricted to layer ℓ, the corresponding taskvector matrix is

Test-time adaptation methods [10], [17], [18] use unlabeled test data to learn merging coefficients. AdaMerging [10] is representative: it learns per-layer scales from test inputs, whereas our solvers determine which spectral directions should be inverted in a data-free manner. These two axes are comple(ℓ) (ℓ) (ℓ) τi := Πℓ (τi ) = Wi − W0 ∈ Rdo ×di , (1) mentary in principle; we can combine AdaMerging-style scaling on top of our closed-form solutions. Dynamic (MoE-style) merging [19]–[22] loads task-specific modules at inference where Πℓ extracts the ℓ-th weight block. When the layer is fixed, we omit ℓ and write τi for clarity. time, which requires router training and increases storage. Merged delta and initial point. The merged delta of a method is τm , and the merged model is Θm = Θ0 + τm . We use (ℓ) C. Model Merging for Multimodal LLMs τ := τm ∈ Rdo ×di to denote the per-layer 2-D variable optimized by WUDI or returned by our closed-form solvers. VL-merging [23] merges modality-specific encoders before The initial merged delta of an iterative or closed-form solver fine-tuning. VisionFuse [24] concatenates visual features and is denoted τinit to distinguish it from the base model Θ0 . applies task arithmetic on the LLM. UnIVAL [25] interpolates 2) The WUDI Loss: WUDI Merging [5] notes that, based between multimodal-task experts. DAMC [26] composes on a linear subspace assumption, the hidden-activation intervision/audio/video MLLMs through parameter decoupling and ference (τ − τi )x for a linear layer with input activations online activation merging. AdaMMS [27] performs unsuperdi ×ns x ∈ R can be effectively approximated by replacing vised hyperparameter search but only merges two MLLMs ⊤ x with τ . This formulation yields a completely data-free i at a time. UQ-Merge [28] uses uncertainty quantification on per-layer loss: unlabeled inputs to determine the merging order, but it treats every fine-tuning subset as a separate task without capabilitylevel categorization. Our prior conference version [6] introduced the first MLLM merging benchmark with a clean separation of training data and evaluation suites for VQA, Geometry, Chart, OCR, and Grounding, and additionally studied modality merging across vision, audio, and video.



min L τ = τ

N X

1 2 (τ − τi ) τi⊤ F . 2 ∥τ ∥ i F i=1

(2)

To minimize this loss, WUDI applies the Adam optimizer for T = 300 steps.

4

3) OptMerge Improvements: OptMerge [6] extends WUDI with three algorithmic refinements. (i) On full fine-tuned models, the task-vector matrix is centered as τ̃i = τi − τ̄ and ⊤ projected onto its top-k singular components U1:k Σ1:k V1:k , PN 1 where the per-layer mean is τ̄ := N j=1 τj . The projected matrix replaces τi⊤ in Eq. (2), denoising the proxy. (ii) On LoRA fine-tuned models, where τi is rank-deficient and the merged vector tends to take “shortcuts” by inflating its Frobenius norm, the optimizer is replaced by SGD with implicit regularization, and a low-rank truncation is applied directly to τi without centering. (iii) The per-layer variable is initialized to τinit = τ̄ , which stabilizes the training trajectory. OptMerge produces the strongest results on the MLLM merging benchmark [6]. Despite these gains, OptMerge retains the 300-iteration optimization loop. The rethinking below argues that this loop is not necessary: it performs an implicit spectral regularization that we can carry out in closed form. B. Rethinking as a Noisy Linear Inverse Problem The two unexplained facts about iterative WUDI/OptMerge (that the exact closed-form minimum is worse than 300-step iterative descent, and that this descent dominates the merging wall-clock time) are explained by a single change of viewpoint: WUDI is a noisy linear inverse problem, and iterative descent on it acts as an implicit spectral regularizer of an ill-posed normal equation. We show that the per-layer WUDI objective in Eq. (2) is a quadratic in τ with a closed-form minimumnorm solution (Proposition 1), explain why that closed form is suboptimal in the presence of proxy noise, and prove that gradient flow induces an exact exponential spectral filter on the closed-form pseudoinverse (Proposition 2). 1) Closed-Form Normal Equation: For a linear layer with task vectors τi ∈ Rdo ×di , define the symmetric operator

Table I. WUDI iteration-count sweep on CLIP-B/32 TA8, including the closed-form WUDI solution. Solver

Iterative WUDI (T steps)

Closed form 1000

DC †

Avg. Acc. (%) 80.08 83.82 84.63 84.82 84.72 84.52

82.33

100

200

300

Optimizing �� �0

�� ��

500

��

�� − ��

��

�� orthogonal

700

�� − �� ��

�0

��

�� − �� ��

Figure 3. Norm-shortcut failure mode of unregularized iterative optimization. When optimizing Eq. (2), τm tends to take a shortcut by inflating its magnitude to make (τm −τi )τi⊤ approximately orthogonal to each task vector, rather than aligning with the signal subspace.

Sketch. By the definition of Ai , each normalized term in Eq. (2) is a trace quadratic in τ − τi ; summing and collecting terms gives Eq. (4), and differentiating gives ∇τ L = 2(τ C − D), hence the normal equation τ C = D. Consistency follows because any z ∈ Null(C) is annihilated by every task vector: τi z = 0 for all i, which implies Ai z = 0 and Dz = 0. Thus Null(C) ⊆ Null(D), equivalently the rows of D lie in Range(C), or D = DC † C. The solutions are therefore τ = DC † + Z(I − CC † ) with arbitrary Z. The second term lies in the null space of C and is orthogonal to DC † , so the unique minimum-Frobenius-norm solution sets it to zero, giving Eq. (6). The eigenform follows from C † = QΛ† Q⊤ .

2) Why Exact Closed Form Is Suboptimal: The closed-form solution is empirically not optimal for downstream performance. Table I makes this gap concrete on CLIP-ViT-B/32: iterative WUDI improves at early steps, peaks at a finite iteration N N count, and then degrades as the trajectory approaches the exact ⊤ X X τ τi Ai := i 2 ∈ Rdi ×di , C := Ai , D := τi Ai . pseudoinverse. The closed-form WUDI solution DC † reaches ∥τi ∥F only 82.33%, below both the 300-step result (84.63%) and i=1 i=1 (3) the best early-stopped result (84.82%). This unimodal pattern C is symmetric positive semidefinite. We let C = QΛQ⊤ is the empirical signature that early stopping regularizes the be the eigendecomposition with eigenvalues λ1 ≥ λ2 ≥ inverse problem, whereas excessive optimization recovers noise· · · ≥ λdi ≥ 0 sorted in descending order and corresponding amplifying tail directions. eigenvectors qk , the columns of Q. The reason is a noise-amplification mechanism familiar from ◦ ◦ Proposition 1 (Closed-form WUDI normal equation). The inverse problems. Decompose D = τ C + E, where τ is an unobserved ideal merged delta and E denotes proxy-induced WUDI objective in Eq. (2) is the quadratic noise from replacing the true input activations xi with the   L(τ ) = tr τ C τ ⊤ − 2 tr τ D⊤ + const, (4) task-vector proxy τ ⊤ . For each eigendirection qk , let τ ◦ := i k ◦ do with gradient ∇τ L(τ ) = 2(τ C − D). Any stationary point τ qk ∈ R . Projecting onto qk gives therefore satisfies the normal equation yk := Dqk = λk τ ◦ + ξk , ξk := Eqk . (7) k

τ C = D.

(5)

The set of stationary points is non-empty: each row of D lies in Range(C), equivalently D = D C † C. Among all stationary points the unique minimum-Frobenius-norm element is τ cf = D C † = D Q Λ† Q⊤ ,

(6)

where Λ† inverts the strictly positive eigenvalues and sets the remaining entries to zero.

For λk > 0, the closed-form pseudoinverse gives τkcf = τk◦ + ξk /λk .

(8)

Small eigenvalues λk correspond to row-space directions weakly supported by any task vector, exactly where the proxyinduced noise ξk is large in magnitude relative to λk τk◦ . The pseudoinverse amplifies this noise. Fig. 4 empirically supports this view: the leading eigendirections explain nearly all proxy reduction, whereas the full pseudoinverse overfits the proxy and

5

(a) Head directions explain proxy reduction SWUDI-A cut retains ≥ 98%

(b) Proxy minimum is not real optimum ̂ real interference I(τ)

remaining WUDI proxy

100

10−1

10−2

10−3 25--75\% band median (72 layers)

proxy minimum

Task Arithmetic

Closed-form

101

SWUDI/SWUDI-A Iterative lower is better

10−4 0.0

0.2

0.4

0.6

0.8

10−2

1.0

retained rank ratio K/d

10−1

WUDI proxy (τ)

Figure 4. The exact pseudoinverse is insufficient. (a) Most WUDI proxy reduction is achieved by the leading eigendirections: the SWUDI-A cut retains at least 98% of the median proxy reduction, indicating that the discarded spectral tail contributes little to the proxy objective. (b) The closed-form pseudoinverse ˆ ) than the regularized alternatives in the lower-left region. This motivates DC † minimizes the WUDI proxy P(τ ) but yields higher real interference I(τ regularized rather than full pseudoinversion.

Proposition 2 (Gradient flow induces an exponential spectral filter). Consider the gradient flow τ̇ (t) = − 12 ∇τ L(τ (t)) for the loss in Eq. (4), started at τ (0) = τinit . The flow is the linear ODE τ̇ (t) = D − τ (t) C, with closed-form solution   τ (t) = τinit + DC † − τinit CC † Q diag hk (t) Q⊤ , (9)

empirical filter ĥ k, n

(a) Adam behaves as a spectral filter 1.00

Adam step step 10 step 50 step 300

0.75 0.50 0.25 0.00

10−1

10−2

eigenvalue λk filter-fit R (Adam)

(b) Filter fit improves after early steps

2

yields higher real interference. A regularized solver therefore replaces the inversion 1/λk by hk /λk for a filter hk ∈ [0, 1] that vanishes (or shrinks) for small λk . A related parameter-level instability surfaces as unconstrained inflation of ∥τ ∥F : when ill-conditioned directions are inverted, an iterative solver of Eq. (4) can drive ∥τ ∥F upward to make (τ − τi )τi⊤ approximately orthogonal to each τi , rather than recovering the underlying signal. The two phenomena are linked (both stem from poorly damped small-λk directions). Fig. 3 illustrates the geometry: when task vectors lie in a narrow subspace, the unique low-loss direction lies far from the origin, so unregularized descent on Eq. (4) keeps inflating the merged-vector norm. The link to theory is made formal in Proposition 4 (Appendix B): the WUDI proxy bounds the real per-layer interference up to a Frobenius slack term proportional to ∥τ −τi ∥2F , so once ∥τ ∥F inflates, the slack term dominates and the proxy ceases to control the real interference. Suppressing small-λk directions therefore plays a dual role: it removes the noise-amplifying inversion of Eq. (8) and keeps ∥τ ∥F controlled, restoring tightness of the proxy bound. 3) Why Iterative Descent Works? Implicit Spectral Filtering:

R 2 = 0.9

1.00 0.75 0.50

∼ 50 steps

0.25 0.00 100

101

102

optimization step

Figure 5. Iterative merging is implicit spectral filtering. (a) Adam’s empirical update filter is well approximated by an exponential spectral filter at three checkpoints: large-eigenvalue directions are fitted earlier than small-eigenvalue directions. (b) Across layers, the median fit quality exceeds R2 = 0.9 after about 50 steps. This supports replacing the iterative loop with the closed-form filter. The exact SGD/Landweber identity is reported in Appendix C.

The discrete Landweber iteration τn+1 = τn + η(D − τn C) n admits the corresponding filter hLW k (n) = 1 − (1 − ηλk ) , stable for 0 < η < 2/λmax . For the optimizer used by WUDI/OptMerge, the Adam trajectory empirically matches where the spectral filter is an exponential filter 1 − e−teff λk up to a non-trivial R2 ∈ hk (t) = 1 − e−λk t , k = 1, . . . , di . (10) [0.92, 0.97] on per-layer fits (see Fig. 5), with Spearman correlation approaching 0.99 for step counts ≥ 50. We therefore Thus, hk (t) → 0 as λk → 0+ , while hk (t) → 1 as λk t → ∞. treat Adam as an empirically early-stopped spectral regularizer. cf † The closed-form pseudoinverse τ = DC is the t → ∞ 4) Implication for Solver Design: The analysis above gives limit on the column space of C (i.e., on the directions where a direct design principle: instead of minimizing the proxy λk > 0), and the trajectory remains at τinit on the null space. objective for hundreds of optimizer steps, compute a spectrally Proof. Let τ̃ (t) := τ (t)Q and D̃ := DQ. The flow decou- regularized solution of the normal equation τ C = D in closed ples into di vector ODEs dτ̃k /dt = D̃k − λk τ̃k (one per form. Concretely, the full pseudoinverse factor 1/λk in DC † eigendirection, each τ̃k , D̃k ∈ Rdo ). For λk > 0 the unique should be replaced by hk /λk , where hk ∈ [0, 1] attenuates solution is τ̃k (t) = τ̃init,k e−λk t + (D̃k /λk )(1 − e−λk t ). For eigendirections that are weakly supported and therefore prone λk = 0, τ̃k (t) ≡ τ̃init,k . Reassembling and identifying the filter to noise amplification. completes the proof. From this, two complementary filter components naturally

6

arise: (1) Soft filtering (hk = 1 − e−tλk ), where a continuous Algorithm 1 Unified closed-form spectral merging time parameter t matches the gradient-flow stopping time; and Require: expert deltas {τi(ℓ) }N i=1 for every linear layer; optional SWUDI parameters (t, r) (2) Hard truncation (hk = 1[k ≤ K]), where a rank cutoff K removes the poorly conditioned spectral tail and can be tuned Ensure: merged delta τm do ×di 1: for each linear layer ℓ with τi ∈ do P PR globally or adapted per layer. The following section integrates ⊤ 2 2: Ai ← τi τi /∥τi ∥F , C ← i Ai , D ← i τi Ai these components into a unified spectral filtering estimator 3: C = Q diag(λ1 , . . . , λdi )Q⊤ , λ1 ≥ · · · ≥ λdi (Eq. (11)), instantiating this approach as a hybrid solver and ▷P C is the spectral operator; D is the right-hand side. an adaptive per-layer truncation rule. 4: τinit ← i τi 5: KS ← ⌈rd ⌉ √ Connection to existing methods. This filter view places several  iP  P (ℓ) 6: K ← ( λk )2 / k λk A data-free merging methods in a common language. Tikhonov k ( 1[k ≤ KS ](1 − e−tλk ), SWUDI, regularization corresponds to the classical filter hk = λk /(λk + hk = (ℓ) 1[k ≤ KA ], SWUDI-A. α). Iso-C [15] and TSV-Merging [14] can be re-read as fixed ▷ hk regularizes each eigendirection. shrinkage or truncation rules on the spectrum of the stacked 7: Ch† ← Q diag(hk /λk )Q⊤ task-vector matrix, equivalently on the square-root spectrum (ℓ) 8: τm ← τinit + (D − τinit C) Ch† of C. Finite-step gradient descent on the WUDI [5] quadratic 9: end for gives the Landweber filter, while the Adam optimizer used 10: Average non-2-D parameters and return τm . in WUDI/OptMerge empirically behaves as an early-stopped spectral regularizer rather than an exact Landweber iteration. More broadly, this comparison clarifies two levels at which unregularized choice h ≡ 1 recovers the pseudoinverse limit k a data-free merging method can intervene: it can denoise the DC † and serves as the unstable reference case. proxy normal equation itself, thereby modifying the estimated This filter perspective dictates the behavior of hk . The operator/right-hand-side pair (C, D), or it can keep the proxy gradient-flow filter s (λ) = 1−e−tλ captures the early-stopping t equation fixed and regularize the inversion of its ill-conditioned regularization provided by finite-step iterative descent across operator. Our OptMerge mainly belongs to the first category the bulk of the spectrum. However, its small-λ behavior must and additionally stabilizes the iterative optimization trajectory. be evaluated in terms of the quantity that actually enters The next section pursues the second route by replacing the C † : although s (λ) → 0 as λ → 0, the effective inverse t h full pseudoinverse C † with closed-form spectral filters that gain gt (λ) := st (λ)/λ converges to t. Consequently, pure suppress noise-amplifying eigendirections. exponential filtering still allows tail noise to propagate with finite gain across many weakly supported directions. Conversely, IV. M ETHODOLOGY hard truncation 1[k ≤ K] completely eliminates the tail by We now turn the spectral view of Sec. III into closed-form, setting gt to zero for truncated indices, but it fails to reproduce data-free merging algorithms. For each layer, we reuse the the gradient-flow regularization on the retained directions. normal-equation quantities C and D from Eq. (3), with eigen- SWUDI combines the two by multiplying them into a single decomposition C = QΛQ⊤ and eigenvalues λ1 ≥ · · · ≥ λdi . two-factor filter that we plug into Eq. (11): hk = mk · st (λk ),

A. SWUDI: Spectrally Regularized WUDI We first cast all closed-form spectral solvers of τ C = D into a single family parameterized by a per-direction filter hk ∈ [0, 1], and then specialize the family to obtain i SWUDI. For spectral filter coefficients {hk }dk=1 , define Ch† := ⊤ Q diag(hk /λk )Q , with the Moore–Penrose convention that the diagonal entry is set to 0 on directions with λk = 0. The unified spectral filtering estimator is τbh = τinit + (D − τinit C) Ch† .

(11)

The core operation is the per-direction filter hk , which controls how strongly each eigendirection of C is inverted; τinit is an optional base point at which the filter is applied, and the τinit = 0 specialization τbh = D Ch† recovers the direct filtered inverse and remains close in accuracy in our experiments. Eq. (11) makes clear that the choice of hk determines the regularization. Two filter behaviors are essential for SWUDI: the exponential filter hk = 1−e−tλk exactly recovers the gradient-flow solution stopped at time t from Proposition 2, transferring the earlystopping effect of iterative WUDI into closed form; the hard filter hk = 1[k ≤ K] yields a rank-K truncated spectral inverse, removing weakly supported tail directions. By contrast, the

mk = 1[k ≤ K],

st (λk ) = 1 − e−tλk , K = ⌈r di ⌉.

(12)

Therefore, SWUDI improves merging quality not through exact proxy minimization, but by preventing the proxy inverse from overfitting to noise-amplifying tail directions. The retained head directions capture most of the transferable task signal while keeping the merged delta norm ∥τ ∥F controlled (Sec. III-B2). Two hyperparameters control this regularizer: a continuous exponential time t ≥ 0, which corresponds to the gradientflow stopping time of WUDI, and a rank ratio r ∈ (0, 1]. The soft factor st applies early-stopping regularization to the retained directions, while the hard mask mk zeroes out the effective inverse gain gt on the long tail of small-λk directions before they can introduce noise-amplifying weights into Ch† . Ultimately, the merged delta τbh is computed using Eq. (11), with the filter hk defined in Eq. (12). B. SWUDI-A: Adaptive Variant The rank ratio r in SWUDI is global. However, spectra differ significantly across layers (e.g., attention q/k/v/o, MLP, embedding) and architectures (e.g., CLIP-ViT, Flan-T5, Llama, MLLMs). The adaptive variant, SWUDI-A, addresses this by

7

choosing Kℓ per layer using a closed-form rank rule based on (ℓ) the eigenvalues {λk }, thereby eliminating the need for a global rank hyperparameter. Within the unified spectral estimator (Eq. (11)), SWUDI-A acts as a hard-truncation specialization: it sets the soft factor to the identity for retained directions and replaces the global rank ratio with a layer-wise spectral rank rule Kℓ , allowing the spectrum itself to dictate the cutoff. Layer-wise rank selection: We provide two parameter-free layer-wise rank rules, each corresponding to a specific spectral regime and computable from the existing eigendecomposition. Both operate on the singular values of the stacked taskN do ×di vector matrix M := . P[τ1 /∥τ1 ∥F ; . . . ; τN /∥τN ∥F ] ∈ R ⊤ Because M M = A = C, the singular values of M are i i √ simply σk = λk , allowing these rank rules to be evaluated directly from the eigenspectrum of C. (i) Participation-square-root rule. q  & P 2 '  P (ℓ) 2 λ σ k k k  Pk 2 Kℓpsqrt = (13) =  P (ℓ)  . σ λ   k k k k This applies a participation-ratio effective-rank estimator [7] to σk , thus measuring the effective column rank of M rather than the squared-energy rank of C. This prevents undue concentration on the largest eigenvalues in heavy-tailed spectra. We utilize this as the default rule when the spectrum decays smoothly without a distinct noise floor. (ii) Marchenko–Pastur Gavish–Donoho rule.  KℓGavish = k : σk > ωGD (β) σ bmed , (14)

Vision MLLM

Video MLLM

Private Datasets

Personal fine-tuning Audio MLLM

VQA MLLM

Geometry MLLM

Chart MLLM

Grounding MLLM

OCR MLLM

License

Release checkpoints

Multi-Task MLLM

Data-Free

Model merging

Omni MLLM

Model merging

Figure 6. Two settings of the MLLM merging benchmark. Capability merging (left) combines task-specialized experts that share the same MLLM backbone into a single multi-task model covering VQA, Geometry, Chart, OCR, and Grounding. Modality merging (right) composes vision-, audio-, and video-language experts that share an LLM backbone but use modality-specific encoders and connectors. Both settings are data-free, enabling the merged model to retain expert capabilities without requiring joint training data.

V. B ENCHMARKS AND E XPERIMENTAL R ESULTS This section details our experimental setup and results. We evaluate five merging scenarios spanning vision, language, and multimodal foundation models, utilizing both LoRA and full fine-tuning settings. Sec. V-A focuses on our proposed MLLM merging benchmark, while Sec. V-B extends the evaluation to four widely adopted model merging benchmarks. Finally, we provide a comprehensive discussion and analysis.

where β = min(N do , di )/ max(N do , di ), ωGD (β) is the A. MLLM Merging Benchmark Gavish–Donoho ratio [9], and σ bmed = mediank (σk ) robustly We evaluate our approach on our MLLM merging benchmark, estimates the noise scale. Under a Marchenko–Pastur spiked briefly summarizing the setup here while deferring comprehenmodel M = M ◦ + Ξ with low-rank M ◦ and i.i.d. noise Ξ, sive details to Appendix D. Fig. 6 illustrates the benchmark’s this rule recovers the spike rank with high probability [8]. We two settings, capability merging and modality merging. apply it to spectra with a clear noise bulk and isolated spikes. Backbones. We consider two MLLMs that cover both fineIn summary, psqrt provides a smooth participation count tuning regimes: InternVL2.5-1B-Instruct [29] (full fine-tuning) that consistently returns a positive rank and tolerates heavy and Qwen2-VL-7B-Base [30] (LoRA fine-tuning). For modality tails, whereas Gavish-Donoho acts as a strict noise-floor merging, we follow [26] and use Vicuna-7B-v1.5 [31] paired test that may return a rank of 0 if no singular value is significant. with CLIP-ViT-L-336px for vision, BEATs-Iter3+ with a QConsequently, the appropriate rule can be selected based on the Former for audio, and LanguageBind for video. spectrum and fine-tuning regime prior to downstream evaluation. Tasks and data. We consider five capabilities (VQA, Geometry, Fig. 7 visualizes these regimes, with detailed per-architecture Chart, OCR, Grounding), each with at least 100K training statistics provided in Appendix C. Algorithm 1 summarizes samples. The dataset table is detailed in Appendix D. the unified closed-form procedure. Evaluation. We use VLMEvalKit [32] and lmms-eval [33] under matched settings. For capability evaluation, we report results on VizWiz [34], GQA [35], MathVista [36], MATHC. Computational Complexity Vision [37], ChartQA [38], TextVQA [39], OCRVQA [40], For a single linear layer, the dominant cost is the symmetric and RefCOCO/+/g [41]. For integrated multimodal QA, we eigendecomposition of C ∈ Rdi ×di , which is O(d3i ) time and report results on MMMU [42], DocVQA [43], ScienceQA [44], O(d2i ) memory. Forming C and D costs O(N do d2i ) FLOPs. AI2D [45], and InfographicVQA [46]. For modality merging, Iterative WUDI/OptMerge performs T matrix multiplications we report results on MUSIC-AVQA [47] and AVQA [48]. of similar shapes per layer plus Adam first/second moments, so the wall-clock speedup is roughly T /c, where c captures implementation-dependent constants. Empirically, we observe B. General Model Merging Benchmarks 28–72× speedups (Sec. VI, Table IX). Because no Adam state Benchmarks. (i) CLIP-ViT TA8. The standard 8-task vision is needed, peak GPU memory is reduced by approximately the benchmark from FusionBench [49] (SUN397, Cars, RESISC45, size of the optimizer state. EuroSAT, SVHN, GTSRB, MNIST, DTD). We evaluate three

8

Table II. Capability merging results on InternVL2.5 (full fine-tuning) across multiple tasks. Best scores are bolded and second-best scores are underlined. Method

VQA

Geometry

Chart

OCR

Grounding

Avg.

VizWiz GQA MathVista MATH-Vision ChartQA TextVQA OCRVQA RefCOCO RefCOCO+ RefCOCOg InternVL2.5-Instruct 29.15 54.62

45.40

18.09

69.48

72.51

41.08

71.69

65.41

67.40

53.48

Weight Average Task Arithmetic TIES Merging TA w/ DARE TIES w/ DARE TSV Merging Iso-C WUDI Merging OptMerge

29.96 30.67 30.63 30.61 30.65 31.15 28.21 31.02 30.85

54.89 56.34 56.48 56.48 56.11 56.67 55.36 56.96 57.05

42.30 40.70 44.10 40.40 44.30 44.90 42.10 44.80 46.90

17.76 17.43 16.78 15.79 18.09 17.76 18.09 15.31 15.79

71.64 72.88 72.28 73.08 72.72 70.56 70.56 69.19 68.80

74.54 76.26 76.29 76.30 76.19 75.66 69.34 75.95 75.98

41.86 43.39 44.01 43.03 43.33 45.38 46.51 46.12 46.35

52.62 74.90 76.01 74.94 75.10 65.19 72.72 76.06 76.09

45.29 68.15 68.45 68.07 68.48 58.51 66.56 70.14 69.82

52.39 72.75 73.65 73.02 73.55 59.17 68.50 74.48 74.18

48.33 55.35 55.87 55.17 55.85 52.50 53.80 56.00 56.18

SWUDI SWUDI-A

31.11 57.04 31.25 56.85

46.60 46.10

18.42 16.45

69.76 70.44

76.04 76.00

46.06 45.90

76.24 76.20

70.18 69.99

74.12 74.08

56.56 56.33

Mixture Training

29.79 61.33

45.00

17.11

70.32

72.96

60.25

72.06

65.93

67.46

56.22

Table III. Capability merging results on Qwen2-VL (LoRA fine-tuning) across multiple tasks. Best scores are bolded and second-best scores are underlined. Method

VQA

Geometry

Chart

OCR

Grounding

Avg.

VizWiz GQA MathVista MATH-Vision ChartQA TextVQA OCRVQA RefCOCO RefCOCO+ RefCOCOg Qwen2-VL-Base

5.52

5.39

54.00

21.05

0.36

20.22

1.07

45.32

37.55

31.26

22.17

Weight Average Task Arithmetic TIES Merging TA w/ DARE TIES w/ DARE TSV Merging Iso-C WUDI Merging OptMerge

41.47 40.52 41.38 40.64 41.63 41.43 12.31 37.19 41.54

57.33 62.31 59.08 62.38 59.96 57.31 13.44 56.45 61.21

57.90 58.40 52.60 58.10 54.50 54.30 49.70 54.70 58.40

25.66 23.68 19.41 23.68 23.03 23.68 20.07 25.66 25.99

59.56 79.67 67.24 79.76 70.68 59.44 2.80 67.84 74.24

81.09 81.09 81.42 81.04 81.53 81.25 30.05 79.92 81.48

57.85 59.50 58.53 59.34 59.63 57.81 6.12 65.56 60.03

80.72 75.96 80.63 75.83 80.73 80.71 53.68 76.25 80.45

65.37 61.33 65.36 61.41 65.65 65.34 38.96 60.72 65.96

77.68 75.85 77.65 75.80 77.77 77.76 41.90 71.99 76.92

60.46 61.83 60.33 61.80 61.51 59.90 26.90 59.63 62.62

SWUDI SWUDI-A

40.31 60.21 40.63 60.80

57.60 57.00

23.03 23.36

70.96 75.32

81.60 81.63

63.96 64.23

80.12 80.22

65.45 65.62

76.07 78.42

61.93 62.72

Qwen2-VL-Instruct 44.09 62.18

57.20

17.43

70.04

78.38

65.42

82.89

77.87

75.63

63.11

CLIP-pretrained backbones [50]: ViT-B/32, ViT-B/16, and ViT- C. Multimodal Model Merging L/14, reporting the mean per-task accuracy. (ii) CLIP-ViTWe first evaluate the proposed solvers on our multimodal B/32 TALL20. A 20-task extension used to evaluate scalability merging benchmark, covering full-parameter InternVL2.5, as the number of merged tasks increases. (iii) Flan-T5-base LoRA-fine-tuned Qwen2-VL, integrated multimodal QA, and on GLUE. Eight GLUE tasks [51] fine-tuned with rankmodality merging across vision, audio, and video experts. Ta16 LoRA on Flan-T5-base [52]. This evaluates our method bles II and III show that the closed-form spectral solvers consison rank-deficient task-vector matrices in the NLP domain. tently match or exceed the strongest iterative WUDI/OptMerge (iv) Llama-3.2-3B. Following MergeBench [53], five domain baselines on capability merging. On InternVL2.5-1B, SWUDI experts (math, code, instruction following, safety, multilingual) obtains the best average accuracy (56.56), while SWUDI-A are merged into a single Llama-3.2-3B model [54]. Evaluation remains close behind (56.33) without method-specific rank uses lm_eval across GSM8K [55], HumanEval/MBPP+ [56], tuning. On Qwen2-VL-7B, SWUDI-A reaches the best average [57], IFEval [58], TruthfulQA [59], MMLU [60], ARC [61], (62.72), slightly above iterative OptMerge (62.62), which and HellaSwag [62]. This setting employs full-parameter deltas demonstrates that adaptive spectral truncation is especially with 3B trainable parameters per expert. useful when LoRA deltas are intrinsically low-rank. Methods. We compare our proposed SWUDI together with The two backbones expose complementary benefits. In the its adaptive variant SWUDI-A, against several model merging full-parameter InternVL2.5 setting, spectral filtering preserves methods: Weight Average [11], Task Arithmetic [3], TIES [12], shared multimodal capabilities while improving the average DARE-TA and DARE-TIES [13], TSV-Merging [14], Iso- over both optimization-based and spectrum-based baselines. In C [15], WUDI Merging [5], and our OptMerge [6]. the Qwen2-VL LoRA setting, the adaptive rank rule prevents Hyperparameters. Following common practice, we search the norm-inflation failure mode of iterative data-free objectives over the global scaling coefficient s ∈ {0.1, 0.2, 0.3, 0.5, 1.0} and retains the compact directions that carry most of the LoRA for all merging methods. For SWUDI, we additionally signal. The multimodal results therefore support the central tune r ∈ {0.55, 0.60, 0.65, 0.70, 0.75, 0.85} and t ∈ claim from two regimes: explicit spectral regularization is not {300, 500, 700, 1000, 1300, 1800}. In contrast, SWUDI-A re- only faster than iterative optimization, but also more stable quires no continuous hyperparameters beyond the global s. when the task-vector geometry is low-rank or noisy. Specifically, SWUDI-A applies the participation-square-root The capability-level wins extend to comparisons against rule in Eq. (13) on CLIP-ViT, Flan-T5, and the MLLMerging the corresponding mixture-trained model, the natural databenchmark, while employing the Gavish–Donoho rule in rich baseline. SWUDI on InternVL2.5-1B slightly exceeds Eq. (14) on the Llama-3.2-3B MergeBench (a full-parameter mixture training on average (56.56 vs. 56.22), and SWUDI-A LLM with spiked-noise spectra). on Qwen2-VL-7B is within 0.4 points of mixture training

9

Table IV. Modality merging results on zero-shot image-audio-video question answering tasks by merging vision-language, audio-language, and video-language models. The “Individual Modalities” columns show baseline performance for each single-modality model. Individual Modalities Datasets

Merging Methods

MUSIC-AVQA 50.77 27.93 AVQA 75.55 47.57 Avg. 63.16 37.75

49.02 79.20 64.11

47.75 69.39 58.57

52.14 78.62 65.38

50.35 75.84 63.10

Table V. Evaluation on general multimodal QA benchmarks. Method

MMMU

DocVQA

SciQA

AI2D

InfoVQA

Avg.

Individual VQA Individual Chart Individual Geometry Individual Grounding Individual OCR

26.00 30.33 33.67 34.22 38.00

62.93 57.13 64.29 65.64 77.67

50.83 40.01 73.25 76.54 63.66

44.59 29.86 62.27 63.24 54.39

39.07 26.02 29.79 33.82 41.97

44.68 36.67 52.65 54.69 55.14

OptMerge SWUDI-A

39.33 39.33

Online Composing

Weight Task TIES TSV WUDI Iso-C OptMerge SWUDI-A NaiveMC DAMC Vision Audio Video Average Arithmetic Merging Merging Merging

84.18 84.14

91.89 93.41

79.44 79.47

56.84 56.57

70.34 70.58

(62.72 vs. 63.11). Reaching this accuracy regime without any joint training data, using only the experts’ parameter deltas and a single eigendecomposition per layer, is the practical case that capability merging makes for production multimodal systems. Next, we examine modality merging, where vision-, audio-, and video-language Vicuna-7B experts are integrated into an Omni-language model [26] (in Table IV). SWUDI-A outperforms both OptMerge and all offline merging baselines. It even surpasses online composition methods that require modalityspecific, inference-time composition. These results demonstrate that the spectral regularization principle generalizes effectively from capability merging to cross-modal composition, a setting where preserving complementary modality information is more critical than optimizing for any single expert. Finally, we assess whether the merged model preserves composite abilities rather than only isolated capabilities. Table V evaluates the InternVL2.5-1B merge on integrated multimodal QA benchmarks. SWUDI-A attains the highest average (70.58), improving over the OptMerge result (70.34) and giving the largest gain on ScienceQA. This pattern is consistent with the noise-amplification analysis in Sec. III-B2: integrated tasks are sensitive to spurious low-eigenvalue updates, so explicitly filtering those directions improves robustness beyond the percapability averages. Across capability, integrated-QA, and modality-merging evaluations, the closed-form solvers match or exceed iterative WUDI/OptMerge in nearly all average multimodal metrics, while reducing the merging cost by over an order of magnitude. The results demonstrate that merged multimodal experts can surpass individual or mixture-trained models when their complementary skills are optimally combined, further indicating that these benefits arise from explicit spectral regularization rather than a costly iterative optimizer.

53.78 80.90 67.34

52.77 77.51 65.14

52.43 76.86 64.65

53.17 80.82 67.00

53.91 81.26 67.59

53.50 80.26 66.88

52.80 80.78 66.79

Table VI. Cross-backbone average accuracy (%) on CLIP-ViT TA8 (B/32, B/16, L/14) and the 20-task extension TALL20 (B/32). Detailed per-task results for all three TA8 backbones are provided in Appendix C. The best and second-best results in each column are bolded and underlined, respectively. The last two rows further apply AdaMerging [10] to the closed-form merged delta using unlabeled test data. TA8

Method

TALL20

B/32

B/16

L/14

B/32

Weight Average Task Arithmetic TIES Merging TA w/ DARE TIES w/ DARE TSV Merging Iso-C τ cf = DC † WUDI Merging OptMerge

66.32 67.55 71.90 67.46 60.96 83.07 80.39 82.33 84.63 84.53

72.33 77.14 77.60 77.15 74.30 87.10 85.07 88.04 89.17 89.49

79.87 80.47 83.83 80.49 74.33 90.57 90.65 91.69 92.16 92.38

61.10 60.62 62.76 60.55 62.22 73.22 70.35 72.54 61.06 61.71

SWUDI SWUDI-A

85.55 85.53

89.57 89.49

92.51 92.52

75.60 75.62

SWUDI → AdaMerging SWUDI-A → AdaMerging

86.08 86.05

89.78 89.81

92.72 92.75

78.12 78.03

SWUDI and SWUDI-A achieve the best or second-best datafree averages across all backbones, with SWUDI-A matching the tuned SWUDI while eliminating the need for methodspecific rank tuning. To understand this performance, we include the unregularized closed-form pseudoinverse τ cf = DC † as a direct ablation. Its performance gap to SWUDI-A shrinks monotonically as the backbone size increases (−3.20 pt on B/32, −1.45 pt on B/16, −0.83 pt on L/14). This aligns with the noiseamplification analysis in Sec. III-B2: larger backbones yield better-conditioned task-vector spectra, meaning division by small λk causes less degradation. However, on TALL20, this gap widens again to −3.08 pt, reflecting the re-emergence of long-tail noise as the number of tasks increases. The TALL20 setting further demonstrates the robustness of our solvers in a more densely populated task space. Here, SWUDI-A outperforms TSV-Merging and significantly exceeds the iterative WUDI and OptMerge. Furthermore, the performance drop from TA8 to TALL20 is substantially smaller for the spectral solvers than for iterative methods, confirming that suppressing noise-amplifying tail directions becomes increasingly critical as more task vectors interact. Finally, when unlabeled test data are available, applying AdaMerging [10] D. Vision Model Merging on top of our spectral anchors yields further improvements: We next evaluate whether the spectral solvers generalize from +0.2 to +0.5 pt on TA8, and a substantial +2.4 to +2.5 pt on multimodal models to vision models. Table VI summarizes TALL20. This demonstrates that closed-form spectral merging the results for CLIP-ViT on the TA8 benchmark across three provides a robust data-free initialization that remains highly backbones, as well as the larger TALL20 setting. On TA8, complementary to test-time adaptation.

10

Table VII. Multi-task performance when merging Flan-T5-base (LoRA fine-tuned) models on all eight tasks. The metric is accuracy except for STSB (Spearman ρ). Method

CoLA

MNLI

MRPC QNLI

QQP

RTE

SST2

STSB

Avg.

Weight Average Task Arithmetic TIES Merging TA w/ DARE TIES w/ DARE TSV Merging Iso-C WUDI Merging OptMerge

69.70 68.84 68.17 68.94 31.16 69.32 69.13 68.65 68.36

59.66 55.18 48.96 55.10 0.43 77.09 57.35 72.18 70.29

78.92 78.68 78.92 78.92 79.90 80.39 76.72 78.43 80.39

90.08 89.79 89.31 89.73 84.44 90.04 88.63 84.64 89.58

83.79 83.67 83.43 83.71 82.23 83.62 82.66 82.70 83.28

80.51 79.06 79.78 79.06 76.90 79.06 80.14 71.48 79.06

91.17 91.51 91.51 91.51 89.56 92.55 91.28 93.00 93.00

72.00 72.38 74.22 72.58 75.94 82.55 63.32 83.82 84.27

78.23 77.39 76.79 77.44 65.07 81.83 76.15 79.36 81.03

SWUDI SWUDI-A

69.22 68.94

82.00 80.00

77.94 83.33

89.80 89.77

83.46 80.87 83.26 80.87

93.00 92.55

85.33 85.10

82.70 82.98

Table VIII. Llama-3.2-3B MergeBench with five experts. We report per-task and average accuracy (%) on eight tasks. Multilingual subtasks follow the fr-only legacy protocol. Best results are shown in bold, and the second-best results are underlined. Method

GSM8K

HE+

MBPP+ IFEval

TQA

MMLUfr

ARCfr

HSwagfr

Avg.

Weight Average Task Arithmetic TIES Merging TA w/ DARE TIES w/ DARE TSV Merging Iso-C WUDI Merging OptMerge

42.76 44.73 42.61 46.70 52.99 55.72 48.22 52.54 53.53

31.71 33.54 30.49 33.54 35.37 36.59 35.37 37.20 34.76

59.52 59.79 57.14 59.26 57.41 56.88 55.56 57.14 58.99

9.24 14.42 7.58 18.67 25.51 20.15 8.13 17.93 25.51

46.02 47.38 44.91 47.80 47.42 46.28 44.83 46.56 45.23

46.35 46.34 48.25 48.12 47.55 48.17 47.49 44.86 46.99

35.76 36.27 37.13 36.44 37.04 37.13 37.04 37.21 37.30

44.53 44.88 44.75 45.56 45.19 45.01 44.29 44.17 44.66

39.49 40.92 39.11 42.01 43.56 43.24 40.12 42.20 43.37

SWUDI SWUDI-A

58.45 58.15

37.80 36.59

57.67 56.35

24.58 22.18

45.02 45.42

47.01 46.85

37.55 37.30

44.69 44.83

44.10 43.46

Table IX. Wall-clock time, peak GPU memory, and speedup for the merging

E. Language Model Merging step. Speedup is measured against the iterative data-free baseline in each We evaluate language-model merging in two contrasting setting (WUDI for CLIP/Llama and OptMerge for MLLMs). Mixture-training rows report the cost of jointly fine-tuning a single multi-task model. regimes: LoRA fine-tuning on Flan-T5 GLUE [52] and fullMethod Time Peak Mem. Speedup parameter fine-tuning on Llama-3.2-3B MergeBench [53]. Setting Table VII shows that the Flan-T5 LoRA setting tightens the 86.3 s 4.82 GB 1.0× WUDI Merging CLIP-B/32 SWUDI / SWUDI-A 2.8 s 3.64 GB 30.8× gap between iterative and closed-form data-free objectives: WUDI and OptMerge reach 79.36% and 81.03%, respectively, WUDI Merging 5126.1 s 42.03 GB 1.0× sitting close to TSV Merging (81.83%) but still trailing Llama-3.2-3B SWUDI / SWUDI-A 70.7 s 39.78 GB 72.5× the proposed solvers, with SWUDI-A achieving the best OptMerge 13608.9 s 21.97 GB 1.0× average (82.98%) and SWUDI the second-best (82.70%). Qwen2-VL-7B SWUDI / SWUDI-A 487.6 s 10.81 GB 27.9× Mixture training 24.56 h 256 GB — The benefit comes from matching the solver to the low-rank structure of LoRA deltas: adaptive truncation preserves the OptMerge 552.3 s 7.03 GB 1.0× 8.0 s 4.16 GB 69.0× informative subspace and avoids the interference that the InternVL2.5-1B SWUDI / SWUDI-A Mixture training 25.38 h 240 GB — iterative quadratic loss leaves under rank-deficient C, where many small eigendirections couple weakly to the proxy gradient and slow convergence. The Llama-3.2-3B benchmark complements this LoRA case A. Efficiency and Accuracy–Cost Trade-off with full-parameter experts spanning math, code, instruction folTable IX reports the raw wall-clock time and peak GPU lowing, safety, and multilingual tasks. As reported in Table VIII, memory of the merging step on representative settings. The SWUDI obtains the best average among the merging methods, closed-form solvers consistently reduce both quantities because and SWUDI-A (with the Gavish–Donoho rank rule appropriate they eliminate optimizer state and per-iteration workspaces and for the spiked-noise spectra of full-parameter LLM deltas) replace hundreds of matrix multiplications with one symmetric remains the strongest tuning-free alternative. The contrast with eigendecomposition per layer. The magnitude of the memory Flan-T5 illustrates why a single rank rule is not sufficient: saving depends on the regime: it is modest when resident model LoRA deltas favor a heavy-tailed low-rank prior, whereas parameters dominate the footprint, as in full-parameter Llama full-parameter LLM deltas are better described by a spiked- merging, but substantial when optimizer states dominate, as in noise spectrum. In both regimes, the same benefit emerges: Qwen2-VL LoRA merging. spectral filtering converts an unstable proxy inversion into a These efficiency gains are most compelling when viewed controlled merge that improves accuracy while avoiding the alongside accuracy. Combining Table IX with the benchmark cost of iterative optimization. results demonstrates that the proposed solvers shift the merging Evaluating 5 to 20 experts across multimodal, vision, process toward the upper left of the accuracy-cost plane. They and language foundation models, we find that SWUDI and achieve comparable or superior accuracy under a substantially SWUDI-A define the high-accuracy end of the closed-form reduced computational budget. Fig. 1 illustrates this trade-off Pareto frontier. Consequently, both solvers transform the across four representative settings. Here, SWUDI-A acts as the implicit regularization, previously obtained through hundreds ideal low-cost, tuning-free solution, while SWUDI provides a of optimizer steps, into an explicit closed-form computation. high-accuracy alternative if hyperparameter tuning is permitted. This yields the accuracy benefits of spectral filtering with Both approaches successfully replace the iterative optimizer substantially lower wall-clock time and memory costs. with explicit spectral regularization. The mixture-training reference rows provide further context VI. E FFICIENCY AND D IAGNOSTIC A NALYSIS for these computational savings. Joint multi-task fine-tuning This section distills the empirical analysis into two high-level of Qwen2-VL-7B and InternVL2.5-1B requires approximately messages. First, replacing iterative optimization with closed- 24 to 25 hours and 240 to 256 GB of aggregate GPU memory. form spectral filtering substantially reduces wall-clock time This represents roughly 180× the wall-clock time and 20× the and memory. Second, spectral diagnostics explain why adaptive peak memory required by SWUDI-A on the same backbones, truncation is needed across architectures. even under the strict assumption that all task-specific training

amplified noise ν̂ k2/λk2

11

VII. C ONCLUSION

(a) Small eigenvalues amplify noise

10−2

smaller λ ⇒ larger amplified noise 10−4 α=2

10−6

10−6

10−4

10−2

100

eigenvalue λk

(b) Different spectra need different cuts 100 KGavish

λk/λ1

10−2

Kpsqrt

CLIP early MLP

10−4 CLIP late MLP

10−6 0.0

0.2

0.4

0.6

0.8

1.0

normalized index k/d

Figure 7. Noise amplification motivates adaptive rank truncation. (a) Small-eigenvalue directions are associated with larger amplified noise ν̂k2 /λ2k under the closed-form pseudoinverse (binned median + 25–75% band, pooled over CLIP-ViT-B/32 TA8 layers). (b) Different layer spectra lead to different SWUDI-A rank-rule cuts: the psqrt rule retains more directions in heavytailed spectra, whereas the Gavish–Donoho rule is more conservative for concentrated spectra. This motivates layer-wise adaptive rank selection.

data are centrally co-located. Therefore, closed-form spectral merging not only reduces the cost of iterative OptMerge by an order of magnitude but also offers a data-free alternative to the natural baseline (joint training) at a mere fraction of its computational and data-governance costs. B. Spectral Diagnostics The spectral diagnostics in Fig. 7 explain why explicit spectral regularization is needed. Small-eigenvalue directions are most vulnerable to pseudoinverse instability: dividing by a small λk amplifies proxy noise, so the spectral tail should not be inverted without regularization. This supports the hard truncation used by SWUDI and SWUDI-A. The same diagnostics also show why the cutoff should be adaptive. Spectra vary substantially across layers and architectures: vision and LoRA merges often exhibit a headand-tail structure, where a participation-style rule [7] preserves the useful subspace, whereas full-parameter LLM merges more often resemble a bulk-plus-spike regime, favoring a more conservative Gavish–Donoho cutoff [8], [9]. The retained ranks reflect the underlying fine-tuning geometry: for example, SWUDI-A-psqrt keeps a mean rank ratio of 0.149 on Qwen2-VL-7B LoRA, versus 0.615 on fully fine-tuned InternVL2.5-1B. These diagnostics are explanatory rather than tuning criteria; they show why layer-wise rank adaptation is preferable to a single global cutoff. Full per-architecture statistics are reported in Table XVII (Appendix C). Together, these results show that WUDI/OptMerge works primarily through implicit spectral regularization. Our closedform solvers make this regularization explicit, preserving accuracy across diverse settings while avoiding hundreds of optimizer steps.

We revisit data-free model merging as a noisy linear inverse problem. While WUDI and OptMerge optimize a quadratic objective over hundreds of steps, this objective has a closed-form pseudoinverse whose small-eigenvalue directions amplify proxy noise. Iterative descent succeeds largely because it implicitly filters these unstable directions. This insight motivates a spectral-filtering estimator and closed-form solvers: SWUDI, which combines an exponential filter with hard rank truncation, and SWUDI-A, which uses layer-wise spectral rules to eliminate rank hyperparameters. Both require only one eigendecomposition per layer, with no training data or optimizer state. We also introduce a multimodal benchmark for capability and modality merging. Across vision, language, LoRA, fullparameter LLM, and MLLM settings, our solvers match or exceed iterative baselines while running 28× to 72× faster and using up to 50% less peak memory. These findings suggest that effective merging need not rely on long optimization trajectories: once the inverse problem is formulated, the central design choice is the spectral filter to impose. R EFERENCES [1] T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz et al., “Huggingface’s transformers: State-of-the-art natural language processing,” arXiv preprint arXiv:1910.03771, 2019. 1 [2] P. Yadav, T. Vu, J. Lai, A. Chronopoulou, M. Faruqui, M. Bansal, and T. Munkhdalai, “What matters for model merging at scale?” arXiv preprint arXiv:2410.03617, 2024. 1 [3] G. Ilharco, M. T. Ribeiro, M. Wortsman, L. Schmidt, H. Hajishirzi, and A. Farhadi, “Editing models with task arithmetic,” in ICLR, 2023. 1, 2, 8, 20 [4] E. Yang, L. Shen, G. Guo, X. Wang, X. Cao, J. Zhang, and D. Tao, “Model merging in llms, mllms, and beyond: Methods, theories, applications and opportunities,” arXiv preprint arXiv:2408.07666, 2024. 1 [5] R. Cheng, F. Xiong, Y. Wei, W. Zhu, and C. Yuan, “Whoever started the interference should end it: Guiding data-free model merging via task vectors,” in ICML, 2025. 1, 3, 6, 8 [6] Y. Wei, R. Cheng, W. Jin, E. Yang, L. Shen, L. Hou, S. Du, C. Yuan, X. Cao, and D. Tao, “Optmerge: Unifying multimodal LLM capabilities and modalities via model merging,” in ICLR, 2026. 1, 3, 4, 8, 19, 27 [7] O. Roy and M. Vetterli, “The effective rank: A measure of effective dimensionality,” in EUSIPCO, 2007. 2, 7, 11 [8] V. A. Marčenko and L. A. Pastur, “Distribution of eigenvalues for some sets of random matrices,” Mathematics of the USSR-Sbornik, vol. 1, no. 4, pp. 457–483, 1967. 2, 7, 11 [9] M. Gavish and √ D. L. Donoho, “The optimal hard threshold for singular values is 4/ 3,” IEEE Transactions on Information Theory, vol. 60, no. 8, pp. 5040–5053, 2014. 2, 7, 11, 19 [10] E. Yang, Z. Wang, L. Shen, S. Liu, G. Guo, X. Wang, and D. Tao, “Adamerging: Adaptive model merging for multi-task learning,” in ICLR, 2024. 2, 3, 9 [11] M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith et al., “Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time,” in ICML, 2022. 2, 8 [12] P. Yadav, D. Tam, L. Choshen, C. A. Raffel, and M. Bansal, “TIESmerging: Resolving interference when merging models,” NeurIPS, 2023. 2, 8 [13] L. Yu, B. Yu, H. Yu, F. Huang, and Y. Li, “Language models are super mario: Absorbing abilities from homologous models as a free lunch,” in ICML, 2024. 2, 8, 21 [14] A. A. Gargiulo, D. Crisostomi, M. S. Bucarelli, S. Scardapane, F. Silvestri, and E. Rodolà, “Task singular vectors: Reducing task interference in model merging,” in CVPR, 2025. 3, 6, 8

12

[15] D. Marczak, S. Magistri, S. Cygert, B. Twardowski, A. D. Bagdanov, and J. van de Weijer, “No task left behind: Isotropic model merging with common and task-specific subspaces,” in ICML, 2025. 3, 6, 8 [16] Y. Wei, A. Tang, L. Shen, C. Yuan, and X. Cao, “Modeling multi-task model merging as adaptive projective gradient descent,” in ICML, 2025. 3 [17] E. Yang, L. Shen, Z. Wang, G. Guo, X. Chen, X. Wang, and D. Tao, “Representation surgery for multi-task model merging,” in ICML, 2024. 3 [18] N. Daheim, T. Möllenhoff, E. Ponti, I. Gurevych, and M. E. Khan, “Model merging by uncertainty-based gradient matching,” in ICLR, 2024. 3 [19] A. Tang, L. Shen, Y. Luo, N. Yin, L. Zhang, and D. Tao, “Merging multi-task models via weight-ensembling mixture of experts,” in ICML, 2024. 3 [20] C. Huang, P. Ye, T. Chen, T. He, X. Yue, and W. Ouyang, “EMRMerging: Tuning-free high-performance model merging,” in NeurIPS, 2024. 3 [21] Z. Lu, C. Fan, W. Wei, X. Qu, D. Chen, and Y. Cheng, “TwinMerging: Dynamic integration of modular expertise in model merging,” in NeurIPS, 2024. 3 [22] L. Shen, A. Tang, E. Yang, G. Guo, Y. Luo, L. Zhang, X. Cao, B. Du, and D. Tao, “Efficient and effective weight-ensembling mixture of experts for multi-task model merging,” IEEE TPAMI, 2025. 3 [23] Y.-L. Sung, L. Li, K. Lin, Z. Gan, M. Bansal, and L. Wang, “An empirical study of multimodal model merging,” in EMNLP, 2023. 3 [24] Z. Chen, J. Hu, Z. Deng, Y. Wang, B. Zhuang, and M. Tan, “Enhancing perception capabilities of multimodal llms with training-free fusion,” arXiv preprint arXiv:2412.01289, 2024. 3 [25] M. Shukor, C. Dancette, A. Rame, and M. Cord, “UnIVAL: Unified model for image, video, audio and language tasks,” TMLR, 2023. 3 [26] C. Chen, Y. Du, Z. Fang, Z. Wang, F. Luo, P. Li, M. Yan, J. Zhang, F. Huang, M. Sun et al., “Model composition for multimodal large language models,” in ACL, 2024. 3, 7, 9, 27 [27] Y. Du, X. Wang, C. Chen, J. Ye, Y. Wang, P. Li, M. Yan, J. Zhang, F. Huang, Z. Sui et al., “AdaMMS: Model merging for heterogeneous multimodal large language models with unsupervised coefficient optimization,” in CVPR, 2025. 3 [28] H. Qu, X. Zhao, J. Peng, K. Lee, B. Dariush, and T. Chen, “UQ-Merge: Uncertainty guided multimodal large language model merging,” in ACL, 2025. 3 [29] Z. Chen, W. Wang, Y. Cao, Y. Liu, Z. Gao, E. Cui, J. Zhu, S. Ye, H. Tian, Z. Liu et al., “Expanding performance boundaries of opensource multimodal models with model, data, and test-time scaling,” arXiv preprint arXiv:2412.05271, 2024. 7, 27 [30] P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge et al., “Qwen2-VL: Enhancing vision-language model’s perception of the world at any resolution,” arXiv preprint arXiv:2409.12191, 2024. 7, 27 [31] L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing et al., “Judging llm-as-a-judge with mt-bench and chatbot arena,” in NeurIPS, 2023. 7, 27 [32] H. Duan, J. Yang, Y. Qiao, X. Fang, L. Chen, Y. Liu, X. Dong, Y. Zang, P. Zhang, J. Wang et al., “VLMEvalKit: An open-source toolkit for evaluating large multi-modality models,” in MM, 2024. 7, 27 [33] K. Zhang, B. Li, P. Zhang, F. Pu, J. A. Cahyono, K. Hu, S. Liu, Y. Zhang, J. Yang, C. Li et al., “Lmms-eval: Reality check on the evaluation of large multimodal models,” arXiv preprint arXiv:2407.12772, 2024. 7, 27 [34] D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham, “VizWiz grand challenge: Answering visual questions from blind people,” in CVPR, 2018. 7, 27 [35] D. A. Hudson and C. D. Manning, “GQA: A new dataset for real-world visual reasoning and compositional question answering,” in CVPR, 2019. 7, 27 [36] P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao, “MathVista: Evaluating mathematical reasoning of foundation models in visual contexts,” in ICLR, 2024. 7, 27 [37] K. Wang, J. Pan, W. Shi, Z. Lu, H. Ren, A. Zhou, M. Zhan, and H. Li, “Measuring multimodal mathematical reasoning with math-vision dataset,” in NeurIPS, 2024. 7, 27 [38] A. Masry, D. X. Long, J. Q. Tan, S. Joty, and E. Hoque, “ChartQA: A benchmark for question answering about charts with visual and logical reasoning,” arXiv preprint arXiv:2203.10244, 2022. 7, 27

[39] A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach, “Towards vqa models that can read,” in CVPR, 2019. 7, 27 [40] A. Mishra, S. Shekhar, A. K. Singh, and A. Chakraborty, “OCRVQA: Visual question answering by reading text in images,” in ICDAR, 2019. 7, 27 [41] S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg, “Referitgame: Referring to objects in photographs of natural scenes,” in EMNLP, 2014. 7, 27 [42] X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun et al., “MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,” in CVPR, 2024. 7, 27 [43] M. Mathew, D. Karatzas, and C. Jawahar, “Docvqa: A dataset for vqa on document images,” in WACV, 2021. 7, 27 [44] P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan, “Learn to explain: Multimodal reasoning via thought chains for science question answering,” in NeurIPS, 2022. 7, 27 [45] A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi, “A diagram is worth a dozen images,” in ECCV, 2016. 7, 27 [46] M. Mathew, V. Bagal, R. Tito, D. Karatzas, E. Valveny, and C. V. Jawahar, “InfographicVQA,” in WACV, 2022. 7, 27 [47] G. Li, Y. Wei, Y. Tian, C. Xu, J.-R. Wen, and D. Hu, “Learning to answer questions in dynamic audio-visual scenarios,” in CVPR, 2022. 7, 27 [48] P. Yang, X. Wang, X. Duan, H. Chen, R. Hou, C. Jin, and W. Zhu, “AVQA: A dataset for audio-visual question answering on videos,” in MM, 2022. 7, 27 [49] A. Tang, L. Shen, Y. Luo, H. Hu, B. Du, and D. Tao, “Fusionbench: A comprehensive benchmark of deep model fusion,” arXiv preprint arXiv:2406.03280, 2024. 7, 22 [50] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable visual models from natural language supervision,” in ICML, 2021. 8, 27 [51] A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman, “GLUE: A multi-task benchmark and analysis platform for natural language understanding,” in ICLR, 2019. 8 [52] H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma et al., “Scaling instruction-finetuned language models,” JMLR, 2024. 8, 10, 21 [53] Y. He, S. Zeng, Y. Hu, R. Yang, T. Zhang, and H. Zhao, “Mergebench: A benchmark for merging domain-specialized llms,” in NeurIPS, 2025. 8, 10 [54] A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan et al., “The Llama 3 herd of models,” arXiv preprint arXiv:2407.21783, 2024. 8 [55] K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano et al., “Training verifiers to solve math word problems,” arXiv preprint arXiv:2110.14168, 2021. 8 [56] M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman et al., “Evaluating large language models trained on code,” arXiv preprint arXiv:2107.03374, 2021. 8 [57] J. Liu, C. S. Xia, Y. Wang, and L. Zhang, “Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation,” in NeurIPS, 2023. 8 [58] J. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y. Luan, D. Zhou, and L. Hou, “Instruction-following evaluation for large language models,” arXiv preprint arXiv:2311.07911, 2023. 8 [59] S. Lin, J. Hilton, and O. Evans, “TruthfulQA: Measuring how models mimic human falsehoods,” in ACL, 2022. 8 [60] D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,” in ICLR, 2021. 8 [61] P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord, “Think you have solved question answering? try ARC, the AI2 reasoning challenge,” arXiv preprint arXiv:1803.05457, 2018. 8 [62] R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi, “HellaSwag: Can a machine really finish your sentence?” in ACL, 2019. 8 [63] H. W. Engl, M. Hanke, and A. Neubauer, Regularization of Inverse Problems, ser. Mathematics and Its Applications. Dordrecht, The Netherlands: Kluwer Academic Publishers, 1996. 17

13

[64] G. Ortiz-Jimenez, A. Favero, and P. Frossard, “Task arithmetic in the tangent space: Improved editing of pre-trained models,” in NeurIPS, 2023. 20 [65] R. M. Gower, N. Loizou, X. Qian, A. Sailanbayev, E. Shulgin, and P. Richtárik, “SGD: General analysis and improved rates,” in ICML, 2019. 20 [66] A. Khaled and P. Richtárik, “Better theory for SGD in the nonconvex world,” TMLR, 2023. 20 [67] L. Li, T. Zhang, Z. Bu, S. Wang, H. He, J. Fu, Y. Wu, J. Bian, Y. Chen, and Y. Bengio, “MAP: Low-compute model merging with amortized pareto fronts via quadratic approximation,” in ICLR, 2025. 21 [68] G. Merlin, V. Nanda, R. Rawal, and M. Toneva, “What happens during finetuning of vision transformers: An invariance based investigation,” in CoLLAs, 2023. 21 [69] C. Wu, T. Wang, Y. Ge, Z. Lu, R. Zhou, Y. Shan, and P. Luo, “π-tuning: Transferring multimodal foundation models with optimal multi-task interpolation,” in ICML, 2023. 21, 24 [70] Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu et al., “MMBench: Is your multi-modal model an allaround player?” in ECCV, 2024. 26 [71] B. Li, Y. Ge, Y. Ge, G. Wang, R. Wang, R. Zhang, and Y. Shan, “Seedbench: Benchmarking multimodal large language models,” in CVPR, 2024. 26 [72] C. Fu, P. Chen, Y. Shen, Y. Qin, M. Zhang, X. Lin, J. Yang, X. Zheng, K. Li, X. Sun et al., “MME: A comprehensive evaluation benchmark for multimodal large language models,” in NeurIPS, 2025. 26 [73] L. Chen, J. Li, X. Dong, P. Zhang, Y. Zang, Z. Chen, H. Duan, J. Wang, Y. Qiao, D. Lin, and F. Zhao, “Are we on the right way for evaluating large vision-language models?” in NeurIPS, 2024. 26 [74] Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the v in vqa matter: Elevating the role of image understanding in visual question answering,” in CVPR, 2017. 27 [75] K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi, “OK-VQA: A visual question answering benchmark requiring external knowledge,” in CVPR, 2019. 27 [76] H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” in CVPR, 2024. 27 [77] W. Wang, Q. Lv, W. Yu, W. Hong, J. Qi, Y. Wang, J. Ji, Z. Yang, L. Zhao, S. XiXuan et al., “CogVLM: Visual expert for pretrained language models,” in NeurIPS, 2024. 27 [78] J. Cao and J. Xiao, “An augmented benchmark dataset for geometric question answering through dual parallel text encoding,” in COLING, 2022. 27 [79] J. Gao, R. Pi, J. Zhang, J. Ye, W. Zhong, Y. Wang, L. Hong, J. Han, H. Xu, Z. Li et al., “G-LLaVA: Solving geometric problem with multimodal large language model,” arXiv preprint arXiv:2312.11370, 2023. 27 [80] K. Kafle, B. Price, S. Cohen, and C. Kanan, “DVQA: Understanding data visualizations via question answering,” in CVPR, 2018. 27 [81] O. Sidorov, R. Hu, M. Rohrbach, and A. Singh, “TextCaps: a dataset for image captioning with reading comprehension,” in ECCV, 2020. 27 [82] G. Kim, T. Hong, M. Yim, J. Nam, J. Park, J. Yim, W. Hwang, S. Yun, D. Han, and S. Park, “Ocr-free document understanding transformer,” in ECCV, 2022. 27 [83] Y. Zhang, R. Zhang, J. Gu, Y. Zhou, N. Lipka, D. Yang, and T. Sun, “LLaVAR: Enhanced visual instruction tuning for text-rich image understanding,” arXiv preprint arXiv:2306.17107, 2023. 27 [84] A. F. Biten, R. Tito, A. Mafla, L. Gomez, M. Rusinol, E. Valveny, C. Jawahar, and D. Karatzas, “Scene text visual question answering,” in ICCV, 2019. 27 [85] S. Svetlichnaya, “DeepForm: Understand structured documents at scale,” 2020. 27 [86] T. Stanisławek, F. Graliński, A. Wróblewska, D. Lipiński, A. Kaliska, P. Rosalska, B. Topolski, and P. Biecek, “Kleister: key information extraction datasets involving long documents with complex layouts,” in ICDAR, 2021. 27 [87] W. Chen, H. Wang, J. Chen, Y. Zhang, H. Wang, S. Li, X. Zhou, and W. Y. Wang, “TabFact: A large-scale dataset for table-based fact verification,” in ICLR, 2020. 27 [88] L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in ECCV, 2016. 27 [89] J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy, “Generation and comprehension of unambiguous object descriptions,” in CVPR, 2016. 27 [90] R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma et al., “Visual genome: Connecting

language and vision using crowdsourced dense image annotations,” IJCV, 2017. 27 [91] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in ICLR, 2022. 27 [92] H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in NeurIPS, 2023. 27 [93] S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, “BEATs: Audio pre-training with acoustic tokenizers,” in ICML, 2023. 27 [94] J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping languageimage pre-training with frozen image encoders and large language models,” in ICML, 2023. 27 [95] X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y. Zou, and W. Wang, “WavCaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,” TASLP, 2024. 27 [96] Y. Gong, H. Luo, A. H. Liu, L. Karlinsky, and J. R. Glass, “Listen, think, and understand,” in ICLR, 2024. 27 [97] A. Panagopoulou, L. Xue, N. Yu, J. Li, D. Li, S. Joty, R. Xu, S. Savarese, C. Xiong, and J. C. Niebles, “X-instructblip: A framework for aligning x-modal instruction-aware representations to llms and emergent crossmodal reasoning,” arXiv preprint arXiv:2311.18799, 2023. 27 [98] B. Zhu, B. Lin, M. Ning, Y. Yan, J. Cui, H. Wang, Y. Pang, W. Jiang, J. Zhang, Z. Li et al., “LanguageBind: Extending video-language pretraining to n-modality by language-based semantic alignment,” arXiv preprint arXiv:2310.01852, 2023. 27 [99] R. Luo, Z. Zhao, M. Yang, J. Dong, D. Li, P. Lu, T. Wang, L. Hu, M. Qiu, and Z. Wei, “Valley: Video assistant with large language model enhanced ability,” arXiv preprint arXiv:2306.07207, 2023. 27 [100] M. Maaz, H. Rasheed, S. Khan, and F. Khan, “Video-ChatGPT: Towards detailed video understanding via large vision and language models,” in ACL, 2024. 27 [101] B. Lin, Y. Ye, B. Zhu, J. Cui, M. Ning, P. Jin, and L. Yuan, “VideoLLaVA: Learning united visual representation by alignment before projection,” in EMNLP, 2024. 27

14

Supplementary Material of Closed-Form Spectral Regularization for Multi-Task Model Merging C ONTENTS Appendix A: Notations A-A Models and Task Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . A-B Spectral and Filter Quantities . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . A-C Rank Rules and Hyperparameters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . A-D Diagnostics and Proxies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

14 14 14 15 15

Appendix B: Theoretical Proofs B-A From the WUDI Loss to the Normal Equation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . B-B Task-Vector Proxy for the Input Subspace . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . B-C Noise Amplification and Spectral Risk . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . B-D Spectral Filters and the Unified Estimator . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . B-E Adaptive Rank Rules . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . B-E1 Participation-Square-Root Rule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . B-E2 Marchenko–Pastur Gavish–Donoho Rule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . B-F Parameter-Drift Bound and Empirical Companion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . B-F1 Notation and Setting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . B-F2 Assumptions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . B-F3 Supporting Lemmas . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . B-F4 Main Theorems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . B-F5 Empirical Fine-Tuning Step Sweep . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

15 16 16 17 17 18 18 19 19 19 20 20 21 22

Appendix C: Additional Analyses C-A Per-Task Vision Results on CLIP-ViT . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . C-B Spectral Diagnostics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . C-B1 Per-Architecture Spectral Statistics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . C-B2 Task-Vector Proxy and Optimizer-Filter Diagnostics . . . . . . . . . . . . . . . . . . . . . . C-C OptMerge Analysis and Rank-Truncation Evidence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . C-C1 Component-Wise Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . C-C2 Evidence for Head-Spectrum Truncation . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

23 23 24 24 25 25 25 25

Appendix D: Our MLLMerging Benchmark D-A Motivation and Benchmark Scope . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . D-B Capability-Merging Tasks and Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . D-C Backbones and Expert Construction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . D-D Modality-Merging Track . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . D-E Evaluation Protocol . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . D-F Answer Extraction Prompt . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

26 26 26 27 27 27 27

A PPENDIX A N OTATIONS This appendix lists the symbols that are used repeatedly across the theory, method, and diagnostic sections. Unless explicitly written, the layer index ℓ is suppressed for per-layer matrices; indices i and k denote experts and eigendirections, respectively. Hats indicate constructed estimators, and the superscript cf denotes a closed-form solution. A. Models and Task Vectors This group fixes the model-level and per-layer objects used throughout the paper. Table X distinguishes full-model parameters from 2-D layer blocks and records the main task-vector quantities used by the closed-form estimator. B. Spectral and Filter Quantities The spectral notation in Table XI is used to express the WUDI objective as a normal equation. The matrix C defines the eigendirections to be inverted, D is the right-hand side, and hk controls how strongly each direction is retained.

15

Table X. Notation: models and task vectors. Symbol

Meaning

Θ0 , Θ i , Θ m (ℓ) (ℓ) W0 , Wi ∈ Rdo ×di (ℓ) τi or τi τm and τ τinit τ cf = DC † τ◦ τbh

base, fine-tuned expert, and merged model parameters ℓ-th 2-D weight block of the base/expert model (ℓ) (ℓ) per-layer task vector Wi − W0 (ℓ) full-model merged delta and its per-layer variable, τ := τm P optional initialization in the spectral filtering estimator; default i τi , or 0 for the direct filtered inverse minimum-norm closed-form solution of the WUDI normal equation unobserved ideal merged delta used only in the inverse-problem analysis spectral filtering estimator in Eq. (11) Table XI. Notation: spectral and filter quantities.

Symbol

Meaning

Ai = τi⊤ τi /∥τi ∥2F P C = i Ai

single-task input-side covariance proxy aggregated input-side covariance proxy right-hand side of the normal equation τ C = D stacked normalized task vectors, satisfying M ⊤ M = C eigendecomposition of C k-th eigenvalue and eigenvector of C k-th singular value of M Moore–Penrose pseudoinverse of C spectral filter coefficient applied to direction qk soft exponential filter (gradient-flow stopping time t) hard top-K truncation mask SWUDI two-factor spectral filter, see Eq. (12) effective inverse gain on direction qk filtered pseudoinverse used in Eq. (11)

P D = i τi A i M = [τ1 /∥τ1 ∥F ; . . . ; τN /∥τN ∥F ] C = QΛQ⊤ λk , q k √ σ k = λk † C hk st (λk ) = 1 − e−tλk mk = 1[k ≤ K] hk = mk st (λk ) gt (λk ) = st (λk )/λk Ch† = Q diag(hk /λk ) Q⊤

Table XII. Notation: rank rules and hyperparameters. Symbol

Meaning

K Kℓ r ∈ (0, 1] t s Kℓpsqrt KℓGavish

retained-rank cutoff per-layer retained rank in SWUDI-A SWUDI global rank ratio, K = ⌈rdi ⌉ SWUDI exponential time parameter global scaling coefficient applied to the merged delta  P √  P participation-square-root rank rule, ( k λk )2 / k λk Gavish–Donoho hard-threshold rank rule for spiked-noise spectra Table XIII. Notation: diagnostics, proxies, and per-direction noise model.

Symbol

Meaning

P(τ ) ˆ ) I(τ E = D − τ ◦C ξk = Eqk

P WUDI proxy, i ∥(τ − τi )τi⊤ ∥2F /∥τi ∥2F calibration-based estimate of real layer-wise interference proxy-noise matrix in the inverse-problem view proxy noise projected onto eigendirection qk

C. Rank Rules and Hyperparameters Table XII collects the quantities that control truncation and final scaling. We keep only the symbols needed to describe SWUDI and SWUDI-A. Other quantities are defined where they are used. D. Diagnostics and Proxies The final notation block, summarized in Table XIII, supports the diagnostic figures and the noise-amplification analysis. It separates the computable WUDI proxy from the unobserved signal/noise quantities used to explain why small-eigenvalue directions should be regularized. A PPENDIX B T HEORETICAL P ROOFS This appendix expands the derivations underlying Sec. III and Sec. IV. We first derive the WUDI normal equation and justify the task-vector proxy for the input subspace; we then analyze the resulting inverse problem, spectral filters, adaptive rank rules,

16

and parameter-drift bound. Throughout, we work per linear layer with task vectors τi ∈ Rdo ×di , and we use the matrix calculus identity ∂X tr(XAX ⊤ B) = BXA + B ⊤ XA⊤ . A. From the WUDI Loss to the Normal Equation Recall that L(τ ) = Expanding term i:

N X

1 ⊤ 2 . 2 (τ − τi )τi F ∥τ ∥ i F i=1

 2 (τ − τi )τi⊤ F = tr (τ − τi ) τi⊤ τi (τ − τi )⊤ .

Let Ai := τi⊤ τi /∥τi ∥2F . Then L(τ ) =

X

 tr (τ − τi )Ai (τ − τi )⊤ .

i

Expanding the quadratic and summing,   L(τ ) = tr τ C τ ⊤ − 2 tr τ D⊤ + const, P P where C = i Ai , D = i τi Ai . Differentiating, ∇τ L = 2(τ C − D). Stationary points satisfy τ C = D. We first verify P solvability: for any z ∈ Null(C), z ⊤ Cz = i ∥τi z∥22 /∥τi ∥2F = 0 forces τi z = 0 for all i, hence Ai z = 0 and Dz = 0. Therefore Null(C) ⊆ Null(D), equivalently each row of D lies in Range(C). The full set of stationary points is parameterized by τ = DC † + Z(I − CC † ) for an arbitrary free matrix Z ∈ Rdo ×di , which spans the null-space component. The minimumFrobenius-norm choice is Z = 0, yielding the closed form τ cf = D C † , computed via the eigendecomposition C = QΛQ⊤ . B. Task-Vector Proxy for the Input Subspace a) Row-space justification: The WUDI loss in Eq. (2) substitutes the transpose of the task vector τi for the input subspace xi . We provide a self-contained justification. For a single linear layer with weight matrix Wl , the per-sample loss gradient has the standard outer-product form ∇Wl Lt,n = gt,n x⊤ t,n , where xt,n is the layer’s input activation at training step t on sample n and gt,n is the back-propagated output-side gradient. Summing the resulting GD updates across iterations gives τi,l = −η

Bs T X X

gt,n x⊤ t,n ,

t=1 n=1

where Bs is the per-step batch size. Each row of τi,l is therefore a g-weighted superposition of input vectors visited during fine-tuning. Equivalently, the row space of τi,l is contained in (or, in finite-trajectory practice, biased toward) the activation subspace span{xt,n }. This containment is the formal content of the WUDI substitution: the proxy operator τi⊤ τi acts on the P same subspace as the unobserved input Gram t,n xt,n x⊤ t,n , up to the gradient-induced reweighting that we collect into the residual term E of Sec. III-B2. In particular, τ C = D in Sec. III-B1 can be read as a noisy linear inverse problem with respect to that activation subspace, which is the foundation on which the spectral filters in Sec. IV are built. b) Covariance-dominance bound: The discussion above establishes a row-space containment between τi,l and the input activation subspace. We now upgrade it to a quantitative inequality that justifies treating the WUDI proxy as a computable upper bound on the real per-layer interference   2 Ii (τ ) := Ex∼Di,l (τ − τi,l ) x 2 = tr (τ − τi,l ) Σi,l (τ − τi,l )⊤ , di ×ni where Σi,l := E[xi,l x⊤ with ni i,l ] is the input second moment. Equivalently, given an empirical activation matrix Xi,l ∈ R 1 1 2 ⊤ samples, Ii (τ ) = ni ∥(τ − τi,l )Xi,l ∥F and Σi,l = ni Xi,l Xi,l .

Assumption 3 (Activation covariance dominated by task-vector Gram). For each task i and each linear layer l, the empirical input second moment Σi,l := E[xi,l x⊤ i,l ] admits the decomposition Σi,l ⪯ ai,l

⊤ τi,l τi,l + Ri,l , ∥τi,l ∥2F

(15)

with constants ai,l > 0 and a residual operator Ri,l ⪰ 0 whose action on any layer-wise delta δi,l := τ − τi,l is bounded by  ⊤ tr δi,l Ri,l δi,l ≤ bi,l ∥δi,l ∥2F , (16) for some bi,l ≥ 0.

17

The decomposition in Eqs. (15)–(16) has two complementary readings. (i) The first term states that the directions on which ⊤ Σi,l has appreciable mass are precisely the directions on which τi,l τi,l has appreciable mass, scaled by the layer-specific constant ai,l . This is the formal version of “task vectors approximate the input subspace they were trained on.” (ii) The residual Ri,l collects every input direction that the gradient trajectory failed to cover (e.g., directions visited only at very early or very late iterations); bi,l controls how much Ri,l can leak into the interference computation. Proposition 4 (Computable upper bound on real interference). Under Assumption 3, the real per-layer interference satisfies  ⊤ 2 Ii (τ ) ≤ ai,l (τ − τi,l ) τi,l ∥τi,l ∥2F + bi,l ∥τ − τi,l ∥2F . F Proof. Write δ := τ − τi,l . By definition, Ii (τ ) = tr(δ Σi,l δ ⊤ ). Substituting Eq. (15) and using Eq. (16), Ii (τ ) ≤ ⊤ ⊤ 2 ⊤ 2 ⊤ ai,l tr(δ τi,l τi,l δ ⊤ )/∥τi,l ∥2F + bi,l ∥δ∥2F . The first term equals ai,l ∥δ τi,l ∥F /∥τi,l ∥2F since ∥δ τi,l ∥F = tr(δ τi,l τi,l δ ⊤ ). The first summand is exactly the WUDI proxy contribution from task i (Eq. (2)); the second summand is a Frobenius-norm regularization on the merged delta. Minimizing the WUDI proxy therefore controls the real interference up to a ∥δ∥F slack term. Two consequences follow. First, the spectral filters in Sec. IV that suppress small-λk directions of C are precisely those that make the proxy a tight bound: discarding low-eigenvalue directions reduces the proxy without inflating ∥δ∥F . Second, the Frobenius-norm-inflation regime documented in Fig. 3 is exactly the failure mode in which iterative WUDI drives the proxy down by inflating ∥δ∥F , leaving the second summand large; closed-form spectral solvers avoid this regime by construction. The empirical validity of Assumption 3 is supported by the capture-gap diagnostics in Fig. 14(a): the task-vector subspaces capture input energy in early and middle layers (gap +0.18–0.43 vs. random subspaces). The last MLP layer is a documented exception (gap ≈ 0); for that layer bi,l is comparable to ai,l and the proxy is correspondingly looser, in line with the observation that the input-subspace assumption is layer-conditional rather than global. C. Noise Amplification and Spectral Risk a) Inverse-problem view: Decompose D = τ ◦ C + E, where τ ◦ is an unobserved ideal merged delta (the model that would minimize the true downstream loss) and E collects the proxy mismatch (replacing xi by τi⊤ , plus the linear-subspace approximation error). Project on qk , the k-th eigenvector of C: yk := Dqk = λk τk◦ + ξk ,

ξk := Eqk .

The closed-form solution gives, for λk > 0, τkcf = yk /λk = τk◦ + ξk /λk . As λk → 0+ , the noise term ξk /λk dominates. Equivalently, defining the residual signal Rk◦ := (τ ◦ − τinit )qk and the residual right-hand side B := D − τinit C, so that bk = hk (Bqk /λk ) Bqk = λk Rk◦ + ξk , any spectral filter hk with hk → 0 as λk → 0 produces a regularized residual estimator R b and merged update τbqk = τinit qk + Rk , in which the noise contribution scales as hk /λk , controllable by the filter shape. This shrinks the residual rather than the absolute estimate, so discarded directions retain the initial point τinit , matching the implementation of SWUDI-A. b) Empirical noise-amplification fit: The inverse-problem view above treats the residual ξk = Eqk as an arbitrary noise vector. Empirically, its squared norm is well described by a power law in λk : ν̂k2 ≈ σ02 + σ12 λα k,

(17)

with three parameters (σ02 , σ12 , α) ≥ 0 fitted per layer in log–log space. Across the 72 linear layers of CLIP-ViT-B/32 TA8, we obtain α with mean 1.871 and median 1.855; 68% of layers fit exponents below 2, and only a small minority fit exponents above. The empirical α ≈ 2 regime means ν̂k2 /λ2k behaves as σ02 /λ2k + σ12 in the small-λk tail, so the closed-form pseudoinverse risk diverges only through the offset σ02 , while the bulk of the spectrum (where λk is large) sees a vanishing noise contribution. This is consistent with the binned-median plot in Fig. 7(a) and motivates suppressing small-λk directions with a spectral filter. D. Spectral Filters and the Unified Estimator a) Gradient flow and Landweber filters: The Frobenius gradient flow is τ̇ (t) = D − τ (t)C,

τ (0) = τinit .

Multiplying both sides on the right by Q and writing τ̃ (t) = τ (t)Q, D̃ = DQ, we obtain di decoupled vector ODEs τ̃˙k (t) = D̃k − λk τ̃k (t) (one per eigendirection, each τ̃k , D̃k ∈ Rdo ), each solved by τ̃k (t) = τ̃init,k +

 D̃k − λk τ̃init,k 1 − e−λk t . λk

Reassembling and identifying the spectral filter,   τ (t) = τinit + DC † − τinit CC † Q diag 1 − e−λk t Q⊤ . n The discrete Landweber iteration τn+1 = τn + η(D − τn C) has filter hLW k (n) = 1 − (1 − ηλk ) , stable for 0 < η < 2/λmax . cf As n → ∞, both filters converge to 1 on λk > 0, recovering τ . We refer to [63] for the classical theory of spectral filters as regularizers for ill-posed linear inverse problems.

18

(a) Filter shapes (layers/0/self_attn, d = 768)

(b) Predicted Bayes risk (diagnostic) lowest-risk filter ≠ best merged accuracy

(c) SNR around the cutoff no sharp SNR gap at the cut (median ρK ≈ 6.7)

0.6

0.4

closed-form (h = 1) 0.2

empirical Wiener

2148

103

337

291

SWUDI-A (K = 241) 0

50

100

150

index k (head, capped at 256)

200

ρ = 1 (theory cut) median ρK = 6.7

SWUDI (K = 500 > 256; no cut in view) 0.0

ρKA + 1

2215

2167

2123

boundary SNR ρ

total predicted Bayes risk (log)

0.8

filter coefficient hk

ρKA

101

1.0

250

rm

-fo

ed

s Clo

I UD IW

t=

30

0

I UD SW

r=

0.6

5

I-A UD SW

λ

en Wi

er

100

all

-

op Dr

0

10

20

30

40

50

60

70

layer index

Figure 8. Spectral-filter diagnostics on CLIP-ViT-B/32 TA8. This figure is diagnostic rather than prescriptive: it shows that simple Wiener/Bayes-risk or SNR-gap criteria do not by themselves pick the best merging filter, which is why SWUDI-A instead relies on the conservative rank rules of Appendix B-E. Panel (a) plots the filter value hk (vertical axis) against the eigen-direction index k sorted by decreasing eigenvalue (horizontal axis), comparing the full pseudoinverse (hk = 1), the empirical Wiener filter, the tuned SWUDI hard cutoff, and the adaptive SWUDI-A cutoff; the tuned SWUDI cut (K = 500) lies beyond the displayed head range, so within view it coincides with the pseudoinverse. Panel (b) compares the total predicted residual (Bayes) risk of these filters, where lower bars indicate smaller risk under the residual-noise model introduced in Sec. III-B2 and detailed in Appendix B-C0b; the Wiener and pseudoinverse filters attain the lowest predicted risk, yet they are not the best on real merged accuracy, so this risk model is a negative diagnostic and not a selection criterion. Panel (c) plots the boundary signal-to-noise ratio ρk at the eigen-direction k (the ratio of retained signal energy to residual-noise energy at that boundary) evaluated at and just after the SWUDI-A cutoff KA across layers, where KA is the per-layer SWUDI-A retained rank. The median boundary SNR stays well above 1 with no sharp drop across the cut, confirming that SWUDI-A is a conservative spectral-rank rule rather than a sharp SNR threshold: it removes the spectral tail while retaining the directions that dominate the proxy reduction.

b) Derivation of the unified spectral filtering estimator: The estimator τbh in Eq. (11) follows directly from the linear ODE solution. Let R(t) := τ (t) − τinit be the deviation from the initial point and recall B = D − τinit C from Appendix B-C. The corresponding ODE Ṙ(t) = B − R(t) C with R(0) = 0 has solution R(t) = B Q diag (1 − e−λk t )/λk λ >0 Q⊤ . Replacing k 1 − e−λk t by an arbitrary filter hk in the eigenbasis yields R = B Ch† , which is exactly Eq. (11) after adding τinit . Setting τinit = 0 gives the direct filtered inverse τbh = D Ch† . Fig. 8 compares the resulting filters—the full pseudoinverse, the Wiener filter, the tuned SWUDI filter, and the adaptive SWUDI-A cutoff—together with their predicted residual risk and the boundary signal-to-noise ratio at the cutoff. E. Adaptive Rank Rules The goal of this section is to explain how SWUDI-A selects a retained rank Kℓ for each layer without using a global rank ratio. We use two complementary rules: a participation-ratio rule for heavy-tailed spectra, and a Gavish–Donoho threshold for spectra with a clearer signal-plus-noise structure. √ 1) Participation-Square-Root Rule: For singular values σk = λk , the participation ratio P ( σk )2 Rpart := Pk 2 k σk counts the effective number of comparable singular components. We set  P √ 2 ( k λk ) P . Kℓpsqrt = ⌈Rpart ⌉ = k λk If the active singular spectrum is flat and supported on exactly K directions, then Rpart = K, so the rule recovers the active rank exactly. More generally, suppose the active singular values P are σk = µ(1 + δk ) for k ≤ K, with empirical mean perturbation close to zero and empirical second moment c2 := K −1 k≤K δk2 . A first-order expansion gives X X σk ≈ Kµ, σk2 ≈ Kµ2 (1 + c2 ), k≤K

k≤K

and therefore (σ)

Rpart ≈

K . 1 + c2

Thus the rule contracts the ideal active rank only according to the relative spread of the active singular values, not their absolute (λ) scale. By contrast, computing the same participation ratio on eigenvalues λk = σk2 gives Rpart ≈ K(1 − 4c2 ) for small c, which is more sensitive to spectral spread and tends to retain too few directions in practice.

19

(a) Rank rules by layer type Kpsqrt ≈ tuned SWUDI; mlp.fc2 most aggressive SWUDI tuned r = 0.65

Kpsqrt

KGavish

(c) Task-vector capture gap task-vec capture − random capture

(b) Optimizer–filter fit SGD matches the filter; Adam grows filter-like 1.0

layer L0

0.8

0.8

dtd

L4

+0.22

+0.00

+0.28 0.4

L7 L11

eurosat filter-fit R 2

mean K/d

0.6

0.4

+0.16

+0.05

0.3

+0.26

0.4

0.2 0.2

mnist

+0.37

+0.02

+0.37

sun397

+0.24

-0.02

+0.23

0.1

0.0 0.2

−0.2

0.0

optimiser Adam (solid)

−0.4

capture gap

0.6

SGD (dotted, ≈1.0)

0.0

attn q

attn k

attn v

attn out

mlp fc1

mlp fc2

0

100

101

optimization step

102

early

(L0)

mid

(L6)

last

(L11

)

Figure 9. Rank-rule and optimizer-trajectory diagnostics on CLIP-ViT-B/32 TA8. Panel (a) plots the mean retained-rank ratio K/di (vertical axis) for different layer types (horizontal axis) under the participation-square-root rule Kpsqrt , the eigenvalue-participation rule Kλ , and the Gavish–Donoho rule KGavish ; the useful rank varies across layers (with mlp.fc2 the most aggressively truncated), so it should be selected adaptively rather than fixed globally. Panel (b) plots how well iterative optimizer updates match a spectral filter: the horizontal axis is the optimization step (symmetric-log, so step 0 is the initialization) and the vertical axis is the filter-fit R2 , with higher values indicating a closer spectral-filter interpretation; SGD (dotted) matches the filter throughout, while Adam (solid) starts unstable in the first few steps and becomes increasingly filter-like with training. Panel (c) reports the capture gap—task-vector capture minus random capture—by task (rows) and layer depth (columns); cell color and the printed value give the gap (warm red for a positive gap, cool blue for a near-zero or negative gap), so every cell carries a value and none are missing. Together, the diagnostics support layer-wise rank adaptation and the view that iterative merging behaves like spectral filtering.

2) Marchenko–Pastur Gavish–Donoho Rule: Stack the normalized task-vector matrices vertically as M ∈ RN do ×di . Since p ⊤ M M = C (Sec. IV-B), the singular values of M are σk = λk (C). When the spectrum resembles a low-rank signal plus random noise, the Marchenko–Pastur law and the Gavish–Donoho threshold give a conservative hard cutoff: retain singular values above min(N do , di ) ωGD (β) σ bmed , β= , max(N do , di ) √ where σ bmed is the empirical median singular value. In the unknown-noise setting, ωGD (β) = λ∗ (β)/ µβ , where µβ is the median of the Marchenko–Pastur distribution and s 8β p λ∗ (β) = 2(β + 1) + (18) (β + 1) + β 2 + 14β + 1 is the known-noise threshold of [9, Eq. 6]. This gives the per-layer rank in Eq. (14). We use this rule as a conservative alternative when the spectrum has a visible noise bulk rather than a long heavy tail. Fig. 9 supports the need for adaptive rank selection: different layer types prefer different retained ranks, while SGD/Adam trajectory diagnostics connect the rank-rule behavior back to the spectral-filtering view.

F. Parameter-Drift Bound and Empirical Companion This subsection reproduces the parameter-drift theorem of the conference version [6]. The bound motivates the expertconstruction protocol used in our merging experiments, namely controlling parameter drift during fine-tuning so that experts remain near a common basin around the base model. 1) Notation and Setting: a) Tasks and losses: For task i, let the loss Li : Rd → R be evaluated at parameters Θ ∈ Rd . b) Task vectors: After T steps of (deterministic) gradient descent (GD) with fixed step size η > 0 from a common initialization Θ, the task vector for task i is T −1 X (i) τi := −η ∇Li (Θt ). t=0

PN c) Merged update: Let τm := j=1 αj τj with nonnegative weights αj ≥ 0. We study the loss of task i at the merged point Θ + τm . d) Norm and inner product: ∥·∥ denotes the Euclidean norm and ⟨·, ·⟩ the Euclidean inner product. For nonzero vectors u, v, cos(u, v) := ⟨u, v⟩ /(∥u∥ ∥v∥).

20

2) Assumptions: Assumption 5 (L-smoothness). Each Li has L-Lipschitz continuous gradients: for all Θ, Θ′ , ∥∇Li (Θ) − ∇Li (Θ′ )∥ ≤ L ∥Θ − Θ′ ∥ , equivalently, for any ∆,

2

Li (Θ + ∆) ≤ Li (Θ) + ⟨∇Li (Θ), ∆⟩ + L2 ∥∆∥ . Assumption 6 (Polyak–Łojasiewicz (PL) condition). Each Li satisfies, for some µ > 0,  2 ∗ 1 2 ∥∇Li (Θ)∥ ≥ µ Li (Θ) − Li , where L∗i := inf Θ Li (Θ). Assumption 7 (Directional similarity). For each i and some κ ∈ (0, 1],  cos − ∇Li (Θ), τi ≥ κ, equivalently, ⟨∇Li (Θ), τi ⟩ ≤ −κ ∥∇Li (Θ)∥ ∥τi ∥ . This ensures τi is a descent direction for task i, with alignment quantified by κ. Assumption 8 (Approximate orthogonality). For all i ̸= j and some ε ∈ [0, 1), cos(τi , τj ) ≤ ε. Prior works [3], [64] show that task vectors are nearly orthogonal in high-dimensional parameter space, which helps explain the success of model merging. A small ε means that tasks are nearly orthogonal in update space, reducing negative transfer. Assumption 9 (Bounded gradients). There exists G > 0 such that for all i and all Θ on the trajectory, ∥∇Li (Θ)∥ ≤ G. This boundedness condition is widely adopted in the optimization literature [65], [66]. 3) Supporting Lemmas: Lemma 10 (Cross-task cosine leakage). Under Assumptions 7–8, with ∇Li (Θ) ̸= 0 and τj ̸= 0, for i ̸= j, p p cos(∇Li (Θ), τj ) ≤ δ, δ := κε + 1 − κ2 1 − ε2 . Proof sketch. Normalize u = −∇Li / ∥∇Li ∥, vi = τi / ∥τi ∥, vj = τj / ∥τj ∥. Assumption 7 gives ⟨u, vi ⟩ ≥ κ and Assumption 8 gives ⟨vi , vj ⟩ ≤ ε. Decomposing u and vj along vi and its orthogonal complement and applying Cauchy–Schwarz yields the stated bound. Lemma 11 (PL convergence under GD). Under Assumptions 5–6 and η ∈ (0, 1/L], the GD iterates for task i satisfy  Li (ΘT ) − L∗i ≤ (1 − ηµ)T Li (Θ0 ) − L∗i . Proof. For one GD step Θt+1 = Θt − η∇Li (Θt ), the L-smooth upper bound (Assumption 5) gives   2 Li (Θt+1 ) ≤ Li (Θt ) − η 1 − Lη 2 ∥∇Li (Θt )∥ . Since η ≤ 1/L, we have 1 − Lη/2 ≥ 1/2, hence 2

Li (Θt+1 ) ≤ Li (Θt ) − η2 ∥∇Li (Θt )∥ . 2

Applying the PL inequality 12 ∥∇Li (Θt )∥ ≥ µ(Li (Θt ) − L∗i ) yields  Li (Θt+1 ) − L∗i ≤ (1 − ηµ) Li (Θt ) − L∗i . Unrolling this recursion over t = 0, . . . , T − 1 gives the claim. PT −1 (j) (j) Lemma 12 (Task-vector norm bound). If τj = −η t=0 ∇Lj (Θt ) and ∇Lj (Θt ) ≤ G for all t, then ∥τj ∥ ≤ ηT G. Proof. By the triangle inequality, ∥τj ∥ ≤ η

T −1 X t=0

(j)

∇Lj (Θt ) ≤ η

T −1 X t=0

G = ηT G.

21

Lemma 13 (Inner-product upper bound). Under Assumptions 5–6 and η ∈ (0, 1/L],   2 ⟨∇Li (Θ), τi ⟩ ≤ − 1 − (1 − ηµ)T Li (Θ) − L∗i + L2 ∥τi ∥ . Proof. L-Lipschitz gradients (Assumption 5) also imply the quadratic lower bound; applying it with ∆ = τi , 2

Li (Θ + τi ) ≥ Li (Θ) + ⟨∇Li (Θ), τi ⟩ − L2 ∥τi ∥ , and rearrange to

2

⟨∇Li (Θ), τi ⟩ ≤ Li (Θ + τi ) − Li (Θ) + L2 ∥τi ∥ . Since Θ + τi = ΘT , Lemma 11 gives Li (ΘT ) − L∗i ≤ (1 − ηµ)T (Li (Θ) − L∗i ), hence   Li (Θ + τi ) − Li (Θ) ≤ − 1 − (1 − ηµ)T Li (Θ) − L∗i , which combined with the rearranged smoothness inequality yields the claim. 4) Main Theorems: Theorem 14 (Finite-step parameter-drift bound). Consider task i trained for T iterations of gradient descent with a fixed step PN size η ∈ (0, 1/L], and let γ := 1 − ηµ ∈ (0, 1). Then the merged update τm = j=1 αj τj satisfies Li (Θ + τm ) ≤ Ci + O(γ T ) + O(δηT ) + O(η 2 T 2 ), where O(γ T ) is the residual error from incomplete convergence on task i, O(δηT ) is the cross-task interference term, and O(η 2 T 2 ) is the curvature term from L-smoothness. Proof. Define the η, T -independent constant  Ci := Li (Θ) − αi Li (Θ) − L∗i . By L-smoothness,

2

Li (Θ + τm ) ≤ Li (Θ) + ⟨∇Li (Θ), τm ⟩ + L2 ∥τm ∥ . Decompose the inner product as ⟨∇Li , τm ⟩ = αi ⟨∇Li , τi ⟩ +

X

αj ⟨∇Li , τj ⟩ .

j̸=i

For the self term, Lemma 13 provides a constant part absorbed into Ci and a residual term of order O(γ T ), plus a curvature correction O(η 2 T 2 ) via Lemma 12. For the cross terms, Lemma 10 together with Assumption 9 gives ⟨∇Li , τj ⟩ ≤ δηT G2 , P so the sum over j ̸= i is O(δηT ). Finally, ∥τm ∥ ≤ ηT G j αj implies the smoothness term is O(η 2 T 2 ). Combining all contributions yields the stated bound. Theorem 15 (Near-convergence regime). Suppose the residual PL error after T steps is below a tolerance ζ > 0:  (1 − ηµ)T Li (Θ) − L∗i ≤ ζ, equivalently,

 ln (Li (Θ) − L∗i )/ζ . T ≥ − ln(1 − ηµ)

Then Li (Θ + τm ) ≤ Ci + O(ζ) + O(δηT ) + O(η 2 T 2 ), with the same Ci as in Theorem 14. Proof. Starting from Theorem 14, replace the residual term O(γ T ) by O(ζ) using the near-convergence assumption. The cross-task and curvature terms are unchanged. Remark 16. At a fixed learning rate, the improvement on the target task (captured by 1 − γ T ) typically outweighs the influence of other task vectors in the early training stage, especially when those vectors are close to orthogonal (small ε, hence small δ). As training approaches convergence, the negative impact from cross-task interference grows linearly in T as O(δηT ), and curvature errors grow quadratically as O(η 2 T 2 ); even when individual single-task losses keep decreasing, the merged loss can worsen due to accumulated interference. Once (1 − ηµ)T (Li − L∗i ) ≤ ζ, the dominant residual terms are interference and curvature; reducing directional leakage (small δ) and limiting ηT are therefore essential for high-quality merging. This motivates the conference benchmark’s choice to fine-tune each MLLM expert for one epoch with a reduced learning rate, and is consistent with prior empirical observations that less intensive fine-tuning often yields stronger merging [13], [67] and that fine-tuned models tend to converge near the base model [52], [68], [69].

22

5) Empirical Fine-Tuning Step Sweep: To illustrate Theorem 14 empirically, we run the standard CLIP-ViT-B/32 merging benchmark following the FusionBench fine-tuning setup [49]. We train each task expert with Adam at learning rate 10−5 for 4,000 steps with batch size 32, and save checkpoints every 500 steps. Across the eight TA8 tasks, single-task accuracy on the corresponding test split typically converges around 3,000 steps (Fig. 10), whereas merged accuracy peaks earlier and then declines as fine-tuning proceeds (Fig. 11).

Task-Specific Accuracy (%)

100.0 90.0 80.0 70.0

DTD EuroSAT GTSRB MNIST RESISC45 Stanford Cars SUN397 SVHN

60.0 50.0 40.0 30.0

0

500

1000

1500

2000

2500

Fine-tuning Steps

3000

3500

4000

Figure 10. Single-task fine-tuning accuracy of CLIP-ViT-B/32 on the eight TA8 tasks as a function of fine-tuning steps. Accuracy converges around 3,000 steps on every task, providing the per-task ground truth against which merging accuracy is measured in Fig. 11.

Task Arithmetic

70.6 70.5 70.4 70.3

Weight Average

67.2

Average Accuracy (%)

Average Accuracy (%)

70.7

67.0 66.8 66.6 66.4 66.2 66.0 65.8

70.2 500

1000

1500

2000

2500

Fine-tuning Steps

3000

3500

500

4000

1000

(a) Task Arithmetic

2000

2500

Fine-tuning Steps

3000

3500

4000

(b) Weight Average

DARE

70.8 70.7 70.6 70.5 70.4 70.3 70.2

TSV Merging

83.6

Average Accuracy (%)

Average Accuracy (%)

1500

83.5 83.4 83.3 83.2

70.1 500

1000

1500

2000

2500

Fine-tuning Steps (c) DARE

3000

3500

4000

500

1000

1500

2000

2500

Fine-tuning Steps

3000

3500

4000

(d) TSV Merging

Figure 11. Average merging accuracy on CLIP-ViT-B/32 TA8 against the fine-tuning step at which each expert was checkpointed. Across all four merging methods, accuracy first rises and then declines as fine-tuning progresses, with the peak occurring well before single-task convergence. This unimodal pattern is the empirical signature of Theorem 14: in the early phase the target-task improvement 1 − γ T dominates; once the loss approaches its minimum, the cross-task interference O(δηT ) and curvature O(η 2 T 2 ) terms grow large enough to outweigh the single-task gains, and merged accuracy decreases. MLLM training is organized in epochs rather than steps, so we fix the number of epochs to 1 and reduce the learning rate, which keeps fine-tuned experts close to the base model in parameter space while still improving on the target task.

23

Table XIV. CLIP-ViT-B/32: per-task accuracy (%) on the 8 vision tasks. Best per column in bold; second-best avg underlined.

Method

SUN397

Cars

RESISC45

EuroSAT SVHN

Weight Average Task Arithmetic TIES Merging TA w/ DARE TIES w/ DARE TSV Merging Iso-C τ cf = DC † WUDI Merging OptMerge

65.44 57.01 67.01 57.06 39.25 67.62 71.66 66.82 68.47 67.16

62.43 55.70 64.15 55.40 43.10 71.65 73.44 70.25 72.68 72.11

70.63 64.75 74.30 64.48 52.65 84.70 84.76 82.48 84.44 85.25

75.74 73.30 74.52 73.30 62.37 93.44 88.04 90.11 95.26 94.85

SWUDI-soft (ablation) SWUDI SWUDI-A

69.29 70.06 69.99

72.50 73.03 72.81

86.35 87.27 87.22

95.44 95.81 95.70

GTSRB MNIST

DTD

Avg.

64.51 77.93 77.74 78.07 81.39 91.90 78.69 93.27 94.90 95.27

54.96 68.50 69.38 68.38 71.48 92.53 84.62 92.83 95.00 95.66

86.28 96.07 94.13 96.06 97.47 98.86 96.69 99.13 99.29 99.33

50.59 47.13 53.99 46.97 39.95 63.83 65.21 63.72 67.02 66.60

66.32 67.55 71.90 67.46 60.96 83.07 80.39 82.33 84.63 84.53

94.53 94.22 94.36

94.76 95.00 95.00

99.27 99.29 99.30

68.62 69.68 69.84

85.10 85.55 85.53

Table XV. CLIP-ViT-B/16: per-task accuracy (%) on the 8 vision tasks. Best per column in bold; second-best avg underlined.

Method

SUN397

Cars

RESISC45

GTSRB MNIST

DTD

Avg.

Weight Average Task Arithmetic TIES Merging TA w/ DARE TIES w/ DARE TSV Merging Iso-C τ cf = DC † WUDI Merging OptMerge

68.74 65.91 70.65 65.88 56.10 73.12 75.13 73.61 75.07 74.91

69.05 68.31 71.23 68.28 61.14 80.74 81.06 79.37 82.10 82.51

75.06 75.49 79.89 75.54 70.68 89.75 90.35 91.65 92.13 92.92

EuroSAT SVHN 83.30 84.52 87.52 84.26 77.59 96.19 94.70 96.93 97.85 97.81

74.98 88.87 83.29 88.96 92.22 94.15 86.21 94.28 95.92 96.15

62.57 81.96 76.29 82.16 85.96 94.10 89.13 96.41 96.65 97.40

93.75 98.08 96.43 98.10 98.76 99.08 97.68 99.32 99.37 99.40

51.17 53.99 55.48 53.99 51.91 69.68 66.28 72.77 74.26 74.84

72.33 77.14 77.60 77.15 74.30 87.10 85.07 88.04 89.17 89.49

SWUDI-soft (ablation) SWUDI SWUDI-A

75.79 76.01 75.91

81.95 82.18 81.97

92.54 92.73 92.81

97.89 97.89 97.78

95.86 95.74 95.84

96.84 97.01 96.86

99.40 99.37 99.34

75.96 75.64 75.43

89.53 89.57 89.49

Table XVI. CLIP-ViT-L/14: per-task accuracy (%) on the 8 vision tasks. Best per column in bold; second-best avg underlined.

Method

SUN397

Cars

RESISC45

EuroSAT SVHN

Weight Average Task Arithmetic TIES Merging TA w/ DARE TIES w/ DARE TSV Merging Iso-C τ cf = DC † WUDI Merging OptMerge

72.53 72.02 74.77 72.09 65.78 78.18 79.78 79.52 80.03 80.15

81.54 79.00 83.16 78.85 69.59 89.79 90.76 90.13 90.71 90.86

82.32 80.57 86.51 80.46 69.29 93.52 94.37 93.70 93.89 94.29

88.52 84.63 89.70 84.48 73.30 96.70 96.48 97.41 98.33 98.33

SWUDI-soft (ablation) SWUDI SWUDI-A

80.55 80.56 80.59

91.00 91.10 91.06

94.17 94.46 94.40

98.52 98.37 98.48

GTSRB MNIST

DTD

Avg.

81.63 87.49 89.67 87.60 87.39 95.58 92.92 96.35 96.93 97.10

74.02 83.48 85.19 83.57 80.69 96.48 95.38 97.84 98.01 98.44

96.62 98.05 97.75 98.03 97.84 99.08 98.77 99.31 99.30 99.32

61.76 58.51 63.88 58.83 50.74 75.27 76.76 79.26 80.11 80.59

79.87 80.47 83.83 80.49 74.33 90.57 90.65 91.69 92.16 92.38

97.03 96.83 97.07

98.08 98.19 98.25

99.33 99.38 99.39

80.74 81.17 80.90

92.43 92.51 92.52

A PPENDIX C A DDITIONAL A NALYSES This appendix provides protocol details for the multimodal benchmark, per-task vision/language results, ablations of the closed-form solvers, spectral diagnostics, and additional MLLM scaling and checkpoint-merging studies.

A. Per-Task Vision Results on CLIP-ViT This subsection expands the CLIP-ViT TA8 summary in Table VI with per-task results for the three evaluated backbones. Per-task accuracy on the eight vision tasks for CLIP-ViT-B/32, B/16, and L/14 is reported in Tables XIV, XV, and XVI.

24

1.0

Density

0.8 0.6

OCR VQA Geometry Chart Grounding

4.0 3.5 3.0 2.5 2.0 1.5 1.0 0.5 0.0 7.0

Density

1.2

0.4 0.2 0.0

6

5

4

3

The magnitude of the parameter (Log10)

2

(a) InternVL2.5-1B (full FT)

OCR VQA Geometry Chart Grounding

6.5

6.0

5.5

5.0

4.5

4.0

3.5

The magnitude of the parameter (Log10)

3.0

(b) Qwen2-VL-7B (LoRA r=8)

Figure 12. Task-vector magnitude distribution on the MLLM benchmark. InternVL2.5 (full fine-tuning) exhibits a right-skewed distribution typical of dense parameter updates, whereas Qwen2-VL (LoRA) displays a multi-modal distribution: the low-rank constraint and LoRA scaling factor restrict deltas to a reduced subspace, causing them to cluster along a few dominant magnitudes. Both backbones show distinct distributions across tasks, supporting layer-wise rather than global rank rules.

OCR VQA Geometry Chart Grounding

Normalized Frobenius Norm

0.0016 0.0014 0.0012 0.0010

OCR VQA Geometry Chart Grounding

0.00008 0.00007 0.00006 0.00005

0.0008

0.00004

0.0006

0.00003

0.0004

0.00002

0.0002 0.0000

0.00009

Normalized Frobenius Norm

0.0018

0.00001

layer1 layer3 layer5 layer7 layer9 layer11 layer13 layer15 layer17 layer20

0.00000

(a) InternVL2.5-1B (full FT)

layer1 layer3 layer5 layer7 layer9 layer11 layer13 layer15 layer17 layer20

(b) Qwen2-VL-7B (LoRA r=8)

Figure 13. Normalized Frobenius norm of task vectors across layers. Norms are divided by the number of parameters of the corresponding linear layer. The Frobenius norm varies substantially across both layers and tasks, and the variation pattern differs by architecture and fine-tuning regime. Layer-wise rank adaptation in SWUDI-A addresses this heterogeneity directly. The small magnitudes in absolute terms (well below 1% of the base-model weight norm) are consistent with the conference-version observation that fine-tuned MLLMs and base models occupy adjacent regions of the loss landscape with linear connectivity [69]. Table P XVII. Per-architecture spectral diagnostics motivating adaptive rank selection. For each benchmark, we summarize the per-layer operators C (ℓ) = i τi⊤ τi /∥τi ∥2F by update magnitude, effective-rank ratio, spectral conditioning, layer-wise variability, and the ranks retained by SWUDI-A-psqrt. λ90 := λ⌈0.9 di ⌉ is the eigenvalue at the 90% rank position from the top under descending eigenvalue order.

Statistic

CLIP-B/32

Flan-T5 LoRA

Llama-3.2-3B

Qwen2-VL-7B (LoRA)

InternVL2.5-1B

# linear layers (ℓ) (ℓ) (ℓ) ∥τi ∥F /∥W0 ∥F (mean over i, ℓ) Effective-rank ratio reff /di (mean) λmax /λmed (median) λ90 /λmax (median) Layer-wise CV of λmax SWUDI-A-psqrt mean K/di SWUDI-A-psqrt min/median/max K

72 0.0173 0.177 85 2.9·10−3 0.601 0.554 26/180/512

72 0.0111 0.0043 1.3·108 4.5·10−10 0.257 0.013 4/12/64

196 0.0093 0.266 80 5.3·10−3 0.720 0.602 16/512/3072

560 0.0059 (LoRA ∆) 0.041 4.2·106 1.7·10−8 0.844 0.149 128/570/4010

168 0.0181 0.408 73 4.1·10−3 0.512 0.615 64/472/2048

B. Spectral Diagnostics This subsection gathers the spectrum-level evidence used to motivate layer-wise adaptive truncation. We first compare architectures and fine-tuning regimes, then include additional diagnostic panels. Spectral Statistics: Per-architecture spectral statistics on the per-layer interference operators C (ℓ) = P1) ⊤Per-Architecture 2 i τi τi /∥τi ∥F , referenced from Sec. VI-B, are summarized in Table XVII. Figs. 12 and 13 provide the empirical hook for the LoRA-vs-full-FT contrast that motivates the dual rank-rule design.

25

(a) Task-vector subspace captures input energy

(b) SGD matches Landweber exactly 1.0

0.8 0.6 0.4 early (L0) mid (L6) last (L11) random subspace

0.2 0.0 0.0

0.2

0.4

0.6

0.8

1.0

empirical filter ĥ k, n

input-energy capture

1.0

0.8

step step 1 step 5 step 50 step 200

0.6 0.4 0.2 0.0 10−1

10−2

retained rank ratio K/d

eigenvalue λk

Average norm of layers

Figure 14. Task-vector proxy and optimizer-filter diagnostics. (a) Task-vector subspaces capture input energy in early and middle CLIP-ViT-B/32 layers (capture gap +0.18–0.43 vs. random subspaces); the last MLP layer is a documented exception (gap ≈ 0). (b) SGD on the WUDI quadratic exactly matches the Landweber spectral filter 1 − (1 − ηλk )neff at all checkpoints (R2 = 1.0000, neff ≈ 2·step), confirming Proposition 2.

0.00027

0.0002 0.00012

0.0001 0.0000 0

50

WUDI Merging Ours

100 150 200 Optimization iterations

250

300

Figure 15. Frobenius-norm trajectory of τm under iterative WUDI/OptMerge on the Qwen2-VL LoRA setting (averaged over linear layers). The unregularized iterative loss inflates ∥τm ∥F throughout optimization (the norm-shortcut behavior of Sec. III-B2; see also Fig. 3). Mean initialization plus the low-rank truncation of τi keep the trajectory norm-bounded while reducing the loss successfully. The closed-form spectral solvers SWUDI/SWUDI-A avoid the inflation altogether by suppressing small-λk directions in the eigenbasis.

2) Task-Vector Proxy and Optimizer-Filter Diagnostics: This subsection presents two per-layer diagnostics evaluated on CLIP-ViT-B/32. Fig. 14(a) illustrates the capture gap between the task-vector subspaces and the input activation subspace, serving as the empirical basis for the task-vector proxy and Assumption 3. Furthermore, Fig. 14(b) demonstrates the exact equivalence between SGD on the WUDI quadratic objective and the Landweber spectral filter. This result confirms Proposition 2 and validates the SGD/Landweber identity deferred from Sec. III-B3. C. OptMerge Analysis and Rank-Truncation Evidence This subsection revisits two pieces of evidence from the OptMerge analysis. The component-wise analysis explains how the iterative baseline was stabilized, while the truncation-ratio sweep motivates the hard spectral cutoff used by SWUDI. 1) Component-Wise Analysis: OptMerge introduces three modifications to the iterative WUDI objective: replacing Adam with SGD, initializing τm with the mean of the task vectors, and applying a low-rank approximation to τi . We analyze the contribution of each component on Qwen2-VL (LoRA capability merging) and Vicuna-7B (modality merging), which are the two settings most affected by the narrow active subspace characteristic of LoRA. Replacing Adam with SGD in isolation is detrimental, as the optimizer struggles to escape the norm-inflation regime documented in Fig. 15. However, incorporating the mean initialization recovers and further improves accuracy, while the low-rank approximation of τi yields an additional marginal gain. Furthermore, these components have a neutral or mildly positive effect in the modality-merging setting. This indicates that they do not degrade performance in regimes for which OptMerge was not explicitly tuned. 2) Evidence for Head-Spectrum Truncation: Table XVIII presents a truncation-ratio sweep (k/di ∈ {0.1, 0.2, 0.3, 0.4, 0.5}) conducted for OptMerge in the InternVL2.5-1B capability merging setting. We include this analysis because the truncation index k plays an identical role in both the hard-truncation factor of SWUDI and the low-rank denoising of τi in OptMerge. Average performance remains essentially stable for k ∈ [0.1, 0.3] (ranging from 56.63% to 57.43%), but declines for k ≥ 0.4 as more low-eigenvalue directions are incorporated into the inversion. This observation aligns with the analysis in Sec. III-B2: the head of the spectrum accounts for almost all the proxy reduction, whereas including tail directions introduces noise rather

26

Table XVIII. OptMerge truncation-ratio sweep on InternVL2.5-1B, used to motivate SWUDI rank choices. Each row reports per-task accuracy (%).

k ratio

VQA

Geometry

Chart

OCR

Grounding

Avg.

VizWiz GQA MathVista MATH-Vision ChartQA TextVQA OCRVQA RefCOCO RefCOCO+ RefCOCOg 10% 20% 30% 40% 50%

30.90 30.97 31.55 31.49 31.37

57.26 57.13 57.15 56.92 56.68

51.49 54.48 54.50 55.77 56.75

18.42 21.05 21.05 25.00 23.68

68.40 68.72 68.72 67.36 68.08

76.10 76.01 76.27 76.06 75.81

46.39 46.35 45.67 45.96 45.02

76.36 75.97 73.63 65.55 61.45

69.99 69.72 66.84 58.40 54.80

73.96 73.94 70.92 59.64 56.19

56.93 57.43 56.63 54.22 52.98

than signal. Consequently, this trend provides empirical justification for setting the default SWUDI configuration to a small truncation ratio. A PPENDIX D O UR MLLM ERGING B ENCHMARK This appendix documents MLLMerging, the benchmark introduced in Sec. V-A for merging multimodal large language models (MLLMs). It isolates the merging algorithm as the only free variable: experts share a backbone, are fine-tuned on capability-aligned data, and are merged purely in parameter space without access to the original training data. The benchmark provides the training suite, full fine-tuning and LoRA expert checkpoints, and a matched evaluation protocol, enabling fair comparison across merging methods. A. Motivation and Benchmark Scope Existing model-merging benchmarks are dominated by vision-only classifiers or text-only language tasks. MLLMs introduce additional complications because a single model must preserve visual perception, language reasoning, grounding, OCR, and modality-specific alignment. MLLMerging targets three gaps that are not fully covered by earlier benchmarks. Training-evaluation mismatch. Public MLLMs are often trained on mixtures of proprietary, licensed, and open-source data, while they are evaluated on standalone suites such as MMBench [70], SEED-Bench [71], MME [72], and MMStar [73]. The same backbone can therefore exhibit very different capability profiles depending on the fine-tuning data. MLLMerging aligns capability-specific training data with capability-specific evaluation suites, so that the merging algorithm, rather than the upstream data composition, is the primary variable. Task expertise versus instruction following. Capability datasets such as VQA, OCR, and grounding provide strong task supervision, but they may not match the broad instruction-following distribution of modern MLLMs. The benchmark therefore reports both per-capability evaluations (Tables II and III) and integrated multimodal QA evaluations (Table V). Capability and modality composition. Beyond combining task-specialized experts that share a full MLLM backbone, the benchmark also studies modality merging: vision-, audio-, and video-language experts share an LLM backbone but use modality-specific encoders and connectors. This setting tests whether a parameter-space merge can preserve complementary sensory channels without online routing or joint retraining. B. Capability-Merging Tasks and Data A core contribution of MLLMerging is the curated capability-merging data suite. Unlike prior studies that rely on a few vision-classification heads or a single text corpus, we construct a large, capability-aligned training pool. This ensures each expert is a true specialist, isolating the merging algorithm as the sole variable during evaluation. The suite encompasses five complementary MLLM capabilities (VQA, Geometry, Chart understanding, OCR, and Grounding), aggregating approximately 1.37M instruction-tuning samples from over twenty public datasets (Table XIX). We deliberately collect at least 100K samples per capability and prioritize source diversity (e.g., incorporating ten OCR datasets ranging from scene text to document and table understanding). This prevents experts from overfitting to a single dataset’s distribution, promoting broad generalization within their respective domains. Two additional design choices ensure the suite is readily reusable as a benchmark. First, all data sources are standardized into a ShareGPT-style instruction-tuning format with a uniform grounding-coordinate convention (Appendix D-C). This guarantees that any backbone can be fine-tuned, and any merging method evaluated, under identical supervision conditions. Second, the suite intentionally mixes English-only and bilingual (English/Chinese) sources. While InternVL2.5-1B utilizes the full multilingual dataset, Qwen2-VL-7B is restricted to the English-only subsets. This design allows us to evaluate merging algorithms on both multilingual full-parameter experts and monolingual low-rank (LoRA) experts within a unified framework. Ultimately, pairing this training suite with the capability-matched evaluation protocol (Appendix D-E) closes the train-evaluation gap often present in earlier MLLM benchmarks (Appendix D-A). Consequently, any changes in downstream accuracy can be confidently attributed to the merging algorithm itself, rather than variations in upstream data composition.

27

Table XIX. Capability training datasets used to construct the MLLMerging expert checkpoints: five capabilities, over twenty public datasets, and ≈ 1.37M instruction-tuning samples in total (≥ 100K per capability).

Capability

Total

Datasets (language)

VQA

588K

Geometry Chart OCR

190K 218K 238K

Grounding

135K

GQA (en) [35], VQAv2 (en) [74], OKVQA (en) [75], LLaVA-Instruct (zh) [76], CogVLM-Singleround (en & zh) [77], CogVLM-Multiround (en & zh) [77] GeoQA+ (zh) [78], G-LLaVA (en) [79] ChartQA (en) [38], DVQA (en) [80] OCRVQA (en) [40], TextCaps (en) [81], SynthDoG (en) [82], LLaVAR (en) [83], ST-VQA (en) [84], TextVQA (en) [39], DocVQA (en) [43], DeepForm (en) [85], KLC (en) [86], TabFact (en) [87] RefCOCO (en) [41], [88], [89], VG (en) [90] Table XX. Modality components and training data for the three single-modality experts.

Modality

Encoder

Connector

Alignment Data

Fine-tuning Data

Reference

Vision

CLIP-ViT-L-336px [50]

MLP

LCS 558K [92]

LLaVA-mixed 665K [76]

LLaVA-1.5 [76]

Audio

BEATs-Iter3+ [93]

Q-Former [94]

WaveCaps 400K [95]

OpenAQA filtered 350K [96]

X-InstructBLIP [97]

Video

LanguageBind [98]

MLP

LCS 558K [92], Valley 702K [99]

Video-ChatGPT 100K [100], LLaVA- Video-LLaVA [101] mixed subset 140K [76]

C. Backbones and Expert Construction Capability merging. We employ two representative MLLM backbones. InternVL2.5-1B-Instruct [29] is fully fine-tuned for one epoch with a learning rate of 4e−5 and a warmup ratio of 3e−2. Qwen2-VL-7B-Base [30] is fine-tuned using LoRA [91] with a rank of r=8, a learning rate of 10−5 , and a warmup ratio of 10−1 . These two configurations allow us to evaluate the merging methods on both dense full-parameter deltas and low-rank LoRA deltas. Data preprocessing. Following standard training practices for InternVL and Qwen2-VL, we utilize only the training splits. We filter out corrupted images and samples where the combined question-answer length exceeds 8192 tokens. The remaining data is then converted into the ShareGPT-style instruction-tuning format. Grounding coordinates are linearly mapped to the [0, 1000) range and enclosed within Qwen2-VL box tokens (e.g., <|box_start|>· · · <|box_end|> [30]). D. Modality-Merging Track For modality merging (Sec. V-C, Table IV), we follow [26] and pair Vicuna-7B-v1.5 [31] with three modality-specific encoder/connector pairs. The vision, audio, and video experts share the same LLM backbone but are trained on different bi-modal data. Table XX lists the modality components. The modality experts are trained in two stages. Stage 1 aligns each modality encoder to the LLM by training only the connector. Stage 2 fine-tunes the connector and the LLM, with LoRA of rank r=128 applied to all linear modules in the LLM. At merging time, the modality-specific encoders and connectors are kept intact, and only the LLM LoRA deltas are merged. The resulting Omni model can process vision, audio, and video inputs while using a single merged LLM backbone. E. Evaluation Protocol Capability evaluation is conducted using VLMEvalKit [32] and lmms-eval [33] with consistent decoding, preprocessing, and answer-extraction settings. The five-capability suite includes VizWiz [34] and GQA [35] for VQA; MathVista [36] and MATH-Vision [37] for Geometry; ChartQA [38] for Chart understanding; TextVQA [39] and OCRVQA [40] for OCR; and RefCOCO/+/g [41], [88], [89] for Grounding. Math geometry subsets. While the original conference paper [6] restricted MathVista to its geometry-related subsets and MATH-Vision to four specific geometry categories, the updated protocol in Tables II and III reports the official overall results for both MathVista and MATH-Vision. Integrated QA and modality evaluation. For integrated multimodal QA, we employ MMMU [42], DocVQA [43], ScienceQA [44], AI2D [45], and InfographicVQA [46]. Modality merging is evaluated on AVQA [48] and MUSIC-AVQA [47], which assess spatio-temporal reasoning across audio-visual scenes. F. Answer Extraction Prompt For MathVista and MATH-Vision, free-form model outputs are normalized by GPT-4o-mini using the prompt below. The template variables {question} and {prediction} denote the original question and the model’s raw response.

28

Please read the following examples. Then extract the answer from the model response and type it at the end of the prompt. Hint: Please answer the question requiring an integer answer and provide the final value, e.g., 1, 2, 3, at the end. Question: Which number is missing? Model response: The number missing in the sequence is 14. Extracted answer: 14 Hint: Please answer the question requiring a floating-point number with one decimal place and provide the final value, e.g., 1.2, 1.3, 1.4, at the end. Question: What is the fraction of females facing the camera? Model response: The fraction of females facing the camera is 0.6, which means that six out of ten females in the group are facing the camera. Extracted answer: 0.6 Hint: Please answer the question requiring a floating-point number with two decimal places and provide the final value, e.g., 1.23, 1.34, 1.45, at the end. Question: How much money does Luca need to buy a sour apple candy and a butter-scotch candy? (Unit: $) Model response: Luca needs $1.45 to buy a sour apple candy and a butterscotch candy. Extracted answer: 1.45 Hint: Please answer the question requiring a Python list as an answer and provide the final list, e.g., [1, 2, 3], [1.2, 1.3, 1.4], at the end. Question: Between which two years does the line graph saw its maximum peak? Model response: The line graph saw its maximum peak between 2007 and 2008. Extracted answer: [2007, 2008] Hint: Please answer the question and provide the correct option letter, e.g., A, B, C, D, at the end. Question: What fraction of the shape is blue? Choices: (A) 3/11 (B) 8/11 (C) 6/11 (D) 3/5 Model response: The correct answer is (B) 8/11. Extracted answer: B {question} Model response: {prediction} Extracted answer:

Record · ID 266189 · SHA-256 23370f67bd02bd2b
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.