ConceptioArchivearXiv CS
arXiv CSopen access

PFAdapter: Hierarchical LoRA Decomposition for Personalized Federated MLLMs

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
distributedsystemsprotocols
networking, internet, protocols, distributed systems

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKS

1

PFAdapter: Hierarchical LoRA Decomposition for Personalized Federated MLLMs Jing Liu ID , Member, IEEE, Kun Yang ID , Yan Wang ID , Dingkang Yang ID , Xiaoshuai Hao ID ,

arXiv:2607.12111v1 [cs.LG] 13 Jul 2026

Wei Zhang ID , Yang Liu ID , Member, IEEE, and Wei Zhou ID Senior Member, IEEE

Abstract—Agentic AI systems are reshaping communications and networking by deploying autonomous intelligent agents capable of collaborative learning while maintaining data privacy at network edges. Within distributed network environments, Multimodal Large Language Models (MLLMs) serve as cognitive engines for edge devices, yet federated fine-tuning faces substantial challenges in balancing global knowledge aggregation with local adaptation under heterogeneous network conditions. Conventional federated protocols typically rely on uniform parameter aggregation, which conflates domain-invariant features with client-specific nuances, thereby resulting in suboptimal personalization and excessive communication overhead. To address these challenges, we propose PFAdapter, a communication-efficient framework introducing hierarchical LoRA decomposition to explicitly separate adapter parameters into global-shared and local-private components. Query and key projections are assigned to global synchronization for capturing universal multimodal semantics across the network, while value and output projections remain localized for edgespecific adaptation. Additionally, orthogonality regularization based on the Frobenius norm enforces strict separation between these components, preventing redundant feature learning. Selective aggregation protocols synchronize only global-shared components across the federated network, preserving local expertise and reducing communication costs by nearly 50%. Experiments on medical VQA and social multimodal benchmarks show that PFAdapter consistently improves over matched federated LoRA baselines while nearly halving synchronized adapter traffic. These results indicate that projection-level decomposition offers a practical path toward communication-efficient agentic MLLM deployment in resource-constrained edge networks. Index Terms—Federated Learning, Edge Intelligence, Multimodal Large Language Models, Communication Efficiency, Agentic AI, Personalization, Parameter-Efficient Fine-Tuning Jing Liu is with the College of Future Information Technology, Fudan University, Shanghai 200433, China, also with the Division of Natural and Applied Sciences, Duke Kunshan University, Suzhou 215316, China, and also with the Department of Electrical and Computer Engineering, The University of British Columbia, BC V6T 1Z4, Canada (e-mail: [email protected]). Kun Yang is with the Ant Group, also with the College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou 310013, China (e-mail: [email protected]). Yan Wang is with the School of Data Science and Engineering, East China Normal University, Shanghai 200062, China (e-mail: [email protected]). Dingkang Yang is with the College of Intelligent Robotics and Advanced Manufacturing, Fudan University & Fysics AI, Shanghai 200433, China (email: [email protected]). Xiaoshuai Hao is with Xiaomi EV, Xiaomi Campus, Anningzhuang Road, Haidian District, 100085, Beijing, China (e-mail: [email protected]). Wei Zhang is with the Information and Communications Technology Cluster, Singapore Institute of Technology (SIT), Singapore 828608 (e-mail: [email protected]). Yang Liu is with the College of Electronic and Information Engineering, Tongji University, Shanghai 201804, China (e-mail: [email protected]). Wei Zhou is with the School of Computer Science and Informatics, Cardiff University, CF24 4AG Cardiff, U.K. (e-mail: [email protected]).

I. I NTRODUCTION GENTIC AI represents a transformative paradigm for modern communication networks, where autonomous agents collaborate to process multimodal information at network edges while keeping raw training data local to each client [1], [2]. Multimodal Large Language Models (MLLMs) serve as cognitive cores for intelligent agents deployed across distributed networking systems, enabling sophisticated reasoning over heterogeneous data modalities including text, images, and sensor data [3]–[5]. Urgent deployment scenarios emerge in manufacturing anomaly detection requiring rapid inference under constrained budgets, autonomous driving processing multisensor data with stringent latency, and medical diagnostics handling privacy-sensitive imaging across distributed facilities [6], [7], where Federated Learning (FL) enables collaborative model training without centralized data aggregation [8], [9]. Integrating FL with MLLMs becomes critical for realizing scalable data-local agentic systems across next-generation communication networks [10]–[12]. We use this wording deliberately: FL keeps raw samples on device, but it does not by itself guarantee resistance to update inversion, gradient leakage, or membership inference [13], [14]. The privacy scope of PFAdapter is therefore limited to decentralized training without centralized raw-data pooling, and the stronger attackmodel discussion is deferred to Sec. V-C. Deploying federated MLLMs across heterogeneous edge networks introduces substantial challenges arising from extreme variations in local data distributions and network conditions. Medical edge devices often process specialized imaging data that varies significantly across equipment manufacturers and patient demographics [15]–[17], while IoT nodes in social sensing networks must handle diverse cultural contexts and evolving linguistic patterns [18], [19]. Standard global models frequently fail to achieve optimal performance on specialized edge tasks due to uniform aggregation strategies that ignore local data characteristics. Consequently, developing personalized federated learning approaches becomes imperative for edge intelligence systems, enabling adaptation to local nuances while benefiting from collective knowledge [20]–[22]. Fundamental technical barriers hinder effective MLLM personalization within distributed communication networks. Balancing generalizable multimodal representations against edge-specific task expertise poses a primary challenge for network-deployed agents [23], [24], where uniform parameter aggregation induces catastrophic weight washing that dilutes critical client-specific patterns through global averaging

A

2

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKS

the adapter space. To address communication and personalization challenges in + VS edge networks, we propose PFAdapter, a resource-efficient ⊥ Weight washing federated learning framework for agentic MLLM deployment Challenge Advantages via hierarchical adapter decomposition. As illustrated in Fig. 1, ✓ Knowledge disent ✓ Local expertise preserved ✗ Weight washing dilutes local knowledge ✓ ~50% comm. reduction ✓ Global semantics shared ✓ ✗ High communication overhead (100%) ✗ our framework explicitly separates adapter parameters into ✓ Superior personalization ✓ +2.4-4.8% accuracy gain ✗ Global-local knowledge entanglement global-shared and local-private sets based on functional roles Fig. 1: Illustration of challenges in federated MLLM fine- of self-attention projections, where query and key projections tuning and our proposed solution. (a) Existing FL methods (q , k ) capture universal multimodal semantics for global p p uniformly aggregate all model parameters, leading to high synchronization while value and output projections (v , o ) p p communication costs, loss of client-specific knowledge, and remain localized for edge-specific adaptation. Orthogonality poor personalization under heterogeneous data distributions. regularization enforces strict separation by minimizing inner (b) PFAdapter introduces hierarchical LoRA decomposition products between global and local parameter matrices, thereby with selective aggregation, where only global-shared adapters reducing interference in respective feature subspaces while are synchronized, while local-private adapters remain on-device. maximizing complementary representation capacity and preventing redundant feature learning. The main contributions of [25], [26]. In manufacturing anomaly detection, extended this work are summarized as follows: local training capturing equipment-specific fault signatures • We introduce a novel architectural decomposition for paradoxically degrades targeted performance after aggregaMLLM adapters, categorizing projection modules into tion while fundamentally undermining model reliability by global and local components based on their functional erasing calibrations essential for safety-critical autonomous roles in capturing multimodal semantics versus edgesystems [27], [28]. Although earlier FL approaches attempted specific features. mitigation through proximal regularization [29], multi-task • We propose Frobenius-norm based orthogonality regularlearning [30], and meta-learning [20], high-dimensional MLLM ization to minimize correlation between global and local parameter spaces frequently lead to insufficient personalization adapter weights, ensuring precise knowledge separation or catastrophic forgetting across the federated network [31]. for distributed agents. Communication overhead represents a critical bottleneck for • We develop a communication-efficient synchronization federated MLLM deployment across bandwidth-constrained strategy that transmits only global-shared parameters, edge networks. Even when employing Parameter-Efficient Finereducing network traffic by nearly 50% while preserving Tuning (PEFT) techniques like Low-Rank Adaptation (LoRA) local edge expertise. [32], MLLM backbones impose prohibitive synchronization • We conduct evaluations across medical VQA and social costs for resource-limited edge devices [3]. Under typical multimodal benchmarks, showing consistent accuracy, industrial edge deployments with limited bandwidth (e.g., 5robustness, and communication-efficiency gains under 20 Mbps in manufacturing facilities or vehicular networks), matched federated LoRA protocols. transmitting hundreds of megabytes of adapter parameters per The remainder of this paper is organized as follows. Sec. II round incurs substantial delays that accumulate across training reviews related work in federated learning and MLLM fineiterations, proving operationally unacceptable for time-sensitive tuning. Sec. III provides the necessary preliminaries and applications like autonomous vehicle fleets or manufacturing problem formulation. Sec. IV details the proposed PFAdapter anomaly detection [33], [34]. Existing solutions frequently framework and its technical components. Sec. V describes overlook redundancy within adapter parameters by transmitting the experimental setup and analyzes the results. Finally, all modules uniformly despite certain components capturing Sec. VI concludes the paper. Additional theoretical analysis, universal features while others remain highly task-specific [10], implementation details, and extended experimental results are [20], consequently wasting scarce network resources. provided in the supplementary material, Secs. II, III and V. Recent advances in federated MLLM fine-tuning have introduced various strategies to mitigate communication chalII. R ELATED W ORK lenges. FedAvg [8] establishes the baseline for collaborative A. Federated learning for large language models learning across distributed networks, whereas FedProx [29] incorporates proximal regularization for handling system and Federated fine-tuning of Large Language Models (LLMs) and statistical heterogeneity. More recently, frameworks including MLLMs has emerged as a critical research direction due to the FlexLoRA [35] and FedMLLM [36] have explored LoRA- growing demand for collaborative learning across institutions based adaptation to reduce trainable parameters. Systematic that cannot pool raw multimodal data centrally [31], [41]. Early analysis reveals that uniform aggregation of all LoRA modules federated learning frameworks such as FedAvg [8] established conflates domain-invariant features with client-specific nuances the foundation for decentralized model training, yet they [37], [38]. Additionally, methods including FedPer [39] and were designed for relatively small-scale neural networks and FedRep [40] attempt model splitting into global and local homogeneous data distributions. The emergence of parameterlayers, yet their application to complex MLLM architectures efficient fine-tuning techniques, particularly LoRA [32], has remains suboptimal. Observed limitations highlight critical enabled practical federated adaptation of billion-scale models gaps in achieving precise knowledge disentanglement within by reducing the number of trainable parameters from billions Problem: Uniform aggregation in FL

+ ( LoRA )

Domain B Domain A Communication: 100% - All LoRA transmitted

Solution: PFedAdapter decomposition

(a)

Global ( )

Domain A

Domain B

(b)

Local

(

)

Orthogonal

Communication: ~50% - Only Q, K transmitted

Optimal balance

AUTHOR et al: PFADAPTER: HIERARCHICAL ADAPTER DECOMPOSITION FOR PFM

to millions. Recent frameworks such as FlexLoRA [35] and split-learning approaches [42] have demonstrated the potential of LoRA-based federated fine-tuning for heterogeneous tasks and resource-constrained environments. However, recent work has identified significant challenges in applying standard FL protocols to MLLMs, including catastrophic forgetting of global knowledge during local updates and the "weight washing" phenomenon where client-specific adaptations are diluted through uniform aggregation [25]. Several approaches have attempted to address these limitations through regularization techniques, multi-stage training protocols [43], and hybrid architectures that separate feature extractors from task-specific heads. Nevertheless, existing methods treat adapter parameters as a monolithic entity without considering the functional heterogeneity within different projection modules of the self-attention mechanism. In contrast, PFAdapter explicitly recognizes that query and key projections tend to capture structural multimodal relationships that are more amenable to global aggregation, whereas value and output projections are inherently more taskspecific and should remain personalized.

B. Personalized federated learning

3

C. Parameter-efficient fine-tuning in federated settings PEFT has become indispensable for adapting large-scale models in resource-constrained environments [47]. Beyond LoRA, several PEFT variants have been developed including adapter layers [48], prompt tuning [49], prefix tuning [50], and Hadamard adapters [51]. Recent work has explored the integration of these techniques into federated learning frameworks to reduce communication overhead and computational burden [10]. Methods such as FedAdapter [52] and pFedLoRA [37] have demonstrated the effectiveness of adapter-based personalization for language models. Building upon these foundations, FloRA [38] introduced heterogeneous low-rank adaptations for federated LLM fine-tuning, while recent work on adaptive LoRA experts and differentially private federated LoRA [13] have further advanced the field. For multimodal applications, approaches like FedDLP [53] have employed dual adapters with selective pruning to balance local specialization and global knowledge sharing, while visuallanguage enhancement systems motivate similarly compact adaptation under edge visual workloads [54]. However, these approaches typically apply uniform aggregation to all adapter parameters, failing to distinguish between domain-invariant and domain-specific components. The FedMLLM framework [36] represents the state-of-the-art in federated MLLM finetuning, yet it relies on full adapter synchronization that conflates global and local knowledge. Recent theoretical analysis has shown that the optimal aggregation strategy should vary across different parameter subsets based on their sensitivity to local data distributions [20]. Motivated by this insight, PFAdapter introduces selective aggregation that only synchronizes the global-shared adapter components, reducing communication costs while preserving local expertise. Table III provides the full structured comparison of representative personalized FL and federated LoRA methods [37], [38], [53]. In addition, it contrasts PFAdapter with the closely related FedMLLM setting [36].

Personalization in federated learning aims to balance the acquisition of global knowledge with the preservation of local expertise, particularly important for non-IID data distributions [10]. Among pioneering approaches, FedPer [39] introduced the concept of maintaining personalized layers at each client while aggregating only the base model parameters. Building upon this foundation, FedRep [40] proposed a representation learning approach where feature extractors are synchronized globally and classifiers remain local. Recent work has explored more sophisticated personalization strategies including prototype-based learning [44], meta-learning frameworks [20], and dual-prompt optimization [22], [35]. Furthermore, FedBABU [45] demonstrated that keeping the model head frozen during server aggregation can significantly improve personalization performance. For multimodal scenarios, [46] III. P RELIMINARIES introduced task-similarity-aware model aggregation for heterogeneous multi-modal clients. However, these methods primarily A. Problem Formulation focus on architectural separation at the layer level rather Consider a federated learning system consisting of a central than parameter-level decomposition within individual modules. server and a set of K heterogeneous clients, where each client Moreover, layer-wise splitting strategies become inefficient for k ∈ {1, . . . , K} possesses a local multimodal dataset Dk . deep transformer architectures where task-specific knowledge Given the input space X comprising image-text pairs and is distributed across multiple layers. PFAdapter addresses the output space Y representing textual responses, the local this limitation by performing fine-grained decomposition at the k dataset is defined as Dk = {(xk,i , yk,i )}ni=1 , where nk denotes projection module level, enabling more nuanced control over the number of local samples. Let fθ : X → Y represent an which aspects of the model are shared versus personalized. MLLM parameterized by θ ∈ Rd . The primary objective in Compared with FedPer and FedRep [39], [40], which personalized federated learning is to find a set of parameters personalize entire layers or heads, PFAdapter moves the {θ1 , . . . , θK } that minimize the aggregate empirical risk: personalization boundary inside each attention block and K X explicitly separates globally aggregated query/key projections nk min E(x,y)∼Dk [ℓ(fθk (x), y)], (1) from locally retained value/output projections. Accordingly, N {θk }K k=1 k=1 the projection-level split is tailored to MLLM adapters, where P structural cross-modal alignment must be shared across clients where N = k nk is the total number of samples across but semantic realization remains strongly client-dependent all clients and ℓ(·, ·) denotes the cross-entropy loss function. under non-IID data. Unlike standard federated learning which seeks a single global

4

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKS

model θ∗ , personalized approaches allow for client-specific variations to account for statistical heterogeneity [20], [21].

through value/output adapters. The supplementary structured comparison summarizes this distinction before the formal derivation. The closest architectural contrast is with FedDLP [53]: both methods separate shared and personalized adaptation paths, but FedDLP decides what to transmit by pruning an auxiliary shared branch, whereas PFAdapter fixes the communication boundary at the projection type itself and aggregates only the query/key adapters. Consequently, the distinction matters in multimodal non-IID settings because it ties the synchronized subspace directly to attention-relation formation rather than to a sparsified duplicate branch.

B. Low-Rank Adaptation LoRA serves as the foundational parameter-efficient finetuning technique for large-scale models. For a pre-trained weight matrix W0 ∈ Rd×k , LoRA represents the weight update ∆W as the product of two low-rank matrices B ∈ Rd×r and A ∈ Rr×k , where the rank r ≪ min(d, k). The forward pass of a modified linear layer is expressed as: h = W0 x + ∆W x = W0 x + BAx.

(2)

K X  nk  G L Lt (Dk ; ΘG , ΘL k ) + λo Lo (Θ , Θk ) L K G N Θ ,{Θk }k=1

(4) During the fine-tuning process, W0 remains frozen while only k=1 A and B are updated. In the context of MLLMs, LoRA is G L typically applied to the projection matrices within the self- where Θ represents the global-shared parameters and Θk attention mechanism, specifically the query (Wq ), key (Wk ), denotes the local-private parameters for client k. The task loss Lt is typically defined as the negative log-likelihood of the value (Wv ), and output (Wo ) projections [51]. target sequence given the multimodal input: min

C. Federated Optimization Standard federated optimization often employs the FedAvg protocol to synchronize model updates across clients. In each communication round t, the server selects a subset of clients St and transmits the current global parameters θt . Each selected client performs E epochs of local stochastic gradient descent (SGD) to obtain updated parameters θkt+1 . The server then aggregates these updates using a weighted average: X n P k θt+1 = θkt+1 . (3) j∈St nj k∈St

However, applying this uniform aggregation to all LoRA parameters in MLLMs often leads to the dilution of local taskspecific knowledge, necessitating a more granular approach to parameter management [37], [38].

n

|yk,i |

k X 1 X Lt (Dk ; Θ) = − log P (yk,i,j | yk,i,<j , xk,i ; Θ) (5) nk i=1 j=1

B. Hierarchical LoRA Decomposition Hierarchical decomposition in PFAdapter formally splits the total set of trainable LoRA parameters Θ into two disjoint sets: ΘG and ΘL . Let M = {q, k, v, o} denote the set of projection types in the transformer layers. The global parameter set ΘG is defined as the union of LoRA weights for query and key projections: ΘG =

L [

{Al,m , Bl,m | m ∈ {q, k}},

(6)

l=1

IV. M ETHOD A. Method Overview Personalized federated learning for MLLMs requires a delicate balance between global knowledge acquisition and local task adaptation. To achieve this balance, PFAdapter introduces a structured decomposition of the adapter parameter space that explicitly separates domain-invariant features from client-specific nuances. Our design philosophy rests on the observation that different projection modules within the self-attention mechanism exhibit varying degrees of taskspecificity. Specifically, the framework partitions the set of LoRA modules into global-shared and local-private subsets based on their functional roles in processing multimodal information. Fig. 2 illustrates the overall architecture, where query and key projections are synchronized globally while value and output projections are maintained locally. The overall optimization objective for the system is formulated as: This design differs from layer-wise personalized FL and fully synchronized federated LoRA [36], [39], [40] because it introduces an intermediate projection-level control point: attention-map formation is shared through query/key adapters, while client-specific representation realization is preserved

where L is the number of transformer layers. Conversely, the local parameter set ΘL contains the weights for value and output projections: L

Θ =

L [

{Al,m , Bl,m | m ∈ {v, o}}.

(7)

l=1

The intuition behind this specific split arises from the functional roles of self-attention modules. Query and key projections encode the attention patterns that determine which tokens attend to each other, capturing structural relationships between multimodal tokens that reflect domain-invariant correspondence patterns [55], [56]. Recent work on adapter decomposition has shown that attention patterns exhibit higher cross-domain transferability than value representations [32], [48]. The attention score matrix S is computed as: S=

(Wq + ∆WqG )x((Wk + ∆WkG )x)⊤ √ . dk

(8)

In contrast, value and output projections encode the semantic content that is transformed into the final representation, making them inherently more task-specific and susceptible to local data variations [57]. The final attended representation V ′ is obtained

AUTHOR et al: PFADAPTER: HIERARCHICAL ADAPTER DECOMPOSITION FOR PFM

Key Features

✓ Orthogonality regularization ✓ Local expertise preserved ✓ ~50% communication ↓ ✓ +2.4~4.8% accuracy

☁️

Server

5

Query projection

Only

,

Legend

Global-shared parameters

Selective aggregation

Global (Sync)

Trainable

Local (Private)

Frozen

Key projection

synchronized

Upload

Global model

Client 1

:

Client K

+

:

Self-attention layer Global

Broadcast

+

Self-attention layer Local

Local data

Global

Local

Local data

Fig. 2: Architecture of PFAdapter. The framework decomposes LoRA adapters into global-shared components (qp , kp ) and local-private components (vp , op ). Orthogonality regularization enforces knowledge disentanglement between these sets. via: V ′ = Softmax(S)(Wv + ∆WvL )x.

(9)

LoRA-based federated learning, ensuring no additional memory overhead during local training.

Consistent with this mechanism, the design aligns with empirical findings that value layers capture more task-specific C. Orthogonality-Driven Knowledge Disentanglement information than query-key pairs in multi-task learning scenarEffective knowledge disentanglement requires that the global ios [30], [58]. Under non-IID multimodal federation [29], [59], [60], the server needs to preserve client-agnostic alignment and local components learn non-redundant features. To enforce cues while avoiding over-averaging client-specific semantics. this separation, PFAdapter incorporates an orthogonality Eqs. 8 and 9 make this separation explicit: query/key adapters regularization term into the local optimization objective. Let perturb the attention logits that determine which visual-textual Wl,m = Bl,m Al,m denote the effective weight update for layer tokens interact, whereas value/output adapters govern what l and module m. The orthogonality loss Lo is formulated using task-specific content is injected after the shared attention map the Frobenius norm of the product between global and local is formed [55], [57]. We therefore assign ΘG in Eq. 6 to weight matrices: the transferable relation-encoding subspace and keep ΘL in L X X X ⊤ Eq. 7 client-resident to absorb label skew, vocabulary bias, Lo = ∥Wl,m Wl,ml ∥2F . (11) g and modality imbalance without washing out local semantics l=1 mg ∈{q,k} ml ∈{v,o} during aggregation. To keep the notation consistent throughout Minimizing this term encourages the column spaces of global the remainder of the paper, ΘL in Eq. 7 denotes the structural and local adapters to be orthogonal, thereby preventing the set of local-private slots, while the actual parameters owned local modules from re-learning information already captured by client k are written as ΘL k . Accordingly, all optimization by the global components. The Frobenius norm ∥ · ∥F for a objectives, algorithmic updates, and convergence statements matrix A ∈ Rm×n is defined as: below use the pair (ΘG , ΘL v k ) when referring to a concrete uX n client state, consistent with personalized FL notation where um X t ∥A∥ = a2ij . (12) F local client states remain distinct from the shared server-side i=1 j=1 model [39], [40]. For a given input x, the forward pass of a decomposed self-attention layer is expressed as: Differentiability of the Frobenius norm allows for efficient ! gradient-based optimization. During local backpropagation, the (Wq + ∆WqG )x((Wk + ∆WkG )x)⊤ √ Attn(x) = Softmax gradients of the orthogonality loss with respect to the global dk and local parameters are computed as: × (Wv + ∆WvL )x X ⊤ ∇ΘG Lo = 2 Wl,ml Wl,m Wl,mg , (13) (10) l l,m ,m where ∆Wm = Bm Am represents the low-rank update for g l X module m. Given the resulting partition, the local model ⊤ ∇ΘL Lo = 2 Wl,mg Wl,m Wl,ml . (14) g for client k is represented as the combination f (·; ΘG , ΘL k ). l,m ,m g l During the training process, only ΘG is subject to federated L aggregation, while Θk remains resident on the client device Consequently, the total local loss function for client k becomes: to preserve personalized features. Moreover, the rank r for Ltotal = Lt (Dk ; ΘG , ΘL (15) k ) + λ o Lo , both global and local adapters is kept consistent to maintain architectural symmetry. Consequently, the total number of where λo is a hyperparameter controlling the strength of the trainable parameters per client remains identical to standard disentanglement constraint. Regularization via orthogonality

6

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKS

ensures that the local adapters focus exclusively on domainspecific nuances that cannot be captured by the global model. Moreover, the orthogonality constraint facilitates more stable federated aggregation by reducing the variance of local updates in the global parameter space. This regularizer also sharpens the global-local interpretation behind the Q/K versus V/O split: when query/key updates already explain a shared attention relation, the penalty discourages value/output adapters from redundantly encoding the same direction, forcing them to capture residual client-specific semantics instead [61]. In gradient terms, Eqs. 13 and 14 project each update away from directions already occupied by its counterpart, thereby reducing the cosine overlap between global-shared and local-private descent steps before aggregation. As a result, the selective aggregation step operates on a subspace whose cross-client bias is controlled, which is the quantity explicitly bounded in our convergence analysis below.

corrupted by highly specialized local features that do not generalize across the client population. E. Algorithm Description The complete training procedure for PFAdapter is detailed in Alg. 1. Initial steps involve the initialization of global parameters ΘG,0 and local parameters ΘL,0 for each client. k In each communication round, selected clients perform local updates using the combined loss function defined in Eq. 15. Following local training, only the global components are synchronized. Iterative synchronization continues until convergence or for a fixed number of rounds T . The algorithm ensures that local expertise is preserved while global knowledge is shared efficiently. For consistency with Eqs. 6 and 7, we use ΘG for the server-synchronized query/key adapters and ΘL k for the client-specific value/output adapters throughout the pseudocode. A client-side gradient step therefore takes the form:

L,t D. Selective Aggregation and Communication Efficiency (ΘG,t+1 , ΘL,t+1 ) = (ΘG,t k k k , Θk ) (18) L,t Selective aggregation protocols in PFAdapter significantly − η∇(ΘG ,ΘL ) Ltotal (ΘG,t k , Θk ). k reduce communication overhead while maintaining high performance. In each round t, the server only collects and averages For cold-start personalization, a newly joined client knew the global-shared parameters ΘG,t+1 from the participating does not participate in the federated rounds used to learn k ΘG,T . After server-side training finishes, the final global clients. The aggregation rule is defined as: X query/key adapters ΘG,T are broadcast to knew , while the n P k ΘG,t+1 = ΘG,t+1 . (16) value/output adapters remain client-private and are initialized k j∈St nj k∈St locally as in the standard training phase. The client is then evaluated at step 0 (zero-shot transfer with no local updates) Meanwhile, the local parameters ΘL k are updated locally and never transmitted to the server. Let Ctotal denote the and after a small number of local adaptation rounds using communication cost of standard federated LoRA tuning, where only its private data, matching the protocol visualized in all adapter parameters are synchronized. The communication Fig. 2(d) and the personalization setting considered in federated adaptation work [36], [37]. Alg. 1 also has four stages explicitly: cost of PFAdapter, denoted as CP F ed , is given by: server initialization, client-side local update, upload of only L X X the global branch, and server aggregation. As a result, the out Pall = r(din l,m + dl,m ), synchronization boundary becomes visually explicit and the l=1 m∈{q,k,v,o} earlier ambiguity about whether private value/output adapters L X X (17) in out are ever transmitted is removed, matching the communicationPPF = r(dl,m + dl,m ), accounting motivation of LoRA-based federated tuning [38]. l=1 m∈{q,k} (round)

CP F ed

= 2bPPF =

2 (round) PPF (round) C ≈ Ctotal , Pall total 4

din l,m

dout l,m

F. Theoretical Analysis

Convergence analysis of PFAdapter can be established unHere r is the LoRA rank [32], and are the input/output dimensions of projection m at layer l, and b is the der standard assumptions of smoothness and bounded variance. transmitted bytes per parameter. Because the four self-attention Let F (ΘG , {ΘL k }) denote the global objective function. Given projections in the deployed MLLM use the same LoRA rank that the orthogonality regularization is a smooth function of the and matched hidden dimensions, synchronizing only {q, k} parameters, the local updates follow a descent direction for the yields the exact parameter-count ratio PPF /Pall = 2/4 = 0.5. regularized objective. Furthermore, the selective aggregation The measured traffic in our implementation is therefore reduced of ΘG can be viewed as a block-coordinate descent step in from 617 MB/round to 315 MB/round, corresponding to the parameter space. We make the non-IID setting explicit 30.85 GB versus 15.75 GB over 50 rounds, with the small through four assumptions: (A1) each local objective Fk is LF deviation from an ideal 50.0% explained by serialization and smooth; (A2) stochastic gradients satisfy E∥gkt − ∇Fk ∥2 ≤ σ 2 and ∥∇F (A3) client heterogeneity is bounded packet rounding overhead. Furthermore, the preservation of ΘL k PKk ∥ ≤ G; G 1 G L 2 2 ensures that the model retains its personalized expertise across by K ∥∇F (Θ , ΘL k k ) − ∇F (Θ , {Θj })∥ ≤ δ ; and k=1 communication rounds, mitigating the negative effects of weight (A4) the selective-aggregation bias and orthogonality gradient washing. Given the massive scale of MLLM backbones, such are bounded as ∥bt ∥ ≤ β and ∥∇Lo ∥ ≤ Ho . Assumption reductions in communication traffic are critical for deployment (A3) does not require IID data; it only requires the crossin resource-constrained environments. Moreover, the selective client drift induced by non-IID partitions to remain bounded, aggregation strategy prevents the global model from being which is the regime probed by the Dirichlet-α experiments in

AUTHOR et al: PFADAPTER: HIERARCHICAL ADAPTER DECOMPOSITION FOR PFM

Algorithm 1: PFAdapter Training Protocol Input:: Local datasets {Dk }K k=1 , rounds T , epochs E, rate η, weight λo K Output:: Personalized parameters {ΘL k }k=1 and global ΘG Server initialization: initialize shared query/key adapters ΘG,0 and each client’s private value/output adapters ΘL,0 k for round t = 0, 1, . . . , T − 1 do Server broadcast: select participating clients St and transmit ΘG,t to all k ∈ St for each client k ∈ St in parallel do Client k local update: set ΘG,t,0 ← ΘG,t and k L,t,0 keep Θk private on device

7

and telescoping derivation, is provided in the supplementary material, Sec. II and Sec. V-A. The assumptions and theorem statement are placed next to the method definition so that the convergence guarantee remains visible where the selectiveaggregation mechanism is introduced, following standard nonIID FL analyses that separate stochastic variance from clientdrift terms [29], [59]. V. E XPERIMENTS A. Experimental Setup

Datasets and evaluation protocols. We evaluate PFAdapter on four diverse multimodal benchmarks to assess its effectiveness across medical imaging and social media domains. VQARAD [15] comprises 3,515 question-answer pairs across 315 for epoch e = 1, . . . , E do radiology images, emphasizing specialized clinical reasoning Sample batch B ∼ Dk and medical domain knowledge. SLAKE [16] provides 14,028 G,t,e−1 L,t,e−1 samples with 642 images in a bilingual medical VQA setting, Compute Lt (B; Θk , Θk ) offering more complex semantic structures for evaluation. HateCompute Lo via Eq. 11 G,t,e G,t,e−1 ful Memes [18] contains 10,000 multimodal entries requiring Θk ← Θk − η∇ΘG (Lt + λo Lo ) joint text-image processing for hate speech detection in social L,t,e L,t,e−1 media contexts. CrisisMMD [19] consists of 16,080 image-text Θk ← Θk − η∇ΘL (Lt + λo Lo ) pairs from disaster scenarios, categorized into humanitarian end G,t+1 G,t,E assistance tasks including damage severity assessment and Θk ← Θk resource needs identification. Performance metrics included L,t+1 L,t,E Θk ← Θk Accuracy and F1-score for all datasets, with Area Under Upload: client k sends only ΘG,t+1 to the k the ROC Curve (AUC) additionally reported for the binary server; ΘL,t+1 is never uploaded k classification task on Hateful Memes. Weighted F1-scores were end employed to account for class imbalance inherent in medical Server aggregation: update datasets. All experiments followed the Aligned modal scenario P ΘG,t+1 ← k∈St P nk n ΘG,t+1 with Dirichlet concentration parameter α = 0.5 to simulate k j j∈St moderate data heterogeneity, matching the evaluation protocol end established in prior federated MLLM work [36]. Detailed Termination: return the final shared ΘG,T and preprocessing, partition reuse, and local-test evaluation protocol L,T K personalized {Θk }k=1 notes are provided in the supplementary material, Sec. V-B, for VQA-RAD [15], SLAKE [16], Hateful Memes [18], and CrisisMMD [19]. Secs. V-A and V-B [29], [59]. Theoretical results indicate that Baseline methods and implementation configuration. To √ the framework achieves a convergence rate of O(1/ T ) for benchmark PFAdapter against established federated optinon-convex objectives, matching the performance of standard mization strategies, we select five representative methods federated learning while providing superior personalization spanning adaptive learning rates and momentum-based agguarantees. gregation. Zero-shot performance of the pre-trained base Theorem 1 (Convergence of√PFAdapter). Assume (A1)–(A4) model served as the lower bound, while Local-only training above and choose ηt = c/ T with 0 < c ≤ 1/LF . For the provided an upper bound for client-specific personalization iterates generated by Alg. 1, the averaged stationarity measure without any knowledge sharing. FedYogi [36] implemented adaptive moment-based federated averaging with per-coordinate satisfies: learning rates, representing the strongest baseline in prior  T −1  2 2 work. FedAdam [62] employed the Adam optimizer [63] 1 X σ C 1 E ∇F (ΘG,t , {ΘL,t ≤ √ + C 2 δ 2 + C3 k }) with server-side momentum accumulation, while FedAvgM T t=0 |St | T [59] combined momentum acceleration with standard FedAvg + C4 λ2o Ho2 + C5 β 2 , updates [8]. FedAdagrad [62] utilized Adagrad’s adaptive (19) learning rate strategy for federated optimization. The base where C1 = 2(F 0 − F ⋆ )/c, C2 = cLF , architecture employed MiniCPM-V-2_6-int4, a quantized C3 = cLF , C4 = 2, and C5 = 2. Consequently, multimodal large language model with Qwen2 backbone. LoRA L,t min0≤t<T E[∥∇F (ΘG,t , {Θk })∥2 ] = O(T −1/2 ) whenever [32] was applied to self-attention projection matrices with rank the non-IID drift δ 2 , orthogonality-gradient magnitude Ho , r = 8 and scaling factor αLoRA = 16. Local training utilized and selective-aggregation bias β remain bounded. AdamW optimizer with learning rate 2 × 10−5 and cosine A detailed proof sketch, including the descent inequality annealing over 50 communication rounds. Each client executed

8

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKS

TABLE I: Performance comparison across datasets on the aligned-modal scenario (α = 0.5). Learned federated baselines are reported as mean ± standard deviation over three seeds, and best results are highlighted in bold. Method

Zero-shot Local FedYogi [36] FedAdam [62] FedAvgM [59] FedAdagrad [62] pFedLoRA [37] PFAdapter

VQA-RAD

SLAKE

Hateful Memes

CrisisMMD

Comm.

Acc (%)↑

F1 (%)↑

Acc (%)↑

F1 (%)↑

Acc (%)↑

AUC (%)↑

Acc (%)↑

F1 (%)↑

(MB/R)

56.98 59.64 60.53±0.88 60.31±1.42 58.98±1.35 60.54±0.55 61.74±0.74 62.83±0.62

52.4 56.8 57.9±0.61 57.5±0.75 55.8±0.67 57.8±0.31 59.1±0.55 60.1±0.44

64.95 61.63 58.67±0.55 56.74±0.52 58.47±1.06 55.83±0.32 59.12±0.43 60.08±0.38

61.2 58.9 55.4±0.33 53.8±0.41 55.1±0.47 52.9±0.60 56.8±0.37 57.6±0.35

66.57 66.39 71.41±0.82 72.56±1.23 72.18±1.71 73.76±0.89 74.38±0.68 75.23±0.51

65.89 67.12 72.48±1.46 73.24±2.43 72.94±1.61 73.34±1.99 74.92±0.72 75.63±0.63

24.20 47.34 60.82±0.95 59.12±0.56 56.87±0.42 60.43±0.60 61.48±0.58 62.49±0.46

22.8 45.2 58.6±0.58 57.1±0.41 54.8±0.53 58.3±0.23 59.4±0.49 60.3±0.40

617 617 617 617 617 315

TABLE II: Accuracy (%) and communication cost (MB/R) of ablation study on component contribution.

performance degradation. Removing the orthogonality regularization loss results in accuracy decreases of 1.5% on VQA-RAD, 1.5% on SLAKE, and 0.9% on Hateful Memes, Configuration VQA-RAD SLAKE Hateful Comm. validating that knowledge disentanglement between global w/o Hierarchical Split 60.53 58.67 72.50 617 and local adapters is crucial for effective personalization. w/o Orthogonality Loss 61.3 58.6 74.3 315 Disabling selective aggregation while maintaining orthogonality w/o Selective Aggregation 60.9 58.2 73.8 617 PFAdapter (Full) 62.8 60.1 75.2 315 constraints leads to performance drops of 1.9%, 1.9%, and 1.4% across the three datasets, respectively, while simultaneously one local epoch with batch size 1 and gradient accumulation doubling communication overhead to 617 MB per round. steps of 16. Federated training sampled 2 clients per round Most significantly, eliminating the hierarchical split entirely from a total population of 10. Orthogonality regularization (equivalent to FedYogi) causes the largest degradation, with weight was set to λo = 0.1 based on validation performance. accuracy decreases of 2.27% on VQA-RAD, 1.43% on SLAKE, Experiments were conducted on a single Nvidia L60 GPU and 2.7% on Hateful Memes. Experimental results confirm with 48GB VRAM, utilizing 8-bit quantization and gradient that: i) explicit disentanglement prevents local adapters from redundantly learning global knowledge, ii) selective aggregation checkpointing for memory efficiency. preserves client-specific expertise through private value and output projections, and iii) hierarchical decomposition enables B. Performance Evaluation more nuanced control over knowledge sharing compared to Main results on aligned modal scenario. Quantitative com- monolithic adapter synchronization. Detailed weight-washing parisons between PFAdapter and state-of-the-art federated diagnostics and cross-client attention-map visualizations are learning baselines are presented in Table I. The proposed provided in the supplementary material, Sec. V-D. method consistently achieves superior performance across all four multimodal datasets while simultaneously reducing Decomposition strategy analysis. Different module assigncommunication overhead by nearly 50%. On the medical VQA- ment strategies lead to varying performance outcomes dependRAD dataset, PFAdapter attains 62.83% overall accuracy, ing on which projections are designated for global versus outperforming the strongest baseline FedYogi by 2.30% and local adaptation. Assigning query (qp ) and key (kp ) projections demonstrating the effectiveness of hierarchical decomposition to the global set while keeping value (vp ) and output (op ) for clinical reasoning tasks. Performance gains are more projections local yields the optimal configuration, achieving pronounced on SLAKE, where PFAdapter achieves 60.08% 62.8% accuracy on VQA-RAD and 60.1% on SLAKE, as accuracy compared to 58.67% for FedYogi, representing a detailed in Table III. Alternative decomposition strategies result 1.41% improvement. For social media multimodal classification, in varying degrees of performance degradation. Specifically, PFAdapter obtains 75.63% AUC on Hateful Memes, sur- assigning qp and vp to global aggregation reduces accuracy passing FedYogi (72.48%) by 3.15%, and achieves 62.49% by 1.3% on VQA-RAD, suggesting that value projections are accuracy on CrisisMMD, outperforming FedYogi (60.82%) inherently more task-specific and should remain personalized. by 1.67%. Communication analysis reveals that PFAdapter Restricting global synchronization to only qp leads to a transmits only 315 MB per round compared to 617 MB for more substantial 2.6% accuracy decrease, indicating that baseline methods, achieving a 48.9% reduction in bandwidth key projections also capture essential cross-client structural requirements through selective aggregation of query and key information. Conversely, assigning three modules (qp , kp , vp ) projection modules only. Table I reports macro-averaged client- to the global set reduces communication less substantially local test scores in the aligned setting together with mean ± and sacrifices 3.1% accuracy, demonstrating the diminishing standard deviation over three seeds for the learned federated returns of excessive global synchronization. The projectionbaselines. Larger-client scaling, fairness, and claim-scope level evidence directly matches the mechanism in Eqs. 8 and 9: details are provided in the supplementary material, Secs. V-C removing kp from the global set degrades cross-client attention and V-D. alignment, while promoting vp to the global set erodes the local Ablation study on component contributions. To assess semantic capacity needed under non-IID supervision [55], [57]. the contribution of individual components, we systematically The best Q/K-global and V/O-local split therefore emerges removed each technical module and measured the resulting not as a heuristic partition, but as the configuration that best

AUTHOR et al: PFADAPTER: HIERARCHICAL ADAPTER DECOMPOSITION FOR PFM

Global

Local

VQA-RAD

Comm.

qp only qp , kp , vp qp , vp qp , kp

kp , vp , op op only kp , op vp , op

60.2 59.7 61.5 62.8

155 469 315 315

FedYogi

64

FedAvgM

PFAdapter

FedYogi

Global Training Loss

VQA-RAD Accuracy (%)

TABLE III: Accuracy (%) and communication cost (MB/R) of different global-local decomposition strategies on VQA-RAD.

62 60 58 56

FedAvgM

10

20

30

40

(a) Accuracy vs. Rounds

50

Full Tuning FedYogi [36] FedAvgM [59] PFAdapter

2.0 1.5 1.0

10

20

30

Train Time GPU-h/R Inf. Time Peak VRAM Total Comm. 45.2 10.3 10.6 10.8

0.753 0.172 0.177 0.180

125.4 14.7 14.2 14.8

42.5 15.7 15.5 15.8

125.4 30.85 30.85 15.75

PFAdapter

2.5

0

TABLE IV: Comparison of train time (min/round), GPUhours per round, peak VRAM (GB), inference time, and total communication (GB) efficiency. Method

0.5 0

9

40

50

(b) Training Loss vs. Rounds

Fig. 3: Convergence analysis on VQA-RAD. (Left) Accuracy vs. communication rounds. (Right) Training loss reduction. preserves relation sharing and client-specific reconstruction simultaneously. Parameter sensitivity and robustness analysis. Comprehensive sensitivity analysis across three critical hyperparameters reveals optimal configuration ranges and robustness characteristics. Fig. 2(a) demonstrates that orthogonality weight λo = 0.1 yields optimal performance for VQA-RAD (62.8%) and Hateful Memes (75.6% AUC), while SLAKE achieves peak accuracy at λo = 0.05 (60.1%). Increasing λo beyond these optimal values to 0.5 or 1.0 causes gradual degradation, as excessive orthogonality constraints restrict local adapters from capturing client-specific knowledge. Fig. 2(b) examines the trade-off between accuracy and communication efficiency across different global-local decomposition ratios, where the 50:50 configuration achieves optimal balance with 62.8% accuracy at 50% communication cost. Fig. 2(c) evaluates robustness under varying data heterogeneity levels, measured by Dirichlet parameter α ranging from 0.1 (extreme non-IID) to 5.0 (nearly IID). Under high heterogeneity (α = 0.1), PFAdapter achieves 61.2% accuracy compared to 57.8% for FedYogi, representing a 3.4% improvement that narrows to 0.9% under low heterogeneity conditions, confirming that hierarchical decomposition provides greater benefits when client distributions diverge more significantly. Fig. 2(d) demonstrates cold-start adaptation capability for newly joined edge devices, wherein PFAdapter achieves 59.8% zero-shot accuracy compared to 56.98% for FedYogi, and reaches 64.8% after only 5 local tuning rounds through effective knowledge transfer from pre-aggregated global components. Detailed heterogeneitytheory interpretation, cold-start protocol, fairness statistics, rank ablation, and local-epoch discussion are provided in the supplementary material, Sec. V-D. The full sensitivity visualization is provided in the supplementary material, Fig. 2. Convergence behavior and training dynamics. Orthogonalitydriven decomposition accelerates training convergence by reducing parameter conflicts during federated aggregation. PFAdapter exhibits significantly faster convergence over 50 communication rounds on VQA-RAD, reaching 61.2% accuracy by round 20 compared to 59.1% for FedYogi

and 58.2% for FedAvgM, as illustrated in Fig. 3(a). Faster convergence can be attributed to the orthogonality constraint, which reduces parameter conflicts between local and global updates and facilitates more stable aggregation. In particular, the gradients in Eqs. 13 and 14 penalize overlap between the global and local update subspaces, so the server aggregates less mutually contradictory information from different clients at each round. Fig. 3(b) demonstrates that PFAdapter achieves consistently lower training loss throughout the optimization process, with final loss of 0.35 compared to 0.82 for FedYogi and 1.15 for FedAvgM, indicating a better-optimized loss landscape and more efficient utilization of the parameter budget. Computational and communication efficiency. Selective aggregation substantially reduces network resource requirements while maintaining computational efficiency comparable to baseline LoRA methods. Compared to full model fine-tuning requiring 45.2 minutes per round and 42.5 GB peak VRAM, LoRA-based methods reduce training time by approximately 75% and memory usage by 63%, as summarized in Table IV. PFAdapter incurs a marginal 5% increase in training time (10.8 vs. 10.3 min/round for FedYogi) due to the additional orthogonality loss computation, but achieves a 48.9% reduction in total communication volume over 50 rounds (15.75 GB vs. 30.85 GB for FedYogi). Communication savings are achieved by transmitting only the global-shared query and key projection adapters (2 of 4 LoRA modules), while preserving clientspecific value and output projections locally. Peak VRAM consumption of 15.8 GB remains comparable to baseline LoRA methods, making PFAdapter suitable for deployment on single-GPU systems without requiring specialized distributed computing infrastructure. Comprehensive heterogeneity analysis. Hierarchical decomposition demonstrates particularly strong advantages when edge devices exhibit severe data distribution mismatches. Under high label skew (α = 0.1), PFAdapter achieves a substantial +3.4% improvement over FedYogi, while the margin decreases to +0.9% under near-IID conditions (α = 5.0), as shown in Table XII, thereby validating that hierarchical decomposition provides greatest benefits precisely when heterogeneity poses the most significant challenges. Missing modal scenarios with higher missing rates (β = 50%) yield +2.9% improvement compared to +2.3% at β = 30%, demonstrating that local adaptation through vp and op parameters effectively compensates for modality-specific distribution shifts. Cross-modal and hybrid scenarios maintain consistent advantages (+1.57% and +1.26%, respectively), confirming that the decomposition strategy generalizes across diverse heterogeneity types. The controlled split-policy note and the full heterogeneity table are provided in the supplementary material, Table XII.

10

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKS

Robustness across multimodal heterogeneity scenarios. Selective parameter aggregation enables PFAdapter to maintain consistent performance advantages across diverse challenging scenarios. In the Missing Modal scenario where 50% of clients lack either image or text modalities, PFAdapter achieves 76.8% AUC on Hateful Memes and 56.5% F1 on CrisisMMD, outperforming FedYogi by 1.68% and 2.72%, respectively, as visualized in Fig. 3(a). Subsequently, Fig. 3(b) evaluates generalization when image-dominant and text-dominant clients coexist (I-5:T-5 split), with PFAdapter maintaining superior performance across all three metrics. Moreover, Fig. 3(c) combines aligned (p = 70%) and missing modal conditions, where the local adaptation capability of vp and op parameters proves particularly beneficial. Finally, Fig. 3(d) demonstrates graceful degradation under increasing Gaussian noise levels, with PFAdapter maintaining 54.6% accuracy at 20% noise compared to 48.5% for FedYogi, achieving a 6.1% absolute advantage. Collectively, experimental results validate that selective aggregation and hierarchical decomposition provide inherent robustness to diverse real-world data distribution challenges.

ploying multimodal large language models as intelligent agents across heterogeneous edge networks. Hierarchical LoRA decomposition was introduced to explicitly separate adapter parameters into global-shared and local-private components based on functional roles of self-attention modules. Query and key projections are assigned to global synchronization across the federated network, whereas value and output projections remain localized for edge-specific adaptation. Orthogonality regularization enforces effective knowledge disentanglement between network-synchronized and edge-retained parameters. Selective aggregation protocols transmit only global-shared components, reducing communication overhead by nearly 50% while preserving edge-specific expertise. Future research directions include exploring dynamic decomposition ratios adapted to network conditions, layer-adaptive splitting strategies for heterogeneous edge devices, and integration with emerging 6G network to further enhance agentic AI deployment across next-generation communication systems.

C. Discussion and Limitations Superior performance of PFAdapter stems from three interrelated factors: i) Structural decomposition recognizes functional heterogeneity within self-attention mechanisms, wherein query and key projections encode cross-attention patterns amenable to global sharing across the network, whereas value and output projections modulate edge-specific representations; ii) Orthogonality regularization prevents redundant learning by enforcing local adapters to capture orthogonal directions in parameter space relative to global knowledge, leading to more efficient parameter budget utilization across distributed agents; and iii) Selective aggregation reduces the weight washing effect common in federated MLLM fine-tuning, wherein clientspecific adaptations become diluted through uniform parameter mixing. Nevertheless, several limitations merit discussion regarding deployment in heterogeneous edge networks: i) Optimal decomposition ratios may vary depending on the degree of local-global divergence across network nodes, with 50:50 split providing the best balance for moderate heterogeneity (α = 0.5) yet potentially requiring adaptation for extreme distribution shifts; ii) Current framework applies uniform decomposition across all transformer layers, whereas layer-specific ratios based on sensitivity analysis could further enhance performance for edge devices with varying computational capabilities; and iii) Although orthogonality regularization effectively disentangles knowledge, Frobenius norm constraints may not fully capture complex non-linear dependencies between global and local parameters in highly dynamic network environments. Future work could explore learnable decomposition ratios, layeradaptive splitting strategies tailored to network topology, and more sophisticated disentanglement metrics based on information-theoretic measures suitable for agentic AI systems.

[1] T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 41, no. 2, pp. 423–443, 2018. [2] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable visual models from natural language supervision,” in ICML. PMLR, 2021, pp. 8748–8763. [3] S. Liu, Y. Shen, J. Yuan, C. Wu, and R. Yin, “Storage-aware joint user scheduling and bandwidth allocation for federated edge learning,” IEEE Trans. Cogn. Commun. Netw., vol. 11, no. 1, pp. 581–593, 2025. [4] B. Mao, Y. Wu, J. Liu, H. Guo, J. Wang, and N. Kato, “Optimizing secrecy rate for federated learning model aggregation with intelligent reflecting surface toward 6g ubiquitous intelligence,” IEEE Trans. Cogn. Commun. Netw., vol. 11, no. 2, pp. 1258–1267, 2025. [5] J. Liu, H. Yu, G. Fang, J. Wu, L. Gong, X. Zhu, Y. Liu, and P. Sun, “Semantic prototype learning for cross-domain activity recognition,” in IEEE ICME, 2026. [6] Q. Yang, Y. Liu, T. Chen, and Y. Tong, “Federated machine learning: Concept and applications,” ACM Trans. Intell. Syst. Technol., vol. 10, no. 2, pp. 1–19, 2019. [7] Y. Gao, Z. Ye, Y. Xiao, M. Xiao, and W. Xiang, “Learner referral for cost-effective federated learning over hierarchical iot networks,” IEEE Trans. Cogn. Commun. Netw., vol. 11, no. 3, pp. 1830–1844, 2025. [8] B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, “Communication-efficient learning of deep networks from decentralized data,” in AISTATS, 2017, pp. 1273–1282. [9] G. Cheng, P. Li, B. Tan, R. Yu, Y. Wu, and M. Pan, “Snowball effect in federated learning: An approach of exponentially expanding structures for optimizing the training efficiency,” IEEE Trans. Cogn. Commun. Netw., vol. 11, no. 3, pp. 1803–1817, 2025. [10] Y. Wu, C. Tian, J. Li, H. Sun, K. Tam, Z. Zhou, H. Liao, Z. Guo, L. Li, and C. Xu, “A survey on federated fine-tuning of large language models,” arXiv preprint arXiv:2503.12016, 2025. [11] W. He, H. Yao, X. Ren, T. Ouyang, Z. Xiong, Y. He, and Y. Liu, “Dual-circulation generative ai for optimizing resource allocation in multi-granularity heterogeneous federated learning,” IEEE Trans. Cogn. Commun. Netw., vol. 11, no. 2, pp. 817–831, 2025. [12] J. Liu, G. Fang, L. Teng, L. Qian, B. Hu, and P. Sun, “Enhancing collaborative learning efficiency via control-theoretic merit gating in federated networks,” in IEEE ICC, 2026. [13] X.-Y. Liu, R. Zhu, D. Zha, J. Gao, S. Zhong, M. White, and M. Qiu, “Differentially private low-rank adaptation of large language model using federated learning,” ACM Trans. Manage. Inf. Syst., vol. 16, no. 2, pp. 11:1–11:24, 2025. [14] Y. Deng, Y. Zhang, and J. Chen, “Federated fine-tuning of large language models: A comprehensive survey,” arXiv preprint arXiv:2312.10845, 2023. [15] J. J. Lau, S. Gayen, A. Ben Abacha, and D. Demner-Fushman, “A dataset of clinically generated visual questions and answers about radiology images,” Sci. Data, vol. 5, no. 1, p. 180251, 2018.

VI. C ONCLUSION In this paper, we presented PFAdapter, a communicationefficient personalized federated learning framework for de-

R EFERENCES

AUTHOR et al: PFADAPTER: HIERARCHICAL ADAPTER DECOMPOSITION FOR PFM

[16] B. Liu, L.-M. Zhan, L. Xu, L. Ma, Y. Yang, and X.-M. Wu, “Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering,” in IEEE ISBI, 2021, pp. 1650–1654. [17] J. Liu, L. Gong, J. Guo, J. Wu, L. Sun, Y. Bi, K. Patwari, B. Chen, L. Zhang, W. Zhou, Y. Liu, X. Zhu, C.-N. Chuah, and B. Rajaratnam, “Multimodal large language models in medicine and nursing: A survey,” 2025. [18] D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, and D. Testuggine, “The hateful memes challenge: Detecting hate speech in multimodal memes,” in NeurIPS, vol. 33, 2020, pp. 2611–2624. [19] F. Alam, F. Ofli, and M. Imran, “Crisismmd: Multimodal twitter datasets from natural disasters,” in ICWSM, vol. 12, no. 1, 2018. [20] A. Fallah, A. Mokhtari, and A. Ozdaglar, “Personalized federated learning with theoretical guarantees: A model-agnostic meta-learning approach,” in NeurIPS, vol. 33, 2020, pp. 3557–3568. [21] C. T. Dinh, N. Tran, and J. Nguyen, “Personalized federated learning with moreau envelopes,” in NeurIPS, vol. 33, 2020, pp. 21 394–21 405. [22] Y. Zhang, K. Guo, Z. Lu, Y. Wang, and J. Liang, “Personalized federated learning via dual-prompt optimization and cross fusion,” arXiv:2506.21144, 2025. [23] G. Wilson and D. J. Cook, “A survey of unsupervised deep domain adaptation,” ACM Trans. Intell. Syst. Technol., vol. 11, no. 5, pp. 1–46, 2020. [24] X. He, H. Huang, B. Xie, C. Wang, R. Li, H. Cui, and Z. Zheng, “Hivefl: Gan-empowered semi-asynchronous federated learning with selfdetermining clients,” IEEE Trans. Cogn. Commun. Netw., vol. 11, no. 2, pp. 805–816, 2025. [25] Y. Cheng, W. Zhang, Z. Zhang, C. Zhang, S. Wang, and S. Mao, “Toward federated large language models: Motivations, methods, and future directions,” IEEE Commun. Surveys Tuts., vol. 27, no. 4, pp. 2733–2764, 2025. [26] Y. Jia, Z. Huang, J. Yan, Y. Zhang, K. Luo, and W. Wen, “Joint optimization of resource allocation and data selection for fast and costefficient federated edge learning,” IEEE Trans. Cogn. Commun. Netw., vol. 11, no. 1, pp. 594–606, 2025. [27] J. Liu, Y. Liu, J. Lin, J. Li, L. Cao, P. Sun, B. Hu, L. Song, A. Boukerche, and V. C. Leung, “Networking systems for video anomaly detection: A tutorial and survey,” ACM Comput. Surv., vol. 57, no. 10, pp. 270:1– 270:37, Apr. 2025. [28] Y. Liu, J. Liu, C. Li, R. Xi, W. Li, L. Cao, J. Wang, L. T. Yang, J. Yuan, and W. Zhou, “Anomaly detection and generation with diffusion models: A survey,” Jun. 2025. [29] T. Li, A. K. Sahu, M. Zaheer, M. Sanjabi, A. Talwalkar, and V. Smith, “Federated optimization in heterogeneous networks,” in MLSys, vol. 2, 2020, pp. 429–450. [30] Y. Zhang and Q. Yang, “A survey on multi-task learning,” IEEE Trans. Knowl. Data Eng., vol. 34, no. 12, pp. 5586–5609, 2021. [31] Y. Yao, T. Li, Z. Wang et al., “Federated large language models: Current progress and future directions,” arXiv preprint arXiv:2409.15723, 2025. [32] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” in ICLR, 2021. [33] W. Y. B. Lim, N. C. Luong, D. T. Hoang, Y. Jiao, Y.-C. Liang, Q. Yang, D. Niyato, and C. Miao, “Federated learning in mobile edge networks: A comprehensive survey,” in IEEE Commun. Surv. Tutor., vol. 22, no. 3. IEEE, 2020, pp. 2031–2063. [34] Y. Liu, D. Yang, Y. Wang, J. Liu, J. Liu, A. Boukerche, P. Sun, and L. Song, “Generalized video anomaly event detection: Systematic taxonomy and comparison of deep models,” ACM Comput. Surv., vol. 56, no. 7, pp. 189:1–189:38, 2024. [35] J. Bai, D. Chen, B. Qian, L. Yao, and Y. Li, “Federated fine-tuning of large language models under heterogeneous tasks and client resources,” in NeurIPS, vol. 37, 2024, pp. 14 457–14 483. [36] B. Xu, X. Shu, H. Mei, G. Xie, B. Fernando, and J. Tang, “Fedmllm: Federated fine-tuning mllm on multimodal heterogeneity data,” arXiv:2411.14717, 2025. [37] L. Yi, H. Yu, G. Wang, X. Liu, and X. Li, “pFedLoRA: Modelheterogeneous personalized federated learning with lora tuning,” arXiv:2310.13283, 2024. [38] Z. Wang, Z. Shen, Y. He, G. Sun, H. Wang, L. Lyu, and A. Li, “Flora: Federated fine-tuning large language models with heterogeneous low-rank adaptations,” in NeurIPS, vol. 37, 2024, pp. 22 513–22 533. [39] M. G. Arivazhagan, V. Aggarwal, A. K. Singh, and S. Choudhary, “Federated learning with personalization layers,” arXiv:1912.00818, 2019. [40] L. Collins, H. Hassani, A. Mokhtari, and S. Shakkottai, “Exploiting shared representations for personalized federated learning,” in ICML, 2021, pp. 2089–2099. [41] J. Liu, Z. Ma, H. Yu, B. Ju, W. Yang, C. Li, B. Hu, and L. Song, “Collaborative adaptive curriculum for progressive knowledge distillation,” in IEEE ICME, 2026.

11

[42] Z. Li, S. Wu, L. Li, and S. Zhang, “Energy-efficient split learning for fine-tuning large language models in edge networks,” IEEE Netw., vol. 7, no. 3, pp. 176–180, 2025. [43] S. Yang, Z. Chen, Y. Lin, X. Chen, G. Cai, H. Yu, P. Wu, and Q. Yang, “A survey on vision-language models for multimodal federated learning tasks,” techRxiv preprint techRxiv:175624545.56457516/v2, 2025. [44] Y. Tan, G. Long, L. Liu, T. Zhou, Q. Lu, J. Jiang, and C. Zhang, “Fedproto: Federated prototype learning across heterogeneous clients,” AAAI, vol. 36, no. 8, pp. 8432–8440, 2022. [45] J. Oh, S. Kim, and S.-Y. Yun, “Fedbabu: Toward enhanced representation for federated personalization,” in ICLR, 2021. [46] Z. Wang, Z. Shen, Y. He, G.-h. Sun, H. Wang, L. Lyu, and A. Li, “Not all clients are equal: Personalized federated learning on heterogeneous multi-modal clients,” in AAAI, 2024. [47] D. Kim and T. Kim, “Missing modality prediction for unpaired multimodal learning via joint embedding of unimodal models,” in ECCV, 2025, pp. 171–187. [48] N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Brunslo, M. Dawson-Haggerty, P. Piponi, V. Vasudevan, A. Kalenichenko, and J. Uszkoreit, “Parameterefficient transfer learning for nlp,” ICML, 2019. [49] B. Lester, R. Al-Rfou, and N. Constant, “The power of scale for parameterefficient prompt tuning,” in EMNLP, 2021. [50] X. L. Li and P. Liang, “Prefix-tuning: Optimizing continuous prompts for generation,” in ACL, 2021. [51] Y. Chen, Q. Fu, G. Fan, L. Du, J.-G. Lou, S. Han, D. Zhang, Z. Li, and Y. Xiao, “Hadamard adapter: An extreme parameter-efficient adapter tuning method for pre-trained language models,” in ACM CIKM, 2023, pp. 276–285. [52] Y. Yan, Q. Yang, S. Tang, and Z. Shi, “Federa:efficient fine-tuning of language models in federated learning leveraging weight decomposition,” arXiv:2404.18848, 2024. [53] D. P. Nguyen, J. P. Muñoz, T. Roosta, and A. Jannesari, “Federated multimodal learning with dual adapters and selective pruning for communication and computational efficiency,” in IEEE CCGrid, 2025, pp. 01–10. [54] J. Wu, S. Zhang, M. Hou, Z. Wang, W. Chen, Z. Tian, F. R. Yu, and V. C. M. Leung, “Clip-ae: A multi-modal unsupervised images enhancement method based on high-order adaptive curve for visual disbalance defects,” IEEE Trans. Multimedia, vol. 27, pp. 4269–4283, 2025. [55] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” NeurIPS, 2017. [56] J. Liu, Y. Du, Y. Liu, Z. Wang, P. Sun, and V. C. M. Leung, “Projecting to consensus: Communication-efficient collaborative learning across heterogeneous networks,” in IEEE ICC, 2026. [57] T. Mickus, E. Ponti, and G. Rätsch, “The role of value projection in multi-task learning,” ICLR, 2024. [58] J. Koo, J. Park, and J. Kim, “Lora-a2: Towards robust and efficient federated low-rank adaptation,” ACL, 2025. [59] T.-M. H. Hsu, H. Qi, and M. Brown, “Measuring the effects of nonidentical data distribution for federated visual classification,” arXiv preprint arXiv:1909.06335, 2019. [60] J. Liu, Z. Guo, Y. Wang, X. Zhu, Y. Du, Z. Wang, and V. C. M. Leung, “Diffusion-guided semantic consistency for multimodal heterogeneity,” in IEEE ICME, 2026. [61] C. Wu, R. Zhang, J. Liu, and L. Zhang, “Towards better orthogonality regularization with disentangled norm,” ICLR, 2023. [62] S. J. Reddi, Z. Charles, M. Zaheer, Z. Garrett, K. Rush, J. Konečný, S. Kumar, and H. B. McMahan, “Adaptive federated optimization,” in ICLR, 2020. [63] D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in ICLR, 2015.

Record · ID 366219 · SHA-256 eaa151873699a742
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.