1
Robust Decentralized Federated Distillation via Multi-Modality Knowledge Collaboration Xiao Maa , Hong Shenb , Hui Tianc , Wei Kea , Wenqi Lyua a
Faculty of Applied Sciences, Macao Polytechnic University, Macao SAR, China School of Engineering and Technology, Central Queensland University, Australia c School of Information and Communication Technology, Griffith University, Australia
arXiv:2609.07230v1 [cs.LG] 7 Sep 2026
b
Abstract—Decentralized federated learning enables clients to collaborate without centralized coordination, but most existing methods assume that all clients use the same model architecture. This requirement limits their use in practical systems where clients may adopt different models because of heterogeneous computation and resource constraints. The problem becomes more challenging under Byzantine attacks, since parameterbased robust aggregation is ineffective for heterogeneous models, especially for defense against receiver-specific attacks where malicious clients may send different manipulated predictions to different honest clients. To address this issue, we propose a robust decentralized federated distillation method that enables clients with heterogeneous models to collaborate through predictions on shared unlabeled public data. In the proposed method, each client first evaluates the received predictions in three modalities of class prediction, boundary decision, and prediction correlation. It then filters unreliable clients, assigns reliability-based weights to the retained clients, and constructs a teacher for each type of knowledge. Finally, the corresponding distillation gradients are validated using a supervised gradient computed from private data. Conflicting prediction and boundary gradients are removed, and conflicting relation gradients are suppressed before the final model update. We prove the convergence of the proposed method by showing stable local optimization for honest clients under Byzantine distillation. Particularly, we show that our method ensures a bounded Byzantine influence on both distillation gradients and individual client private gradients after cross-modality fusion, thereby enabling stable local optimization for honest clients under Byzantine distillation. Extensive experiments on CIFAR-10 and CIFAR-100 demonstrate that the proposed method improves the prediction accuracy of heterogeneous models of clients under non-IID data and Byzantine attacks. As the booming demands of federated learning in decentralized environments such as edge computing and mission-oriented UAV collaborations, our method has a great potential for adoption of DFL in unreliable real-world scenarios where clients are exposed to receiver-specific Byzantine messages of malicious predictions. Index Terms—Decentralized Federated Learning, Robustness, Knowledge Distillation
I. I NTRODUCTION EDERATED learning (FL) enables multiple clients to train models collaboratively while retaining their raw data locally. Conventional FL, however, still relies on a central server to distribute a global model and aggregate client updates[1]. This coordinator can become a communication bottleneck and a single point of failure, and it may be unavailable or undesirable in collaborations among autonomous organizations or peer devices. Decentralized federated learning
F
(DFL) removes this dependency by allowing clients to communicate directly over a graph. Decentralized stochastic optimization can distribute the communication load across neighboring nodes[2], while consensus-based FL demonstrates that serverless model optimization can be implemented through cooperation among devices[3]. These properties make DFL attractive when no trusted server exists or no participant should control a shared global model. Most DFL algorithms nevertheless inherit a restrictive assumption from parameter-aggregation FL: all clients optimize models with the same architecture and an aligned parameter space[4]. This assumption is difficult to maintain in practice because clients may differ in computation, memory, energy, latency, and hardware support. Consequently, they may deploy models with different depths, widths, operators, and parameter dimensions. Under such model heterogeneity, averaging parameters or gradients is undefined when their dimensions differ. Even when two parameter vectors have the same dimension, coordinate-wise comparison or Euclidean distance need not represent functional similarity because distinct architectures encode knowledge differently. Prior heterogeneous FL studies have therefore identified gradient misalignment and architectural dependence as central limitations of parameterspace aggregation[5], [6]. The same limitation affects many robustness rules that judge a client by the distance or coordinates of its transmitted model. Knowledge distillation (KD) provides an alternative communication space. Instead of transferring internal parameters, KD trains a student model to match the predictive behavior of a teacher through soft outputs[7]. This principle allows models with different architectures to exchange task knowledge as long as they share a common input and label space. FedMD applies prediction matching on public data to heterogeneous federated models[8]. In serverless networks, consensus-based multi-hop federated distillation exchanges predictions on shared samples to approximate collaboration in function space[9], while SD-Dist uses model outputs on a shared reference dataset to select useful peers[10]. Compared with transmitting a parameter vector whose modality depends on the local architecture, transmitting logits on a shared public reference dataset provides a compact, architecture-agnostic interface for decentralized collaboration. Moving communication from parameter space to output space does not remove adversarial risk; it changes the object
2
that an adversary can manipulate. A Byzantine client may transmit arbitrary logits, reverse class preferences, amplify selected entries, or imitate plausible but misleading predictions. Recent work on federated distillation has shown that attacks tailored to logits can substantially disrupt knowledge aggregation and that defenses designed for parameter updates do not directly transfer to this setting[11], [12]. The threat is more subtle in a peer-to-peer network. Without a server that produces one common aggregate, a Byzantine client can send different messages to different clients during the same training round. Such equivocation is permitted by standard decentralized Byzantine models and can drive honest receivers toward different references [13]. Thus, an honest client must assess the predictions available at its own receiver rather than rely on a globally consistent view of the network. Existing research addresses only parts of this combined problem. Byzantine-robust DFL methods typically screen or transform client model vectors in a shared parameter space. BRIDGE performs coordinate-wise filtering before decentralized mixing[14]; UBAR combines parameter distance with local performance evaluation[15]; and ClippedGossip and remove-then-clip aggregation bound the effect of suspicious model differences[13], [16]. These methods are effective when the exchanged vectors have compatible semantics, but their coordinate and distance operations cannot be directly applied to models with different architectures. Conversely, decentralized distillation methods support heterogeneous models by exchanging outputs, yet they are primarily designed for benign knowledge transfer rather than receiver-specific Byzantine manipulation[9], [10], [6]. Robust output aggregation has been studied in server-coordinated collaborative learning and federated distillation[17], [11], [12], but a central aggregate does not capture the local and potentially inconsistent message sets observed by different clients in a serverless network. This leaves an unresolved gap: how to achieve robust decentralized knowledge collaboration when honest clients may use heterogeneous models and observe non-IID data, while Byzantine clients can manipulate the prediction messages delivered to different receivers. To address this gap, we propose a robust decentralized federated distillation method for collaborative learning with heterogeneous local models. Instead of exchanging model parameters across heterogeneous architectures, clients communicate only their predictions on a shared unlabeled public dataset. Each honest client independently evaluates the received knowledge from three complementary modalities: class prediction, boundary decision, and prediction correlation. For each modality, unreliable clients are filtered and weighted to construct a teacher. Before distillation, the corresponding gradients are further validated using a supervised gradient computed from private data, and conflicting directions are projected or suppressed. This enables heterogeneous models to collaborate under reduced influence of Byzantine predictions. Our method has a great application value for decentralized federated learning in unreliable real-world scenarios where clients are exposed to receiver-specific Byzantine messages of malicious predictions. The main contributions of this work are summarized as
follows: 1) We propose a robust decentralized federated distillation method with heterogeneous models under Byzantine attacks. By exchanging predictions on shared public data rather than model parameters, our method enables clients with different model architectures to collaborate without a central server while allowing each honest client to independently defend against malicious predictions. 2) We design a multi-modality knowledge collaboration mechanism with private-gradient validation. Neighbor selection, reliability weighting, and teacher construction at each modality reduce the influence of unreliable clients, while gradient validation prevents conflicting external knowledge from being directly absorbed into the local model. 3) We prove that our proposed method ensures a bounded Byzantine influence on both distillation gradients and private gradients of individual clients after crossmodality fusion, thereby enabling stable local optimization for honest clients under Byzantine distillation. 4) We conduct experiments under heterogeneous decentralized settings with different model architectures, non-IID data, and multiple Byzantine attacks. The results show that the proposed method maintains effective collaboration among heterogeneous client models under adversarial prediction manipulation and yields an improved local model prediction accuracy, thanks to the multi-modality knowledge collaboration and private-gradient validation. The rest of the paper is organized as follows: Section II summarizes related work; Section III gives problem formulation; Section IV describes the methodology and algorithm; Section V provides theoretical analysis; Section VI presents experimental results and Section VII concludes the paper. II. R ELATED W ORK A. Decentralized Federated Learning Decentralized federated learning (DFL) removes the centralized parameter server and lets clients exchange information over a communication graph to protect privacy. Representative decentralized stochastic optimization methods, such as D-PSGD, can reduce the communication bottleneck at the busiest node[2] via mixing a client’s local model state with the states of its neighbors. Most gossip- or consensus-based DFL methods nevertheless exchange model parameters or gradients. They therefore assume a common architecture and an aligned parameter space, which makes direct aggregation undefined when clients use models with different structures and parameter sizes. Knowledge transfer in the output space offers a natural way to relax this requirement. FedMD showed that independently designed client models can collaborate by matching predictions on a public dataset[8]. In fully decentralized networks, consensus-based multi-hop federated distillation (CMFD) averages neighbors’ predictions on shared samples to approximate consensus in function space[9]. SD-Dist further uses soft predictions on a common reference dataset to identify similar peers and transfer knowledge among heterogeneous
3
personalized models[10]. Other approaches preserve model autonomy through a shared proxy model, as in ProxyFL[6], or learn personalized collaboration weights from distillationbased statistical distances, as in KD-PDFL[18]. These studies establish that function-space communication can support model heterogeneity, but methods such as CMFD [9] and SDDist [10] mainly focus on benign knowledge collaboration and do not consider Byzantine manipulation of exchanged predictions. B. Robust Decentralized Federated Learning Byzantine-robust DFL has mainly been studied in a shared parameter space. BRIDGE applies coordinate-wise screening before decentralized model mixing[14], while UBAR combines model-distance filtering with local performance evaluation[15]. ClippedGossip bounds the influence of neighborhood model differences on arbitrary communication graphs[13], and remove-then-clip aggregation couples client filtering with clipping to improve robustness under heterogeneous data[16]. More recent personalized DFL work allocates aggregation weights using critical parameter indices to suppress Byzantine clients[19]. Although these methods address malicious participants and, in some cases, non-IID data, their screening rules compare coordinates, distances, or subsets of model parameters. Such comparisons require compatible parameter representations and may confuse architectural differences with malicious deviations. More recently, DFLDual [20] combines data-domain and model-domain distances with trust bootstrapping to identify Byzantine clients, while BALANCE [21] uses each client’s local model as a similarity reference to evaluate received client models. However, these methods still rely on comparing model updates, model states, or model-domain distances, which are not directly applicable when clients use heterogeneous architectures. Robustness has also been investigated directly in the prediction space. Cronus applies robust statistics to black-box predictions and supports heterogeneous local models[17]. FedTGD studies top-k and impersonation attacks against federated distillation and filters suspicious logits using density clustering and cosine similarity[11]. Roux et al. analyze the Byzantine resilience of distillation-based FL and introduce history-aware client weighting to strengthen robust prediction aggregation[12]. Robust Federated Inference further studies adversarial aggregation of predictions from heterogeneous proprietary models and robustifies nonlinear DeepSets-based aggregators[22]. These methods demonstrate the value of output-space defenses, but they are primarily server-coordinated or inference-oriented and do not directly address receiver-specific Byzantine messages during iterative serverless distillation. III. P ROBLEM S TATEMENT We consider a decentralized federated learning system with N clients, denoted by N = {1, . . . , N }, communicating over a fully connected peer-to-peer network. Each client i owns a private labeled dataset Di and a local classifier fi (·; θi ) : X → RC . Clients may use different model architectures and
capacities, but all models share the same C-class output space. Therefore, their model parameters need not have compatible dimensions or semantics. All clients additionally have access to an unlabeled public dataset Dpub . Its labels are not used for reliability estimation, teacher construction, or client optimization. Let A ⊂ N and H = N \ A denote the Byzantine and honest client sets, respectively. Let Ni = N \ {i} denote i’s neighbors and Ai = A ∩ Ni its Byzantine neighbors. During each round of collaborative distillation, while an honest client broadcasts its genuine logits (predictions) on the public data, a Byzantine client may send arbitrary, receiver-specific logits to its neighbors to corrupt the distillation process. At training round t, each honest client initializes its local model as θit,0 = θit and performs K local SGD steps on its private data: θit,k+1 = θit,k − ηgi θit,k ; ξit,k , k = 0, . . . , K − 1, and obtains θ̄it = θit,K . After local training steps, client i performs S public distillation steps, starting from the locally trained model θ̄it . Let t,s Bpub denote the public mini-batch used at distillation step s, where s = 1, . . . , S. Logits sent from client j are genuine and generated by j’s current local model if j is honest, and arbitrary messages if j is Byzantine. IV. M ETHODOLOGY A. Framework Overview The proposed method achieves Byzantine-robust decentralized FL among heterogeneous models over T training rounds. In each training round t, each client i first performs supervised training on its private data to obtain the locally updated model θ̄it . It then sequentially processes distillation on S randomly sampled mini-batches of public data, updating the model from θeit,0 = θ̄it to θeit,S through the following three steps: 1) Logit Exchange and Standardization: Each client generates logits by applying its local model to a randomly sampled mini-batch of shared data and exchanges it with other clients. Since different model architectures may produce logits with different scales, each logit vector is centered and rescaled before comparison. 2) Robust Distillation within Each Modality: Each client compares the received logits with a median reference across three modalities of class prediction, boundary decision and prediction correlation. For each modality, clients with large discrepancies are filtered out, while the retained clients are assigned reliability weights. Their knowledge is then aggregated to construct a modalityspecific teacher, from which the corresponding distillation gradient is obtained. 3) Cross-Modality Fusion: The reliability of each modality is first estimated to determine its adaptive weight. A supervised gradient computed from private data is then used to validate the three distillation gradients: conflicting prediction and boundary gradients are removed,
4
and conflicting relation gradients are suppressed. Finally, the validated gradients are adaptively weighted and combined to update the local model. The detailed procedure of the proposed method is presented in Algorithm 1 in Section IV-E.
mb (A) = Ab,(1) − Ab,(2) ,
B. Logit Exchange and Standardization t,s B Let Bpub = {xt,s b }b=1 denote the s-th mini-batch of public data used at distillation step s of training round t. On the s-th mini-batch, honest client j evaluates its current model θejt,s−1 and obtains t,s et,s−1 Zjt,s = fj Bpub ; θj ∈ RB×C .
For receiver i, the logit attributed to client j is denoted by t,s Zj→i . An honest client sends its genuine logits, t,s = Zjt,s , Zj→i
whereas a Byzantine client may send an arbitrary finite logits and may transmit different messages to different receivers. Because heterogeneous model architectures may produce logits with different offsets and scales, each logit is centered by subtracting its mean and rescaled by its standard deviation across the C classes. The resulting logits have approximately zero mean and scale at most one, making outputs from heterogeneous models directly comparable while preserving their relative class ordering. For z ∈ RC , define S(z) =
z − µ(z)1C , σ(z) + ϵ PC 1 and c=1 zc C
(1)
where µ(z) = σ(z) = h P i1/2 C 2 1 , and ϵ > 0 prevents division c=1 (zc − µ(z)) C by zero. This operation is applied row-wise to obtain t,s t,s Zbj→i = S(Zj→i ). C. Robust Distillation within Each Modality Let Vi = Ni ∪ {i} denote client i and its neighbors. Client i first constructs a coordinate-wise median reference Mi = CoordM edianj∈Vi Zbj→i .
This discrepancy measures the difference in class preferences between client j and the median reference. For boundary decision knowledge on logit matrix A, letting Ab,(1) and Ab,(2) denote the largest and second-largest entries of row Ab , respectively, we define
(2)
Pb (A) = Sof tmax(Ab /Tkd ),
where Tkd > 0 denotes the distillation temperature used to smooth the class probabilities, with a larger Tkd producing a smoother distribution over classes, and Φbd (A)b = [mb (A); Pb (A)] . The boundary discrepancy is then Φbd Zbj→i − Φbd (Mi ) 1 dbd Zbj→i , Mi = . B(C + 1)
(4)
This discrepancy measures the difference in boundary decision and class probabilities. For prediction correlation knowledge on logit matrix A, we define Ab , R(A) = H(A)H(A)⊤ . Hb (A) = ∥Ab ∥2 + ϵ The relation discrepancy is defined by 2 1/2 P b ′ ′ b̸=b′ R(Zj→i )b,b − R(Mi )b,b bj→i , Mi = h drel Z . i1/2 P 2 R(M ) max , ϵ i b,b′ b̸=b′ (5) This discrepancy measures the difference in pairwise relations among different public samples. Therefore, each client is evaluated against the same median reference from class prediction, boundary decision, and prediction correlation perspectives. 2) Robust Neighbor Filtering and Weighting : For each modality, client i first filters out clients whose logits deviate too much from the median reference, and then assigns reliability weights to the remaining clients before aggregation. Client i first determines an adaptive threshold to identify clients whose logits deviate excessively from the median reference. Specifically, cqi = M edianj∈Vi \{i} dq Zbj→i , Mi .
The median is computed independently for each public sample and class coordinate, and is used only to evaluate the received logits. Based on this reference Mi , client i then compares the logits from each client across three modalities: class prediction, boundary decision, and prediction correlation. 1) Multi-Modality Logit Comparison : For the three modalities Q = {pred, bd, rel},
The corresponding median absolute deviation is bj→i , Mi − cq . M ADiq = M edianj∈Vi \{i} dq Z i
client i measures the difference between each client’s stanbj→i and the median reference Mi . dardized logits Z For class prediction knowledge, we define D E bj→i,b , Mi,b B Z X 1 1 − . dpred Zbj→i , Mi = B b Zj→i,b ∥Mi,b ∥2 + ϵ b=1 2 (3)
where κ controls the filtering tolerance and ϵMAD > 0 prevents a degenerate threshold when MAD is zero. Using the median and MAD makes the threshold less sensitive to extreme Byzantine logits and allows it to adapt to the current discrepancy distribution at each modality. Client i retains itself and all clients whose discrepancies are below the threshold: n o Siq = {i} ∪ j ∈ Ni : dq Zbj→i , Mi ≤ ξiq . (7)
The filtering threshold is ξiq = cqi + κ (M ADiq + ϵMAD ) ,
(6)
5
Algorithm 1 Robust Multi-Modality Decentralized Federated Distillation Require: Graph G = (V, E), private datasets {Di }i∈V , public dataset Dpub , local steps K, training rounds T , distillation −1 steps S, and learning rates η, ηd and distillation strength {λt }Tt=0 1: Initialize local models {θi0 }i∈V 2: for t = 0, 1, . . . , T − 1 do // Each client i ∈ V performs in parallel 3: Perform K local SGD steps on Di to obtain θ̄it , and set θeit,0 = θ̄it 4: for s = 1, 2, . . . , S do // Public logit exchange t,s 5: Evaluate θeit,s−1 on Bpub to obtain Zit,s t,s 6: Send Zit,s to clients and receive {Zj→i }j∈Ni t,s t,s 7: Transform each Zj→i into Zbj→i according to Eq. (1) // Robust Distillation within Each Modality bt,s }j∈V according to Eq. (2) 8: Construct the median reference Mit,s from {Z i j→i 9: for q ∈ {pred, bd, rel} do bt,s with M t,s at modality q according to Eqs. (3), (4), and (5) 10: Compare Z i j→i t,s 11: Filter clients using dq (Zbj→i , Mit,s ) and the adaptive threshold ξiq,t,s according to Eqs. (6) and (7), obtaining q,t,s Si q,t,s q,t,s 12: For each j ∈ Siq,t,s , compute rij and wij according to Eqs. (8) and (9) q,t,s 13: Aggregate the retained client knowledge using {wij } to construct Qq,t,s i et,s−1 and obtain g q,t,s according to Eqs. (10), (11), and (12) 14: Distill Qq,t,s into θ i i i 15: Compute the reliability sq,t,s of modality q according to Eq. (13) i 16: end for // Cross-Modality Fusion 17: Compute the adaptive modality weights {ωiq,t,s }q∈Q according to Eq. (14) 18: Sample a private labeled mini-batch and compute gisup,t,s according to Eq. (15) 19: Validate the class prediction and boundary decision gradients according to Eq. (16), and validate the prediction correlation gradient according to Eq. (17) 20: Combine the validated gradients according to Eq. (18) 21: Update θeit,s according to Eq. (19) 22: end for 23: Set θit+1 = θeit,S 24: end for 25: Output: Local models {θiT }i∈V
Filtering removes clearly unreliable clients, but some output, i.e., Byzantine clients may still remain. Therefore, client i further Zistu,t,s = Zit,s . assigns a reliability score to each retained client: Its standardized logits and normalized probabilities are q bj→i , Mi +βlocal dq Zbj→i , Z bi , rij = αmed dq Z j ∈ Siq . Zbistu,t,s = S Zistu,t,s , (8) Here, the first term measures the difference from the median stu,t,s bstu,t,s /Tkd . P = Sof tmax Z i i reference, and the second measures the difference from client i’s own prediction. The retained clients are then weighted as a) Class Prediction Knowledge: The class prediction q teacher is obtained by weighted aggregation: exp(−r /T ) w ij q wij =P . (9) q ℓ∈Siq exp(−riℓ /Tw ) X pred b Qpred =S wij Zj→i , i Thus, more reliable clients receive larger weights in the pred j∈Si subsequent aggregation. 3) Modality-Specific Teacher Construction and Distillation Pipred = Sof tmax Qpred /T kd . i : After filtering and weighting, client i constructs one teacher for each modality and distills the corresponding knowledge For each public sample xb , define its teacher label and into its local model. weighted voting confidence as At distillation step s of training round t, the current logit pred ybi,b = arg max Qpred matrix Zit,s defined in Section IV-B is used as the student i,b,c , c
6
vi,b =
h i pred bj→i,b,c = ybpred . 1 arg max Z wij i,b
X
c
j∈Sipred
Let ai,b = vi,b 1[vi,b ≥ τconf ], where τconf ∈ [0, 1] is the confidence threshold for retaining the teacher label. The class prediction distillation loss is B 2 X Tkd pred,t,s pred stu,t,s = Li DKL Pi,b ∥Pi,b (10) B b=1 PB stu,t,s pred , ybi,b b=1 ai,b CE Zi,b +µ . PB b=1 ai,b + ϵ The corresponding distillation gradient is gipred,t,s = ∇θi Lpred,t,s . i b) Boundary Decision Knowledge.: Similarly, the boundary decision teacher is X bd bd Qi = S wij Zbj→i , j∈Sibd
D. Cross-Modality Fusion 1) Adaptive Modality Weighting : The three knowledge modalities may have different reliability at different clients and distillation steps. We therefore estimate the overall reliability of modality q from its retained clients as X q q sqi = wij exp(−rij ). (13) j∈Siq
A larger sqi indicates that the retained clients at modality q have smaller reliability costs and provide more consistent knowledge. P 0 0 0 Let ω 0 = ωpred , ωbd , ωrel , and q∈Q ωq0 = 1, denote the initial modality weights, and define 1X q si . s̄i = 3 q∈Q
The adaptive modality weight is exp log ωq0 + γω (sqi − s̄i )/Tω q ωi = X . exp log ωℓ0 + γω (sℓi − s̄i )/Tω
Pibd = Sof tmax Qbd i /Tkd .
ℓ∈Q
bd = arg maxc Qbd For sample xb , let ybi,b i,b,c . The teacher and student margins are bd bd mbd i,b = Qi,b,b y bd − max Qi,b,c , bd c̸=y bi,b
i,b
bstu,t,s bstu,t,s . mstu − max Z i,b = Zi,b,b i,b,c y bd bd c̸=y bi,b
i,b
The boundary decision distillation loss is B
Lbd,t,s = i
1 X stu mi,b − mbd i,b B
(11)
b=1
B
ζT 2 X stu,t,s bd + kd DKL Pi,b ∥Pi,b . B b=1
The corresponding gradient is gibd,t,s = ∇θi Lbd,t,s . i c) Prediction Correlation Knowledge.: For inter-sample prediction correlation knowledge, client i directly aggregates the relation matrices of the retained clients and takes this as its teacher: X rel Qrel wij R Zbj→i . i = j∈Sirel
The student relation matrix is Ristu,t,s = R
Thus, ωiq ≥ 0,
X
ωiq = 1.
q∈Q
A modality whose retained clients are more reliable than the current average receives a larger weight in the final public update. It is worth distinguishing the two levels of q weighting used in the method.The client weight wij controls the contribution of client j within modality q, whereas ωiq controls the contribution of modality q in the final multimodality update. 2) Private-Gradient Validation and Model Update : At distillation step s of training round t, honest client i samples a private labeled mini-batch Bisup,t,s ⊂ Di and computes the supervised gradient at the same parameter point used for public distillation: gisup,t,s = ∇θi LCE Bisup,t,s ; θeit,s−1 . (15) This private gradient is used only to validate the external distillation directions and does not introduce an additional supervised update during the public phase. For q ∈ {pred, bd}, if a distillation gradient conflicts with the private supervised gradient, its conflicting component is removed: D E gisup,t,s , giq,t,s g q,t,s − gisup,t,s , i sup,t,s 2 g eiq,t,s = g i 2 q,t,s gi ,
bstu,t,s . Z i
The prediction correlation distillation loss is 2 X stu,t,s 1 Ri,b,b′ − Qrel , Lrel,t,s = i,b,b′ i B(B − 1) ′
(14)
(12)
b̸=b
and relation distillation gradient is girel,t,s = ∇θi Lrel,t,s . i Therefore, the three modalities produce three distillation gradients, gipred,t,s , gibd,t,s , and girel,t,s , which are subsequently validated and combined.
D
gisup,t,s , giq,t,s
gisup,t,s
2
E
< 0,
> 0,
.
otherwise, (16)
The projected gradients therefore satisfy gisup,t,s , geiq,t,s ≥ 0,
q ∈ {pred, bd}.
Prediction correlation knowledge is treated more conservatively. Its gradient is retained only when it agrees with the private supervised direction: hD E i geirel,t,s = 1 gisup,t,s , girel,t,s > 0 girel,t,s . (17)
7
At distillation step s of training round t, the validated multimodality gradient is X q,t,s q,t,s giMG,t,s = ωi gei . (18)
at θeit,s−1 and validated using a private supervised gradient, producing the validated multi-modality gradient giMG,t,s . The public update is θeit,s = θeit,s−1 − ηd λt giMG,t,s .
q∈Q
(20)
Since all modality weights are nonnegative and every validated modality-wise gradient has a nonnegative inner product with the private supervised gradient, D E gisup,t,s , giMG,t,s ≥ 0.
After S distillation steps, θit+1 = θeit,S . The private supervised gradient is used only for validating the public distillation directions and does not introduce an additional model update.
Client i then performs the public update
A. Assumptions
θeit,s = θeit,s−1 − ηd λt giMG,t,s ,
(19)
where ηd > 0 is the distillation learning rate and λt ≥ 0 controls the overall strength of public knowledge transfer. After processing all S public mini-batches, θit+1 = θeit,S . The resulting procedure first controls unreliable knowledge within each modality through client filtering and reliability weighting, and then controls the interaction among the resulting distillation gradients through private-gradient validation and adaptive modality fusion. E. Algorithm Description The complete procedure implementing the proposed method is summarized in Algorithm 1. Specifically, each client first performs local supervised training and then exchanges logits with its neighbors on the shared public data. The received logits are transformed to reduce architecture-dependent offset and scale differences. For each modality, the client compares the received logits with the median reference, filters unreliable clients, assigns reliability weights to the retained clients, and constructs a teacher to obtain the corresponding distillation gradient. Finally, the three distillation gradients are adaptively weighted according to their modality reliability, validated using a private supervised gradient, and combined to update the local model. V. T HEORETICAL A NALYSIS We analyze the optimization behavior of each honest client separately. Since heterogeneous clients may use different model architectures and therefore have incompatible parameter spaces, convergence to a common model parameter is not welldefined. Instead, we study whether each honest client’s own supervised objective remains stable under Byzantine public distillation. Specifically, we bound the residual Byzantine influence on the multi-modality distillation gradient, analyze the safety of private-gradient validation, and then establish bounded average stationarity for each honest client. Within training round t, honest client i first performs K local supervised SGD steps from θit,0 = θit and obtains θ̄it = θit,K . The public distillation phase then starts from θeit,0 = θ̄it and processes S public mini-batches sequentially. At distillation step s, the class prediction, boundary decision, and prediction correlation distillation gradients are computed
For each honest client i, let Fi (θ) denote its local supervised objective. We make the following assumptions for the convergence analysis. Assumption 1 (Smooth and lower-bounded local objective): For every honest client i, Fi is L-smooth and lower bounded. Specifically, for any θ and θ′ , Fi (θ′ ) ≤ Fi (θ) + ⟨∇Fi (θ), θ′ − θ⟩ +
L ′ 2 ∥θ − θ∥2 , 2
(21)
and Fi (θ) ≥ Fiinf ,
(22)
where Fiinf is a finite lower bound of Fi . Assumption 2 (Unbiased stochastic gradients with bounded second moment): For every honest client i, let gi (θ; ξ) denote a stochastic gradient computed from private data at model parameter θ. We assume Eξ [gi (θ; ξ)] = ∇Fi (θ),
(23)
and that there exists a finite constant Gg > 0 such that h i 2 Eξ ∥gi (θ; ξ)∥2 ≤ G2g . (24) Assumption 3 (Bounded sensitivity of distillation gradients to teacher representations).: For each modality q ∈ {pred, bd, rel}, let Γt,s i,q (Q) denote the distillation gradient of honest client i induced by the teacher representation Q, which is obtained after client selection, reliability weighting, and teacher construction on the s-th mini-batch of the t-th training round. There exists a finite constant Lq > 0 such that t,s ′ ′ Γt,s i,q (Q) − Γi,q (Q ) 2 ≤ Lq ∥Q − Q ∥F .
(25)
This condition states that a bounded change in the teacher representation induces a bounded change in the corresponding distillation gradient. Assumption 4 (Bounded honest distillation gradients).: Along the optimization trajectory of every honest client i, there exists a finite constant GH i > 0 such that X q,t,s ωiq,t,s gi,H ≤ GH ∀t, s, (26) i , q∈{pred,bd,rel}
2
q,t,s where gi,H denotes the distillation gradient induced by the honest teacher at modality q. This condition bounds the weighted magnitude of the honest distillation gradients along the optimization trajectory.
8
B. Main Result We first state the main result and then establish the supporting properties in the subsequent lemmas. Let γ = ηK, and α = supt ηd λt , where γ is the localupdate step size and α is an upper bound on the publicdistillation step size. Let δi denote the error bound of the private validation gradient and Gi denote the bound on the validated multi-modality distillation gradient, which will be established below. Theorem 1. (Local convergence under Byzantine distillation) Suppose Assumptions 1--4 hold. Consider any honest client i that follows Algorithm 1 for T training rounds, where each round consists of K local supervised SGD steps followed by S validated public distillation steps. Let γ = ηK and α = 1 , then the sequence of local models supt ηd λt . If 0 < γ ≤ 4L t T {θi }t=0 generated by Algorithm 1 satisfies T −1 i 2 F (θ0 ) − F inf 1 X h i i i t 2 E ∇Fi (θi ) 2 ≤ T t=0 γT 5 + L2 G2g η 2 K 2 + LγG2g . 8 2Sαδi Gi LSα2 G2i + + γ γ The theorem shows that, for every honest client, the average squared gradient norm of its local supervised objective remains bounded under Byzantine public distillation. The first term decreases as O(1/T ), while the remaining terms characterize the residual effects of local-update drift, stochastic gradients, private-gradient validation, and public distillation. Corollary 2. Under the conditions of the theorem, √ choose γ = T −1/2 and α ≤ T −1 , equivalently η = 1/(K T ) and ηd λt ≤ 1/T , with T ≥ 16L2 . Then T −1 1 X E ∥∇Fi (θit )∥22 = O(T −1/2 ). T t=0
Thus a uniformly sampled local iterate converges to stationarity in expectation. Proof: Substituting γ = T −1/2 , ηK = T −1/2 , and α ≤ T −1 into the theorem gives orders T −1/2 , T −1 , T −1/2 , T −1/2 , and T −3/2 for its five terms, respectively. The step-size condition follows from T −1/2 ≤ 1/(4L). Proof sketch: To prove Theorem 1, we analyze how the public distillation stage affects the descent provided by local supervised training. The proof proceeds in three steps: 1) We first derive the decrease of the local supervised objective after the K local SGD updates. This part follows the same analysis as our previous work [23]. 2) We then analyze the Byzantine knowledge after filtering and reliability weighting within each knowledge modality. Byzantine clients are not necessarily completely removed, but their remaining contribution to each distillation gradient is bounded by the total weight assigned to the retained Byzantine clients. Consequently, the magnitude of the Byzantine distillation gradients is bounded by Gi , as established in Lemma 3.
3) We next analyze the cross-modality fusion process. The adaptive modality weights determine the contribution of the three distillation gradients, while private-gradient validation projects or suppresses gradients that conflict with the local supervised direction. After validation and fusion, the final multi-modality gradient remains bounded by Gi , and its possible adverse effect on the local objective is further bounded by δi Gi , as shown in Lemma 4. Combining the local descent with these two bounds shows that Byzantine public distillation introduces only bounded perturbation terms into the local optimization. Therefore, the descent achieved by local SGD cannot be unboundedly disrupted by the Byzantine knowledge, which leads to the convergence bound in Theorem 1. C. Bounded Residual Byzantine Distillation The public-distillation terms in Theorem 1 depend on Gi , which bounds the multi-modality distillation gradients. We next show that Gi consists of a bounded honest component and a residual Byzantine component determined by the weights assigned to retained Byzantine sources. Let Q = {pred, bd, rel}. For each q ∈ Q, let Siq,t,s and q,t,s wij denote the retained source set and the corresponding normalized weights. Define the residual Byzantine weight as X q,t,s βiq,t,s = wij . j∈Ai ∩Siq,t,s
Since the trusted local source of an honest receiver is always retained, 0 ≤ βiq,t,s < 1. Let Qq,t,s j→i denote the source representation at modality q, and define X q,t,s q,t,s Qq,t,s = wij Qj→i , i j∈Siq,t,s
and its normalized honest component as X 1 q,t,s q,t,s wij Qj→i . Qq,t,s i,H = q,t,s 1 − βi q,t,s j∈Hi ∩Si
The corresponding distillation gradients are q,t,s giq,t,s = Γt,s ), i,q (Qi
q,t,s q,t,s gi,H = Γt,s i,q (Qi,H ).
Lemma 3. (Bounded residual Byzantine distillation). Suppose Assumptions~3--4 hold. For every honest client i, modality q ∈ Q, training round t, and distillation step s, the difference between the actual distillation gradient and the gradient induced by the retained honest sources satisfies q,t,s giq,t,s − gi,H ≤ 2Lq Mq βiq,t,s , (27) 2 (√ BC, q ∈ {pred, bd} where Mq = . Consequently, B, q = rel. X q,t,s q,t,s ωi gi ≤ Gi , (28) 2 q∈Q
9
+ εByz where P Gi = GH i i q,t,s q,t,s supt,s q∈Q 2ωi Lq Mq βi .
and
εByz i
=
Proof: We first note that the normalized weights are positive and sum to one over Siq,t,s . Since the trusted local source i of an honest receiver is always retained, X q,t,s q,t,s βiq,t,s = 1 − wij ≤ 1 − wii < 1. j∈Hi ∩Siq,t,s
We next bound the representation contributed by each retained source. For a standardized logit vector zb ∈ RC , PC (zc − µ(z))2 Cσ 2 (z) 2 ∥b z ∥2 = c=1 = ≤ C. (σ(z) + ϵ)2 (σ(z) + ϵ)2 Therefore, for a public mini-batch of size B, √ t,s ≤ BC. Zbj→i F
t,s For the prediction correlation modality, each entry of R(Zbj→i ) is the inner product of two row-normalized vectors and therefore has magnitude at most one. Since the relation matrix has size B × B, bt,s ≤ B. R Z j→i F
Thus, the corresponding source representations satisfy Qq,t,s j→i F ≤ Mq .
(29)
For βiq,t,s > 0, define the normalized Byzantine component as X 1 q,t,s q,t,s wij Qj→i , Qq,t,s i,A = q,t,s βi q,t,s j∈Ai ∩Si
and the normalized honest component as X 1 q,t,s q,t,s Qq,t,s wij Qj→i . i,H = q,t,s 1 − βi q,t,s The retained-source representation can then be decomposed as q,t,s q,t,s Qq,t,s = 1 − βiq,t,s Qq,t,s Qi,A . i i,H + βi
(30)
q,t,s Since Qq,t,s i,H and Qi,A are convex combinations of representations satisfying (29),
≤ Mq ,
Qq,t,s i,A
F
≤ Mq .
Therefore, Qq,t,s − Qq,t,s i i,H
F
≤
X
q,t,s ωiq,t,s gi,H
q,t,s = βiq,t,s Qq,t,s i,A − Qi,H
F
≤ 2Mq βiq,t,s .
(31)
If βiq,t,s = 0, no Byzantine source is retained and the same bound holds trivially. By Assumption~3, q,t,s t,s q,t,s q,t,s = Γt,s Q − Γ Q giq,t,s − gi,H i,q i i,q i,H 2
2
− Qq,t,s ≤ Lq Qq,t,s i i,H ≤ 2Lq Mq βiq,t,s ,
F
≤GH i +
X
+ 2
X
q,t,s ωiq,t,s giq,t,s − gi,H
q∈Q
2
2ωiq,t,s Lq Mq βiq,t,s
q∈Q H ≤Gi + εByz = Gi , i
(32)
where the second inequality follows from Assumption~4 and (27). This proves (28). D. Bound of Private-Gradient Validation The next lemma analyzes the safety of private-gradient validation in the cross-modality fusion process. It shows that the validated multi-modality gradient remains bounded by Gi , while its possible adverse effect on the local objective is bounded by δi Gi . Lemma 4. (Safety of private-gradient validation) Under Assumption 2 and Lemma 3, for every honest client i, training round t, and distillation step s, the validated multi-modality gradient satisfies giMG,t,s ≤ Gi . (33) 2
Moreover, there exists a finite constant δi with δi2 ≤ G2g such that hD E i E ∇Fi θeit,s−1 , giMG,t,s | Ft,s ≥ −δi Gi , (34) where Ft,s denotes the information available before sampling the private mini-batch used to compute gisup,t,s .
j∈Hi ∩Si
F
q∈Q
q∈Q
Hence, 0 ≤ βiq,t,s < 1.
Qq,t,s i,H
which proves (27). Hence, the effect of the retained Byzantine sources on each distillation gradient is controlled by their total retained weight. Finally, by the triangle inequality, X q,t,s q,t,s gi ωi 2
Proof: We first characterize the estimation error of the private supervised gradient. At distillation step s, gisup,t,s is computed from a private mini-batch at the current model θeit,s−1 . Conditioning on Ft,s fixes θeit,s−1 , and the remaining randomness comes only from the sampled private mini-batch. By the unbiasedness condition in Assumption 2, E gisup,t,s | Ft,s = ∇Fi θeit,s−1 . (35) Using this unbiasedness together with the bounded second moment in Assumption 2 gives 2 sup,t,s t,s−1 e E gi − ∇Fi θi | Ft,s 2 i 2 h 2 =E gisup,t,s 2 | Ft,s − ∇Fi θeit,s−1 2
≤G2g .
(36)
We denote a uniform upper bound on this mean-square estimation error by δi2 , so that 2 sup,t,s t,s−1 e E gi − ∇Fi θi | Ft,s ≤ δi2 , δi2 ≤ G2g . 2
(37)
10
We next consider the validated public gradient. By the exact projection used for the class prediction and boundary decision gradients, their components opposing the private supervised gradient are removed. The relation gradient is retained only when it has a positive inner product with the private supervised gradient. Therefore, the validation mechanism defined in the method section guarantees E D (38) gisup,t,s , giMG,t,s ≥ 0. The validation operations do not increase the norm of any modality-wise gradient. For q ∈ {pred, bd}, the exact projection removes only the component parallel and opposite to gisup,t,s and hence geiq,t,s 2 ≤ giq,t,s 2 .
(39)
For the prediction correlation modality, geirel,t,s is either girel,t,s or the zero vector, and therefore geirel,t,s
2
≤ girel,t,s
.
(40)
2
Since all modality weights are non-negative, the triangle inequality and Lemma 3 yield giMG,t,s
=
X
≤
X
2
ωiq,t,s geiq,t,s
q∈Q
Thus, although the private supervised gradient is only a stochastic estimate of the true local gradient, the possible adverse effect of the validated public-distillation direction is bounded by δi Gi . This completes the proof. E. Proof of the Main Result We now prove Theorem 1 by combining the descent achieved during local supervised training with the bounded effect of the subsequent public-distillation stage. The localtraining analysis gives the local-update drift and stochasticgradient terms, while Lemmas 3 and 4 control the two publicdistillation terms. Proof: The proof consists of three steps. (A) We use the local-SGD properties established in our previous work [23] to bound the change of the local supervised objective after the K private updates. (B) Lemma 3 provides the bound Gi on the Byzantine-affected distillation gradients, and Lemma 4 uses this bound to control the subsequent public-distillation updates. (C) We combine the two stage-wise bounds over one training round and telescope the resulting inequality over T rounds. (A) For the local supervised stage, define the average local stochastic gradient as
2
K−1
ωiq,t,s geiq,t,s 2
uti =
q∈Q
≤
X
(45)
k=0
ωiq,t,s giq,t,s 2
so that
q∈Q
≤ Gi .
θ̄it = θit − γuti ,
(41)
This proves Eq. (33). Finally, we relate the validated public gradient to the true local gradient. We write D E ∇Fi θeit,s−1 , giMG,t,s D E D E = gisup,t,s , giMG,t,s + ∇Fi θeit,s−1 − gisup,t,s , giMG,t,s . (42) Using Eq.(38) and the Cauchy--Schwarz inequality, D E ∇Fi θeit,s−1 , giMG,t,s ≥ − ∇Fi θeit,s−1 − gisup,t,s giMG,t,s 2 2 ≥ − Gi ∇Fi θeit,s−1 − gisup,t,s . 2
γ = ηK.
with LGg ηK, 2
i 2 uti − E[uti | θit ] 2 | θit ≤ G2g . (48) The bias bound follows from L-smoothness and E∥θit,k − θit ∥2 ≤ ηkGg , averaged over k = 0, . . . , K − 1; Jensen’s inequality and Assumption 2 give the second-moment bound. By the L-smoothness of Fi , bti 2 ≤
h
E
(43)
2
(44)
(46)
Under Assumptions~1--2, the local-SGD result established in [23] gives E uti | θit = ∇Fi (θit ) + bti , (47)
Fi (θ̄it ) ≤ Fi (θit ) − γ ∇Fi (θit ), uti +
Taking conditional expectation and applying Jensen’s inequality together with Eq.(37), we obtain hD E i E ∇Fi θeit,s−1 , giMG,t,s | Ft,s h i ≥ − Gi E ∇Fi θeit,s−1 − gisup,t,s | Ft,s 2 s 2 ≥ − Gi E ∇Fi θeit,s−1 − gisup,t,s | Ft,s ≥ − δi Gi .
1 X t,k t,k gi θi ; ξi , K
Lγ 2 t 2 ui 2 . 2
(49)
Taking expectation and using (48), together with |⟨a, b⟩| ≤ 1 2 2 4 ∥a∥2 + ∥b∥2 and Lγ ≤ 1/4, gives i γ h 2 E Fi (θ̄it ) ≤ E Fi (θit ) − E ∇Fi (θit ) 2 (50) 2 5γ 2 2 2 2 Lγ 2 2 + L Gg η K + Gg . 16 2 (B) We now analyze the public-distillation stage. Let θeit,0 = θ̄it , and for s = 1, . . . , S let θeit,s = θeit,s−1 − αt giMG,t,s ,
αt = ηd λt .
(51)
11
The model at the beginning of the next training round is θit+1 = θeit,S . By the L-smoothness of Fi ,
Fi θeit,s ≤ Fi θeit,s−1
D E − αt ∇Fi θeit,s−1 , giMG,t,s (52)
+
Lαt2 2
giMG,t,s
2
. 2
Dividing both sides by γT /2 gives T −1
2i 1 X h E ∇Fi θit 2 T t=0 2 Fi θi0 − Fiinf 5 ≤ + L2 G2g η 2 K 2 γT 8 2Sαδi Gi LSα2 G2i + LγG2g + + . γ γ
By Lemma 3, the multi-modality distillation gradients are bounded by Gi . Lemma 4 further gives E i hD E ∇Fi θeit,s−1 , giMG,t,s | Ft,s
This is the bound stated in Theorem 1. Moreover, since Gi = Byz , smaller residual Byzantine weights reduce εByz GH i + εi i and tighten the Byzantine-dependent terms in the bound.
giMG,t,s
VI. E XPERIMENTS
≥ − δi Gi ,
2
≤Gi . Taking conditional expectation in Eq. (52) therefore gives h i Lαt2 2 E Fi θeit,s | Ft,s ≤ Fi θeit,s−1 +αt δi Gi + Gi . (53) 2 Applying this bound successively for s = 1, . . . , S and using the tower property, together with θeit,0 = θ̄it and θeit,S = θit+1 , yields LSαt2 2 Gi . (54) E Fi θit+1 ≤ E Fi θ̄it + Sαt δi Gi + 2 (C) Combining (50) and (54), and using αt ≤ α, gives 2i γ h E Fi θit+1 ≤E Fi θit − E ∇Fi θit 2 2 5γ 2 2 2 2 Lγ 2 2 + L Gg η K + Gg 16 2 LSα2 2 + Sαδi Gi + Gi . (55) 2 Summing (55) over t = 0, . . . , T − 1 yields T −1 2i γ X h E ∇Fi θit 2 2 t=0 ≤Fi θi0 − E Fi θiT 5γT 2 2 2 2 Lγ 2 T 2 + L Gg η K + Gg 16 2 LT Sα2 2 + T Sαδi Gi + Gi . 2
Since Fi (θ) ≥ Fiinf by Assumption~1, T −1
2i γ X h E ∇Fi θit 2 2 t=0 ≤Fi θi0 − Fiinf 5γT 2 2 2 2 Lγ 2 T 2 + L Gg η K + Gg 16 2 LT Sα2 2 + T Sαδi Gi + Gi . 2
A. Experimental Setup 1) Datasets: We evaluate the proposed method on CIFAR10 and CIFAR-100. For each dataset, 10% of the official training set is used as an unlabeled public dataset, 5% is used for validation, and the remaining samples are partitioned among clients as private labeled data. The official test set is used only for final evaluation. Labels of the public samples are not used during teacher construction or distillation. 2) Federated Setting: We simulate a fully connected decentralized system with N = 10 clients on a single machine. No central server performs model, gradient, or prediction aggregation. The coordinator only schedules computation and routes messages between clients. Among the ten clients, three are Byzantine: A = {7, 8, 9}, and the remaining clients are honest: H = {0, 1, 2, 3, 4, 5, 6}. Only honest clients are included in the reported evaluation metrics. 3) Data and Model Heterogeneity: We use label-skewed Dirichlet partitioning to generate non-IID private data. Unless otherwise specified, the concentration parameter is set to α = 0.5. Each client receives at least 100 private training samples. To introduce model heterogeneity, clients use CIFARcompatible variants of ResNet-18[24], MobileNetV2[25], ShuffleNetV2[26], and VGG-11[27]. Different architectures are assigned to different clients as shown in Table I. TABLE I H ETEROGENEOUS MODEL ARCHITECTURES ACROSS CLIENTS . Client Client 0 & Client 4 & Client 8 Client 1 & Client 5 & Client 9 Client 2 & Client 6 Client 3 & Client 7
Architecture ResNet-18 MobileNetV2 ShuffleNetV2 VGG-11
4) Training and Distillation Settings: Each experiment runs for 120 training rounds. In each round, every client first performs one epoch of supervised training on its private data. The local optimizer is SGD with momentum 0.9, batch size 64, initial learning rate 0.05, and weight decay 5 × 10−4 . The
12
learning rate is multiplied by 0.1 at 50% and 75% of the total training rounds. The proposed method then processes two public minibatches per round, each with batch size 256. Public logits are standardized row-wise before reliability estimation. Class prediction, boundary decision, and prediction correlation discrepancies are used for client selection and weighting, followed by the construction of three distillation teachers. The three distillation gradients are validated using a private supervised gradient before they are combined for the public update. The principal distillaiton hyperparameters are listed in Table II. TABLE II H YPERPARAMETERS OF D ISTILLATION S ETTING . Hyperparameters Tkd ηkd τ λ ϵ
Value 2.0 0.01 0.5 1.0 10−8
5) Byzantine Attacks: We evaluate four Byzantine attacks together with a benign setting. In the Gaussian attack, malicious clients replace their outgoing messages with random Gaussian values. The Sign-flip attack reverses the sign of the transmitted values. The Targeted attack increases the preference for a selected target class, while the Bias attack adds a fixed class-wise bias to the outgoing messages. For the proposed method and other output-space methods, the attacks are applied to the exchanged logits. Byzantine clients may send different manipulated logits to different honest clients. 6) Evaluation Metrics: Test performance is reported only for honest clients. For honest client i, let Acci denote its test accuracy. We report 1 X Accmean = Acci , Accworst = min Acci . (56) i∈H |H| i∈H
Accmean measures the overall performance of honest clients, while Accworst measures the performance of the most affected honest client. The latter is particularly useful when Byzantine clients send different messages to different receivers. 7) Compared Methods: The primary comparison includes public-distillation methods that support the same heterogeneous model assignment, private-data partition, public batches, training schedule, and logit-space Byzantine attacks: • FedMD[8]: prediction-based distillation using the arithmetic mean of the available source logits; • FedDF[28]: public ensemble distillation using the arithmetic mean of source probabilities without parameter averaging; • Ours: the proposed robust decentralized federated distillation method with multi-modality knowledge collaboration. B. Main Comparison under Heterogeneous Models 1) CIFAR-10 Results: Table III reports the CIFAR-10 test results. Ours achieves the highest mean accuracy under None,
Gaussian, and Sign-flip, while FedDF and FedMD obtain the highest mean accuracy under Targeted and Bias, respectively. More importantly, Ours achieves the highest worst-client accuracy in all five distillation conditions. TABLE III CIFAR-10 ACCURACY (%) WITH HETEROGENEOUS MODELS AND D IRICHLET NON -IID DATA . ”M EAN ” AND ”W ORST ” DENOTE MEAN AND WORST HONEST- CLIENT ACCURACY. Attack None Gaussian Sign-flip Targeted Bias Average
FedMD Mean Worst 51.28 32.43 48.96 33.94 49.09 35.27 51.57 33.69 53.44 39.51 50.87 34.97
FedDF Mean Worst 52.08 39.71 50.31 34.92 49.72 35.94 52.66 38.33 51.94 36.89 51.34 37.16
Ours Mean Worst 54.79 45.66 51.92 42.13 50.01 39.33 51.50 39.90 52.57 42.23 52.16 41.85
Averaged over the five conditions, Ours improves mean accuracy by 0.82 percentage points over FedDF, the stronger primary baseline on this aggregate metric. Its average worstclient accuracy is 41.85%, which is 4.69 points above FedDF and 6.88 points above FedMD. The larger gain in worstclient accuracy indicates that receiver-specific robust teacher construction primarily benefits vulnerable clients rather than only the best-performing model. 2) CIFAR-100 Results: Table IV presents the CIFAR100 results. Ours obtains the highest mean and worst-client accuracy under every evaluated attack. Averaged across all five conditions, it improves mean accuracy by 0.96 points over FedMD and improves worst-client accuracy by 1.79 points over FedMD, which is the stronger baseline according to the corresponding aggregate metrics. TABLE IV CIFAR-100 ACCURACY (%) WITH HETEROGENEOUS MODELS AND D IRICHLET NON -IID DATA . Attack None Gaussian Sign-flip Targeted Bias Average
FedMD Mean Worst 25.13 20.56 24.47 21.07 23.73 20.21 25.46 21.44 25.45 21.02 24.85 20.86
FedDF Mean Worst 25.24 20.84 24.09 18.20 24.08 19.48 24.87 18.35 25.24 20.62 24.70 19.50
Ours Mean Worst 26.01 22.87 25.92 22.98 25.46 22.79 25.87 22.74 25.80 21.85 25.81 22.65
Figure 1 visualizes the main comparison. The separation between Ours and the two baselines is more consistent for worst-client accuracy than for mean accuracy, especially on CIFAR-10. C. Robustness across Attack Types On CIFAR-10, the no-attack mean and worst-client accuracies of Ours are 54.79% and 45.66%. Gaussian, Sign-flip, Targeted, and Bias reduce mean accuracy by 2.87, 4.78, 3.29, and 2.22 points, respectively. The corresponding worst-client reductions are 3.53, 6.33, 5.76, and 3.43 points. Sign-flip is therefore the strongest tested attack against Ours on CIFAR10. The CIFAR-100 results are considerably more stable across attacks. Relative to the no-attack result, the mean-accuracy
13
TABLE V P ER - CLIENT O URS ACCURACY (%) ON CIFAR-10. C0 AND C4 USE R ES N ET-18; C1 AND C5 USE M OBILE N ET V2; C2 AND C6 USE S HUFFLE N ET V2; C3 USES VGG-11. Attack None Gaussian Sign-flip Targeted Bias
C0 83.45 81.15 81.36 82.34 80.53
C1 46.94 43.85 39.33 43.72 42.23
C2 49.51 47.86 42.15 49.13 49.65
C3 50.62 47.15 44.87 48.02 49.75
C4 54.46 49.32 49.19 48.32 50.69
C5 45.66 42.13 40.02 39.90 42.68
C6 52.89 52.00 53.16 49.08 52.44
Mean 54.79 51.92 50.01 51.50 52.57
Worst 45.66 42.13 39.33 39.90 42.23
Mean 26.01 25.92 25.46 25.87 25.80
Worst 22.87 22.98 22.79 22.74 21.85
TABLE VI P ER - CLIENT O URS ACCURACY (%) ON CIFAR-100. Attack None Gaussian Sign-flip Targeted Bias
C0 26.31 26.73 26.23 26.42 25.90
C1 22.87 22.98 22.79 22.74 21.85
C2 26.32 25.97 25.66 25.58 26.24
C3 27.08 26.94 26.15 27.21 27.17
C4 27.19 27.44 26.51 27.28 27.49
C5 27.06 26.07 25.98 27.26 27.03
C6 25.21 25.33 24.87 24.57 24.90
D. Client-Level Performance Tables V and V I report each honest client’s accuracy. These results expose behavior that is hidden by a single aggregate mean.
Fig. 1. Comparison of FedMD, FedDF, and Ours under the benign setting and four Byzantine attacks. The top and bottom rows show mean and worst-client accuracy, respectively. Each panel uses an independently scaled vertical axis to expose relative differences; exact values are reported in Tables III and IV .
Fig. 2. Per-client Ours accuracy under different attacks. Model abbreviations are R18 (ResNet-18), MBV2 (MobileNetV2), SNV2 (ShuffleNetV2), and VGG11 (VGG-11). Each dataset uses its own color scale.
TABLE VII VARIANT COMPARISON IN MEAN / WORST HONEST- CLIENT ACCURACY (%). “ATTACK AVG .” AVERAGES THE FOUR B YZANTINE ATTACKS . Dataset
changes under Gaussian, Sign-flip, Targeted, and Bias are −0.08, −0.55, −0.14, and −0.21 points. The worst-client changes are +0.11, −0.08, −0.13, and −1.02 points, respectively. The small positive Gaussian difference is within normal single-run variation and should not be interpreted as an improvement caused by the attack. The validation-selected rounds provide an additional indication of training stability. On CIFAR-10, the selected rounds are 120, 90, 85, 90, and 120 for None, Gaussian, Sign-flip, Targeted, and Bias, respectively. In contrast, all CIFAR-100 runs select round 120. The earlier CIFAR-10 checkpoints under the three more disruptive attacks suggest that late-stage public supervision can accumulate attack-dependent noise, even though validation-based model selection limits its effect on the reported test result.
CIFAR-10
CIFAR-100
Method Ours-Base Ours-MM Ours-MM+CA Ours Ours-Base Ours-MM Ours-MM+CA Ours
None Mean Worst 50.08 37.48 49.95 34.69 50.61 35.66 54.79 45.66 24.46 21.28 24.79 21.06 24.96 21.68 26.01 22.87
Attack avg. Mean Worst 49.97 36.73 50.95 37.22 50.88 35.56 51.50 40.90 24.39 20.82 24.59 20.52 24.80 20.39 25.76 22.59
CIFAR-10 exhibits substantial client heterogeneity. Across the five conditions, C0 averages 81.77%, whereas C1 and C5 average 43.21% and 42.08%. Because C0 and C4 both use ResNet-18 but obtain markedly different accuracies, architecture alone cannot explain the gap; the Dirichlet private partitions and client-specific optimization trajectories are also major factors. On CIFAR-100, client averages across attacks lie between 22.65% and 27.18%, producing a substantially narrower client-performance range.
14
TABLE VIII F INAL - ROUND O URS DIAGNOSTICS WITHOUT ATTACK . Dataset CIFAR-10 CIFAR-100
Pseudo cov. 67.49% 22.72%
Conf. 0.739 0.693
pred 0.745 0.734
The same pattern is visible in Figure 2: CIFAR-10 performance is dominated by both client identity and private-data heterogeneity, whereas CIFAR-100 is more uniform across honest clients.
bd 0.209 0.219
rel 0.046 0.047
P roj.pred. 35.71% 78.57%
P roj.bd. 42.86% 50.00%
the three knowledge modalities. Figure 3 further highlights that the most pronounced advantage of Ours lies in worst-client accuracy rather than in attack-averaged mean performance. F. Mechanism Diagnostics
E. Variant Analysis Table V II compares four versions under identical heterogeneous settings. “Attack average” denotes the macro-average over Gaussian, Sign-flip, Targeted, and Bias, excluding None. This is a variant comparison rather than a one-factor-at-a-time ablation because successive versions modify more than one mechanism. Ours-Base denotes the original single-modality robust distillation method. Ours-MM introduces multi-modality knowledge transfer. Ours-MM+CA further incorporates class-aware boundary modeling and modality-wise gradient validation. Ours denotes the complete method with confidence-aware dual targets and conflict-preserving gradient projection.
Table V III summarizes final-round diagnostics for Ours in the no-attack runs. Pseudo-label coverage is the fraction of public examples whose weighted vote confidence exceeds τconf = 0.5. Projection frequency is averaged over honest clients and the two public updates in the final round. The adaptive modality weights remain close to the configured prediction-dominant prior, while retaining nonzero boundary and correlation contributions. CIFAR-100 has substantially lower pseudo-label coverage than CIFAR-10, which is expected in the more difficult 100-class label space. The high CIFAR-100 prediction-projection frequency also shows that private gradient validation is active rather than merely adding an unused safeguard. VII. C ONCLUSIONS
Fig. 3. Comparison of Ours-Base and the three multi-modality variants. “Attack average” is the macro-average over Gaussian, Sign-flip, Targeted, and Bias. All bar charts start at zero.
On CIFAR-10 without attack, Ours improves the mean and worst-client accuracies over Ours-MM+CA by 4.18 and 10.00 percentage points, respectively. Averaged over the four Byzantine attacks, the corresponding gains are 0.62 and 5.34 percentage points. On CIFAR-100, Ours surpasses Ours-MM+CA by 0.96 and 2.20 percentage points in attack-averaged mean and worst-client accuracy, respectively. These results show that the complete design primarily improves robustness for the most vulnerable honest client, consistent with its confidence-aware dual-target supervision and prediction-preserving gradient projection. Because Ours and Ours-MM+CA are cumulative designs rather than one-factor variants, we complement this comparison with dedicated ablations of gradient validation and
In this paper, we proposed a robust decentralized federated distillation method for heterogeneous models under Byzantine attacks. The key motivation is that clients with different model architectures cannot directly exchange or compare model parameters, while Byzantine clients may manipulate the predictions sent to different honest clients. Therefore, instead of performing collaboration in the parameter space, the proposed method exchanges predictions on shared unlabeled public data. Specifically, each client first evaluates the received predictions in three modalities of class prediction, boundary decision, and prediction correlation. Unreliable clients are then filtered, and the retained clients are weighted to construct three teachers in the three modalities for distillation. Finally, a supervised gradient computed from private data is used to validate the three distillation gradients. Conflicting prediction and boundary gradients are removed, and conflicting relation gradients are suppressed before the final model update. Theoretically, we showed that our algorithm achieves a bounded Byzantine influence on both distillation gradients of all modalities and final client private gradients after crossmodality fusion, thereby ensuring stable local optimization for honest clients under Byzantine distillation. Empirically, experiments on CIFAR-10 and CIFAR-100 for clients with heterogeneous model architectures, non-IID data, and multiple Byzantine attacks demonstrated that the proposed method provides robust collaboration among clients and achieves improved local model prediction accuracy, showing that our method can reduce the impact of different malicious predictions received by different clients. Across all evaluated attacks, the method achieves the highest worst client-accuracy on both
15
datasets, demonstrating its application value for decentralized federated learning in unreliable real-world scenarios where clients are exposed to receiver-specific Byzantine messages of malicious predictions. ACKNOWLEDGMENT This work is supported by the Science and Technology Development Fund of Macao (FDCT) (Project #0015/2023/RIA1), Queensland State Department of Environment and Science under the Quantum Challenges 2032 Program (Project #Q2032001). The corresponding author is Hong Shen. R EFERENCES [1] B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y. Arcas, “Communication-efficient learning of deep networks from decentralized data,” in Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, ser. Proceedings of Machine Learning Research, A. Singh and J. Zhu, Eds., vol. 54. PMLR, 2017, pp. 1273–1282. [Online]. Available: https://proceedings.mlr.press/v54/ mcmahan17a.html [2] X. Lian, C. Zhang, H. Zhang, C.-J. Hsieh, W. Zhang, and J. Liu, “Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent,” in Advances in Neural Information Processing Systems, vol. 30, 2017, pp. 5330–5340. [Online]. Available: https://proceedings.neurips.cc/paper/ 2017/hash/f75526659f31040afeb61cb7133e4e6d-Abstract.html [3] S. Savazzi, M. Nicoli, and V. Rampa, “Federated learning with cooperating devices: A consensus approach for massive IoT networks,” IEEE Internet of Things Journal, vol. 7, no. 5, pp. 4641–4654, 2020. [4] M. Ye, X. Fang, B. Du, P. C. Yuen, and D. Tao, “Heterogeneous federated learning: State-of-the-art and research challenges,” ACM Computing Surveys, vol. 56, no. 3, pp. 1–44, 2023. [5] Y. Tan, G. Long, L. Liu, T. Zhou, Q. Lu, J. Jiang, and C. Zhang, “FedProto: Federated prototype learning across heterogeneous clients,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 8, 2022, pp. 8432–8440. [6] S. Kalra, J. Wen, J. C. Cresswell, M. Volkovs, and H. R. Tizhoosh, “Decentralized federated learning through proxy model sharing,” Nature Communications, vol. 14, no. 1, p. 2899, 2023. [7] G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” arXiv preprint arXiv:1503.02531, 2015. [8] D. Li and J. Wang, “FedMD: Heterogenous federated learning via model distillation,” arXiv preprint arXiv:1910.03581, 2019. [9] A. Taya, T. Nishio, M. Morikura, and K. Yamamoto, “Decentralized and model-free federated learning: Consensus-based distillation in function space,” IEEE Transactions on Signal and Information Processing over Networks, vol. 8, pp. 799–814, 2022. [10] G. Ye, H. Yin, and T. Chen, “A decentralized collaborative learning framework across heterogeneous devices for personalized predictive analytics,” arXiv preprint arXiv:2205.13705, 2022. [11] W. Li, H. Gu, S. Wan, Z. Luan, W. Xi, L. Fan, Q. Yang, and B. Chen, “Byzantine robust aggregation in federated distillation with adversaries,” in 2024 IEEE 44th International Conference on Distributed Computing Systems, 2024, pp. 881–890. [12] C. Roux, M. Zimmer, and S. Pokutta, “On the byzantine-resilience of distillation-based federated learning,” in International Conference on Learning Representations, 2025. [Online]. Available: https://proceedings.iclr.cc/paper_files/paper/2025/ hash/6b40b312ce4933f4f2d1ac0b0b6b07f7-Abstract-Conference.html [13] L. He, S. P. Karimireddy, and M. Jaggi, “Byzantine-robust decentralized learning via ClippedGossip,” in International Conference on Learning Representations, 2023. [Online]. Available: https://openreview.net/ forum?id=qxcQqFUTIpQ [14] C. Fang, Z. Yang, and W. U. Bajwa, “BRIDGE: Byzantine-resilient decentralized gradient descent,” IEEE Transactions on Signal and Information Processing over Networks, vol. 8, pp. 610–626, 2022. [15] S. Guo, T. Zhang, H. Yu, X. Xie, L. Ma, T. Xiang, and Y. Liu, “Byzantine-resilient decentralized stochastic gradient descent,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 6, pp. 4096–4106, 2022.
[16] C. Yang and J. Ghaderi, “Byzantine-robust decentralized learning via remove-then-clip aggregation,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 19, 2024, pp. 21 735–21 743. [17] H. Chang, V. Shejwalkar, R. Shokri, and A. Houmansadr, “Cronus: Robust and heterogeneous collaborative learning with black-box knowledge transfer,” arXiv preprint arXiv:1912.11279, 2019. [18] E. Jeong and M. Kountouris, “Personalized decentralized federated learning with knowledge distillation,” in ICC 2023 - IEEE International Conference on Communications, 2023, pp. 1982–1987. [19] A. Zhang, P. Zhao, W. Lu, and G. Zhang, “Personalized decentralized federated learning: A privacy-enhanced and byzantine-resilient approach,” IEEE Transactions on Computational Social Systems, vol. 12, no. 5, pp. 3206–3217, 2025. [20] P. Sun, X. Liu, Z. Wang, and B. Liu, “Byzantine-robust decentralized federated learning via dual-domain clustering and trust bootstrapping,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 24 756–24 765. [21] M. Fang, Z. Zhang, Hairi, P. Khanduri, J. Liu, S. Lu, Y. Liu, and N. Gong, “Byzantine-robust decentralized federated learning,” in Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security, 2024, pp. 2874–2888. [22] A. Dhasade, S. Farhadkhani, R. Guerraoui, N. Gupta, M. Jacovella, A.-M. Kermarrec, and R. Pinot, “Robust federated inference,” in The Fourteenth International Conference on Learning Representations, 2026. [Online]. Available: https://openreview.net/forum?id=47eKYCaBIV [23] X. Ma, H. Shen, H. Tian, W. Lyu, and W. Ke, “Robust decentralized personalized federated learning via prediction-constrained neighborhood collaboration,” https://github.com/MPUMxX/R-PDFL, 2026, gitHub repository. [24] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770–778. [25] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 4510–4520. [26] N. Ma, X. Zhang, H.-T. Zheng, and J. Sun, “Shufflenet v2: Practical guidelines for efficient cnn architecture design,” in Proceedings of the European Conference on Computer Vision, 2018, pp. 116–131. [27] K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in International Conference on Learning Representations, 2015. [28] T. Lin, L. Kong, S. U. Stich, and M. Jaggi, “Ensemble distillation for robust model fusion in federated learning,” in Advances in Neural Information Processing Systems, 2020.