1
PoCQ: Proof of Contribution Quality as a Lightweight Blockchain Consensus for Secure Federated Learning
arXiv:2606.05642v1 [cs.DC] 4 Jun 2026
Sudad Abed, Nasser Sabar, Abdun Mahmood, and Mohammad Jabed Morshed Chowdhury
Abstract—Decentralized Federated Learning (FL) removes reliance on centralized coordinators but remains vulnerable to model poisoning, unreliable validation, and high validation overhead. This paper introduces Proof of Contribution Quality (PoCQ), a blockchain-based consensus framework designed to secure decentralized FL through reputation-aware validation and aggregation. PoCQ evaluates client updates using cryptographic commitments and lightweight norm-based validation, enabling efficient detection of malicious contributions while limiting validation cost. A reputation-driven consensus mechanism dynamically adjusts the influence of participants based on their historical contribution quality, while the blockchain stores only compact audit metadata to preserve scalability. Extensive experiments under poisoning scenarios across three benchmark datasets demonstrate that PoCQ outperforms the strongest state-of-the-art methods, achieving accuracy gains of 34.1% on challenging medical datasets in highly non-iid settings and an 11% improvement in global average accuracy. In addition, PoCQ reduces validation time by 21.27% on average per round, highlighting its effectiveness in jointly enhancing robustness and efficiency for fully decentralized federated learning. Index Terms—Decentralized Federated Learning, Blockchain Consensus, Reputation Systems, Model Poisoning, Byzantine Fault Tolerance.
I. I NTRODUCTION
T
HE digitization of healthcare has generated vast repositories of sensitive patient data, ranging from medical imaging to electronic health records (EHRs). While these datasets hold immense potential for training diagnostic Artificial Intelligence (AI) models, strict privacy regulations (e.g., HIPAA, GDPR) creates ”data silos,” preventing institutions from pooling their information centrally [1]. Federated Learning (FL) has emerged as a critical solution to this dilemma, enabling hospitals and medical centers to collaboratively train global models by exchanging only parameter updates, while patient data remains securely within the institution’s firewall [2]. However, the traditional FL architecture relies on a central aggregation server. In a healthcare context, this centralized S. Abed is with the Department of Computer Science and Information Technology, La Trobe University, Bundoora, VIC 3083, Australia, and also with the Electronic Computer Center, University of Anbar, Ramadi, 31001, Iraq. E-mail: [email protected] N. Sabar, A. Mahmood, and M. J. M. Chowdhury are with the Department of Computer Science and Information Technology, La Trobe University, Bundoora, VIC 3083, Australia. © 2026 IEEE. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible.
design presents a high-risk Single Point of Failure (SPoF) and a prime target for cyberattacks. To mitigate the problem of SPoF, recent research has pivoted toward Decentralized Federated Learning (DFL) and optimized methods to speed the learning process [3], [4]. In addition, recent research has used blockchain technology to replace the central authority with a peer-to-peer consensus network to improve resilience and trust [5], [6]. While decentralization eliminates the single point of failure, it introduces severe adversarial dynamics in a trustless environment. Malicious participants may launch model poisoning attacks, injecting corrupted gradients to degrade the diagnostic accuracy of the global model [7]. Furthermore, within a competitive healthcare environment, certain institutions may perform free-rider attacks by appearing to participate in the federated learning process while ultimately aiming to acquire the final global model without contributing any local data or computational effort [8]. In addition, the consensus mechanism itself is vulnerable to bad-mouthing attacks, where dishonest validators deliberately reject valid updates from rival clients to damage their reputation [9]. Beyond deliberate attacks, real world medical data is rarely uniform or Independent and Identically Distributed (IID) [10]. This uneven data distribution causes model drift and reduces overall accuracy. To address this, recent decentralized frameworks, such as Proof of Interpretation and Selection (PoIS) [11], use model interpretation techniques, such as Shapley values, to fairly evaluate the true quality of each client’s contribution. However, computing these values requires massive processing power, making interpretation based methods too slow and impractical for lightweight validation. Existing Blockchain-based FL (BCFL) frameworks have attempted to secure this infrastructure, though often with limitations that affect their viability in clinical settings. For instance, LBFL [12] employs a committee-based consensus to improve throughput and storage efficiency. While LBFL effectively addresses the scalability of storing model updates, its primary focus is on system performance rather than cryptographic defense against different attacks or the sophisticated manipulation of reputation scores by colluding validators. In addition, LBFL reputation score calculations depends not only on the improvement of global model based on the updates of a client but also on the number of samples that the client owns and the number of epochs performed, which is hard to validate. Conversely, VBFL [13] prioritizes security by introducing a validator role to verify the quality of updates before aggregation. However, VBFL relies on an accuracy-
2
based validation mechanism, requiring validators to re-train received updates on their local validation datasets. This process imposes a significant computational overhead and latency, which can bottleneck the training of models on large-scale medical images. To address these challenges, we propose Proof of Contribution Quality (PoCQ), a blockchain-based approach with reputation-aware consensus mechanism designed to secure DFL. Unlike VBFL’s computationally intensive re-training, PoCQ utilizes a geometric L2 -norm analysis to rapidly detect statistical anomalies in model updates. Furthermore, unlike LBFL, PoCQ integrates a strict cryptographic protocol to prevent manipulation of the reputation score, ensuring that credit for clinical contributions is correctly attributed. The specific contributions of this paper can be highlighted as follows: • Introducing PoCQ, a blockchain consensus for decentralized FL that outperforms state-of-the-art methods across three benchmark datasets. It achieves a 34.1% accuracy gain over VBFL in extreme non-iid settings and delivers an 11% improvement in global average accuracy when evaluated against the highest performing state of the art frameworks. • Reducing validation time by 21.27% per round on average compared to LBFL method via lightweight normbased verification and scalable metadata blockchain storage. • Mitigating poisoning and bad-mouthing attacks through a dynamic reputation mechanism. PoCQ achieves a 100 percent adversary detection rate with under a 7 percent false positive rate. Conversely, VBFL unfairly bans 30 percent of honest clients, and LBFL misses 22 percent of attackers while isolating 43 percent of honest nodes. The remainder of this paper is structured into four primary areas. Section II reviews the current state-of-the-art advancements in secure federated learning. Section III comprehensively details the Proof of Contribution Quality consensus phases and the underlying system architecture. Section IV presents the experimental setup alongside a rigorous analysis of the performance results. Finally, Section V summarizes the core findings of the study and outlines promising directions for future research. II. BACKGROUND AND R ELATED W ORK This section reviews the trajectory of Federated Learning (FL) from centralized topologies to blockchain-enabled decentralized architectures. We analyze the adversarial dynamics in trustless healthcare environments and evaluate the limitations of existing consensus and reputation mechanisms in addressing sophisticated attack vectors such as bad-mouthing and freeriding. A. Decentralized Federated Learning in Healthcare Federated Learning (FL), first proposed by McMahan et al. [2], represented a paradigm shift in training Deep Neural Networks (DNNs) by decoupling model optimization from direct data access. This architecture is particularly critical in
the healthcare domain, where stringent regulatory frameworks create localized data silos that prevent the physical aggregation of patient records (e.g., medical imaging, EHRs) [1]. As noted in a recent survey by Nezhadsistani et al. [14], while FL facilitates compliance by keeping data local, the conventional Client-Server architecture relies on a central aggregator. This centralization introduces a distinct Single Point of Failure (SPoF) and a bottleneck for scalability, rendering the global model vulnerable to server-side denial-of-service attacks or corruption. To mitigate these risks, recent literature has pivoted toward Decentralized Federated Learning (DFL). Yuan et al. [15] highlight that DFL distributes the aggregation workload among peer nodes via gossip protocols or structured topologies, enhancing system resilience and utilizing the idle computational power of edge devices. While DFL improves robustness against server failure, it fundamentally alters the trust model, necessitating new mechanisms to verify the integrity of contributions in the absence of an authoritative coordinator. B. Adversarial Dynamics: Poisoning and Consensus Attacks The transition to a trustless, decentralized environment exposes the learning process to severe adversarial threats. Recent studies classify these threats into two primary categories: performance degradation and reputation manipulation [16], [17]. First, model poisoning attacks involve malicious clients injecting mathematically corrupted gradients to prevent the global model from converging or to implant targeted backdoors [16], [18]. In high-dimensional spaces typical of medical imaging, such attacks can be subtle; adversaries may scale their updates to bypass simple magnitude checks while still misdirecting the optimization trajectory. Conversely, in competitive cross institutional settings, the system faces free-rider attacks. As demonstrated by Fraboni et al. [8], selfish participants may download the global model and upload random or repeated weights to simulate participation. This behavior drains network bandwidth and dilutes the global model quality without contributing legitimate data. A more insidious threat in decentralized consensus is the bad-mouthing attack [9]. As detailed by Zhang et al. [19] in their taxonomy of trust attacks, malicious validators may collude to cast negative votes against valid updates submitted by honest nodes. The objective is to artificially degrade the reputation of rival institutions, effectively ejecting them from the consensus committee to monopolize mining rewards or influence. Unlike poisoning, which attacks the model, badmouthing attacks the consensus mechanism itself, requiring defenses that can objectively verify validator honesty. C. Blockchain-Based Consensus Mechanisms: State of the Art and Limitations To enforce trust and auditability in DFL, researchers have integrated blockchain technology, creating Blockchain-Based Federated Learning (BCFL). The immutable ledger provides a tamper-proof history of model updates and participant behavior, yet existing frameworks exhibit distinct trade-offs between
3
security and efficiency. Early versions, such as Blockchained On-Device FL [5], relied on Proof-of-Work (PoW) consensus. While secure, PoW imposes prohibitive computational costs and latency, making it unsuitable for the iterative, high-frequency communication required by deep learning. Consequently, the field has shifted toward committee-based mechanisms to improve throughput, as seen in the work of Li et al. [6], where a randomly selected subset of nodes validates updates. However, random selection does not guarantee the honesty of the committee, leaving the system vulnerable to collusive attacks. To explicitly address the quality of updates, Chen et al. [13] proposed VBFL, which mandates that validators re-train received updates on their local validation datasets to verify accuracy improvement. While this ”Proof-of-Retraining” effectively detects poisoning, it incurs a massive computational overhead; the validation time scales linearly with model complexity, creating a severe bottleneck for large medical models where re-training a single update can take minutes. Addressing this efficiency gap, Qiao et al. [12] introduced LBFL, a lightweight framework employing a Proof-of-Contribution (PoC) consensus that defines contribution not only by retraining the model for one epoch by validators but also based on data volume and number of training epochs by a worker node. LBFL uses one miner only in each iteration, which receives votes for trained workers’ models from a selected validation committee to reduce the communications. Similarly, Zhao et al. [20] proposed a ”Long-Term Proof-of-Contribution” algorithm designed to incentivize sustained participation over time rather than immediate utility. Most recently, Işler et al. [21] introduced FedPoP, utilizing cryptographic proofs to verify that local training physically occurred, thereby effectively preventing free-riding. However, these efficiency focused frameworks exhibit a critical limitation: they largely define contribution quantitatively rather than qualitatively. Frameworks like LBFL and FedPoP lack rigorous mechanisms to verify the semantic quality of updates against intelligent poisoning. For instance, FedPoP ensures a client performed work, but cannot distinguish between work done on legitimate data versus corrupted data. Furthermore, LBFL relies on majority voting without specific defenses against bad-mouthing attacks where validators manipulate consensus scores. In addition, the reputation score mechanism in LBFL relies on metrics that are inherently difficult to verify, such as the client’s local dataset size and the number of training epochs performed. D. Reputation Systems and the Non-IID Challenge To move beyond simple voting, reputation-aware systems have been developed to penalize poor performance and ensure fairness [22], [23]. However, a profound limitation of current decentralized reputation frameworks is their inability to efficiently manage statistically heterogeneous, or non-iid, data. Real-world medical data is rarely uniform [10]. Training on highly non-iid data inherently causes model drift and uneven local updates, which traditional reputation metrics frequently misclassify as malicious poisoning. Recent advancements have
attempted to distinguish between adversarial behavior and natural statistical divergence. For example, Kasyap et al. [11] proposed the Proof of Interpretation and Selection (PoIS) consensus, utilizing model interpretation techniques (Shapley values) to fairly evaluate a client’s contribution across specific data distributions. While interpretation based mechanisms like PoIS provide deep, data-aware security, they introduce a prohibitive computational burden. The calculation of feature attributions scales exceptionally poorly with high-dimensional deep neural networks, rendering it practically infeasible for the rapid, lightweight validation required in clinical edge environments. Consequently, the current literature lacks a consensus mechanism capable of providing the semantic security of robust frameworks while maintaining the low-latency throughput of lightweight models. To resolve this fundamental limitation, we propose PoCQ, a reputation-aware consensus mechanism that leverages lightweight geometric auditing to objectively verify update quality. By bypassing the need for heavy retraining or deep model interpretation, PoCQ simultaneously neutralizes poisoning, mitigates bad-mouthing, and maintains robust performance in highly non-iid healthcare environments. III. T HE P ROPOSED M ETHOD In this paper, we propose Proof of Contribution Quality (PoCQ), a reputation-aware consensus mechanism designed to secure decentralized Federated Learning (FL) against model poisoning and bad-mouthing attacks. Unlike traditional FL, which relies on a central server [2], PoCQ operates on a peerto-peer network secured by a blockchain. The system utilizes a Public Key Infrastructure (PKI) for identity verification, where every node n possesses a key pair {P Kn , SKn }. The P Kn also represents the ID of each client. The consensus process is executed in five synchronized phases per communication round t. The workflow of PoCQ is illustrated in Fig.1. To facilitate the presentation of the mathematical formulations, the key notations and symbols used throughout the formal description of PoCQ are summarized in Table I. A. Phase 1: Local Training and Cryptographic Commitment In each round t, a participating worker node i trains the global model Wgt on its private local dataset Di using Stochastic Gradient Descent (SGD). To capture the node’s contribution, we compute the update vector ∆Wit : ∆Wit = Wit+1 − Wgt
(1)
To ensure data integrity and prevent free-rider attacks where adversaries replicate legitimate updates, the worker must cryptographically commit to the update before transmission. Given the high dimensionality of the parameter space, SGD training yields statistically unique gradient vectors for every node; thus, a digital signature effectively binds the specific update values to the worker’s identity, preventing plagiarism. The worker runs a hash function (SHA-256) to generate a hash Hi of the update and digitally signs it using its private key SKi : Hi = Hash(∆Wit )
(2)
4
Local Training at Round t
Model Exchange
Validator selection
Local model
Validation & Reputation Update Norm-based validation and votes casting
Leader Election & Block Mining Reputation-based leader election [k-3]-th block
Votes aggregation Local model
Validator integrity check
Sending signed hashed model
Global model aggregation
[k-1]-th block
Reputation score update
Local model
Verification of digital signature
Blacklisting module
[k-2]-th block
Block construction
Fig. 1. Workflow of PoCQ.
from ∆Wit for verification:
TABLE I S UMMARY OF K EY N OTATIONS
Pi = {∆Wit , σi }
Notation
Description
t n, i, j P Kn , SKn Wgt ∆Wit Di Hi σi Pi Vi kmin , kmax
Communication round index Node indices (e.g., worker i, validator j) Public and Secret (private) key of node n Global model weights at round t Local model update vector of worker i at round t Private local dataset owned by node i SHA-256 Hash Digital signature of worker i’s update Payload tuple broadcasted by worker i Set of assigned validators for worker i Minimum and maximum size bounds of the validator set L2 -norm discrepancy ratio between updates of i and j Validation threshold Signed vote (-1.0 or 1.0) cast by validator j for worker i Reputation weighted consensus score for worker i Final consensus decision for worker i (1: Accepted, 0: Rejected) Accumulated reputation score of node n ∈ [0, 1] Reputation reward/penalty assigned to validator j Decay factor Number of initial warm-up rounds before blacklisting begins Minimum reputation threshold for node blacklisting Probability of node n being elected as the round leader Number of malicious nodes in the network
rij τ vj→i Sw,i Ψw,i Rn rv,j β Twarmup Rmin Pelect (n) M
σi ← SignSKi (Hi )
(3)
Immediately upon generating the signature, the worker encapsulates the raw update and the signature into a payload tuple Pi and broadcasts it to the assigned validators. While the raw model remains visible to peers, this strict coupling ensures that any attempt by an intermediary to claim the update as their own results in a detectable hash collision, allowing validators to reject duplicates. Note that Hi is recomputed by validators
(4)
B. Phase 2: Distributed Validator Selection To eliminate the single point of failure inherent in centralized client selection, we employ a Distributed Invitation Protocol. At the start of a round, nodes broadcast “Invitation Requests” to their peers. A Worker node i accepts incoming invitations to form a validator set Vi . To ensure fault tolerance and prevent network congestion, the size of Vi is constrained by lower and upper bounds: kmin ≤ |Vi | ≤ kmax
(5)
C. Phase 3: Norm-Based Validation Upon receiving Pi , a validator j ∈ Vi first verifies the digital signature σi against Hi′ using the worker’s public key P Ki and ensures the calculated hash of the received weights Hi′ = Hash(∆Wit ) matches Hi . If the signature is valid, the validator performs a geometric check based on the L2-Norm (Euclidean magnitude) of the gradient vectors. The selection of the L2 -norm for the validation phase is primarily driven by its superior computational efficiency. In distributed architectures, compelling validators to assess model updates via semantic evaluation on local datasets or through partial retraining incurs prohibitive processing costs. In contrast, calculating the L2 -norm is a highly lightweight mathematical operation with a linear time complexity of O(d) for a model with d parameters. This allows the validation check to be executed near-instantaneously, making it highly scalable and ideal for resource-constrained peer-to-peer environments. Recent studies, such as DeFL [18], suggest that malicious updates (e.g., from noise injection) often exhibit statistical anomalies in their gradient norms compared to benign updates. We quantify this using the L2 -norm discrepancy ratio rij : rij =
||∆Wit ||2 ||∆Wjt ||2
(6)
5
where ||∆Wjt ||2 is the norm of the validator’s own local update, serving as a reference. The validator casts a vote vj→i based on a geometric threshold τ . Selecting an appropriate value for this threshold is crucial, particularly in highly heterogeneous (non-iid) environments, to prevent honest nodes with naturally divergent data from being mistakenly identified as malicious. A properly calibrated τ accommodates benign data drift while still detecting the extreme magnitude shifts caused by deliberate poisoning attacks. The optimal τ value is detailed in our experimental evaluation (Section IV-A3). Our system assigns a signed score based on this threshold: ( 1(Valid) if rij ≤ τ (7) vj→i = −1(Invalid) if rij > τ To prevent vote tampering and replay attacks, the vote is encapsulated in a signed vote transaction that includes a timestamp: Vtx = SignSKj ({IDworker , IDval , vj→i , timestamp})
(8)
D. Phase 4: Consensus and Reputation Update The network aggregates votes to determine the consensus status of the worker. Unlike simple majority voting, PoCQ employs reputation weighted consensus to prioritize votes from trustworthy validators. Let Vi be the set of validators for worker i, and Rj be the current reputation of validator j. The weighted consensus score Sw,i is calculated as: P j∈Vi vj→i · Rj P (9) Sw,i = j∈Vi Rj The final consensus decision Ψw,i is determined by a validity threshold: ( 1(Accepted) if Sw,i > 0 Ψw,i = (10) 0(Rejected) otherwise Simultaneously, we update the reputation score Ri of every node. This phase incorporates a specific defense against badmouthing attacks, where malicious validators intentionally vote Invalid against honest workers to damage their reputation [9]. 1) Validator Integrity Check: To defend against badmouthing, we verify if a validator’s vote aligns with the global consensus. A validator j receives a reward rv,j based on three conditions: if Ψw,j = 0 (Malicious Worker) 0 rv,j = 1 if vj→i = Ψw,i (Honesty Reward) −1 if vj→i ̸= Ψw,i (Bad-Mouthing Penalty) (11) Crucially, if a node submits an invalid local update and fails its own worker task (Ψw,j = 0), its validation reward is automatically neutralized to 0. This structural dependency effectively prevents reputation farming, ensuring that a malicious node cannot offset the penalties of model poisoning by accumulating trust through superficially honest validation behavior.
2) Reputation Update: The reputation score Ri is updated using an Exponential Moving Average (EMA) with a decay factor β [24]. This temporal update mechanism is essential for establishing sustained trust within the network. It acts as a historical memory that protects consistently honest nodes from severe penalties due to a single anomalous round, such as one caused by natural gradient variance in non-iid data, while simultaneously preventing malicious actors from instantly accumulating influence. To ensure mathematical stability, updates are applied sequentially for each role. Initially, the node’s reputation is updated to reflect its performance as a worker, utilizing the weighted average validation score Sw,i : Ri′ = max(0, min(1, βRit + (1 − β)Sw,i )) Next, we apply the validator reward rv,i : Rit+1 = max(0, min(1, βRi′ + (1 − β)rv,i )) where max(0, min(1, x)) ensures the reputation score remains bounded between 0 and 1. To permanently remove malicious nodes, PoCQ enforces a strict blacklisting protocol. However, model updates can be highly unstable during the early rounds of federated training, especially with non-iid data distributions. To avoid unfairly penalizing honest nodes during this phase of natural gradient variance, PoCQ incorporates an initial warm-up period of Twarmup rounds. During these early rounds (t ≤ Twarmup ), reputation scores are continuously updated, but the blacklisting mechanism is temporarily suspended. Once the network stabilizes and the warm-up phase concludes, any node whose reputation falls below a predefined critical threshold Rmin is permanently blacklisted from the active network, its identity (P Ki ) is added to a blacklist on the ledger. Once blacklisted, the network severs all ties with the node. It can no longer submit model updates, cast consensus votes, or participate in leader elections, ensuring the system remains secure over time. E. Phase 5: Leader Election and Block Mining To ensure the blockchain is maintained by trustworthy nodes, the round leader is elected via a verifiable lottery mechanism. Crucially, we do not deterministically select the node with the absolute highest reputation. Doing so would allow a single high-performing node to monopolize the mining process, effectively centralizing the network. To maintain decentralization while ensuring security, we employ a probabilistic selection strategy. The probability Pelect (n) of a node n being chosen as the leader is directly proportional to its accumulated reputation score Rn : Rn (12) Pelect (n) = P k∈Nactive Rk To ensure transparency, the lottery utilizes a Verifiable Random Function (VRF) seeded by the hash of the previous block. The winning node broadcasts a cryptographic proof of its selection, allowing peers to independently verify the election and prevent fraudulent leadership claims.
6
Algorithm 1 PoCQ: Reputation-Aware Consensus Mechanism Require: Set of nodes N , Dataset {Di }, Initial Model Wg0 , Rounds T , Validation Threshold τ , Decay Factor β. Ensure: Final Global Model WgT . 0 1: Initialize: ∀n ∈ N , Rn ← 0.5 2: for round t = 1 to T do 3: // Phase 1: Local Training & Commitment 4: for each worker i ∈ N in parallel do 5: ∆Wit ← SGD(Wgt−1 , Di ) − Wgt−1 6: Compute Hash Hi ← SHA256(∆Wit ) 7: Broadcast Pi = {∆Wit , SignSKi (Hi )} to Vi 8: end for 9: // Phase 2: Norm-Based Validation 10: for each validator j ∈ Vi do 11: if Verify(σi , P Ki ) then 12: rij ← ||∆Wit ||2 /||∆Wjt ||2 13: if rij ≤ τ then 14: vj→i ← 1.0 {Valid} 15: else 16: vj→i ← −1.0 {Invalid} 17: end if 18: end if 19: end for 20: // Phase 3: Consensus & Worker Update 21: for each worker i ∈ N do P vj→i ·Rj P i 22: Sw,i ← j∈V j∈Vi Rj 23: Ψw,i ← I(Sw,i > 0) {1 if accepted, else 0} 24: Ri ← Clamp(βRi + (1 − β)Sw,i , [0, 1]) 25: end for 26: // Phase 4: Validator Reputation Update & Blacklisting 27: for each validator j ∈ N do 28: if Ψw,j == 0 then 29: rv,j ← 0.0 {Probation} 30: else if vj→i == Ψw,i then 31: rv,j ← 1.0 {Reward} 32: else 33: rv,j ← −1.0 {Bad-Mouthing Penalty} 34: end if 35: Rj ← Clamp(βRj + (1 − β)rv,j , [0, 1]) 36: if t > Twarmup and Rj < 0.1 then 37: Blacklist(j) {blacklist node after warm-up phase} 38: N ← N \ {j} {Permanently remove from active network} 39: end if 40: end for 41: // Phase 5: Leader Election & Aggregation P 42: Select Leader L via Lottery: P (L) ∝ RL / k∈N Rk 43: L mines block BP t containing IDs {i | Ψw,i = 1} 44: Wgt ← Wgt−1 + i∈Valid PRRi k ∆Wit 45: end for 46: return WgT
While our current implementation utilizes the full active set, for larger-scale deployments we propose restricting the eligibility pool to the top η% (e.g., top 10%) of nodes ranked by reputation. This hybrid approach ensures that only highly trusted nodes can mine blocks, yet the specific winner remains unpredictable to preventing censorship or control. 1) Global Model Aggregation: The Global Model is updated using Reputation-Weighted Federated Averaging [23]. Unlike standard FedAvg, which typically weights contributions by data size, PoCQ weights updates based on the trustworthiness of the source. This ensures that high reputation nodes
exert greater influence on the global model convergence: X Ri P Wgt+1 = Wgt + ∆Wit (13) R k k∈Valid i∈Valid
2) Lightweight Block Construction: Finally, the elected Leader generates a new block. To address scalability, the block body strictly excludes model gradients. Instead, it encapsulates a lightweight set of audit metadata: • Block Index (idx): A sequential integer identifying the communication round. • Transaction Registry (transactions): A list containing exclusively the unique identifiers (IDs) of the workers whose contributions were validated (Ψw,i = 1). • Miner Identity (mined by): The unique ID of the round leader. • Cryptographic Link (previous hash): The SHA-256 hash of the preceding block’s header. • Timestamp: The precise system time of block generation. By decoupling model parameters from the consensus record, PoCQ maintains a lean blockchain state, significantly reducing synchronization latency. A comprehensive step-by-step summary of the PoCQ consensus mechanism is presented in Algorithm 1. IV. E XPERIMENTS A ND D ISCUSSION A. Experiments Setup To rigorously validate the efficacy and resilience of the proposed Proof of Contribution Quality (PoCQ) consensus mechanism, we simulated a comprehensive decentralized federated learning environment. Our empirical evaluation benchmarks the proposed architecture against state-of-the-art paradigms under varying degrees of data heterogeneity and adversarial threat models. 1) Datasets and Algorithmic Architectures: Our experiments are conducted across three diverse image classification datasets to evaluate the framework’s performance on both standard benchmarks and complex real world tasks. The baseline image recognition task utilizes MNIST, a widely used 10-class dataset of handwritten digits [25]. The medical data consisted of two distinct datasets sourced from MedMNIST v2, which is a collection of biomedical datasets encompassing twelve 2D datasets and six 3D datasets [26]. Specifically, we extracted and employed OrganAMNIST (an 11-class dataset of abdominal CT scans) and PathMNIST (a 9-class histological dataset of colon pathology). To accurately simulate the inherently unbalanced data distributions present in real-world federated networks, we evaluate the models under both Independent and Identically Distributed (IID) and Non-IID configurations. Statistical heterogeneity in the non-iid scenarios is induced by partitioning the labels across clients using a Dirichlet distribution, Dir(α) [11], [27]. We test two distinct imbalance severities: α = 0.5 (representing moderate heterogeneity) and α = 0.1 (representing extreme heterogeneity), see Fig.2. The local model employed on each client is a convolutional neural network (CNN) comprising two convolutional layers
7
Class_0Class_1Class_2Class_3Class_4Class_5Class_6Class_7Class_8Class_9Class_10
0 5 51 103 2372 0 1970 0 551 7 0 1451 0 168 1250 50 0 27 1
10 58 141 0 57 63 0 0 5702 1341 423 17 0 0 53 2566 1750 0 1
2750 0 543 0 1542 10 21 0 735 74 0 328 653 4 286 79 848 0 11 2
0 0 973 0 502 60 0 15 279 10 693 2 0 1143 1 0 17 108 5579 19
0 1 0 0 1119 3 4 0 1 57 0 5 82 0 0 3243 0 2 3314 5054
Class_0 Class_1 Class_2 Class_3 Class_4 Class_5 Class_6 Class_7 Class_8
Class
169 797 29 4 454 1 1324 40 112 226 474 103 252 48 466 5 4087 113 643 19
0 54 79 243 2908 1951 299 6 2 122 864 64 16 4 643 0 474 640 1062 78
1281 12 624 656 226 1 1 81 1495 139 43 70 1814 355 774 1332 790 84 43
0 224 18 3229 347 170 317 500 4 2335 27 650 72 174 552 11 132 2 338
416 88 1587 50 764 213 7 1295 87 487 1174 19 8 0 231 352 457 80 349
6 50 1133 631 241 26 319 1728 1355 0 130 2314 62 863 21 1651 281 14 1022
97 719 1007 195 678 430 1178 161 61 1 773 12 588 2 219 479 199 667 1
192 28 350 553 3335 357 203 146 0 2 1445 278 34 183 162 291 549 6 1
131 182 2150 2748 34 8 8 5 1023 862 12 367 581 69 625 120 101 1112 626
Class
( 286= 1.0) 292 297 315 293 319 259 276 277 279 270 252 310 269 283 269 275 273 260 260 251 265 271 276 254
283 298 299 305 305 300 329 315 291 278 304 265 280 275 287 284 314 292 317
318 330 336 336 302 305 274 335 307 322 301 321 290 329 286 324 311 322 301
292 280 295 285 285 282 293 271 280 304 326 281 293 298 289 302 277 310 315
313 281 281 285 284 294 315 306 307 295 281 286 286 311 319 303 284 295 304
4000
Number of Samples
294 292 304 267 283 308 300 279 317 278 275 349 295 283 293 269 286 301 283
3000 2000 1000 0
100 106 90 103 102 85 92 95 110 88 99 104 105 107 87 90 102 106 90 95
61 72 87 61 74 65 84 72 69 75 73 68 66 78 73 47 73 57 66 69
64 58 65 71 66 47 59 70 64 80 56 77 66 71 73 66 71 73 81 79
81 82 64 65 73 58 81 76 89 65 75 75 71 66 85 88 65 70 67 78
(206 =1821.0) 304 195 195 154 186 203 181 186 184 210 207 209 197 192 200 201 193 198 219 202 202 182 188 202
208 190 197 209 198 185 184 190 189 186 179 189 181 215 193 190 186 180 186
294 306 290 340 328 268 313 318 307 288 307 311 304 280 327 298 348 299 334
188 204 190 181 183 219 184 182 198 223 202 192 208 186 193 210 200 214 167
185 184 221 192 227 207 200 194 177 202 214 187 195 181 187 205 179 212 185
159 173 168 140 151 148 159 137 166 150 133 157 140 149 141 149 141 160 156
173 184 176 167 176 178 166 178 191 176 168 191 180 180 194 163 186 171 177
3000 2500
Number of Samples
Device
Class_0 Class_1 Class_2 Class_3 Class_4 Class_5 Class_6 Class_7 Class_8
294 295 315 315 299 328 312 326 296 297 297 299 291 335 333 334 297 290 295 283
2000 1500 1000 500 0
Class
Device
0 779 0 1 0 3705 9 0 5690 5 0 41 0 0 110 0 0 60 1
(b) ( = 0.5) 539 1299 342 335 419 1286 2121
285 287 324 278 297 288 308 255 297 301 310 315 300 309 283 305 318 309 286 303
Class_0Class_1Class_2Class_3Class_4Class_5Class_6Class_7Class_8Class_9Class_10
20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1
1328 0 0 248 1750 0 377 429 9 0 0 0 1 0 0 2965 60 652 1858 683
Device
0 0 0 0 10 5340 2655 26 0 0 1 676 0 0 0 333 0 138 329 1
168 0 206 2 75 80 40 108 104 2 31 43 15 562 12 56 43 8 81 3 786 355 787 146 655 43 0 86 458 25 198 153 230 838 1157 2 30 49 378 75 530 1390 11 180 7 0 24 106 9 335 149 2 36 407 0 45 3 100 921 85 3 213 48 11 71 641 541 127 60 1 371 29 195 0 43 425 47 0 312 3 0 19 331 241 0 1 1 263 1 1 5 4 581 40 604 603 4 83 161 9 92 10 725 9 116 111 52 4 238 10 4 20 460 14 178 8 244 47 25 0 270 2 76 21 3 119 101 78 114 7 171 67 190 2 203 174 2 174 517 374 122 37 247 1 104 18 226 88 155 229 31 4 1586 11 1 40 4 522 2 339 59 120 6 48 173 321 1 1 191 414 45
Class
20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1
Device
20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1
0 313 0 0 43 0 1131 0 0 2 52 0 3 0 38 48 0 7576 0 160
0 8 11 35 426 285 44 0 1 17 127 9 2 1 94 0 69 94 155 12
Class_0Class_1Class_2Class_3Class_4Class_5Class_6Class_7Class_8Class_9Class_10
Class
(0 =0 0.1)0
35 166 6 1 95 0 277 8 24 47 99 21 53 10 97 1 854 23 135 4
326 328 337 331 327 345 371 307 325 332 357 314 345 351 328 331 342 371 331 343
Class
20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1
Device
0 0 12 0 0 0 0 3 0 0 0 0 0 0 0 110 2 18 424 405 0 0 23 0 32 0 26 44 0 0 0 0 120 1 229 0 51 0 1206 210 341 309 0 781 0 0 1174 18 8 25 1 0 0 388 50 525 0 20 16 0 1 23 0 4 56 1 975 0 0 6 0 111 0 0 1 0 0 0 575 116 1 0 0 0 0 807 272 1786 58 4 17 137 10 0 0 1 4 421 0 289 0 0 0 98 0 0 0 132 256 1 2 0 0 0 0 5 718 5 511 0 25 524 38 0 0 0 0 0 2 477 0 137 0 0 0 0 84 0 224 0 0 0 0 49 389 16 618 17 62 0 988 3 155 0 7 0 25 804 662 7 0 0 3103 20 86 0 0 548 1 45 1 0 0 48 243 8 13 0 8 2326 1011 0 99 1 90 1 1 1 2 8 1541 1787 1
(a) ( = 0.5) 70 184 169 105 328 536 646 4 472
293 331 267 284 320 310 268 291 307 285 290 310 289 301 300 305 296 287 292 297
Class_0 Class_1 Class_2 Class_3 Class_4 Class_5 Class_6 Class_7 Class_8 Class_9
Class
20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1
Device
20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1
0 65 0 0 9 0 236 0 0 1 10 0 1 0 8 10 0 1582 0 34
Device
Class_0 Class_1 Class_2 Class_3 Class_4 Class_5 Class_6 Class_7 Class_8 Class_9
Class
( 0 =0 0.1) 0 174 0 2149 0
8 128 59 211 19 83 85 233 976 84 368 1248 2272 2223 15 353 238 4 71 135 4 418 97 2 727 0 465 7 2 391 3 962 6 161 186 166 102 23 264 17 121 31 6 108 284 4 194 55 485 366 46 6 4 504 116 1 285 813
72 539 756 147 509 322 884 121 46 1 580 9 441 1 165 359 149 501 1
447 450 461 481 463 455 457 496 490 447 475 486 449 468 469 497 489 468 467 450
502 520 487 497 498 461 466 479 454 485 468 452 459 457 470 480 463 455 503 450
517 484 511 481 501 516 529 540 494 527 536 516 498 536 533 511 536 529 513 551
( =3881.0)618 403 475 664
485 585 513 488 515 536 501 504 514 485 512 495 529 537 538 537 517 522 540 548
368 403 424 393 398 390 419 382 409 400 427 414 385 410 380 409 392 414 399
611 610 602 643 653 636 547 603 619 630 625 629 566 573 585 608 611 612 598
399 386 393 405 389 385 393 406 381 381 404 389 419 383 405 384 397 384 399
438 487 503 464 447 458 473 487 481 466 492 454 475 463 481 462 470 438 484
644 641 630 617 644 677 648 669 665 631 602 678 656 660 623 631 655 628 620
Class_0 Class_1 Class_2 Class_3 Class_4 Class_5 Class_6 Class_7 Class_8
7000 6000 5000 4000 3000 2000 1000 0
Number of Samples
Class_0 Class_1 Class_2 Class_3 Class_4 Class_5 Class_6 Class_7 Class_8 Class_9
( = 0.5)315 857 963
107 0 310 765 249 149 504 38 736 0 304 3 18 56 7 132 64 22 3 172 359 11 1158 504 287 2062 377 1904 37 281 0 1384 131 204 557 107 838 212 0 100 156 12 25 4 0 187 5 142 71 1 47 295 945 768 142 87 860 2 64 603 301 612 80 1377 354 1 65 46 25 16 857 57 159 11 40 383 14 1030 30 3 1043 42 6 27 295 456 204 103 0 385 3 0 446 325 168 9 2585 336 765 6 257 735 71 454 455 78 334 125 407 752 48 2 58 6 12 56 25 199 255 455
20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1
Device
20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1
Device
20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1
( = 0.1)
0 0 764 0 0 0 2063 0 0 0 197 0 0 0 0 4 0 0 0 0 0 0 0 459 3 26 408 648 0 0 0 0 142 0 38 63 0 0 0 0 28 7 1007 0 75 0 1157 335 508 607 0 3786 0 0 1731 25 8 40 2 0 715 1883 216 2185 0 28 16 0 2 46 0 18 247 5 1437 0 0 10 0 218 0 0 5 0 0 0 551 186 0 0 1 0 0 3354 402 2538 56 6 26 268 33 0 0 3 5 596 0 462 0 0 0 480 0 0 0 189 246 1 2 0 2 0 1 24 1059 7 490 0 38 1028 0 0 0 0 0 0 3 762 0 269 24 0 0 0 123 0 214 1 0 0 31 236 1705 65 912 24 59 0 1472 5 0 0 35 0 36 1142 637 11 0 0 4791 98 374 0 0 778 0 72 1 1 0 233 1069 35 20 0 9 3718 1505 0 101 1 393 1 1 1 1 13 2295 3507
Class
(c) Fig. 2. Class distribution across 20 devices for three datasets under different heterogeneity levels (α = 0.1, α = 0.5, α = 1.0). Lower α values indicate higher data heterogeneity among devices. (a) MNIST. (b) OrganAMNIST. (c) PathMNIST
with 32 and 64 filters of size 5×5, each followed by ReLU activation and 2×2 max pooling, and two fully connected layers of sizes 512 and the number of classes of each dataset. The model accepts single- or multi-channel input images of size 28×28, making it compatible with all datasets used in our experiments. 2) Benchmark Algorithms and Threat Configurations: To rigorously contextualize the performance enhancements of the PoCQ mechanism, we benchmark it against three representative distributed learning paradigms. The foundational control baseline is Vanilla Federated Learning (VFL) [2], which utilizes the standard Federated Averaging (FedAvg) protocol without cryptographic provenance, establishing the theoretical lower bound for Byzantine resilience. To evaluate against contemporary blockchain-integrated defenses, we compare against Validation-based Blockchained Federated Learning (VBFL) [13], which evaluates local updates via a distributed Proof-ofStake (PoS) consensus. The third method we compared with is Lightweight Blockchained Federated Learning (LBFL) [12], a resource-efficient framework that isolates malicious updates through a dedicated proof-of-contribution committee.
To systematically evaluate Byzantine fault tolerance, the simulated network is subjected to two discrete operational scenarios. The Benign Network Topology serves as a collaborative baseline featuring entirely honest clients (M = 0), establishing the maximum achievable accuracy and optimal convergence trajectory. Conversely, the Byzantine Network Topology introduces M = 4 malicious nodes specifically engineered to aggressively subvert the global model. To simulate a severe and persistent model poisoning attack, these malicious nodes intentionally corrupt their local gradient calculations prior to network broadcast by injecting targeted Gaussian noise into their updated weights, with the adversarial noise variance strictly constrained to σ 2 = 1.0. 3) Computational Setup and Experiments Configuration: The simulated federated network, cryptographic consensus protocols, and deep learning models were implemented using Python 3.11 and the PyTorch framework. All experiments were executed on an NVIDIA T400 GPU with total of 12 GB of RAM. The experimental federated network consists of 20 participating nodes, and all experiments were executed for a total of
8
0.9906 0.9903 0.9907 0.9900
0.9886 0.9844 0.1008 0.9883
0.8402 0.8437 0.8436 0.8412
0.7992 0.7542 0.1116 0.8381
0.7969 0.7997 0.7994 0.8007
0.7541 0.6102 0.1735 0.7897
Global Accuracy
IID = 1.0 Clean Att.
M: 0 | Dataset: organamnist 0.8
0.96
0.7
M: 0 | Dataset: pathmnist 0.8 0.7 0.6
0.94
0.6
0.92
0.5
0.5
0.90
0.4
0.4
0.88
0.3
M: 4 | Dataset: mnist
1.0
M: 4 | Dataset: organamnist 0.8
0.8
0.7
0.6
0.5
0.4
20
40
60
Round
80
0.5
0.4
0.4
0.3
0.3
0.2
0.2
100
LBFL PoCQ VBFL VFL
0.6
0.1
0.1 0
Method
M: 4 | Dataset: pathmnist
0.8 0.7
0.6
0.2
0
20
40
60
Round
80
100
0.0
0
20
40
60
Round
80
100
(a)
Note. IID = Independent and Identically Distributed; Att. = Attack (4 malicious clients); Clean = 0 malicious clients; LBFL = Lightweight Blockchain FL; VBFL = Validation-based Blockchain FL; VFL = Vanilla FL. Underlined values indicate the highest accuracy.
M: 0 | Dataset: mnist
1.00
Global Accuracy
IID = 0.1 IID = 0.5 Method Clean Att. Clean Att. MNIST LBFL 0.9630 0.9664 0.9869 0.9859 VBFL 0.9860 0.9453 0.9886 0.9877 VFL 0.9345 0.0686 0.9774 0.0984 PoCQ (Ours) 0.9885 0.9875 0.9902 0.9898 OrganAMNIST LBFL 0.7338 0.2125 0.8262 0.6914 VBFL 0.7832 0.5805 0.8225 0.6764 VFL 0.4644 0.1141 0.8293 0.1120 PoCQ (Ours) 0.8059 0.8027 0.8363 0.8364 PathMNIST LBFL 0.4027 0.2038 0.7827 0.5210 VBFL 0.5797 0.3192 0.7777 0.5967 VFL 0.2696 0.1720 0.7706 0.1720 PoCQ (Ours) 0.7344 0.6602 0.7996 0.7843
M: 0 | Dataset: mnist 0.98
Global Accuracy
TABLE II G LOBAL ACCURACY P ERFORMANCE ACROSS F EDERATED L EARNING M ETHODS AND DATASETS
0.8
0.90
0.7
0.85 0.80 0.75 0.70 0.65
Global Accuracy
0.6
0.6
0.5
0.5
0.4 0.3 0.2
M: 4 | Dataset: mnist
M: 4 | Dataset: organamnist 0.8
0.5
0.5
0.4 0.2
0.4
0.4
0.3
0.3
0.2
0.2 0.1
0.1 0
20
40
60
Round
80
100
LBFL PoCQ VBFL VFL
0.6
0.6 0.6
Method
M: 4 | Dataset: pathmnist
0.8 0.7
0.7
0.8
0
20
40
60
Round
80
100
0.0
0
20
40
60
Round
80
100
(b) M: 0 | Dataset: mnist
1.0
M: 0 | Dataset: organamnist 0.8
Global Accuracy
0.9
0.5
0.6
0.5
0.4
0.5
0.4
0.3
0.4
0.3
0.2
0.3
0.1
0.2
0.2
M: 4 | Dataset: mnist
1.0
Global Accuracy
0.6
0.6
0.7
M: 0 | Dataset: pathmnist 0.7
0.7
0.8
M: 4 | Dataset: organamnist 0.8
0.8 0.6 0.4
0.6
0.6
0.5
0.5
0.4
0.4
0.3
0.3
0.2
0
20
40
60
Round
80
100
LBFL PoCQ VBFL VFL
0.1
0.1
B. Results and Discussion
Method
M: 4 | Dataset: pathmnist
0.7
0.7
0.2
0.2
To comprehensively evaluate the proposed Proof of Contribution Quality (PoCQ) framework, this section presents a detailed empirical analysis of its performance across several critical dimensions. We explore the global accuracy and robustness of the model under varying degrees of data heterogeneity and targeted adversarial model poisoning. Furthermore, we investigate the underlying mechanisms that drive this resilience by analyzing the dynamic threat isolation and blacklisting behavior of the network during active attacks. Finally, we assess the practical viability of the framework by quantifying its computational overhead and validation efficiency relative to existing blockchain based solutions, demonstrating its suitability for scalable federated learning deployments. 1) Accuracy and Robustness: We first established the baseline convergence of the global model in a benign, trusted
0.7
0.3
0.60
M: 0 | Dataset: pathmnist 0.8
0.4
1.0
100 communication rounds. To optimize validation efficiency and limit communication overhead, the validator committee size is bounded between Kmin = 4 to ensure redundancy and Kmax = 6. Empirical tuning established a validation threshold (τ ) of 4 to strictly filter anomalous mathematical updates. The dynamic reputation system utilizes a decay factor of 0.8, ensuring that the framework remains robust when honest nodes hold highly skewed data while still rapidly adapting to sudden adversarial behavior. Finally, to avoid prematurely excluding honest nodes during early training volatility, participants become exposed for blacklisting only after a warmup period of Rwarm = 5 rounds. Following this phase, any node whose reputation score drops below the 0.1 blacklisting threshold is permanently isolated from the network. To ensure full transparency and reproducibility of our empirical findings, the complete source code and simulation configurations are publicly available at https://github.com/sudad/PoCQ.git.
M: 0 | Dataset: organamnist
0.95
0.0 0
20
40
60
Round
80
100
0
20
40
60
Round
80
100
(c) Fig. 3. Global accuracy over 100 communication rounds on MNIST, OrganAMNIST, and PathMNIST and heterogeneity levels ((a) iid with α = 1.0, (b) Non iid with α = 0.5, and (c) Non iid with α = 0.1) under clean (top row) and adversarial (bottom row) settings.
network environment where M = 0. As shown in Table II and the convergence plots for iid, moderate non-iid distributions, and extreme non-iid distributions in figures 3a, 3b, and 3c respectively, PoCQ consistently matches or exceeds the test accuracy of the established baselines and state of the art (VFL, VBFL, and LBFL) across all three datasets. A primary challenge in federated learning is managing weight divergence
9
0.7
Global Accuracy
0.6 0.5
method LBFL PoCQ VBFL
0.4 0.3 0.2 4
5
6 7 8 Malicious Count (M)
9
10
6 6
= 0.1
= 0.5
= 1.0
6 6
6 6 7
6 6
6 7
6 6 6 6
1 4 7101316192225283134374043464952555861646770737679828588919497100
Round
Honest (Active)
1 4 7 10 13 16 19 22 25 28 31 34 37 40 43 46 49 52 55 58 61 64 67 70 73 76 79 82 85 88 91 94 97 100 1 4 7 10 13 16 19 22 25 28 31 34 37 40 43 46 49 52 55 58 61 64 67 70 73 76 79 82 85 88 91 94 97 100
Device
20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1
Fig. 4. Global accuracy under an escalating malicious presence on PathMNIST dataset (α = 0.5).
Round
Honest (Blacklisted)
Round
Malicious (Active)
Malicious (Blacklisted)
= 0.1
= 0.5
6 6
= 1.0
6
6
6 6 91
7
6 6
96
90
6
6 100
6 6
1 4 7101316192225283134374043464952555861646770737679828588919497100
Round
Honest (Active)
1 4 7 10 13 16 19 22 25 28 31 34 37 40 43 46 49 52 55 58 61 64 67 70 73 76 79 82 85 88 91 94 97 100 1 4 7 10 13 16 19 22 25 28 31 34 37 40 43 46 49 52 55 58 61 64 67 70 73 76 79 82 85 88 91 94 97 100
Device
20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1
(a)
Round
Honest (Blacklisted)
Malicious (Active)
Round
Malicious (Blacklisted)
(b) = 0.1 100 6 6
6
= 0.5
= 1.0 6 6
6 6
Device
6 6
6 6 1 4 7101316192225283134374043464952555861646770737679828588919497100
Round
Honest (Active)
6
1 4 7 10 13 16 19 22 25 28 31 34 37 40 43 46 49 52 55 58 61 64 67 70 73 76 79 82 85 88 91 94 97 100 1 4 7 10 13 16 19 22 25 28 31 34 37 40 43 46 49 52 55 58 61 64 67 70 73 76 79 82 85 88 91 94 97 100
20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1
caused by non-iid data distributions. The results show that PoCQ effectively handles this statistical variance. For example, under extreme non-iid conditions (α = 0.1) on the complex PathMNIST dataset (Fig. 3c), PoCQ achieves a clean accuracy of 0.7344. This performance substantially outpaces Vanilla FL (0.2696) and LBFL (0.4027), while maintaining a clear margin over the second best performer, VBFL (0.5797). This performance gap suggests that the quality-weighted aggregation based on the contribution score in PoCQ filters out subpar local updates from honest nodes more efficiently than standard validation mechanisms, ultimately stabilizing global optimization. The structural resilience of the framework becomes much more apparent under the adversarial configuration of M = 4. When targeted with Gaussian noise attacks, the unprotected Vanilla FL baseline completely collapses, degrading to roughly 10-17% accuracy across all datasets, which equates to random guessing. On the other hand, PoCQ demonstrates strong Byzantine fault tolerance. While competing blockchain defenses like LBFL and VBFL successfully protect the simpler MNIST dataset, their validation logic struggles significantly when dealing with the combined difficulty of high data heterogeneity and complex medical images. On the OrganAMNIST dataset with extreme non-iid data (α = 0.1) and 4 attackers, LBFL and VBFL drop to accuracies of 0.2125 and 0.5805, respectively. This vulnerability is even more pronounced on the highly complex PathMNIST dataset under identical conditions, where LBFL and VBFL plummet to 0.2038 and 0.3192. While VBFL and LBFL do not claim to be effective in non-iid settings, it is noticeable that VBFL is slightly better than LBFL at handling these skewed distributions. Their reliance on strict consensus thresholds often leads them to reject skewed but legitimate honest updates while occasionally accepting welldisguised malicious ones. Under these identical conditions, PoCQ filters the adversarial noise effectively, sustaining a global accuracy of 0.8027 on OrganAMNIST and 0.6602 on PathMNIST, closely mirroring their clean baselines. To test the theoretical limits of this fault tolerance, we increased the number of malicious nodes from M = 4 to M = 10 on the PathMNIST dataset (α = 0.5), representing a severe scenario where half of the network is compromised. This stress test highlights the vulnerabilities of traditional consensus mechanisms. As the volume of poisoned gradients grows, LBFL and VBFL experience a linear deterioration in accuracy, dropping toward 0.20 as their validators are mathematically overwhelmed, see Fig. 4. PoCQ, however, shows a distinct non-linear resilience. By evaluating both the mathematical quality of incoming updates and the historical reputation of the sending nodes, PoCQ successfully protects the global model and keeps accuracy above 0.70, even when 35% of the participating nodes are malicious. This sustained performance relies directly on the system’s ability to identify and permanently blacklist malicious actors before their accumulated noise corrupts the global state, a mechanism explored further in the next section. 2) Threat Isolation and Blacklisting Dynamics: The core mechanism that enables PoCQ to maintain high accuracy is its dynamic reputation and isolation protocol. To better under-
Round
Honest (Blacklisted)
Malicious (Active)
Round
Malicious (Blacklisted)
(c) Fig. 5. Blacklisting dynamics and threat isolation speed across the three datasets under varying Dirichlet distributions (α = 0.1, 0.5, and 1.0)(a) MNIST. (b) OrganAMNIST. (c) PathMNIST
stand how the framework preserves performance during active poisoning, we tracked the network’s blacklisting behavior across different Dirichlet distributions (with α ∈ 1.0, 0.5, 0.1), as illustrated in Fig.5. The data reveals a clear relationship between statistical
10
True Node Identity
PoCQ
VBFL
LBFL
Actual Honest
135
9
Actual Honest
101
43
Actual Honest
82
62
Actual Malicious
0
36
Actual Malicious
0
36
Actual Malicious
8
28
Kept Active Blacklisted
Kept Active Blacklisted
System Action
Kept Active Blacklisted
System Action
System Action
Fig. 6. Confusion matrices illustrating the blacklisting efficacy of PoCQ, VBFL, and LBFL across all experimental configurations.
integrity. To comprehensively evaluate these temporal dynamics, we plotted the cumulative rate of blacklisting events over the progression of the training rounds in Fig. 7. Analyzing the time to detection for malicious actors reveals that PoCQ achieves instantaneous and deterministic isolation. Immediately following the initial five round warm-up phase, the cumulative TP trajectory for PoCQ spikes vertically to the absolute maximum of 36 isolated instances at precisely Round 6. In contrast, VBFL exhibits a heavily delayed response curve that allows malicious updates to compromise the global model for an average of 25 rounds before achieving full isolation. Conversely, LBFL plateaus early and fails to ever reach the maximum detection target. Furthermore, the cumulative FP trajectory highlights the superior fault tolerance of the proposed framework. While LBFL aggressively purges honest nodes within the first ten rounds and VBFL accumulates unfair bans continuously throughout training, PoCQ maintains a near zero FP rate for the vast majority of the simulation. The few honest nodes that are ultimately isolated by PoCQ only cross the blacklisting threshold in the extreme terminal stages of training, ensuring their valuable local data remains integrated into the global model. Cumulative Malicious Nodes Blacklisted (True Positives) Cumulative Honest Nodes Blacklisted (False Positives)
40 30
50
25
40
20
30
15 PoCQ VBFL LBFL Total Injected (36)
10 5 0
PoCQ VBFL LBFL
60
35
Cumulative Count
heterogeneity and the speed of threat isolation. In a balanced iid setting (α = 1.0), local data is uniformly distributed. When malicious nodes inject Gaussian noise, their updates stand out sharply against the global consensus. As a result, PoCQ flags these extreme statistical outliers almost immediately, isolating and permanently blacklisting all Byzantine actors within the first few warm-up communication rounds across every dataset tested. However, as data heterogeneity increases at α = 0.5 and particularly at α = 0.1, the isolation process becomes more complex. In highly non-iid environments, the natural skew of local datasets means that mathematically valid updates from honest nodes can frequently look like statistical anomalies. This variance provides a layer of cover that attackers can use to partially hide their poisoned gradients. The blacklisting graphs for the extreme α = 0.1 setting show that PoCQ requires slightly more communication rounds to build a reliable historical reputation score before confidently executing a network ban. Importantly, the consensus mechanism balances absolute network security with high precision. In highly skewed environments, such as the extreme non-iid setting of α = 0.1, the natural statistical variance of certain local datasets is severe enough that a small number of honest nodes are mistakenly blacklisted. This occurrence of false positives highlights an inherent challenge in federated learning where distinguishing between mathematically disguised malicious noise and legitimate but highly unusual data, such as a hospital with a rare patient demographic, remains deeply complex. While the false positive rate can be effectively decreased by expanding the initial warm-up rounds (Rwarm ) or by reducing the blacklisting threshold (Rmin ), the current configuration prioritizes absolute defense. Consequently, despite these occasional false positives under extreme data heterogeneity, the PoCQ framework consistently achieves a 100 percent True Positive (TP) isolation rate across all tested configurations. To systematically evaluate this precision and recall against existing literature, we constructed aggregate confusion matrices for PoCQ, VBFL, and LBFL. Evaluating the frameworks across three datasets and three data distributions with four malicious nodes injected per experiment yields a total of 36 distinct malicious node instances. The visualization in Fig.6 clearly maps the system actions against these true node identities. While PoCQ achieved absolute isolation efficacy by correctly blacklisting all 36 malicious node occurrences yielding zero False Negatives (FN) and only 9 False Positives (FP), competing methods struggled to balance security with client retention. VBFL matched the perfect detection rate by identifying all 36 malicious actors but severely penalized the network by erroneously banning 43 honest nodes. LBFL exhibited the weakest overall performance, recording 8 FN along with a substantial 62 FP. These matrices confirm that the norm based thresholding and dynamic reputation decay inherent to PoCQ vastly outperform state of the art frameworks in balancing aggressive Byzantine isolation with the preservation of honest client diversity. Beyond raw isolation accuracy, the speed and timing of blacklisting events are critical for maintaining global model
0
20
40
60
Communication Round
80
100
20 10 0
0
20
40
60
Communication Round
80
100
Fig. 7. Temporal Dynamics of Blacklisting (Time-to-Isolation).
3) Computational Overhead and Validation Efficiency: A major challenge with adding blockchain to federated networks is that the security checks take a lot of time and computer power, such as proof of work (PoW). While it is crucial to protect the network from attacks, the system still needs to be fast enough for real-world use. To evaluate this efficiency, we measured the average validation time required per communication round for the different methods, see Fig.8. The empirical results confirm that PoCQ operates as a highly optimized architecture. Under identical hardware conditions, PoCQ averages a validation time of 570.5 seconds per round. This offers a clear efficiency advantage over existing blockchain frameworks, running roughly 21%
11
faster than LBFL (724.7 seconds) and 40% faster than the computationally heavy VBFL (952.5 seconds).
Average Validation Time (s)
1000
952.5
800 600
724.7 570.5
400 200
Method
VB FL
LB FL
Po C
Q
0
research. The current consensus mechanism strictly prioritizes absolute network security. Under extreme non-iid conditions where highly skewed legitimate data mathematically resembles adversarial noise, this strict prioritization occasionally results in a small number of false positives. Although this false-positive rate can be effectively minimized by manually increasing the initial warm-up rounds (Rwarm ) or tuning the validation threshold (τ ) and blacklisting threshold (Rmin ), such adjustments currently require domain-specific calibration prior to deployment. Future work will focus on developing adaptive machine learning driven threshold mechanisms capable of automatically calibrating these validation parameters in real time. This dynamic calibration will further optimize the delicate balance between high security and maximum honest participation across entirely unknown data distributions.
Fig. 8. Average Validation Time per Round Across All Datasets.
This drop in validation time comes from two main improvements in how PoCQ is built. First, frameworks like VBFL are slow because they force validators to test incoming updates by training them on their own local data for one epoch. This local training step takes a huge amount of time and effort. PoCQ avoids this completely by using a lightweight, normbased check. Calculating these mathematical norms requires far less computing power and completely removes the need for extra training cycles. Second, PoCQ controls the network workload by limiting the size of the validator committee. It uses set minimum and maximum limits, written as Kmin and Kmax . By keeping the committee size within these strict boundaries, the system prevents major delays even when the network grows. Furthermore, once an attacker is permanently blacklisted, the system completely ignores them in future rounds. This naturally speeds up the network over time. Ultimately, our findings prove that PoCQ delivers strong security and handles complex data effectively, all without the massive slowdowns that usually affect blockchain-based federated learning. V. C ONCLUSION AND F UTURE W ORK This study introduced the Proof of Contribution Quality (PoCQ) framework, a blockchain-based approach designed to resolve the critical vulnerabilities of federated learning when deployed in highly skewed data environments under active Byzantine attacks. By integrating a lightweight normbased validation process with a dynamic reputation weighting system, PoCQ successfully detects malicious actors without imposing the severe computational bottlenecks typically associated with traditional blockchained federated learning models. The empirical results systematically demonstrate that the framework achieves high global accuracy and robust fault tolerance across multiple complex medical imaging datasets. It maintains a perfect true-positive isolation rate, effectively purging large-scale network compromises while seamlessly managing the statistical variance inherent to extreme non-iid data distributions. Despite this strong architectural resilience, the framework presents certain limitations that offer clear avenues for future
R EFERENCES [1] M. Joshi, A. Pal, and M. Sankarasubbu, “Federated learning for healthcare domain-pipeline, applications and challenges,” ACM Transactions on Computing for Healthcare, vol. 3, no. 4, pp. 1–36, 2022. [2] B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, “Communication-efficient learning of deep networks from decentralized data,” in Artificial intelligence and statistics. PMLR, 2017, pp. 1273– 1282. [3] W. Liu, L. Chen, and W. Zhang, “Decentralized federated learning: Balancing communication and computing costs,” IEEE Transactions on Signal and Information Processing over Networks, vol. 8, pp. 131–143, 2022. [4] S. Abed, N. Sabar, and A. Mahmood, “T-flash: Topology-flexible latency-aware scheduling for hierarchical decentralized federated learning,” Information Fusion, p. 103858, 2025. [5] H. Kim, J. Park, M. Bennis, and S.-L. Kim, “Blockchained on-device federated learning,” IEEE Communications Letters, vol. 24, no. 6, pp. 1279–1283, 2019. [6] Y. Li, C. Chen, N. Liu, H. Huang, Z. Zheng, and Q. Yan, “A blockchainbased decentralized federated learning framework with committee consensus,” IEEE Network, vol. 35, no. 1, pp. 234–241, 2020. [7] G. Xia, J. Chen, C. Yu, and J. Ma, “Poisoning attacks in federated learning: A survey,” Ieee Access, vol. 11, pp. 10 708–10 722, 2023. [8] Y. Fraboni, R. Vidal, and M. Lorenzi, “Free-rider attacks on model aggregation in federated learning,” in International conference on artificial intelligence and statistics. PMLR, 2021, pp. 1846–1854. [9] C. Lewis, V. Varadharajan, and N. Noman, “Attacks against federated learning defense systems and their mitigation,” Journal of Machine Learning Research, vol. 24, no. 30, pp. 1–50, 2023. [10] Z. Lu, H. Pan, Y. Dai, X. Si, and Y. Zhang, “Federated learning with non-iid data: A survey,” IEEE Internet of Things Journal, vol. 11, no. 11, pp. 19 188–19 209, 2024. [11] H. Kasyap, A. Manna, and S. Tripathy, “An efficient blockchain assisted reputation aware decentralized federated learning framework,” IEEE Transactions on Network and Service Management, vol. 20, no. 3, pp. 2771–2782, 2022. [12] S. Qiao, Y. Jiang, N. Han, W. Hua, Y. Lin, S. Min, and X. Wu, “Lbfl: A lightweight blockchain-based federated learning framework with proofof-contribution committee consensus,” IEEE Transactions on Big Data, 2024. [13] H. Chen, S. A. Asif, J. Park, C.-C. Shen, and M. Bennis, “Robust blockchained federated learning with model validation and proof-ofstake inspired consensus,” arXiv preprint arXiv:2101.03300, 2021. [14] N. Nezhadsistani, N. S. Moayedian, and B. Stiller, “Blockchain-enabled federated learning in healthcare: Survey and state-of-the-art,” IEEE Access, 2025. [15] L. Yuan, Z. Wang, L. Sun, P. S. Yu, and C. G. Brinton, “Decentralized federated learning: A survey and perspective,” IEEE Internet of Things Journal, vol. 11, no. 21, pp. 34 617–34 638, 2024. [16] E. Hallaji, R. Razavi-Far, M. Saif, B. Wang, and Q. Yang, “Decentralized federated learning: A survey on security and privacy,” IEEE Transactions on Big Data, vol. 10, no. 2, pp. 194–213, 2024.
12
[17] E. T. M. Beltrán, M. Q. Pérez, P. M. S. Sánchez, S. L. Bernal, G. Bovet, M. G. Pérez, G. M. Pérez, and A. H. Celdrán, “Decentralized federated learning: Fundamentals, state of the art, frameworks, trends, and challenges,” IEEE Communications Surveys & Tutorials, vol. 25, no. 4, pp. 2983–3013, 2023. [18] G. Yan, H. Wang, X. Yuan, and J. Li, “Defl: Defending against model poisoning attacks in federated learning via critical learning periods awareness,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, 2023, pp. 10 711–10 719. [19] C. Zhang, S. Lan, L. Wang, L. Liu, and J. Ren, “Trust attacks and defense in the social internet of things: Taxonomy and simulation-based evaluation,” Sensors, vol. 25, no. 24, p. 7513, 2025. [20] Y. Zhao, Y. Qu, Y. Xiang, F. Chen, and L. Gao, “Long-term proofof-contribution: an incentivized consensus algorithm for blockchainenabled federated learning,” IEEE Transactions on Services Computing, vol. 17, no. 5, pp. 2558–2570, 2024. [21] D. İşler, E. van Kempen, S. Hwang, and N. Laoutaris, “Fedpop: Federated learning meets proof of participation,” arXiv preprint arXiv:2511.08207, 2025. [22] W. Zhu, P. Wang, K. Li, and Y. Zhang, “Trustworthy blockchainassisted federated learning: Decentralized reputation management and performance optimization,” IEEE Internet of Things Journal, vol. 12, no. 3, pp. 2890–2905, 2025. [23] S. Barkatsa, M. Diamanti, P. Charatsaris, and S. Papavassiliou, “Fair and robust federated learning via reputation-aware incentives and model aggregation,” in 2025 IEEE 31st International Symposium on Local and Metropolitan Area Networks (LANMAN). IEEE, 2025, pp. 1–6. [24] D. M. Brotons, T. Vogels, and H. Hendrikx, “Exponential moving average of weights in deep learning: Dynamics and benefits,” Transactions on Machine Learning Research Journal, 2024. [25] L. Deng, “The mnist database of handwritten digit images for machine learning research,” IEEE Signal Processing Magazine, vol. 29, no. 6, pp. 141–142, 2012. [26] J. Yang, R. Shi, D. Wei, Z. Liu, L. Zhao, B. Ke, H. Pfister, and B. Ni, “Medmnist v2-a large-scale lightweight benchmark for 2d and 3d biomedical image classification,” Scientific Data, vol. 10, no. 1, p. 41, 2023. [27] H. Reguieg, M. El Hanjri, M. El Kamili, and A. Kobbane, “A comparative evaluation of fedavg and per-fedavg algorithms for dirichlet distributed heterogeneous data,” in 2023 10th International Conference on Wireless Networks and Mobile Communications (WINCOM). IEEE, 2023, pp. 1–6.
Sudad Abed is currently a PhD candidate at La Trobe University. He received his Master’s degree in Computer Science from Ball State University, USA, in 2016. His research interests include machine learning, deep learning, and federated learning, with a focus on improving the efficiency, privacy, and fairness of distributed learning systems.
Nasser Sabar is currently a Senior Lecturer with the Computer Science and IT Department, La Trobe University, Australia. He has published more than 76 papers in international journals and peer-reviewed conferences. His current research interests include hyper-heuristic frameworks, evolutionary computation, and hybrid algorithms, with a specific interest in big data optimization problems, cloud computing, dynamic optimization, and scheduling problems.
Abdun Mahmood (Senior Member, IEEE) received the B.Sc. degree in applied physics and electronics and the M.Sc. (Research) degree in computer science from the University of Dhaka, Dhaka, Bangladesh, in 1999 and 1997, respectively, and the Ph.D. degree from the University of Melbourne, Melbourne, VIC, Australia, in 2008. He has been an academic career in University since 2000, working with the University of Dhaka, RMIT University,Melbourne, VIC, UNSW Canberra, Canberra, ACT, Australia. He is currently with La Trobe University as an Associate Professor (Reader). Dr. Mahmood leads a group of researchers focusing on machine learning and cybersecurity including anomaly detection in smart grid, scada security, memory forensics, and false data injection. Dr. Mahmood has been successful to attract over a $1M+ in grant funding as a CI, including two ARC Linkage Projects.
Mohammad Jabed Morshed Chowdhury (Senior Member, IEEE) received the master’s degree in information security from Norwegian University of Science and Technology, Norway, the master’s degree in mobile computing from the University of Tartu, Estonia, and the Ph.D. degree from the Swinburne University of Technology, Melbourne, Australia. He is currently an Associate Lecturer of Cyber Security Program, La Trobe University, Melbourne. He is also working with Security, Privacy, and Trust. He has published his research in top venues, including TrustComm, HICSS, and REFSQ. His research interests include data sharing, privacy, and blockchain in different top venues.