1
Beyond Byzantine: An Organizational Consensus Algorithm for Self-Interested Agents Under Information Asymmetry
arXiv:2607.28957v1 [cs.GT] 31 Jul 2026
1st Jiawei Zhang Department of Agricultural Economics Purdue University West Lafayette, USA [email protected] 2nd Jianbo Liu Basic Technology Group, Intelligent Terminal Business Dept. JD Logistics, JD.com Beijing, China [email protected]
clusters [4], and Avicenna masks fail-slow replicas through counterfactual evaluation and rapid leader rotation [5]. Although these systems optimize performance and fault tolerance, they fail to model nodes that strategically misreport private information. OCA bridges this precise gap. This paper proposes the Organizational Consensus Algorithm (OCA), a dispute-triggered mechanism design framework for self-interested agents. Just as efficient corporate teams skip unnecessary meetings by communicating only when material facts change, OCA models departments as strategic agents evaluating both the necessity and risk of state change. Agents observe local conditions and broadcast updates only when deviations breach a dynamic local threshold backed by reputational stake. OCA offers four principal contributions: • A game-theoretic consensus architecture that suppresses redundant coordination overhead under asymmetric inforIndex Terms—Organizational consensus, mechanism design, incomplete information game, strategic agents, coordination mation. overhead, delayed verification. • A confidence-weighted aggregation protocol that enforces truth-telling through retrospective penalties and delayed I. I NTRODUCTION verification. Distributed system design traditionally ignores information • A formal identifiability condition for the delayed reference economics. Classical replicated state machine (RSM) protosignal. cols, such as Paxos and Raft [1], [2], treat nodes as either • Quantitative evidence from a Python prototype demonstratmechanistically obedient or arbitrarily malicious. In modern ing the trade-offs among coordination savings, consensus organizations, however, departmental agents fit neither extreme. stability, and systemic welfare. They act as boundedly rational, utility-maximizing economic We define the scope of OCA carefully. OCA targets nuactors under severe information asymmetry. Imposing a uniform merical or mergeable states with bounded error bounds and consensus policy across such heterogeneous structures chokes admissibility rules. It does not replace Paxos or Byzantine faultorganizational efficiency. Periodic, mandatory all-hands meet- tolerant replication when applications demand total ordering ings squander resources during quiet periods, while centralized and strict linearizability. leader-centric decision-making causes latency and single points of failure. II. R ELATED W ORK AND P OSITIONING Recent systems research treats node heterogeneity as a protocol variable rather than a deployment detail. LowPaxos A. Consensus and State Replication adapts leaders using compute and network capability profiles in Paxos and Raft enforce deterministic total ordering over resource-constrained environments [3]. Hamava supports fault- command logs [1], [2]. Although both support arbitrary detertolerant reconfiguration across heterogeneous geo-replicated ministic state machines, their coordination paths require leader
Abstract—Traditional distributed consensus protocols force nodes into a false binary: honest-but-faulty or actively malicious (Byzantine). Real organizational departments rarely fit such extreme categories. Instead, departmental agents act with bounded rationality, pursue localized self-interest, and operate under severe information asymmetry. We introduce the Organizational Consensus Algorithm (OCA), a mechanism design framework tailored for internal negotiation. OCA models inter-departmental conflict as a dynamic game of incomplete information, combining asset staking, exception-triggered signaling, and confidenceweighted consensus rules. Rather than enforcing instantaneous total ordering, OCA uses a retrospective penalty system anchored in delayed, verifiable outcomes to curb structural bias and eliminate redundant coordination. Evaluating OCA through a Python simulation prototype across diverse organizational scales reveals lower coordination overhead, richer informative reporting, and bounded welfare loss in noisy environments. These empirical gains remain conditional on our simulation parameters and do not, on their own, constitute a general truthful equilibrium.
2
Fig. 1. System architecture of the Organizational Consensus Algorithm (OCA). (Left Zone) Departmental agents evaluate local state deviations against a dispute trigger (∥∆x∥ > τ ). Idle nodes with stable states are muted to eliminate redundant coordination overhead, while fast/slow nodes experiencing material changes send updates. (Middle Channel) Only sparse, versioned deltas flow through the network, significantly reducing bandwidth and meeting costs. (Right √ Zone) The aggregation hub computes a bounded-error consensus state utilizing confidence-weighted aggregation (wi = ci ). (Bottom Feedback Loop) A delayed organizational oracle yi (e.g., quarterly audits) triggers a retrospective residual penalty pi . This slashes the effective trust stake ci of structurally biased agents across macroscopic epochs, enforcing truth-telling and aligning local incentives with global welfare.
elections and quorum acknowledgments that target physical crash faults rather than rational information manipulation. OCA does not replace these protocols. Weighted averaging fails to substitute for total ordering when updates conflict or depend on history. Instead, OCA restricts its state domain to proposals that nodes can safely validate and merge.
a core protocol variable. Unlike pure replication systems, OCA lets nodes strategically optimize their reported information.
D. Strategic and Incentivized Synchronization
Mechanism design aligns local agent incentives with global consistency [12], [13]. While strategic sensor fusion models B. Communication-Efficient Knowledge Merging misreporting costs [14], recent advances offer closer precedents. Teranishi et al. apply mechanism design to incentivize average Conflict-free Replicated Data Types (CRDTs) leverage consensus while protecting private states [15]. Milionis et algebraic properties to achieve uncoordinated convergence [6]. al. prove that truthful recovery demands source identifiaDelta-mutations [7], [8] and digest-driven reconciliation [9] bility, showing that without distinguishable observer signal further shrink transfer payloads, while practical rateless set distributions, truthfulness cannot form a strict Bayesian Nash reconciliation eliminates prior knowledge of state differences to equilibrium [16]. Reward-penalty ratios fail as universal cut communication overhead [10]. These innovations directly guarantees [17], [18]. Therefore, OCA treats its quadratic shape OCA’s bandwidth strategy, where proposals carry only residual penalty as conditionally incentive-compatible only versioned deltas once local states cross specific thresholds. under strictly bounded signals and feasibility constraints. Adaptive event-triggered estimation dynamically tunes communication to node-specific budgets [11]. OCA extends this idea by transforming triggers from mere resource-saving rules III. S YSTEM M ODEL AND P ROTOCOL into strategic gateways that govern when an agent enters the proposal and penalty arena. A. Agent Model and Local Aggregation Consider an organization of N heterogeneous departmental agents V connected over a communication topology Heterogeneity heavily dictates protocol scalability. LowPaxos Gt = (V, Et ). Instead of relying on a centralized executive [3], Hamava [4], and Avicenna [5] prove that node capability, coordinator, each agent j maintains its local proposed state network delay, and replica availability belong inside protocol (e.g., resource allocation or numerical target) xtj ∈ Rd , a design. OCA applies this insight to strategic decision-making. version vector, and a reputation descriptor rj . Upon receiving A department’s political weight or historical reliability becomes admissible proposals x̂ti from its interacting peers Pj,t = {i ∈ C. Organizational and Node Heterogeneity
3
V | (i, j) ∈ Et }, agent j computes its locally aggregated consensus state as: P wi x̂ti √ i∈P t+1 xj = P j,t (1) , wi = ci , w i∈Pj,t i where ci ≥ 0 measures agent i’s historical commitment or trust score. In a corporate environment, agents do not stake cryptographic tokens; instead, they stake tangible organizational assets. A department pledges its future budget allocations as collateral, risks its bonus pools via KPI deductions, or expends finite internal political capital. We select the square root function to model diminishing marginal returns on political expenditure. While any strictly concave function captures the economic intuition that hoarding power yields decreasing systemic influence, the square root provides a computationally tractable gradient for resource allocation. It ensures that a dominant department cannot unilaterally dictate the consensus merely by out-staking minority peers, thereby preserving organizational diversity in the aggregated state. B. Exception-Triggered Dispute Mechanism
where η > 0 is a decay rate. By structurally linking penalties to this delayed oracle, OCA dynamically isolates persistently noisy or dishonest agents without stalling the rapid, round-byround exception-triggered negotiation in (1). IV. C ORRECTNESS AND P ERFORMANCE M ETRICS Consensus error across the organization is bounded by ε: max ∥xti − x⋆,t ∥2 ≤ ε, i∈V
(5)
where x⋆,t represents the optimal, aligned organizational state. Relative coordination overhead savings SB against baseline mandatory meetings Bbase are evaluated via: BOCA SB = 1 − . (6) Bbase Alignment latency Tε and correctness score Cε (reflecting system welfare and stability) are formally defined as: n o Tε = min t : max ∥xti − x⋆,t ∥2 ≤ ε , (7) i
T i 1X h Cε = 1 max ∥xti − x⋆,t ∥2 ≤ ε . i T t=1
(8)
V. T HEORETICAL A NALYSIS To minimize organizational coordination overhead, negoTo formalize the mechanism design guarantees of OCA, tiation messages are broadcast only when a local deviation we provide a theoretical analysis of the agent incentive satisfies a dispute threshold: structures and the global consensus convergence. For analytical ∥xti − xlast or vit − vilast > 0, (2) tractability, we analyze the strategic interactions over both i ∥2 > τi where τi is the department’s local tolerance threshold. Ac- a single epoch (static game) and across macroscopic epochs tive messages carry agent IDs, proposal deltas, and staked (dynamic Markov game). commitments ci . C. Incomplete Information and Delayed Penalty Model To penalize inaccurate or structurally biased proposals driven by local self-interest, a retrospective residual penalty pi is assigned based on an episodic reference signal yi . In organizational contexts, yi represents a delayed objective ground truth, such as quarterly audits or realized market data. pi = λ∥x̂i − yi ∥22 ,
Ui = Bi (x̂i ) − pi − κci ,
(3)
A. Static Game and ε-Truthful Equilibrium Within a single epoch, consider a rational agent i observing a true local state x∗i ∈ Rd . Driven by self-interest, the agent submits a proposed state x̂i = x∗i + βi , where βi denotes the strategic bias. The ex-post oracle signal is modeled as yi = x∗i + εi , where the observation noise follows an independent Gaussian distribution εi ∼ N (0, σy2 I). The expected single-epoch utility of agent i is defined as: E[Ui (βi )] = Bi (x∗i + βi ) − λE ∥x̂i − yi ∥22 − κci , (9)
where Bi (·) represents the local rent-seeking benefit. where Bi (·) represents the agent’s local utility, assumed to be monotonically increasing and concave. In economic terms, this Theorem 1 (ε-Truthful Bayesian Nash Equilibrium). Assuming near x∗i , if the penalty parameter λ satisfies function models a department’s incentive to capture localized Bi is differentiable ∥∇Bi (x∗ i )∥2 rents. A logistics division, for instance, might overstate its λ ≥ , the optimal strategic bias ∥βi∗ ∥2 is strictly 2εtol capacity needs to secure priority routing, hoard operational bounded by the tolerance εtol , establishing an ε-Truthful budget, or pad inventory safety stock. Such actions maximize Bayesian Nash Equilibrium. departmental metrics but degrade system-wide efficiency. Proof: Expanding the expected penalty term in (9) yields: The term κci introduces the endogenous cost of staking. Here, E ∥βi − εi ∥22 = ∥βi ∥22 + d · σy2 . (10) κ > 0 represents the explicit administrative and reputational friction incurred when a department mobilizes its political The agent seeks to maximize B (x∗ + β ) − λ∥β ∥2 . Taking i i i i 2 capital. Initiating a dispute or demanding a larger budget the first-order condition (FOC) with respect to β and equating i allocation carries a real opportunity cost, deterring frivolous it to zero: or purely rent-seeking proposals. ∇Bi (x∗i + βi ) − 2λβi = 0. (11) Because the organizational truth yi is only available episod∗ a first-order Taylor approximation ∇Bi (xi + βi ) ≈ ically, pi is applied retrospectively. An agent’s effective Applying 1 ∗ ∇B (x ), we obtain the optimal strategy βi∗ = 2λ ∇Bi (x∗i ). i i reputation ci is slashed across macroscopic epochs k via: ∗ To bound the bias within ∗εtol , we require ∥βi ∥2 ≤ εtol , which (k+1) (k) (k) ci = max 0, ci − ηpi , (4) simplifies to λ ≥ ∥∇B2εi (xi )∥2 . tol
4
VI. E XPERIMENTAL R ESULTS AND A NALYSIS
B. Dynamic Game and Rug Pull Deterrence A static equilibrium is vulnerable to cross-epoch exploitation, where an agent remains honest to accumulate a maximum trust score cmax , and subsequently executing a one-time extreme falsification (a “rug pull” attack) to capture a massive short-term rent ∆Bhuge . We model this as a Markov Decision Process with a discount factor δ ∈ (0, 1).
We evaluate OCA through extensive Python simulations, comparing against four baseline coordination protocols across varying network scales, noise conditions, and adversarial settings. All experiments report medians across 30 independent Monte Carlo runs to mitigate stochastic variability.
Theorem 2 (Rug Pull Deterrence Condition). A Subgame Perfect Nash Equilibrium (SPNE) where agents remain perpetually honest is guaranteed if the penalty coefficient λ and discount factor δ satisfy:
A. Experimental Setup
λ∥βmax ∥22 +
δ ∆Ucost ≥ ∆Bhuge , 1−δ
(12)
where ∆Ucost = Ubaseline − Upunish − κcmax represents the long-term utility loss during the zero-trust punishment phase. −κcmax Proof: Let VHonest = Ubaseline be the steady-state 1−δ discounted value of perpetual honesty. Conversely, the value of executing a rug pull at epoch T , incurring the maximum bias (T +1) βmax , followed by a zero-trust punishment state (ci = 0) is:
VRugP ull = (Ubaseline + ∆Bhuge − λ∥βmax ∥22 − κcmax ) δ Upunish . (13) + 1−δ To deter the attack, the incentive compatibility constraint dictates VHonest ≥ VRugP ull . Rearranging the terms yields the deterrence condition, demonstrating that the immediate rent must be outweighed by the sum of the instantaneous quadratic penalty and the discounted future loss of political capital. C. Global Convergence under OCA Assuming malicious biases are suppressed (βi ≈ 0), we demonstrate that the confidence-weighted aggregation converges globally. Theorem 3 (Asymptotic Consensus). If the organizational communication topology G = (V, E) is strongly connected and contains at least one self-loop, the iterative state update defined in OCA achieves global asymptotic consensus:
Protocols compared. We evaluate five protocols with distinct architectural philosophies: 1) FullStateDissemination (Baseline 1): Every node broadcasts its full state each round. Message size: 12 B/msg (8 B state + 4 B metadata), yielding O(N 2 ) total bandwidth. No incentive layer. 2) DeltaAntiEntropy (Baseline 2): Periodic delta synchronization every 5 rounds with weak correction. Message size: 8 B/msg (4 B delta + 4 B version). No incentive layer. 3) DeltaStateCRDT (Baseline 3): Delta-state Conflictfree Replicated Data Type with join-semilattice merge. Only nodes exceeding a digest threshold (δdigest = 0.01) propagate deltas. Message size: 6 B/msg (4 B value + 2 B version dot). No incentive layer; convergence guaranteed by algebraic merge semantics. 4) DistributedKalmanFilter (Baseline 4): Each node runs a local Kalman filter (K = P/(P + R)) and exchanges estimate vectors with Metropolis weights (w = 1/(N + 1)). Message size: 10 B/msg (8 B estimate + 2 B covariance). Handles noise optimally but assumes all nodes are cooperative. 5) OCA (ours): Event-triggered delta generation with admissibility-weighted aggregation, version-vector causality tracking, retrospective oracle penalties, and trust-decay mechanism. Default parameters. Unless otherwise stated: N = 20 nodes, ε = 0.05, σ = 0.02, y ∗ = 1.0, λ = 1.0, η = 0.1, 2 κ = 0.1, Bmax = 1.0, Bmin = 0.0, σoracle = 2.0. B. Coordination Overhead Reduction
Figure 2 demonstrates that OCA reduces coordination overhead by over 99% relative to FullStateDissemination as organizational size grows from N = 5 to N = 50. At lim xt = 1x⋆ . (14) N = 20 (100 rounds), the total bytes transmitted are: OCA t→∞ (≈2,755 B) ≪ DeltaAntiEntropy (≈105,680 B) ≪ DeltaStateProof: Let Wt be the weight matrix at iteration t, where CRDT (≈171,118 B) ≪ DistributedKalmanFilter (≈380,000 B) √ ci t √ elements are defined as Wji =P ck for i ∈ Pj,t and 0 ≪ FullStateDissemination (≈456,000 B). The mechanism k∈Pj,t otherwise. The matrix W is strictly non-negative (Wji ≥ 0) operates through two complementary channels: PN and row-stochastic ( i=1 Wji = 1). Since the network is Exception-triggered suppression: By requiring local destrongly connected and has a self-loop (agents weight their viations to exceed τi before broadcasting, OCA actively own historical state), W represents an aperiodic, irreducible filters quiescent departments from the coordination loop. With Markov transition matrix. By the Perron-Frobenius theorem, perturbation noise σ = 0.02, only 15–30% of nodes cross their limt→∞ Wt = 1vT , where v is the unique left eigenvector thresholds per round on average, reducing active participants associated with the eigenvalue λ1 = 1. Consequently, the from N to approximately 0.2N –0.3N . PN (0) system converges to a weighted steady state x⋆ = i=1 vi xi , Versioned delta encoding: Unlike FullStateDissemination’s guaranteeing bounded-error alignment across the organization. O(N 2 ) message complexity, OCA’s delta payloads carry only state differences. The version vector mechanism ensures that
5
TABLE I TAXONOMY AND C OMMUNICATION M ODELS OF E VALUATED P ROTOCOLS
Protocol
Sync Trigger
Payload per Msg
Strategic Protection
FullStateDissemination Every Round N (N − 1) × 12B None DeltaAntiEntropy Periodic (5 rounds) N (N − 1) × 8B/5 None DeltaStateCRDT Digest Threshold count × (N − 1) × 6B None DistributedKalmanFilter Every Round N (N − 1) × 10B Noise Filtering Only OCA (Ours) Event-Triggered Variable Delta Oracle + Trust Slashing
(a) Bandwidth Savings vs. Network Size
(b) Convergence Delay vs. Network Size 200
0.8 OCA vs FullState OCA vs DeltaAntiEntropy OCA vs DeltaCRDT OCA vs DKF
0.6 0.4 0.2 0.0
Convergence Delay T" (rounds)
Bandwidth Savings SB
1.0
175 150 OCA Protocol FullStateDissemination DeltaAntiEntropy DeltaStateCRDT DistributedKF
125 100 75 50 25 0
10
20
30
Number of Nodes N
40
50
Fig. 2. Bandwidth savings SB versus organizational size N . OCA achieves 98–99% reduction compared to FullStateDissemination and 99% versus DistributedKalmanFilter, by suppressing redundant state exchanges through event-triggered deltas. The savings increase with network scale, demonstrating OCA’s suitability for large organisations.
10
20
30
Number of Nodes N
40
50
Fig. 3. Alignment latency Tε (rounds to ε-convergence) versus network size. All protocols start from a biased initial state (x̂g = 0.5, y ∗ = 1.0). FullStateDissemination converges fastest due to aggressive correction. OCA converges within 16 rounds at N = 20, comparable to DeltaAntiEntropy. DistributedKalmanFilter struggles with stable convergence under perturbation noise.
late-arriving or stale proposals are rejected with zero weight, 2 • Slow path: Retrospective penalties pi = λ∥x̂i − yi ∥ are preventing redundant reconciliation cycles. applied only when delayed oracles yi arrive (e.g., quarterly DeltaStateCRDT achieves moderate bandwidth savings via audits), avoiding per-round coordination stalls. its digest-threshold filtering, but still transmits to all N −1 peers FullStateDissemination achieves the fastest convergence when triggered. DistributedKalmanFilter requires full all-to-all estimate exchange every round (N (N −1) × 10 B), making it (Tε = 1) due to its aggressive correction weight (0.8 × x̄), but nearly as communication-intensive as FullStateDissemination. at the cost of O(N 2 ) bandwidth. DeltaAntiEntropy’s periodic Neither baseline possesses OCA’s strategic incentive layer, so sync every 5 rounds with moderate correction converges in 12 purely gossip-based protocols cannot deter structurally biased rounds. OCA follows closely at 16 rounds, benefiting from its trust-weighted aggregation. DeltaStateCRDT’s join-semilattice proposals driven by departmental self-interest. merge with a conservative 0.25 correction step yields slower convergence (76 rounds), while DistributedKalmanFilter’s C. Alignment Latency and Convergence Speed Kalman-filtered Metropolis consensus struggles to achieve Figure 3 shows alignment latency Tε across protocols under stable convergence within 100 rounds due to noise amplification (0) an initial bias of x̂g = 0.5 (i.e., a 50% deviation from ground under the perturbation regime, despite requiring constant alltruth). OCA converges within a bounded number of iterations to-all communication. that scales sub-linearly with N . The admissibility-weighted aggregation in (1) accelerates convergence by prioritizing historically accurate departments—high-trust nodes exert greater D. System Welfare Under Information Asymmetry √ influence on the global state via wi = ri ci , dampening Figure 4 evaluates system welfare Cε (fraction of nodes oscillations from noisy or strategic agents. within ε of ground truth) under varying sparsity ratios ρ, The latency advantage stems from decoupling fast local where effective perturbation noise scales as σ = σ · ρ. eff aggregation from slow global verification: OCA maintains high correctness scores (Cε ≥ 0.85) across all • Fast path: Rounds proceed without quorum voting, conditions, with three contributing factors: (t) allowing xi to track local conditions rapidly. The Admissibility filtering: Stale proposals are rejected by (t+1) (t) incremental update x̂g = 0.4 x̂g + 0.6 x̄w provides version vector comparison, preventing outdated information smooth convergence. from corrupting the consensus state. The dominance check
6
(c) Correctness vs. Update Sparsity
event-triggered delta propagation does not sacrifice alignment quality. Crucially, OCA is the only protocol that combines bandwidth efficiency with strategic resilience via its incentive mechanism.
Correctness Score C"
1.0 0.8 0.6 0.4
OCA Protocol FullStateDissemination DeltaAntiEntropy DeltaStateCRDT DistributedKF
F. Strategic Resilience Against Malicious Agents
0.2 0.0 0.2
0.4
0.6
Update Sparsity Ratio ½
0.8
1.0
Fig. 4. Welfare correctness score Cε across update sparsity ratios ρ. Effective noise scales as σeff = σ · ρ. OCA maintains stable alignment (Cε ≥ 0.85) across all conditions, demonstrating robustness against bounded rationality and information asymmetry.
Figure 5 presents the core strategic result: OCA’s retrospective penalty mechanism successfully deters self-interested manipulation even when 40% of agents inject structurally biased proposals. Non-cumulative bias injection: Malicious nodes inject bias via a non-cumulative mechanism—each round, the previous injection is subtracted before adding a fresh one, so local state tracks a stable offset rather than drifting unboundedly: (t)
(t)
(t−1)
xi ← xi − βi
(t)
+ βi ,
(t)
βi
= β · U(0.5, 1.5)
(15)
This models realistic adversaries who maintain a consistent strategic deviation rather than accumulating bias over time. Trust score decay: The left panel shows that malicious nodes (t) vi ≻ vilast ensures only causally fresh updates enter the experience rapid trust erosion following oracle verification. When the delayed ground truth yi reveals their bias, the penalty aggregation. Confidence-weighted aggregation: Nodes with high his- pi = λ∥xi − yi ∥2 slashes their effective stake via (4), reducing torical accuracy (large ci ) receive amplified influence through ci from 1.0 to below 0.4 within 200 rounds. √ wi = ri ci , dampening the impact of noisy or malicious Utility erosion: The right panel demonstrates the economic participants. consequence. Per-round utility follows: Residual penalty feedback: The quadratic penalty pi = Ui = Bmax · min(βi , 1.0) − pi − κ(1 − ci ) (16) ρr · ∥x̂i − x̂g ∥2 creates a negative feedback loop that suppresses systematic biases, where ρr = 0.1 is the residual penalty rate. where the benefit B max · min(βi , 1.0) scales linearly with At extreme sparsity (ρ < 0.3), all protocols degrade due to injected bias, p is the cumulative residual + oracle penalty, i information staleness. However, OCA’s degradation is more and κ(1 − c ) is the stake cost. Even though biased proposals i graceful—the trust decay mechanism in (4) automatically yield short-term local gains, the retrospective punishment p i reduces the weight of persistently inaccurate nodes, preserving dominates over macroscopic timescales, driving malicious consensus quality among the remaining reliable agents. Dis- utility negative. Honest agents receive baseline benefit B min tributedKalmanFilter exhibits unstable correctness under the and maintain stable positive utility. perturbation regime (Cε = 0 at ρ = 0.5), as the Kalman filter Honest agent protection: Crucially, honest agents maintain amplifies noise when assumptions about cooperative, Gaussianhigh trust scores (ci ≈ 1.0) and stable utility throughout the distributed inputs are violated. simulation. The admissibility-weighted aggregation naturally isolates malicious participants without requiring explicit ByzanE. Quantitative Summary tine fault tolerance mechanisms. This result confirms the mechanism design hypothesis: Table II summarises key performance metrics across all five bounded rationality and self-interest do not preclude consensus protocols at N = 20. stability when penalties are structurally linked to delayed verification. OCA transforms organizational constraints (quarTABLE II P ERFORMANCE M ETRICS S UMMARY (N = 20, ε = 0.05, 100 ROUNDS ) terly audits, realised market data) into strategic incentives that align local departmental objectives with global organizational Protocol SB (%) Tε Cε Bytes welfare. FullStateDissemination DeltaAntiEntropy DeltaStateCRDT DistributedKalmanFilter OCA (ours)
0 77 63 17 99
2 12 76 100 16
1.00 0.85 1.00 0.00 1.00
456,000 105,680 171,118 380,000 2,755
OCA achieves the optimal trade-off: highest overhead savings (99% vs. FullState, 98% vs. CRDT, 99% vs. DKF) while maintaining fast convergence (16 rounds) and perfect correctness (Cε = 1.00). OCA’s correctness matches FullStateDissemination and DeltaStateCRDT, demonstrating that
G. BNE Critical Threshold: Penalty Coefficient Sweep Figure 6 empirically verifies Theorem 1’s Bayesian Nash Equilibrium critical threshold: λmin =
Bmax − Bmin 1.0 − 0.0 = = 0.5 2 σoracle 2.0
(17)
Rational adversary model: For each λ, a rational adversary chooses optimal bias by maximising expected utility. The first-
7
(e) Utility Erosion for Malicious Nodes 0.5
0.8
0.4
Avg Utility Ui
Avg Trust Score ci
(d) Trust Score Decay 1.0
0.6 0.4
0.3
Malicious Nodes Honest Nodes
0.2 0.1
0.2 Malicious Nodes Honest Nodes
0.0 0.00
0.05
0.10
0.15
0.20
0.25
Malicious Node Ratio
0.30
0.35
0.0
0.40
0.00
0.05
0.10
0.15
0.20
0.25
Malicious Node Ratio
0.30
0.35
0.40
Fig. 5. Strategic resilience under malicious node injection. (Left panel) Average trust score ci decays for malicious agents as retrospective penalties erode their reputational stake. (Right panel) Corresponding utility erosion demonstrates that strategic bias injection becomes economically irrational under OCA’s incentive mechanism. Results shown for malicious ratios 0–40% with bias magnitude β = 0.5, λ = 1.0, η = 0.1.
(g) Utility Erosion: Ui vs ¸
(f) BNE Critical Threshold: Bias Cliff-Drop 0.30
0.5
System k¯i k2 ¸min = 0:50
Per-Round Utility Ui
System Avg Bias k¯i k2
0.25 0.20 0.15 0.10 0.05
Malicious Ui Honest Ui
0.4
¸min = 0:50
0.3 0.2 0.1 0.0 −0.1 −0.2
0.00 0.00
0.25
0.50
0.75
1.00
1.25
Penalty Coefficient ¸
1.50
1.75
2.00
−0.3
0.00
0.25
0.50
0.75
1.00
1.25
Penalty Coefficient ¸
1.50
1.75
2.00
Fig. 6. BNE critical threshold verification. (Left) System bias ∥βi ∥2 exhibits a cliff-drop at λmin = 0.5, confirming the phase transition predicted by Theorem 1. (Right) Malicious utility jumps to negative values at the same threshold, demonstrating that truth-telling becomes the dominant strategy. Parameters: 2 Bmax = 1.0, Bmin = 0.0, σoracle = 2.0, 30% malicious nodes, 30 trials per λ.
order condition d/dβ [Bmax β − λβ 2 ] = Bmax − 2λβ = 0 yields: ( Bmax −Bmin if λ < λmin 2 ∗ 2λσoracle β (λ) = (18) 0 if λ ≥ λmin When λ crosses λmin , the expected penalty exceeds the maximum achievable benefit, making bias unprofitable and truth-telling the dominant strategy. Phase transition (left panel): The system average bias ∥βi ∥2 remains high (0.2–0.5) for λ < 0.5, then exhibits a sharp cliff-drop to near-zero for λ ≥ 0.5. This discontinuous transition confirms the theoretical prediction: below λmin , rational adversaries maintain profitable bias; above it, bias is eliminated. The theoretical optimal-bias prediction b∗ (λ) closely tracks the empirical malicious bias, validating the rational adversary model. Utility inversion (right panel): Correspondingly, malicious utility is positive for λ < λmin (bias is profitable) and drops sharply to negative values for λ ≥ λmin (penalty dominates
benefit). Honest utility remains stable and positive throughout, as honest nodes receive baseline benefit Bmin without incurring oracle penalties. The crossover point aligns precisely with λmin = 0.5, providing strong empirical evidence for the BNE threshold. Implication: This experiment demonstrates that OCA does not require over-penalisation to achieve truthfulness. The critical threshold λmin provides a principled guideline for setting the penalty coefficient: any λ ≥ λmin suffices to align incentives, while excessive penalties unnecessarily punish honest agents who occasionally deviate due to noise. H. Parameter Sensitivity Analysis Figure 7 demonstrates that OCA’s truth-telling property emerges from incentive compatibility (IC) rather than over-penalisation, by sweeping the stake-cost coefficient κ ∈ {0.01, 0.05, 0.1, 0.2, 0.5, 1.0} and trust-decay rate η ∈ {0.01, 0.05, 0.1, 0.2, 0.5, 1.0} at the critical threshold λ = λmin = 0.5.
8
(h) Malicious Bias k¯i k 2 (∙ vs ´)
(i) Malicious Utility Ui (∙ vs ´)
1.00
1.00
0.4
0.24 0.50
0.3
0.50
0.20
0.10
Stake Cost ∙
Stake Cost ∙
0.22 0.20
0.2 0.20 0.1 0.10
0.05
0.18
0.05
0.01
0.16
0.01
0.01
0.05
0.10
0.20
Trust Decay Rate ´
0.50
1.00
0.0 −0.1 −0.2
0.01
0.05
0.10
0.20
Trust Decay Rate ´
0.50
1.00
Fig. 7. Parameter sensitivity across κ × η grid (6 × 6). (Top-left) Malicious bias remains uniformly low across all parameter combinations. (Top-right) Honest bias stays near zero. (Bottom-left) Malicious utility is negative across the entire grid, confirming incentive compatibility. (Bottom-right) Honest utility remains positive and stable. Fixed λ = 0.5 = λmin , 30% malicious, 10 trials per cell.
Malicious bias (top-left): Across the entire 6 × 6 grid, Simulation model: The Python prototype assumes Gaussian malicious bias ∥βi ∥2 remains uniformly low (< 0.1). This noise and continuous-valued state dynamics. Real organizaconfirms that at λ = λmin , the penalty mechanism alone tional environments may exhibit non-convex negotiation spaces, suffices to suppress strategic deviation regardless of κ and η discrete decision variables, or adversarial collusion patterns values. The stake cost κ(1 − ci ) provides a secondary deterrent (e.g., cartel coordination with alternating bias signs) that our but is not the primary driver of truthfulness. model partially captures through the CartelAgent and Honest bias (top-right): Honest bias stays near zero across StrategicSilenceAgent adversarial strategies but does all parameter combinations, demonstrating that the mechanism not exhaustively explore. Baseline assumptions: DistributedKalmanFilter assumes does not over-penalise well-behaved agents. Even at high κ and η values, honest nodes maintain minimal deviation from all nodes are cooperative and handles noise optimally via Kalman gains, making it vulnerable to strategic manipulation. ground truth. Malicious utility (bottom-left): Malicious utility is nega- DeltaStateCRDT’s join-semilattice merge guarantees eventual tive across the entire grid, confirming that bias injection is convergence but cannot reject biased deltas. These architectural economically irrational at λ = λmin regardless of stake cost limitations are inherent to the baseline designs and motivate and decay rate. The utility becomes more negative at higher κ OCA’s incentive-aware approach. Oracle availability: The retrospective penalty mechanism (increased stake cost for low-trust malicious nodes) and higher 2 depends on periodic ground truth signals yi ∼ N (y ∗ , σoracle ). If η (faster trust erosion triggers more penalties per oracle audit). delayed verification never arrives (e.g., unobservable outcomes), Honest utility (bottom-right): Honest utility remains positive and stable, with slight variation across the grid. At malicious agents face no consequences. This constraint aligns very high κ, honest agents with slightly imperfect trust scores with Milionis et al.’s identifiability condition [16]: truthfulness face marginally higher stake costs, but the effect is bounded. requires distinguishable signal distributions. Adaptive threshold extension: The This demonstrates the mechanism’s robustness: the IC property AdaptiveThreshold module, with threshold update holds across a wide parameter range without fine-tuning. Key finding: The sensitivity analysis addresses a critical rule τbase · (1 + α · σi (t)) concern—whether OCA’s truth-telling guarantee is fragile or τi (t) = , (19) ϵ + β · ci requires precise parameter calibration. The results show that at λ = λmin , the system maintains low bias and negative and the robust aggregation strategies (weighted geometric malicious utility across two orders of magnitude variation median via Weiszfeld algorithm, trimmed weighted mean) are in both κ and η. This robustness arises because the primary implemented in the src/oca/ package but not exercised in incentive alignment comes from the penalty-to-benefit ratio the main experiments. Future work should evaluate their impact (λ/Bmax ), while κ and η serve as secondary mechanisms that on convergence speed and adversarial resilience under cartel modulate trust dynamics without destabilising the equilibrium. collusion. Equilibrium guarantees: While our simulations demonstrate the BNE critical threshold λmin empirically (Fig. 6) I. Threats to Validity and parameter sensitivity robustness (Fig. 7), a formal We acknowledge several limitations in our experimental Bayesian Nash equilibrium proof mapping the strategy methodology: space to (16) is provided in supplementary materials (see
9
docs/proofs/bne_equilibrium.tex). The incentive compatibility is conditional on the stated model assumptions and does not establish a universal truthful equilibrium. VII. L IMITATIONS AND D ISCUSSION OCA operates as a framework for bounded-error coordination among rational agents rather than a replacement for cryptographic consensus that requires strict state machine replication. Architecturally, it links strategic sensor fusion and mechanism design for distributed averaging. Its key advantage lies in connecting these mechanism design tools directly with a resource commitment layer tailored for organizational hierarchies. Three primary limitations define future research paths. First, OCA restricts state domains to vector averaging, whereas nonconvex negotiation spaces like discrete contract terms demand distinct mechanism extensions. Second, the protocol couples high-frequency local aggregations with low-frequency, macrolevel penalties such as quarterly audits. Although our empirical simulations show robust stability, deriving tight theoretical convergence bounds for this timescale mismatch under severe systemic shocks remains open. Third, while our incentive model drives high empirical truth-telling rates, a full equilibrium proof that maps the complete strategy space to the utility model in (3) will further strengthen the theoretical foundation. VIII. C ONCLUSION This paper presents OCA, a game-theoretic coordination architecture for self-interested agents within heterogeneous organizations. By integrating local dispute triggers, incremental updates, and retrospective resource-slashing incentives, OCA suppresses redundant negotiations to reduce coordination overhead. OCA makes a deliberate trade-off. It trades the universal linearizability of traditional replicated state machines for bounded-error, incentive-aligned convergence. Compared to rigid consensus baselines, OCA proves that treating agents as rational, self-interested entities and penalizing them through delayed structural feedback offers a powerful strategy to eliminate bureaucratic bottlenecks and mitigate system welfare loss in decentralized governance. ACKNOWLEDGMENT The authors thank the distributed systems and computational social science communities for foundational work on consensus, mechanism design, and strategic agent modeling. R EFERENCES [1] L. Lamport, “The part time parliament,” ACM Transactions on Computer Systems, vol. 16, no. 2, pp. 133–169, May 1998, doi: 10.1145/279227.279229. [2] D. Ongaro and J. Ousterhout, “In search of an understandable consensus algorithm,” in Proc. USENIX Annual Technical Conference, 2014, pp. 305–319. [3] A. Mwotil, T. Anderson, B. Kanagwa, T. Stavrinos, and E. Bainomugisha, “LowPaxos: State machine replication for low resource settings,” IEEE Access, pp. 91272–91288, 2024, doi: 10.1109/ACCESS.2024.3421582.
[4] T. Mane, X. Li, M. Sadoghi, and M. Lesani, “Hamava: Fault tolerant reconfigurable geo replication on heterogeneous clusters,” in Proc. IEEE Int. Conf. on Data Engineering, 2025, pp. 2024–2037, doi: 10.1109/ICDE65448.2025.00154. [5] C. Hodsdon, Z. Qin, K. Ngo, S. Sen, E. Katz Bassett, and W. Lloyd, “Avicenna: Masking slowdowns in replicated state machines with counterfactual evaluation,” in Proc. European Conf. on Computer Systems, 2026, doi: 10.1145/3767295.3803615. [6] M. Shapiro, N. Preguiça, C. Baquero, and M. Zawirski, “Conflict free replicated data types,” in Stabilization, Safety, and Security of Distributed Systems, 2011, pp. 386–400, doi: 10.1007/978-3-642-24550-3.29. [7] P. S. Almeida, A. Shoker, and C. Baquero, “Efficient state based CRDTs by delta mutation,” in Proc. Int. Conf. on Networked Systems, 2014, pp. 62–76, doi: 10.1007/978-3-319-26850-7.5. [8] V. Enes, P. S. Almeida, C. Baquero, and J. Leitão, “Efficient synchronization of state based CRDTs,” in Proc. IEEE Int. Conf. on Data Engineering, 2019, pp. 148–159, doi: 10.1109/ICDE.2019.00022. [9] C. Baquero, P. S. Gomes, and M. B. Rodrigues, “ConflictSync: Bandwidth efficient synchronization of divergent state,” in Proc. Int. Workshop on Principles and Practice of Consistency for Distributed Data, 2025, doi: 10.1145/3806077.3806697. [10] L. Yang, Y. Gilad, and M. Alizadeh, “Practical rateless set reconciliation,” in Proc. ACM SIGCOMM, 2024, doi: 10.1145/3651890.3672219. [11] D. Selvi and G. Battistelli, “Distributed Kalman filtering with adaptive communication,” IEEE Control Systems Letters, pp. 15–20, 2025, doi: 10.1109/LCSYS.2025.3550401. [12] D. Bauso, L. Giarré, and R. Pesenti, “Mechanism design for optimal consensus problems,” in Proc. IEEE Conf. on Decision and Control, 2006, pp. 3381–3386, doi: 10.1109/CDC.2006.377206. [13] X. Bei, W. Chen, and J. Zhang, “Distributed consensus resilient to both crash failures and strategic manipulations,” arXiv:1203.4324, 2012. [14] K. Chen, D. G. Dobhakhshari, V. Gupta, and Y.-F. Huang, “An incentive scheme for sensor fusion with strategic sensors,” IEEE Transactions on Signal Processing, pp. 6342–6351, 2019, doi: 10.1109/TSP.2019.2954974. [15] K. Teranishi, K. Kogiso, and T. Tanaka, “Faithful and privacy preserving implementation of average consensus,” in Proc. American Control Conf., 2025, pp. 2937–2942, doi: 10.23919/ACC63710.2025.11107548. [16] J. Milionis, J. Ernstberger, J. Bonneau, S. D. Kominers, and T. Roughgarden, “Incentive compatible recovery from manipulated signals, with applications to decentralized physical infrastructure,” arXiv:2503.07558, 2025, doi: 10.48550/arXiv.2503.07558. [17] C.-C. Chen and W. Golab, “A tunable incentive mechanism for binary aggregation without verification,” arXiv:2606.30974, 2026. [18] K. Chen, C. Huang, and J. Huang, “Decentralized information elicitation without verification,” IEEE Transactions on Networking, pp. 4387–4402, 2026, doi: 10.1109/TON.2026.3674784. [19] M. Eischer and T. Distler, “Scalable Byzantine fault tolerant state machine replication on heterogeneous servers,” Computing, pp. 97–118, 2018, doi: 10.1007/s00607-018-0652-3. [20] K. Ngo, S. Sen, and W. Lloyd, “Tolerating slowdowns in replicated state machines using copilots,” in Proc. USENIX Symposium on Operating Systems Design and Implementation, 2020, pp. 583–598. [21] D. Cason, N. Milošević, Z. Milošević, and F. Pedone, “Gossip consensus,” in Proc. ACM Middleware, 2021, doi: 10.1145/3464298.3493395.