Conceptio › Archive › arXiv CS
arXiv CSopen access

Adaptive Bayesian Partner Selection for Federated Clinical Centers

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

Adaptive Bayesian Partner Selection for Federated Clinical Centers

arXiv:2609.16446v1 [cs.LG] 15 Sep 2026

Navid Seidi Department of Computer Science Missouri University of Science and Technology Rolla, MO, USA [email protected]

Satyaki Roy Department of Mathematical Sciences University of Alabama in Huntsville Huntsville, AL, USA [email protected]

Sajal K. Das Department of Computer Science Missouri University of Science and Technology Rolla, MO, USA [email protected]

Abstract Federated learning (FL) in healthcare is challenged by pronounced heterogeneity and temporal concept drift across clinical centers, where evolving patient populations and care practices shift data distributions. Existing approaches rely on persistent global communication, incurring substantial bandwidth overhead while risking negative transfer from poorly aligned peers. We address this by proposing A DAPTIVE BAYESIAN PARTNER S ELECTION (ABPS), a peer-to-peer framework that governs who collaborates, when, and at what cost. Each center maintains a Beta–Bernoulli posterior over prospective peers’ Shapley marginal utility, ranks candidates using an Upper Confidence Bound (UCB) criterion, and forms collaborations through a lightweight propose–reject mechanism, with the option to abstain from communication when mutually beneficial interaction fails. It admits a stochastic decision interpretation, yielding finite-sample concentration guarantees and O(κ log T ) regret in partner selection, along with conditions under which intentional isolation is optimal under negative transfer. Lightweight extensions, head personalization, bfloat16 quantized communication, and tunable active-set cardinality further improve efficiency, while a goal-aware metadata filter enables institution-specific collaboration strategies. Evaluations on binary in-hospital mortality prediction over the first 24 hours of an ICU stay, with 230 non-IID clinical centers drawn from MIMIC-IV, show that the full ABPS-X variant matches the strongest federated baseline (FedDyn, AUROC 0.758) at 0.09× the communication cost of FedAvg, with reduced variability. A diversity-driven configuration activates intentional isolation for a substantial fraction of centers, highlighting the role of selective collaboration under heterogeneity. These results show that adaptive, utility-aware collaboration reduces communication without sacrificing accuracy when centers are numerous and small, providing a scalable paradigm for real-world healthcare FL systems.

1

Introduction

Federated Learning (FL) has emerged as the dominant paradigm for training predictive models across decentralized centers without exposing sensitive patient data McMahan et al. [2017], sidestepping Health Insurance Portability and Accountability Act (HIPAA)-style data-sharing constraints. Realizing FL in clinical settings is nevertheless gated by two persistent obstacles: statistical heterogeneity (i.e., non-IID biomedical data) across institutions Hsu et al. [2019], and concept drift as patient demographics and clinical protocols evolve Gama et al. [2014], Rahimli et al. [2024] (see Appendix A for an instance of distributional drift across admission eras in routinely charted vital signs). Preprint.

This heterogeneity is structural rather than incidental. The healthcare ecosystem comprises three institution types with divergent key performance indicators (KPIs) and data profiles. Integrated delivery networks (Kaiser Permanente, HCA Healthcare) aggregate large but internally siloed cohorts. Academic medical centers see complex, rare, and trial-driven cases at low volume. Community hospitals generate the majority of high-volume routine encounters. No single ecosystem captures the full distribution required for generalizable modeling, yet forcing central FL aggregation across them frequently induces negative transfer that degrades local performance. Peer-to-peer (P2P) alternatives Guha Roy et al. [2019], Hegedűs et al. [2021] avoid the single point of failure but still face the question of which peers a center should collaborate with and at what bandwidth cost. Contributions. To achieve predictive accuracy under concept drift, we introduce A DAPTIVE BAYESIAN PARTNER S ELECTION (ABPS), a serverless P2P federated learning framework in which each center adaptively selects collaborators based on a Bayesian estimate of their utility. ABPS models each candidate peer’s Shapley marginal contribution via a Beta–Bernoulli posterior, ranks peers using an ϵ-greedy UCB policy, and employs a propose–reject protocol with an explicit rest action to mitigate negative transfer. We establish three guarantees: posterior concentration (Theorem 1), sublinear regret of order O(κ log T ) (Theorem 2), and the Bayes-optimality of isolation (Lemma 1). The core framework comprises head personalization, bfloat16 Kalamkar et al. [2019] quantized communication, and a tunable active-set cardinality κ. In addition, a goal-aware metadata pre-filter allows each center to prioritize homogeneity, diversity, or KPI alignment objectives. We denote the maximally-extended configuration as ABPS-X, which composes all of the above with deeper 2-layer head personalization, server-side exponential moving average (EMA) momentum, and validation-AUROC early stopping (full hyperparameters in Sec. 6.4). On MIMIC-IV v3.1, with n=230 non-IID careunit-by-year centers, ABPS-X matches the strongest federated baseline (FedDyn) while using only 0.09× the bandwidth of FedAvg. In contrast, applying the same extensions to FedAvg degrades performance by 6.4 AUROC points, leaving ABPS-X ahead of FedAvg-X by 6.7 points (Table 1), thereby isolating the gains to Bayesian partner-selection.

2

Related Work

Our research intersects with several active domains in decentralized machine learning. We organize the related work below in the order in which the corresponding open Question is answered in the body of the paper, ranging from divergence-aware partner selection to peer-to-peer (P2P) federated architectures, concept-drift adaptation, and communication-efficient FL. 2.1 Divergence Metrics for Non-IID Collaboration Information-theoretic divergences (Kullback-Leibler, hereafter KL, and Wasserstein) quantify interclient distributional disparity in FL Németh et al. [2025]. They appear as regularization penalties that align local models without raw data exposure Németh et al. [2025], as client-selection utilities that favor either stabilizing or diversity-injecting clients Düsing and Cimiano [2026], Rahad et al. [2025], and as theoretical tools for bounding FL generalization under heterogeneity. This philosophy extends to the wire-format level: each center can be summarized by a non-PHI descriptor of its distribution (e.g., positive-class rate, training-set size, optional Charlson-comorbidity prevalence), and peers can be admitted via a goal-dependent monotone function of pairwise similarity. A key limitation is that divergence is treated as a single scalar objective fixed network-wide, without allowing individual centers to declare their own collaboration goal before model exchange. Yet healthcare institution types (Integrated Delivery Networks, Academic Medical Centers, Community Hospitals) require different collaboration objectives, including homogeneity, diversity, or KPI alignment Li et al. [2025]. We label this open problem Q1 (goal-heterogeneous collaboration) and address it through the goal-aware metadata pre-filter of Sec. 4.1, where a center selects one of three admission rules, namely, cosine, anti-cosine, and KPI matching, at runtime without risking exposing model parameters. 2.2 P2P Federated Learning and Partner Selection Peer-to-peer (P2P) FL decentralizes aggregation, eliminating single points of failure and reducing communication bottlenecks Hegedűs et al. [2021], Zhou et al. [2024]. However, identifying the appropriate collaboration topology under non-IID heterogeneity remains an open challenge. Early approaches relied on gossip-based protocols Hegedűs et al. [2021], while more recent work studies the security of P2P training against backdoor attacks Syros et al. [2024] and robustness mechanisms against free-riding and collusion Augello et al. [2024], Ranjan et al. [2022]. Within the multi-armed bandit (MAB) framework, CS-UCB Xia et al. [2020] introduced UCB-style scheduling in servermediated FL, and subsequent work extended this idea to decentralized settings with non-stationary MAB formulations over peer groups Listo Zec et al. [2024]. Complementary approaches based on Shapley value estimate client contributions for server-side selection Singhal et al. [2024], Yang 2

et al. [2024]. A second limitation persists across P2P, MAB-, and Shapley-based methods: all assume mandatory participation in each round. Even decentralized variants enforce collaboration once the topology is established, despite evidence that aggregation across heterogeneous institutions can degrade local performance below the no-collaboration baseline Crowson et al. [2022], Li et al. [2025]. We define this as Q2 (selective collaboration under heterogeneity) and address it through the Bayesian propose–reject mechanism (discussed in Sec. 4.3), allowing each center to abstain from participation when all posterior UCB scores fall below the acceptance threshold τacc . 2.3 Concept Drift and Bayesian Adaptation Clinical data distributions evolve due to changes in patient demographics, treatment protocols, and clinical infrastructure, leading to concept drift Rahimli et al. [2024]. Bayesian methods quantify the resulting uncertainty, with two dominant approaches in prior work. The first models uncertainty over model parameters, where Bayesian neural networks in FL aggregate posterior distributions rather than point estimates Saile et al. [2024], Rahman et al. [2025], often incorporating hierarchical or personalized updates to adapt to local data. The second leverages uncertainty for update filtering, where client contributions are selectively incorporated based on confidence measures, such as credibleinterval thresholds Iglesias Jr. et al. [2024]. Despite their effectiveness, both approaches share a key limitation: abstaining from collaboration is not treated as a principled decision with formal guarantees. In practice, concept drift can render collaboration rounds detrimental, yet existing methods rarely allow nodes to opt out in a theoretically grounded manner Crowson et al. [2022], Gama et al. [2014], Rahimli et al. [2024]. We define this as Q3 (adaptive isolation under drift) and address it through an explicit rest action in Sec. 4.3, supported by a Bayes-optimal isolation criterion (Lemma 1) that provides a closed-form condition under which abstention strictly outperforms collaboration. 2.4 Personalization, Quantization, and Communication-Efficient FL To achieve personalization and bandwidth minimization, it is necessary to keep a subset of model parameters local per client: FedPer and FedRep retain the classifier head locally Arivazhagan et al. [2019]. pFedHN and Ditto provide other per-client adaptations. Quantized transmission compresses the wire format: LLM.int8 Dettmers et al. [2022], QSGD, and signSGD demonstrate single-digitpercent accuracy loss with 2-16× bandwidth reduction. Server momentum (FedAvgM) stabilizes aggregation under heterogeneity. These ideas are orthogonal to partner selection, and in principle, they can be composed with any partner-selection core to push the accuracy-bandwidth frontier further. A fourth limitation runs across all three of these families: bandwidth-saving mechanisms have been studied in isolation from which peers a center should engage with. Real-world multi-institutional medical FL still incurs prohibitive bandwidth Haripriya et al. [2025], and communication-efficient extensions such as knowledge distillation Wu et al. [2022] and metaheuristic aggregation Abdolmaleki and Farahani [2026] treat compression and topology choice as separate problems. We label this open problem Q4 (communication overhead, even in P2P) and answer it through the additive headpersonalization, bfloat16-quantization, and tunable active-set cardinality κ, optimized jointly on the accuracy-bandwidth Pareto frontier (Sec. 6.4, with the row-by-row ablation in Sec. 6.5).

3

Problem Formulation

Goal. We seek to learn, for each of n federated centers, a sequence of local predictors that maximizes predictive accuracy on a center’s evolving data distribution while minimizing the cumulative communication overhead of inter-center collaboration, under continuous concept drift in patient demographics, treatment protocols, and equipment. Predictive accuracy is task-dependent and enters the framework through a per-center utility functional Ui ∈ [0, 1]. Ui is the per-center Area Under the Receiver Operating Characteristic curve (AUROC) for binary in-hospital mortality classification. At each round t ∈ {1, . . . , R}, every center i ∈ {1, . . . , n} holds local data drawn from an unknown, time-varying joint distribution Pt(i) (X, Y ) shaped by the center’s patient demographics, comorbidity profile, and care protocols, observed only through patient-level samples, fits local parameters wi,t ∈ R|W | , and may exchange them with a per-round active peer set Pi,t ⊆ {1, . . . , n} \ {i}. We write Ai (wi,t ) := E(x,y)∼P (i) [AUROC(wi,t ; x, y)] for the expected accuracy of i’s model and treat t Pi,t = ∅ as a legal first-class action (intentional rest, formalized in Lemma 1). ABPS jointly chooses, for every center and every round, the active peer set and the local update rule, solving n

R

1X 1 X Ai (wi,t ) {Pi,t , wi,t } n R t=1 i=1 | {z } max

subject to

h i (i) (i) Ctotal ≤ B , supm,k D Pm,k ∥ Pm,k−1 ≤ ϵm . | {z } | {z }

bandwidth budget

(1)

drift bound at scale m

predictive accuracy

The maximand is the across-center, across-round mean accuracy. Ctotal ≤ B caps cumulative communication (Eq. 2 below). The drift constraint bounds the per-window divergence by a scalespecific tolerance ϵm at each temporal scale m ∈ {1, 2, 3}. When a realized shift exceeds ϵm , A3 3

(Sec. 5) breaks at scale m, the Bayesian posterior decays toward the prior, the UCB rankings re-order, and the propose-reject mechanism re-forms Pi,t+1 . Equation 1 thus makes the three competing forces explicit: accuracy as the maximand, bandwidth as the budget, and drift as the constraint. Communication cost. Let Pt = {(i, j) : j ∈ Pi,t } denote the active peer-pair set in round t, |Wi | the size of center i’s parameter vector, and cegress the per-GB egress cost (e.g., $0.09/GB on AWS Amazon Web Services [2024]). It is worth mentioning here that continuously broadcasting parameters across 100,000+ global healthcare centers typically incurs a prohibitive bandwidth footprint McMahan et al. [2017], and recent medical-FL benchmarks Haripriya et al. [2025] confirm that one 50-round sweep of a VGG-16-class model exceeds 276,000 MB of cross-client traffic. Distillation Wu et al. [2022] and bandwidth-aware aggregation Abdolmaleki and Farahani [2026] cut these figures, but neither couples compression to peer choice. The cumulative cost across R rounds is Ctotal = cegress

R X

X

 |Wi | + |Wj | .

(2)

t=1 (i,j)∈Pt

In Eq. 2, the summand accounts for the bidirectional transfer of parameter aggregation. Under the P uniform-architecture specialization |Wi |=|W |, this reduces to Ctotal = 2 cegress |W | t |Pt |, the form used in Theorem 2 and Lemma 1. Hierarchical concept drift. The drift constraint of Eq. 1 bounds, at three temporal scales m ∈ {1, 2, 3} (short-term operational, mediumterm seasonal, long-term demographic), the interwindow divergence (i)

(i)

D[Pm,k ∥ Pm,k−1 ] ≤ ϵm

admission year (MIMIC-IV v3.1 privacy-shifted → 2-year windows = centers) 2110

2130

2160

2188

MICU (pos ≈ 15.4%) Medical/Surgical ICU (pos ≈ 14.1%) inter-careunit divergence (level h3 )

CCU (pos ≈ 13.0%) TSICU (pos ≈ 10.8%) SICU (pos ≈ 10.6%) CVICU (pos ≈ 3.3%) intra-careunit temporal drift (level h2 ), bounded by ϵ2 cell shade ∝ per-center mortality rate: 3%

8%

13%

18%

level h1 (within-window noise, ϵ1 ): per-round local-SGD stochasticity, not depicted.

between consecutive windows of length ∆tm Figure 1: Hierarchical temporal-shift model, illustrated on MIMIC-IV v3.1 as one concrete **(Figure 1)** Gama instantiation. The construction is dataset-agnostic and applies to any longitudinal multi-center cohort. In this instantiation each cell is one of n=230 federated centers (careunit × 2-year window). Cell et al. [2014], Rahimli shade encodes in-hospital mortality rate, and rows sort six careunits by mean mortality. Horizontal et al. [2024], Cuturi and variation within a row illustrates the level-h2 within-group term bounded by ϵ2 ; because the windows are shifted-year buckets rather than calendar years, on this axis that variation is of the order of sampling Blondel [2017], Düsing variation (Sec. 6). Vertical shade jumps show level-h3 inter-group divergence that goal-aware pre-filter and Cimiano [2026], (Sec. 4.1) reasons about without exchanging model parameters. Rahad et al. [2025]. The window size ∆tm admits an online empirical estimator, e.g., a Welch t-test q 2 /Nt+k stays below tcrit at α=0.05. (Note that expands the window while |µt − µt+k |/ σt2 /Nt + σt+k that during experimental validation, we fix ∆t2 =2 years and leave adaptive estimation of {∆tm } to future work, since online window adaptation is orthogonal to the partner-selection contribution we focus on here, and Theorem 1 already bounds the within-window cost explicitly through ϵm , so the headline accuracy-bandwidth claim is unaffected by any reasonable choice of ∆t2 .)

4

Methodology: Adaptive Bayesian Partner Selection

2. Topology ABPS models inter-center collaboration as a prob- 1. Goal-Aware Filter ϵ-greedy UCB all reject f (v , v ) ≥ τ propose/reject abilistic, temporally adaptive process built from Rest P = ∅ Bayesian-core components, a goal-aware meta3. Utility 4. Beta Update data pre-filter (Sec. 4.1), a Beta-Bernoulli poste- Beta(α+ Shapley ϕ ϕ̄, β+1−ϕ̄) rior over each peer’s marginal utility (Sec. 4.2), an ϵ-greedy UCB propose-reject topology rule Figure 2: ABPS round-level pipeline. Centers prune candidates a goal-aware metadata filter (Eq. 3), propose to peers ranked by (Sec. 4.3), and a rest action when no proposal via UCB on Beta belief, observe Shapley marginal utility, and update the is accepted, and three communication-efficiency belief. When all proposals are rejected, the center rests with Pi = ∅ extensions that compose additively on top of the (Lemma 1). core, namely, head personalization, bfloat16 quantization, and tunable κ, are all discussed hereafter and analyzed in Sec. 6.4). Figure 2 summarizes the round-level pipeline, and the full pseudocode is given in Algorithm 1 in Appendix B. 4.1 Goal-Aware Metadata Pre-Filter To address Q1 on goal-heterogeneous collaboration from Sec. 2, before invoking the expensive Shapley-UCB stage, each center i computes a d-dimensional summary metadata vector vi ∈ Rd of non-sensitive aggregates (positive-class rate, log ntrain , normalized mean year of admission, and for goal-aware variants a 12-dim Charlson-comorbidity prevalence). Every entry is a cohort-level scalar, not a per-patient feature, so vi carries no Protected Health Information. No model parameters cross i

i

j

sim

i

j→i

4

the network at this stage, consistent with HIPAA and inter-institutional data-sharing constraints. Goal-aware admission rule. Real healthcare federations are not uniformly similarity-seeking: the three institution types in Sec. 1 have distinct objectives. A community hospital prefers homogeneity, collaborating with demographically similar peers to improve transfer learning for under-represented cohorts. An academic medical center seeks diversity, leveraging exposure to heterogeneous case-mix. In contrast, an Integrated Delivery Network emphasizes alignment, selecting peers with matching target KPIs (e.g., mortality rates) to meet system-level objectives. We generalize the admission rule into a per-center goal-scoring function  fi : Rd × Rd → R: fi (vi , vj ) =

 goal = homogeneity cos(vi , vj ), − cos(vi , vj ), goal = diversity   −|vi,0 − vj,0 |, goal = alignment

(3)

In the above equation, v·,0 is the pos-rate coordinate used as the KPI marker. Peer j is admitted to Ci iff fi (vi , vj ) ≥ τsim . Each formulation lives on a different scale, so τsim is auto-calibrated per goal to retain a fixed fraction (we use 25%) of candidate pairs: this makes goal choices directly comparable. The threshold can equivalently be learned online via an exponential moving average (EMA) on the acceptance rate. The static calibration suffices for our experiments. Complexity and theoretical preservation. The pre-filter is O(nd) per round and transmits only vi (a few dozen bytes) per center. All downstream Shapley evaluations operate on the filtered arm set, reducing the expected per-round bandwidth by the same fraction. Theorems 1 and 2 transfer verbatim once the arm set is restricted to Ci . The only added bias is whether the true best peer survives the filter, which we bound in Appendix C via a standard top-k coverage argument for descriptor-based recall. Empirically (Sec. 6.4), the diversity goal activates Lemma 1’s |Pi | = 0 regime in ∼ 40% of rounds on the biomedical dataset, the first empirical observation of intentional isolation in our study. 4.2 Bayesian Belief Updating on Marginal Utility Rather than maintaining a belief over a peer’s explicitly shared raw parameters, center i evaluates the actual marginal utility (performance gain) a peer j provides. We model this utility using the Shapley value to objectively allocate credit across historical collaboration subsets S ⊆ Pi \ {j}: ϕj→i =

X S⊆Pi \{j}

|S|! (|Pi | − |S| − 1)! (Ui (S ∪ {j}) − Ui (S)) , |Pi |!

(4)

In the above equation, Ui (·) ∈ [0, 1] is the local model’s normalized held-out validation performance (we use AUROC). Because Ui is bounded in [0, 1], the per-peer Shapley contribution ϕj→i is bounded in [−1, +1]. We map it to the unit interval via the affine clip  ϕ̄j→i = clip[0,1]

ϕj→i − ϕmin ϕmax − ϕmin

 ,

(5)

with truncation thresholds ϕmin , ϕmax chosen so that a single anomalous round cannot saturate the posterior (we use [−0.1, +0.1] throughout this paper). Modeling ϕ̄j→i as a soft Bernoulli observation, the conjugate Beta prior ϕ̄j→i ∼ Beta(αi→j , βi→j ) admits the closed-form update (k+1)

αi→j

(k)

← αi→j + ϕ̄j→i ,

(k+1)

βi→j

(k)

← βi→j + (1 − ϕ̄j→i ).

(6)

The Beta-Bernoulli formulation offers three properties. (i) The posterior support [0, 1] matches the bounded co-domain of ϕ̄j→i , ruling out the model mis-specification that would otherwise arise when Shapley contributions are negative. (ii) The posterior variance αβ/[(α+β)2 (α+β+1)] admits the Hoeffding-style concentration exploited in Theorem 1. (iii) The upper-confidence bound(in Sec. 4.3) for partner ranking takes the standard UCB1 form and inherits sublinear regret (Theorem 2). 4.3 Collaborative Topology Formulation Recall Q2 (on selective collaboration under heterogeneity) and Q3 (on adaptive isolation under drift) from Sec. 2. The propose-reject protocol below operationalizes both, with the explicit rest action carrying through to the optimality result of Lemma 1. The global time horizon T is partitioned into discrete interaction intervals {t1 , . . . , tK }. Collaboration decisions among the centers are governed by a propose-and-reject protocol approximating a stable-marriage solution, augmented with an ϵ-greedy Upper Confidence Bound (UCB) strategy Auer et al. [2002]. During each period tk , center i computes a utility-based UCB score for every candidate peer j : s UCBj→i (tk ) = µ̂i→j (tk ) + γ

2 log tk , ni→j (tk ) + 1

(7)

(tk ) (tk ) (tk ) Here, µ̂i→j (tk ) = αi→j / αi→j + βi→j is the posterior mean of the Beta belief (6), ni→j (tk ) is the number of past observations from j , and γ > 0 is the exploration weight. With probability 1 − ϵ, center i sequentially proposes collaboration to peers sorted by descending UCBj→i . With probability ϵ, it purposefully queries a uniformly random unproven neighbor. A receiving peer j evaluates the



5

proposal and accepts only if its own UCB on i exceeds the threshold τacc . If j rejects, i extends the proposal to the next highest-ranked peer until either κ acceptances accrue or the candidate list is exhausted. Critically, if all proposals are rejected, center i rests: its active peer set Pi collapses to ∅ and no parameters are exchanged this round. This isolation mechanism is a deliberate design feature, not a system failure. In highly heterogeneous networks, forcing connections with dissimilar nodes degrades local performance via negative transfer. Choosing isolation whenever the expected gain falls below the threshold helps protect local utility and conserves bandwidth (Lemma 1). Resting is per-round and self-correcting. The next round re-runs the propose-reject loop with an updated posterior and a UCB exploration radius that grows in t, and the ϵ-greedy swap admits a uniformlyrandom peer with probability ϵ regardless of UCB ranking, so collaboration resumes the moment any peer’s UCB rises above τacc or drift makes a previously-low-utility peer informative again.

5

Theoretical Guarantees and Complexity Analysis

We analyze ABPS along posterior concentration (Theorem 1), partner-selection regret (Theorem 2), and Bayes-optimality of isolation (Lemma 1). Proofs are in Appendix C. The analysis rests on three assumptions. A1 bounded utility, Ui ∈ [0, 1], so ϕ̄j→i ∈ [0, 1] via the affine clip of Eq. 5 (see Sec. 4.2). A2 conditional independence of the ϕ̄ observations given µ⋆i→j , justified by independent local stochastic gradient descent draws and disjoint mini-batches per round. A3 quasi-stationary drift (k) |E[ϕ̄j→i ] − µ⋆i→j | ≤ ϵm within any window ∆tm , binding the theory to the hierarchical formulation of Sec. 3. For the filtered-arm-set refinement of Theorem 2, we additionally assume A4 that the metadata descriptor vi is L-Lipschitz informative about utility, |µ⋆i→j − µ⋆i→j ′ | ≤ L ∥vj − vj ′ ∥ (see Appendix C.2 for the precise statement and use). 5.1 Convergence of the Beta Posterior (K) Theorem 1 (Beta-belief concentration). Let ϕ̄(1) j→i , . . . , ϕ̄j→i ∈ [0, 1] be the clipped Shapley observations collected by center i for peer j across K rounds inside a single drift window, with empirical P (k) α0 +K µ̄K mean µ̄K = K1 K k=1 ϕ̄j→i and posterior mean µ̂K = α0 +β0 +K . Under A1-A3, for any δ ∈ (0, 1) with probability at least 1 − δ r µ̂K − µ⋆i→j

≤

α0 + β 0 log(2/δ) + + ϵm . 2K α0 + β 0 + K {z } | |{z} | {z } Hoeffding

prior decay

(8)

drift

p

Hence µ̂K − → µ⋆i→j as K → ∞ with ϵm → 0. √ The proof (Appendix C.1) decomposes the error into Hoeffding noise (O(1/ K)), prior decay (O(1/K)), and within-window drift bounded by ϵm , the last of which motivates the Welch-t window selection of Sec. 3. 5.2 Regret of the UCB Partner Selection Bandit We bound the regret of the per-center partner-selection problem viewed as a κ-armed banditLattimore and Szepesvári [2020] over the candidate pool Ci . Let µ⋆(1) ≥ µ⋆(2) ≥ . . . denote the ordered true utilities of i’s candidate peers and define the suboptimality gaps ∆j = µ⋆(1) − µ⋆(j) . Theorem 2 (Sublinear partner-selection regret). Under assumptions A1-A2 and a√stationary window (ϵm = 0), the cumulative regret of ABPS with exploration parameter γ = 2 and ϵ-greedy exploration probability ϵ ∈ [0, 1) over T rounds satisfies R(T ) ≤ κ

  X  8 log T π2 + ∆j 1 + + ϵ T · Ej∼Unif(Ci ) [∆j ]. ∆j 3 j:∆ >0

(9)

j

√

Setting ϵ = O(1/ T ) recovers the standard O(log T ) regret rate of UCB1 up to the factor κ. The proof (Appendix C.2) bounds the greedy phase via canonical UCB1 analysis Auer et al. [2002] scaled by κ and the exploratory phase via the ϵT · E[∆j ] term, with a filtered-arm-set refinement under A4 that improves the constants when the goal-aware pre-filter is active. In stationary windows, ABPS therefore matches an oracle that always proposes to the κ best peers. 5.3 Optimality of Intentional Isolation Recall Q3 (adaptive isolation under drift) from Sec. 2. The lemma below is the formal optimality guarantee that the rest of the action of Sec. 4.3 promised. Lemma 1 (Intentional isolation dominates forced collaboration). Let Virest denote the expected one-step utility of center i when it rests (Pi = ∅) and Vicoll (P) the expected utility when it forcibly collaborates with peer set P ̸= ∅. Suppose the per-peer expected marginal utility satisfies µ⋆i→j < cneg 6

for every j ∈ P , where cneg is the negative-transfer threshold defined by Vicoll ({j}) = Virest when µ⋆i→j = cneg . Then rest coll Vi

> Vi

(P),

(10)

Cirest = 2|P| · cegress · |W |

and the bandwidth saved by resting is exactly per round (cf. eq. (1)). The proof (Appendix C.3) combines Shapley efficiency with the affine clip of Eq. 5, which fixes cneg = −ϕmin /(ϕmax −ϕmin ) = 0.5 for (ϕmin , ϕmax ) = (−0.1, +0.1), exactly τacc . ABPS approximates this oracle by resting whenever every UCB falls below τacc , and Theorem 1 guarantees convergence to the oracle as K → ∞. 5.4 Computational Complexity Let ρ ∈ (0, 1] denote the expected top-k coverage of the goal-aware pre-filter (Sec. 4.1), i.e. the fraction of peers that survive the τsim threshold. When τsim is calibrated to retain a prescribed keepfraction, ρ equals that target (we use ρ = 0.25). Define the post-filter candidate size |Ci | ≈ ρ(n − 1). The per-round computational cost at each center is strictly bounded and decomposes into four components. (1) Goal-aware filter: Evaluating fi (vi , vj ) for each of the n − 1 peers using d-dimensional metadata incurs a cost of O(nd), where d ≤ 15 in our implementation (3 base features and 12 comorbidity indicators); (2) UCB ranking and propose–reject: Ranking candidates in the filtered set |Ci | ≈ ρn and traversing the ordered list requires O(ρn log(ρn)) time requires a ρ-factor reduction compared to the unfiltered case; (3) Shapley credit assignment: The exact computation of this involves evaluating 2|Pi | coalitions. Since |Pi | ≤ κ by design, the cost is O(2κ ), independent of n. For κ > 5, we instead employ Truncated Monte Carlo (TMC) Shapley Ghorbani and Zou [2019], which achieves an ε-accurate estimate in O(κ log κ/ε2 ) samples; and (4) Beta–Bernoulli update: Posterior updates require only closed-form scalar operations, yielding O(1) complexity. The total per-round computation at center i is therefore O(nd + ρn log(ρn) + 2κ ). For the default ρ = 0.25 and κ = 1 this collapses to O(nd), independent of the number of model parameters |W |. Communication. Each round transmits (i) the metadata vector vi once at filter time (a few dozen bytes per center, no model weights) and (ii) the active-peer aggregation for the κ peers that survive both the filter and the UCB threshold, costing 2κ|W | per active center per round. With personalization (Sec. 6.4), only the shared trunk is transmitted. With quantization, each |W | is reduced to its bfloat16 footprint. When |Pi | = 0 (intentional isolation, empirically realized under the diversity goal), the communication collapses to zero for that center-round.

6

Experimental Evaluation

We evaluate ABPS along four axes: (i) predictive accuracy under realistic non-IID partitioning of MIMIC-IV, a publicly available database sourced from the Beth Israel Deaconess Medical Center (BIDMC) electronic health record Johnson et al. [2023], (ii) cumulative transmitted bytes as a direct proxy for communication cost, (iii) the empirical realization of the intentional-isolation property of Lemma 1, and (iv) an ablation decomposing the contribution of each framework extension. Code, sbatch scripts, and per-seed per-method result JSONs will accompany the camera-ready submission. 6.1 Dataset and Federation Setup The task at hand is the binary in-hospital mortality prediction over the first 24 hours of an ICU stay Johnson et al. [2023]. After applying the standard age filter (18 ≤ age ≤ 95), the cohort contains roughly 76,000 ICU stays. The patient features: demographics (age, gender, race), admission context (admission type, location, insurance), Charlson-style comorbidity binaries, discussed in Sec. 6.3, derived from ICD-10/ICD-9 codes, and aggregated first-24h vitals (heart rate, systolic/diastolic/mean BP, respiration rate, SpO2 , temperature) summarized as {mean, min, max}. Continuous features are standardized per center on the local training split to respect federated isolation. Our experiments rely on a careunit-by-year partitioning of MIMIC-IV, where a center is defined as the Cartesian product of a care unit and a consecutive two-year admission window. The care units include CVICU, CCU, MICU, Medical/Surgical ICU, SICU, and TSICU, yielding n = 230 centers spanning the shifted temporal range 2110 to 2191 in MIMIC-IV v3.1. MIMIC-IV de-identifies dates by a single random offset per subject, applied uniformly to all of that subject’s events Johnson et al. [2023], so within-patient order is exact, but the shifted-year windows are not aligned with calendar time: across the 71,008 stays in the partition, Cramér’s V between the assigned window and the published anchor_year_group field is 0.045, and mean window purity is 0.335 against a chance value of 0.336. We therefore treat the windows as a partitioning device rather than a calendar axis, and the empirical evidence of drift over calendar time in Appendix A uses the published era field. What the construction provides is a deterministic partition into many small centers whose outcome distributions differ: the mean pairwise Jensen-Shannon divergence between the Bernoulli mortality distributions of centers from different care units is 0.0066 nats, against 0.0009 nats within a care 7

unit, and mortality rates vary across units (approximately 3% in CVICU versus 15% in MICU), aligning with the IDN, AMC, and Community Hospital framing in Sec. 1. The pipeline also supports Dirichlet(α) label-skew partitions Hsu et al. [2019] and uniform IID splits, but the careunit-by-year setting is used for all primary results. A FedProx-Synthetic(α, β) generator Li et al. [2020] matching the MIMIC feature schema drives implementation tests but is not used for headline numbers. 6.2 Baselines We compare against ten methods grouped by purpose. Reference anchors: Centralized pools all data into one Multi-Layer Perceptron (MLP) (non-federated upper bound, in the Table 1 caption), Local-only trains each center independently (no-collaboration lower bound), and FedAvg McMahan et al. [2017] averages weights every round. Centralized non-IID FL: FedProx Li et al. [2020] adds a proximal regularizer, FedDyn Acar et al. [2021] aligns local objectives with the global stationary point, and MOON Li et al. [2021] maximizes agreement between local and global representations. Decentralized P2P FL: DeceFL Yuan et al. [2023] provably converges to the centralized optimum, DeFTA Zhou et al. [2024] is a plug-and-play decentralized FedAvg with trust-based reweighting (the closest peer-selection competitor), and WPFed Ye et al. [2024] uses Locality-Sensitive Hashing (LSH) similarity filtering plus weighted neighbor selection, mirroring the metadata pre-filter of Sec. 4.1 but lacking the Bayesian posterior update. Bayesian FL: BNN+FL Saile et al. [2024] replaces the MLP with a Bayesian neural network and aggregates posterior moments (ABPS’s novelty is Bayesianizing the utility signal rather than the weights). BrainTorrent Guha Roy et al. [2019], Gossip Learning Hegedűs et al. [2021], KL-FedDis Rahad et al. [2025], and Peer-Driven Reputation FL Seidi et al. [2025] are conceptual antecedents discussed in Sec. 2 but not run as baselines. 6.3 Implementation and Metrics The local model is a two-hidden-layer MLP (128 → 64, ReLU, dropout 0.2) optimized with Adam (η = 10−3 , weight decay 10−5 ). Each round runs Klocal = 2 local epochs with batch size 128 over K = 50 rounds (or up to K = 100 for the ABPS-X variant√with validation-AUROC early stopping at patience 10). Base ABPS uses κ = 3, ϵ = 0.1, γ = 2, τacc = 0.5, the 3-dim base metadata vector [positive-class rate, log ntrain , year/2030], and τsim = 0 (no filtering). The goal-aware variants enrich metadata with a 12-dim per-center prevalence vector over Charlson Comorbidity Index (CCI) Charlson et al. [1987] chronic-condition categories (binary presence per patient, averaged over local training cohort, no per-patient leakage) and auto-calibrate τsim to retain 25% of peer pairs per goal. Shapley contributions are computed exactly when |Pi | ≤ 5 and via truncated Monte-Carlo with B = 8 permutations otherwise. We report mean AUROC across centers (across-seed std over 5 seeds in {11, 22, 33, 44, 55}) and cumulative transmitted bytes counted once per undirected edge. Experiments were run on a Simple Linux Utility for Resource Management (SLURM) Yoo et al. [2003]-managed High-Performance Computing (HPC) cluster, each SLURM task on a single NVIDIA Tesla V100-SXM2 GPU (32 GB) with 8 CPU cores and 32 GB RAM, under PyTorch 2.5.1 (CUDA 12.1) and Python 3.11. The n=230-center sweep (60 array tasks: 12 configurations × 5 seeds) finishes in 1 to 3 wall-clock hours, and the headline ABPS-X sweep (10 tasks at 100 rounds with early stopping) finishes in 3 to 9 minutes per task. Reproduction of table cells and figures requires a single sbatch of the two sweep scripts released with the supplementary material and JSON outputs. 6.4 Headline Result: Accuracy vs. Bandwidth 1: In-hospital mortality prediction on MIMIC-IV (n=230 Recall Q4 (communication overhead, even in Table centers, 50 rounds, 5 seeds). AUROC is the mean across centers P2P) from Sec. 2. Table 1 and the Pareto frontier (mean ± standard deviation). Bandwidth is a cumulative parameter of Figure 3 are the empirical answer. Table 1 exchange relative to FedAvg = 1.00×. As a non-federated upperreports final-round mean AUROC over 5 seeds, bound reference (not a fair federated comparator). AUROC (↑) Bandwidth and Figure 3 plots the same data as an accuracy- Method Local-only (lower bound) 0.587 ± 0.013 0.00× bandwidth Pareto frontier. All federated methods FedAvg McMahan et al. [2017] 0.755 ± 0.019 1.00× gain ∼ 15 AUROC points over the local-only FedProx Li et al. [2020] 0.753 ± 0.019 1.00× lower bound (0.587 ± 0.013) and close most of FedDyn Acar et al. [2021] 0.758 ± 0.017 1.00× MOON Li et al. [2021] 0.750 ± 0.020 1.00× the gap to the non-federated Centralized upper DeceFL Yuan et al. [2023] 0.696 ± 0.011 1.00× bound (0.827 ± 0.007). DeFTA Zhou et al. [2024] 0.725 ± 0.013 1.00× Base ABPS (with κ=3 and no personaliza- WPFed Ye et al. [2024] 0.694 ± 0.012 1.00× BNN+FL Saile et al. [2024] 0.749 ± 0.014 2.00× tion or quantization) already achieves AUROC ABPS (base, κ=3) 0.753 ± 0.015 1.50× 0.753 ± 0.015, statistically indistinguishable from ABPS+P (personalize) 0.757 ± 0.015 1.49× 0.753 ± 0.015 0.75× FedAvg (0.755 ± 0.019), FedProx (0.753 ± 0.019), ABPS+Q (bfloat16) ABPS+P+Q+κ=1 0.748 ± 0.012 0.25× and FedDyn (0.758 ± 0.017). It transmits 1.50× FedAvg+P+Q 0.758 ± 0.017 0.50× FedAvg’s bandwidth because a κ=3 mesh has ABPS+P+Q+κ=1 (hom.) 0.742 ± 0.013 0.25× 8

ABPS+P+Q+κ=1 (div.) ABPS+P+Q+κ=1 (align.) ABPS-X (ours, full) FedAvg-X (fair comp., full)

0.694 ± 0.015 0.739 ± 0.010 0.758 ± 0.010 0.691 ± 0.008

0.16× 0.25× 0.09× 0.31×

more unique edges than FedAvg’s star, matching the theoretical accounting in Sec. 5.4. Three extensions independently move the Pareto frontier: (P) head personalization Arivazhagan et al. [2019] boosts AUROC to 0.757 ± 0.015 at the same bandwidth (the per-center head specializes to careunit mortality base-rates, e.g., CVICU ∼ 3% vs MICU ∼ 15%). (Q) bfloat16 Dettmers et al. [2022] halves bandwidth with no accuracy loss (0.753 ± 0.015 at 0.75×). (κ=1) collapsing the active set to a single partner halves the mesh edges again and, combined with (P) and (Q), yields the ABPS+P+Q+κ=1 row: 0.748 ± 0.012 at 0.25×, already Pareto-dominating federated baselines. Adding the goal-aware pre-filter, 2-layer personalization, server momentum, and validation-AUROC early stopping yields full ABPS-X: 0.758 ± 0.010 (matching FedDyn) at 0.09× FedAvg. Fair comparator and ABPS-X. Applying P+Q to FedAvg (FedAvg+P+Q) yields 0.758 ± 0.017 at (a) Accuracy ranking (federated methods) Centralized

FedDyn ABPS+P FedAvg FedProx ABPS+Q ABPS (base, =3) MOON BNN+FL

ABPS+P+Q+ =1

ABPS+P+Q+G[hom] ABPS+P+Q+G[ali] DeFTA DeceFL ABPS+P+Q+G[div] WPFed FedAvg-X (fair comp.) Local-only

(b) Accuracy-bandwidth Pareto (zoomed)

(non-federated) 0.827

0.696 0.694 0.694 0.691

0.587

0.60

0.65

0.70

0.758 0.758 0.758 0.757 0.755 0.753 0.753 0.753 0.750 0.749 0.748 0.742 0.739 0.725

Centralized (off-axis) = 0.827

Mean AUROC across centers ( better)

FedAvg+P+Q

ABPS-X (ours, full)

Methods

ABPS-X (ours, Tier 1+2)

0.76

+P, =1

ABPS+P+Q+ =1 (ours)

0.74

+P +Q

0.72 0.70 0.68

0.75

Mean AUROC across centers ( better)

102

FedAvg FedDyn FedProx MOON FedAvg+P+Q FedAvg-X (fair comp.) DeceFL DeFTA WPFed BNN+FL ABPS+P ABPS+Q ABPS (base, =3) ABPS+P+Q+ =1 ABPS-X (ours, full) ABPS+P+Q+G[ali] ABPS+P+Q+G[div] ABPS+P+Q+G[hom]

103

Cumulative bytes transmitted (MB, log scale)

Figure 3: Results on MIMIC-IV careunit-by-year (n = 230 centers, 5 seeds). (a) Accuracy ranking: mean AUROC with across-seed std, color-coded by method family (blue: star-topology FL, green: decentralized, orange: ABPS+P/Q, red: ABPS-X, purple/brown: goal-aware, olive: local-only). The dashed line denotes the non-federated centralized upper bound (off-axis). (b) Accuracy vs. bandwidth Pareto (log-scale bandwidth): ABPS-X (dark-red star) matches FedDyn’s AUROC (0.758) at 0.09× FedAvg bandwidth. Red arrows show the ablation path ABPS→+P→+Q→+P+Q+κ=1→ABPS-X, with each step improving the trade-off.

0.50×, a modest improvement, but cannot reach κ=1 because the star topology has no mechanism

for selecting a single best peer. The full ABPS-X variant adds 2-layer personalization, server-side EMA momentum β=0.5, 15-dim goal-aware metadata under the homogeneity goal, and validationAUROC early stopping (patience 10 of 100), reaching 0.758 ± 0.010 at 0.09× FedAvg with tighter variance than FedDyn. The same-extension FedAvg-X fair comparator drops to 0.691 ± 0.008, 6.7 AUROC points behind ABPS-X: 2-layer personalization over-adapts each private head when the star topology averages across all 230 peers, whereas ABPS’s κ=1 Bayesian selection supplies the one-peer constraint under which personalization helps. Empirical isolation does not activate here (E[|Pi |] ≈ 1.00 when κ=1), but the goal-aware filter of the next subsection engages Lemma 1’s rest regime in up to 41% of rounds. 6.5 Empirical Validation of Lemma 1 (Intentional Isolation) The diversity-goal experiment below is the empirical activation of the rest regime promised by Lemma 1 to address Q3 (adaptive isolation under drift).. The isolation lemma predicts that when every peer’s true expected utility falls below the negative-transfer threshold cneg , the optimal action is Pi = ∅ and the resulting bandwidth is zero. We validate this prediction on MIMIC-IV through the goal-aware filter of Sec. 4.1: by varying the admission rule, we directly control which peers survive to the UCB-Shapley stage, which in turn controls whether the UCB thresholds admit proposals. We run ABPS+P+Q+κ=1 on the full n=230 federation under each of the three goals in Eq. 3 with τsim auto-calibrated to retain 25% of peers per center, and record the per-round fraction of centers with |Pi |=0. Homogeneity and alignment admit peers with similar (or KPI-matched) marginals, so isolation rarely activates (E[|Pi |]=0.99), giving AUROC 0.742 and 0.739 at 0.25× bandwidth. Diversity admits dissimilar peers, many crossing the negative-transfer threshold, so the UCB falls below τacc and rest activate for 41% of centers (E[|Pi |]=0.59). Bandwidth drops to 0.16× FedAvg, the lowest observed, while AUROC decreases to 0.694. The bandwidth reduction tracks isolation within p 5%, the collaborate-to-rest transition is sharp at the Hoeffding rate log(2/δ)/(2K), and this is the first setting where Lemma 1’s |Pi |=0 regime activates at scale on real clinical data. Ablations: Reading Table 1: ABPS+P improves AUROC by +0.4 at the same bandwidth, while +Q maintains performance at 0.75×. The combined ABPS+P+Q+κ=1 setting preserves AUROC at just 0.25× bandwidth. In contrast, FedAvg+P+Q reaches 0.758 ± 0.017 at 0.50× bandwidth but cannot operate at κ=1, indicating that the additional 2× reduction is due to Bayesian Shapley-UCB selection. The three goal-aware rows (Eq.3) ablate the pre-filter (Sec.6.5). Hyperparameters ϵ, γ , and truncated-MC B are fixed from preliminary synthetic runs. 9

7

Conclusion and Future Work

We presented ABPS, an adaptive Bayesian P2P federated-learning framework that replaces raw parameter aggregation with Shapley-based marginal-utility evaluation, formalizes intentional isolation as a Bayes-optimal action under negative transfer, and composes with three communication-efficiency extensions (head personalization, bfloat16 quantization, tunable κ) plus a goal-aware metadata pre-filter. On MIMIC-IV with n=230 careunit-by-year centers, the full ABPS-X variant matches the strongest federated baseline (FedDyn) at 0.09× FedAvg bandwidth, while applying the same extensions to FedAvg drops 6.4 AUROC points (isolating the gain to the Bayesian selection itself), and the diversity goal activates Lemma 1’s rest regime for 41% of centers per round, a first on real clinical data. Limitations and future work. The evaluation is on a single dataset and binary clinical task, the regret bound degrades with unmeasured ϵm , and we provide no formal (ϵ, δ)-DP, survival, or Byzantine guarantees. Natural follow-ups include (i) a Gaussian-mechanism DP wrapper on vi for certifiable privacy, (ii) multi-modal pipelines fusing EHR with medical imaging, (iii) online Welch-t estimation of {∆tm } together with a drifting-bandit regret analysis, and (iv) porting to time-to-event outcomes (matching the Cox-style framing of Seidi et al. [2025]). Scope of the accuracy claim. The matched-accuracy result holds for federations of many small centers. On a 40-center partition of the same cohort built from the published anchor_year_group eras, with a median of roughly 2,000 stays per center, FedAvg and FedDyn reach 0.813 and 0.805 AUROC after 100 rounds while ABPS-X peaks at 0.793, although ABPS-X reaches each intermediate AUROC target (0.750, 0.770, 0.785) with four to six times fewer bytes. Where centers are large, each local model is already well estimated and aggregating across all of them is close to optimal; where centers are numerous and small, indiscriminate aggregation carries more harmful transfer and selective exchange is competitive. Broader impact. Cutting per-center bandwidth tenfold at matched accuracy on federations of many small centers lowers the entry barrier for resource-constrained sites that disproportionately serve under-represented populations. Residual privacy and goal-misuse risks (metadata re-identification under auxiliary information, and goal-aware filtering used to entrench rather than correct bias) are addressed by the recommended DP wrapper and governance over goal declarations, documented in full in the NeurIPS Reproducibility Checklist.

References Mostafa Abdolmaleki and Bahar Farahani. Sync-GWO: Highly private and bandwidth-efficient federated learning with a case study in healthcare. IEEE Journal of Biomedical and Health Informatics, 30(3):1939–1946, 2026. doi: 10.1109/JBHI.2025.3567913. Durmus Alp Emre Acar, Yue Zhao, Ramon Matas Navarro, Matthew Mattina, Paul N. Whatmough, and Venkatesh Saligrama. Federated learning based on dynamic regularization. In International Conference on Learning Representations (ICLR), 2021. URL https://openreview.net/ forum?id=B7v4QMR6Z9w. Amazon Web Services. Amazon EC2 on-demand pricing: Data transfer. https://aws.amazon. com/ec2/pricing/on-demand/, 2024. Accessed 2024. Manoj Ghuhan Arivazhagan, Vinay Aggarwal, Aaditya Kumar Singh, and Sunav Choudhary. Federated learning with personalization layers. arXiv preprint arXiv:1912.00818, 2019. URL https://arxiv.org/abs/1912.00818. Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer. Finite-time analysis of the multiarmed bandit problem. Machine Learning, 47(2–3):235–256, 2002. doi: 10.1023/A:1013689704352. Andrea Augello, Ashish Gupta, Giuseppe Lo Re, and Sajal K. Das. Tackling selfish clients in federated learning. In ECAI 2024 – 27th European Conference on Artificial Intelligence, volume 392 of Frontiers in Artificial Intelligence and Applications, pages 1888–1895. IOS Press, 2024. doi: 10.3233/FAIA240702. Mary E. Charlson, Peter Pompei, Kathy L. Ales, and C. Ronald MacKenzie. A new method of classifying prognostic comorbidity in longitudinal studies: development and validation. Journal of Chronic Diseases, 40(5):373–383, 1987. doi: 10.1016/0021-9681(87)90171-8. 10

Matthew G. Crowson, Dana Moukheiber, Aldo Robles Arévalo, Barbara D. Lam, Sreekar Mantena, Aakanksha Rana, Deborah Goss, David W. Bates, and Leo Anthony Celi. A systematic review of federated learning applications for biomedical data. PLOS Digital Health, 1(5):e0000033, 2022. doi: 10.1371/journal.pdig.0000033. Marco Cuturi and Mathieu Blondel. Soft-DTW: a differentiable loss function for time-series. In Proceedings of the 34th International Conference on Machine Learning (ICML), volume 70 of Proceedings of Machine Learning Research, pages 894–903. PMLR, 2017. URL https: //proceedings.mlr.press/v70/cuturi17a.html. Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. LLM.int8(): 8-bit matrix multiplication for transformers at scale. In Advances in Neural Information Processing Systems (NeurIPS), volume 35, 2022. URL https://proceedings.neurips.cc/paper_files/ paper/2022/hash/c3ba4962c05c49636d4c6206a97e9c8a-Abstract-Conference.html. Christoph Düsing and Philipp Cimiano. Distribution-controlled client selection to improve federated learning strategies. In Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD 2024 Workshops), Communications in Computer and Information Science, pages 299–313. Springer, 2026. doi: 10.1007/978-3-032-25314-9_21. Preprint: arXiv:2509.20877. João Gama, Indrė Žliobaitė, Albert Bifet, Mykola Pechenizkiy, and Abdelhamid Bouchachia. A survey on concept drift adaptation. ACM Computing Surveys, 46(4):1–37, 2014. doi: 10.1145/2523813. Amirata Ghorbani and James Zou. Data Shapley: Equitable valuation of data for machine learning. In Proceedings of the 36th International Conference on Machine Learning (ICML), volume 97 of Proceedings of Machine Learning Research, pages 2242–2251. PMLR, 2019. URL https: //proceedings.mlr.press/v97/ghorbani19c.html. Abhijit Guha Roy, Shayan Siddiqui, Sebastian Pölsterl, Nassir Navab, and Christian Wachinger. BrainTorrent: A peer-to-peer environment for decentralized federated learning. arXiv preprint arXiv:1905.06731, 2019. doi: 10.48550/arXiv.1905.06731. Rahul Haripriya, Nilay Khare, and Manish Pandey. Privacy-preserving federated learning for collaborative medical data mining in multi-institutional settings. Scientific Reports, 15(1):12482, 2025. doi: 10.1038/s41598-025-97565-4. István Hegedűs, Gábor Danner, and Márk Jelasity. Decentralized learning works: An empirical comparison of gossip learning and federated learning. Journal of Parallel and Distributed Computing, 148:109–124, 2021. doi: 10.1016/j.jpdc.2020.10.006. Wassily Hoeffding. Probability inequalities for sums of bounded random variables. Journal of the American Statistical Association, 58(301):13–30, 1963. doi: 10.1080/01621459.1963.10500830. Tzu-Ming Harry Hsu, Hang Qi, and Matthew Brown. Measuring the effects of non-identical data distribution for federated visual classification. arXiv preprint arXiv:1909.06335, 2019. URL https://arxiv.org/abs/1909.06335. Cristovão Iglesias Jr., Sidney Alves de Outeiro, Claudio Miceli de Farias, and Miodrag Bolic. Two students: Enabling uncertainty quantification in federated learning clients. In NeurIPS 2024 Workshop on Bayesian Decision-making and Uncertainty, 2024. URL https://openreview. net/forum?id=eS9xH4vHEe. Alistair E.W. Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Shammout, Steven Horng, Tom J. Pollard, Sicheng Hao, Benjamin Moody, Brian Gow, et al. MIMIC-IV, a freely accessible electronic health record dataset. Scientific Data, 10(1):1, 2023. doi: 10.1038/s41597-022-01899-x. Dhiraj Kalamkar, Dheevatsa Mudigere, Naveen Mellempudi, Dipankar Das, Kunal Banerjee, Sasikanth Avancha, Dharma Teja Vooturi, Nataraj Jammalamadaka, Jianyu Huang, Hector Yuen, Jiyan Yang, Jongsoo Park, Alexander Heinecke, Evangelos Georganas, Sudarshan Srinivasan, Abhisek Kundu, Misha Smelyanskiy, Bharat Kaul, and Pradeep Dubey. A study of BFLOAT16 for deep learning training. arXiv preprint arXiv:1905.12322, 2019. URL https://arxiv.org/abs/1905.12322. 11

Tor Lattimore and Csaba Szepesvári. Bandit Algorithms. Cambridge University Press, 2020. doi: 10.1017/9781108571401. Ming Li, Pengcheng Xu, Junjie Hu, Zeyu Tang, and Guang Yang. From challenges and pitfalls to recommendations and opportunities: Implementing federated learning in healthcare. Medical Image Analysis, 101:103497, 2025. doi: 10.1016/j.media.2025.103497. arXiv:2409.09727. Qinbin Li, Bingsheng He, and Dawn Song. Model-contrastive federated learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10713– 10722, 2021. doi: 10.1109/CVPR46437.2021.01057. Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith. Federated optimization in heterogeneous networks. In Proceedings of Machine Learning and Systems (MLSys), volume 2, pages 429–450, 2020. URL https://proceedings.mlsys.org/paper_ files/paper/2020/hash/1f5fe83998a09396ebe6477d9475ba0c-Abstract.html. Edvin Listo Zec, Johan Östman, Olof Mogren, and Daniel Gillblad. Efficient node selection in private personalized decentralized learning. In Proceedings of the 5th Northern Lights Deep Learning Conference (NLDL), volume 233 of Proceedings of Machine Learning Research, pages 244–250. PMLR, 2024. URL https://proceedings.mlr.press/v233/zec24a.html. arXiv:2301.12755. Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. Communication-efficient learning of deep networks from decentralized data. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS), volume 54 of Proceedings of Machine Learning Research, pages 1273–1282. PMLR, 2017. URL https: //proceedings.mlr.press/v54/mcmahan17a.html. Gergely D. Németh, Eros Fanì, Yeat Jeng Ng, Barbara Caputo, Miguel Ángel Lozano, Nuria Oliver, and Novi Quadrianto. FedDiverse: Tackling data heterogeneity in federated learning with diversitydriven client selection. In 2025 3rd International Conference on Federated Learning Technologies and Applications (FLTA), pages 432–440. IEEE, 2025. doi: 10.1109/FLTA67013.2025.11336421. arXiv:2504.11216. Md. Rahad, Ruhan Shabab, Mohd. Sultan Ahammad, Md. Mahfuz Reza, Amit Karmaker, and Md. Abir Hossain. KL-FedDis: A federated learning approach with distribution information sharing using Kullback-Leibler divergence for non-IID data. Neuroscience Informatics, 5(1): 100182, 2025. doi: 10.1016/j.neuri.2024.100182. Leyla Rahimli, Feras M. Awaysheh, Sawsan Al Zubi, and Sadi Alawadi. Federated learning drift detection: An empirical study on the impact of concept and data drift. In 2024 2nd International Conference on Federated Learning Technologies and Applications (FLTA), pages 241–250. IEEE, 2024. doi: 10.1109/FLTA63145.2024.10839814. Iftekhar Rahman, Nisal Hemadasa, Dominik Kaaser, Pierre-Alexandre Murena, and Stefan Schulte. Detect, adapt, overcome: Mitigating concept drift in federated learning. In 2025 3rd International Conference on Federated Learning Technologies and Applications (FLTA), pages 17–24. IEEE, 2025. doi: 10.1109/FLTA67013.2025.11336319. Priyesh Ranjan, Ashish Gupta, Federico Corò, and Sajal K. Das. Securing federated learning against overwhelming collusive attackers. In GLOBECOM 2022 – 2022 IEEE Global Communications Conference, pages 1448–1453. IEEE, 2022. doi: 10.1109/GLOBECOM48099.2022.10000830. arXiv:2209.14093. Finn Saile, Julius Thomas, Dominik Kaaser, and Stefan Schulte. Client-side adaptation to concept drift in federated learning. In 2024 2nd International Conference on Federated Learning Technologies and Applications (FLTA), pages 71–78. IEEE, 2024. doi: 10.1109/FLTA63145.2024.10840058. Navid Seidi, Satyaki Roy, and Sajal K. Das. Enhancing federated survival analysis through peerdriven client reputation in healthcare. arXiv preprint arXiv:2505.16190, 2025. doi: 10.48550/ arXiv.2505.16190. 12

Pranava Singhal, Shashi Raj Pandey, and Petar Popovski. Greedy Shapley client selection for communication-efficient federated learning. IEEE Networking Letters, 6(2):134–138, 2024. doi: 10.1109/LNET.2024.3363620. Georgios Syros, Gokberk Yar, Simona Boboila, Cristina Nita-Rotaru, and Alina Oprea. Backdoor attacks in peer-to-peer federated learning. ACM Transactions on Privacy and Security, 28(1):1–28, 2024. doi: 10.1145/3691633. arXiv:2301.09732. Chuhan Wu, Fangzhao Wu, Lingjuan Lyu, Yongfeng Huang, and Xing Xie. Communication-efficient federated learning via knowledge distillation. Nature Communications, 13(1):2032, 2022. doi: 10.1038/s41467-022-29763-x. Wenchao Xia, Tony Q. S. Quek, Kun Guo, Wanli Wen, Howard H. Yang, and Hongbo Zhu. Multiarmed bandit-based client scheduling for federated learning. IEEE Transactions on Wireless Communications, 19(11):7108–7123, 2020. doi: 10.1109/TWC.2020.3008091. Mengwei Yang, Ismat Jarin, Baturalp Buyukates, Salman Avestimehr, and Athina Markopoulou. Maverick-Aware Shapley valuation for client selection in federated learning. arXiv preprint arXiv:2405.12590, 2024. doi: 10.48550/arXiv.2405.12590. Guanhua Ye, Jifeng He, Weiqing Wang, Zhe Xue, Feifei Kou, and Yawen Li. WPFed: Web-based personalized federation for decentralized systems. arXiv preprint arXiv:2410.11378, 2024. URL https://arxiv.org/abs/2410.11378v1. version 1. Andy B. Yoo, Morris A. Jette, and Mark Grondona. SLURM: Simple Linux Utility for Resource Management. In Job Scheduling Strategies for Parallel Processing (JSSPP), pages 44–60. Springer Berlin Heidelberg, 2003. doi: 10.1007/10968987_3. Ye Yuan, Jun Liu, Dou Jin, Zuogong Yue, Tao Yang, Ruijuan Chen, Maolin Wang, Lei Xu, Feng Hua, Yuqi Guo, Xiuchuan Tang, Xin He, Xinlei Yi, Dong Li, Wenwu Yu, Hai-Tao Zhang, Tianyou Chai, Shaochun Sui, and Han Ding. DeceFL: A principled fully decentralized federated learning framework. National Science Open, 2(1):20220043, 2023. doi: 10.1360/nso/20220043. Yuhao Zhou, Minjia Shi, Yuxin Tian, Qing Ye, and Jiancheng Lv. DeFTA: A plug-and-play peer-topeer decentralized federated learning framework. Information Sciences, 670:120582, 2024. doi: 10.1016/j.ins.2024.120582.

13

A

Empirical Evidence of Concept Drift on MIMIC-IV

To support the concept-drift premise empirically, we compared the distribution of two routinely charted ICU vital signs, heart rate and respiratory rate, between two non-overlapping calendar eras of MIMIC-IV v3.1. The era of each stay is the published anchor_year_group field of the patient, which is the true three-year admission era and is unaffected by the per-subject date shift; we use the earliest era, 2008 to 2010, and the latest, 2020 to 2022. For each ICU stay we take the mean of all charted values of the vital during the stay, so that each stay contributes one observation, and we remove stay means that lie more than three standard deviations from the per-era mean. Heart rate has 29,786 stays in the earlier era and 10,740 in the later one, with means of 84.9 and 83.9 beats per minute; respiratory rate has 29,885 and 10,742 stays, with means of 19.3 and 19.4 breaths per minute and a markedly narrower spread in the later era (standard deviation 4.2 against 5.7). Welch two-sample t-tests give t = 5.45, p = 5 × 10−8 for heart rate and t = −2.07, p = 0.038 for respiratory rate. The shifts are small in the mean but significant at this sample size, and the change in spread for respiratory rate indicates a change in charting or patient mix rather than a mean shift (i) alone, which is the kind of distributional drift the per-window treatment of Pt in Sec. 3 is designed to absorb. (a) Heart rate

0.030

(b) Respiratory rate 2008 to 2010 2020 to 2022

0.025

0.10 density

density

0.020 0.015 0.010

0.08 0.06 0.04

0.005 0.000

2008 to 2010 2020 to 2022

0.12

0.02 50

60

70 80 90 100 110 120 130 stay mean (beats per minute)

0.00

15

20 25 30 stay mean (breaths per minute)

Figure 4: Histograms of per-stay mean heart rate (panel a) and respiratory rate (panel b) for ICU stays in two non-overlapping eras of MIMIC-IV v3.1, 2008 to 2010 and 2020 to 2022, defined by the published anchor_year_group field. Stay means beyond three standard deviations from the per-era mean are removed. Vertical dashed lines mark the per-era means. Welch t-tests give t = 5.45, p = 5 × 10−8 for heart rate and t = −2.07, p = 0.038 for respiratory rate, evidencing distributional shift across eras in routinely collected measurements and (i) motivating the per-window adaptive treatment of Pt in Sec. 3.

B

Algorithm Pseudocode

The complete round-level pseudocode for ABPS is given in Algorithm 1, deferred from the main text to save space. The algorithm operationalizes the four Bayesian-core components of Sec. 4 (goal-aware metadata pre-filter, ϵ-greedy UCB ranking, propose-reject topology with explicit rest, and the Beta-Bernoulli posterior update), and is referenced from the regret proof in Appendix C.2 (lines 13 and 16).

C

Full Proofs

This appendix gives the full proofs of Theorems 1 and 2 and Lemma 1. Throughout, we take Assumptions A1-A3 of Sec. 5 as given: bounded utility (ϕ̄j→i ∈ [0, 1]), conditional independence of (k) the per-round observations given µ⋆i→j , and quasi-stationary drift |E[ϕ̄j→i ] − µ⋆i→j | ≤ ϵm within any window ∆tm . C.1

Proof of Theorem 1 (Beta-belief Concentration) PK (k) Let SK = k=1 ϕ̄j→i and µ̄K = SK /K. The Beta posterior after K updates from prior Beta(α0 , β0 ) is Beta(α0 + SK , β0 + K − SK ), with posterior mean µ̂K =

α0 + SK α0 + K µ̄K = . α0 + β0 + K α0 + β0 + K 14

(11)

Algorithm 1 Adaptive Bayesian Partner Selection (ABPS) Require: centers {1, . . . , n}, rounds {t1 , . . . , tK }, rank threshold κ, exploration weight γ, exploration probability ϵ, acceptance threshold τacc , similarity threshold τsim , Beta prior (α0 , β0 ) Ensure: Updated local parameters θi and posterior beliefs (αi→j , βi→j ) 1: Initialization: 2: for each center i do 3: Initialize local model θi and metadata vector vi 4: for each center j ̸= i do 5: (αi→j , βi→j ) ← (α0 , β0 ), ni→j ← 0 6: end for 7: end for 8: for each round k = 1, . . . , K do 9: for each center i do 10: Ci ← {j : fi (vi , vj ) ≥ τsim } ▷ goal-aware pre-filter, eq. (3) 11: Compute UCBj→i (tk ) via eq. (7) for all j ∈ Ci 12: With prob. ϵ, swap top-ranked candidate with a uniformly random one in Ci 13: Pi ← ∅ 14: for j in descending UCB order do 15: if |Pi | ≥ κ then break 16: end if 17: if UCBi→j (tk ) ≥ τacc and |Pj | < κ then 18: Pi ← Pi ∪ {j}, Pj ← Pj ∪ {i} 19: end if 20: end for 21: end for 22: for each center i with Pi ̸= ∅ do 23: θi ← FedAvg({θi } ∪ {θj : j ∈ Pi }) ▷ exchange; optionally personalize head and quantize wire-format (Sec. 6.4) 24: Compute Shapley marginals {ϕj→i }j∈Pi via eq. (4) 25: for each j ∈ Pi do  26: ϕ̄j→i ← clip[0,1] (ϕj→i − ϕmin )/(ϕmax − ϕmin ) 27: (αi→j , βi→j ) ← (αi→j + ϕ̄j→i , βi→j + 1 − ϕ̄j→i ) 28: ni→j ← ni→j + 1 29: end for 30: end for 31: Centers with Pi = ∅ rest this round (intentional isolation; cf. Lemma 1) 32: end for

Apply the triangle inequality with the empirical mean µ̄K as pivot: |µ̂K − µ⋆i→j | ≤ |µ̂K − µ̄K | + |µ̄K − E[µ̄K ]| + |E[µ̄K ] − µ⋆i→j | . | {z } | {z } {z } | (I) prior pull

(II) statistical noise

(12)

(III) drift bias

Bounding (I).

From (11), α0 (1 − µ̄K ) − β0 µ̄K α0 + K µ̄K − (α0 + β0 + K)µ̄K = . µ̂K − µ̄K = α0 + β0 + K α0 + β0 + K Because µ̄K ∈ [0, 1], the numerator is bounded in absolute value by max(α0 , β0 ) ≤ α0 + β0 . Hence α0 + β0 . (13) (I) ≤ α0 + β0 + K This is the deterministic “prior decay” term in (8): it shrinks at rate Θ(1/K) regardless of randomness in the data. Bounding (II). By A1 every observation lies in [0, 1], and by A2 the observations are conditionally independent given µ⋆i→j . Hoeffding’s inequality Hoeffding [1963] for the mean of K independent [0, 1]-valued variables gives    Pr |µ̄K − E[µ̄K ]| ≥ t ≤ 2 exp −2Kt2 . (14) p Setting the right-hand side equal to δ and solving for t yields t = log(2/δ)/(2K), so with probability at least 1 − δ r log(2/δ) (II) ≤ . (15) 2K Bounding (III).

By A3, every per-round expectation is within ϵm of µ⋆i→j :

|E[µ̄K ] − µ⋆i→j |

=

K

K

k=1

k=1

1 X 1 X (k) (k) E[ϕ̄j→i ] − µ⋆i→j ≤ E[ϕ̄j→i ] − µ⋆i→j ≤ ϵm . K K

Hence (III) ≤ ϵm . 15

Substituting (I), (II), (III) into (12) yields, with probability at least 1 − δ, r log(2/δ) α0 + β0 ⋆ + + ϵm , |µ̂K − µi→j | ≤ 2K α0 + β0 + K which is exactly (8). Convergence in probability follows: as K → ∞, (I) and (II) tend to 0, so p µ̂K − → µ⋆i→j provided the drift ϵm → 0. Combining.

C.2

Proof of Theorem 2 (Sublinear Partner-Selection Regret)

Inside a stationary window (ϵm = 0), the per-center partner-selection problem is a stochastic multiarmed bandit over the post-filter candidate set Ci ⊆ {1, . . . , n − 1} in which ABPS pulls κ arms per round (the κ acceptances) rather than one. The goal-aware pre-filter of Sec. 4.1 restricts the arm set before the bandit sees it, and we analyze this case directly below. We bound the cumulative regret   T X X X ⋆  R(T ) = T µ⋆j − E µj  t=1 j∈P (t)

j∈top-κ

i

by analyzing the greedy and exploratory phases separately. Greedy phase (probability 1 − ϵ). With probability 1 − ϵ ABPS ranks candidates by their UCB1 score s 2 log t UCBj→i (t) = µ̂i→j (t) + γ , ni→j (t) + 1 and proposes greedily down the list until κ acceptances accrue. Decompose the per-round greedy regret as the sum of κ single-arm regrets, indexed by the rank position r = 1, . . . , κ of each accepted proposal. Each rank-r slot is a single-arm UCB1 problem played against the residual candidate pool. By the canonical UCB1 analysis Auer et al. [2002], the cumulative regret of UCB1 with confidence √ radius γ = 2 over T rounds satisfies   X  8 log T π2 + ∆j 1 + , (16) RUCB1 (T ) ≤ ∆j 3 j:∆j >0

where ∆j = µ⋆(1) − µ⋆(j) are the suboptimality gaps. Summing κ such bounds gives   X  8 log T π2 Rgreedy (T ) ≤ κ . + ∆j 1 + ∆j 3

(17)

j:∆j >0

Two refinements only decrease this bound and so are absorbed into (17). (a) The receiver-side acceptance check UCBi→j (t) ≥ τacc in the inner-if of Algorithm 1 cannot create new pulls of suboptimal arms, only suppress them. (b) The Beta posterior is sharper than the empirical mean used by vanilla UCB1 (it shrinks toward the prior at rate 1/K, matching term (I) in Theorem 1), so the exploration radius is in fact tighter than (16) assumes. Exploratory phase (probability ϵ). With probability ϵ ABPS replaces the top-ranked candidate with a uniformly random peer (the ϵ-greedy swap step in Algorithm 1). Its expected per-round regret contribution is at most Ej∼Unif(Ci ) [∆j ], so summing over T rounds Rexplore (T ) ≤ ϵ T · Ej∼Unif(Ci ) [∆j ]. Combining.

(18)

Adding (17) and (18) yields the bound (9) of the main text:   X  8 log T π2 R(T ) ≤ κ + ∆j 1 + + ϵT · Ej [∆j ]. ∆j 3 j:∆j >0

√ Recovering the O(log T ) rate. Two annealing schedules suffice. (a) A constant ϵ = c/ T leaves √ an O( T ) exploratory residual that is sublinear but slower than the greedy term. (b) A time-varying PT schedule ϵt = min(1, c/t), in the spirit of Auer et al. [2002], gives t=1 ϵt = O(log T ) and recovers the full O(κ log T ) rate. In either case, the regret is sublinear in T , so the average per-round regret tends to zero. 16

Filtered-arm-set refinement. When the goal-aware filter of Sec. 4.1 restricts Ci to the top-ρ quantile of peers under fi , the bandit plays on C˜i = Ci ∩ {j : fi (vi , vj ) ≥ τsim }. Two changes propagate to the bound (9). (i) Reduced explore-greedy regret. The sums over {j : ∆j > 0} and Ej∼Unif(Ci ) [∆j ] are replaced by sums over C˜i , which is a strict subset, and both terms can only decrease. (ii) Optimal-arm coverage bias. If the oracle-best peer j ⋆ fails to clear fi (·, ·) ≥ τsim , the bandit plays on a suboptimal set and incurs an additive bias of ∆j ⋆ · T against the oracle baseline. This worst-case term is controlled by a standard top-k coverage guarantee: if the descriptor v is L-Lipschitzinformative about true utility (i.e. |µ⋆j − µ⋆j′ | ≤ L∥vj − vj ′ ∥), then the probability P[j ⋆ ∈ / C˜i ] ≤ exp(−cρ|Ci |) for some c > 0 depending on L, so the expected additional regret is O(T · exp(−cρn)), which is negligible for reasonable ρ and n. In our experiments ρ = 0.25, n = 230 gives ρn ≈ 58, the filter recovers the optimal arm with probability > 1 − 10−25 under any non-trivial descriptor-to-utility Lipschitz constant. The net effect is that (9) remains valid with |Ci | replaced by |C˜i | ≈ ρn. The O(κ log T ) asymptotic rate is preserved, and the constants improve proportionally to ρ. C.3

Proof of Lemma 1 (Optimality of Intentional Isolation)

Let Vi (P) denote i’s expected one-step utility (e.g., held-out AUROC) after the round, conditioned on its active peer set being P. Set Virest := Vi (∅) and Vicoll (P) := Vi (P) for P ̸= ∅. Decomposition via Shapley efficiency. The Shapley value (4) satisfies the efficiency axiom: for any coalition P, X ϕj→i = Ui (P) − Ui (∅), (19) j∈P

where Ui (·) is the validation utility used to compute the Shapley value. Since Vi and Ui coincide in expectation under our protocol (both are held-out AUROC of the post-aggregation model), taking expectations in (19) gives X Vicoll (P) − Virest = E[ϕj→i ]. (20) j∈P

Translating to the clipped scale. Within the clipping range, ϕ̄j→i relates to ϕj→i via the affine map (5), ϕj→i − ϕmin ϕ̄j→i = ⇐⇒ ϕj→i = (ϕmax − ϕmin ) ϕ̄j→i + ϕmin , ϕmax − ϕmin so E[ϕj→i ] = (ϕmax − ϕmin )µ⋆i→j + ϕmin . The negative-transfer threshold cneg is the value of µ⋆i→j at which one-peer collaboration breaks even, Vicoll ({j}) = Virest , i.e. E[ϕj→i ] = 0. Solving, −ϕmin . (21) ϕmax − ϕmin For the paper’s choice (ϕmin , ϕmax ) = (−0.1, +0.1) this gives cneg = 0.5, which coincides with τacc used in the algorithm. The map µ⋆i→j < cneg ⇐⇒ E[ϕj→i ] < 0 is therefore exact. cneg =

Strict dominance of resting. Suppose µ⋆i→j < cneg for every j ∈ P. Then E[ϕj→i ] < 0 for all j ∈ P, so by (20), X Vicoll (P) − Virest = E[ϕj→i ] < 0, j∈P

which is exactly the claimed inequality (10), Virest > Vicoll (P). Refinement under approximate Shapley. When the Shapley values are estimated via truncated Monte-Carlo (TMC, the default for κ > 5, cf. Sec. 5), the per-peer estimator carries an additive bias bounded by some η > 0. Equation (20) then becomes X Vicoll (P) − Virest = E[ϕj→i ] − g(|P|), j∈P

17

with |g(|P|)| ≤ η|P|. Strict P dominance survives whenever the true expected gain margin exceeds the approximation slack, j E[ϕj→i ] < −η|P|, which is satisfied with margin η to spare under the strict inequality µ⋆i→j < cneg . This is the form quoted in the proof sketch of Sec. 5.3. Bandwidth saving. Specializing the cost equation (2) to a single round with Pirest = ∅ gives a bandwidth of 0, while collaborating with |P| peers costs 2|P| cegress |W | (one upload and one download per peer per round). The saving is therefore exactly Cirest = 2|P| cegress |W | per round. Connection to Theorem 1. Lemma 1 reasons about the oracle threshold cneg , but ABPS only sees the noisy posterior estimate µ̂K . Combining the lemma with (8): with probability ≥ 1 − δ, the UCB-induced rule rests whenever q log K µ̂K + γ n2i→j +1 < τacc = cneg , where ni→j + 1 is the per-arm pull count from Eq. (7), which itself grows with K under any nonp trivial selection rule. The condition above is implied by µ⋆i→j < cneg − log(2/δ)/(2K) + (α0 + p  β0 )/(α0 + β0 + K) + ϵm + γ 2 log K/(ni→j + 1) . As K → ∞ and ϵm → 0, every term in the bracket vanishes (the exploration term shrinks because ni→j + 1 grows at rate Θ(K) for any peer ever pulled by UCB), and the algorithmic rule converges to the oracle rule of Lemma 1.

18

Acknowledgments This work was supported by NSF grants OAC-2609072 (CHAI), SFS-2335969 (MASTER), OAC2104076 (CANDY), and SATC-2030624 (TAURUS).

19

Record · ID 919337 · SHA-256 aee0ff7cfec17d59
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.