ConceptioArchivearXiv CS
arXiv CSopen access

Aggregation of Statistical Evidence under Exchangeability

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

Aggregation of Statistical Evidence under Exchangeability Antonin Schrab1

Rajen Shah2

Arthur Gretton3

Ilmun Kim4,∗

1 Department of Computer Science and Technology, University of Cambridge, UK 2 Statistical Laboratory, University of Cambridge, UK

3 Gatsby Computational Neuroscience Unit, University College London, UK 4 Department of Mathematical Sciences, KAIST, South Korea

July 20, 2026

arXiv:2607.15823v1 [stat.ME] 17 Jul 2026

Abstract We study aggregation of statistical evidence under unknown and potentially complex dependence using group-invariance. Building on permutation-based constructions that treat transformed datasets as exchangeable units, we aggregate evidence across statistics for each transformed dataset and calibrate the resulting aggregates across transformations. We develop a finite-sample power and adaptivity theory for this framework, together with extensions to sequential and data-dependent aggregation that preserve validity. For single-batch aggregation, which uses one collection of transformed datasets for both standardization and calibration, we show that the critical values uniformly improve on deterministic calibrations valid under arbitrary dependence, including Bonferroni correction, while adapting to the unknown dependence structure. We also introduce a sequential alpha-spending version that permits early rejection when evidence is strong, and a two-batch extension that separates standardization from calibration to accommodate learned aggregation rules and reduce computation. Applications to adaptive nonparametric testing and conformal prediction illustrate how these results sharpen existing aggregation methods. Keywords: aggregation, conformal prediction, exchangeability, multiple testing, permutation test.

1

Introduction

Modern nonparametric testing and distribution-free inference routinely generate collections of test statistics or p-values whose dependence structure is unknown and often highly non-trivial. Such multiplicity arises naturally when aggregating over tuning-parameter grids, combining multiple statistics to capture complementary aspects of the alternative, splitting data repeatedly for stability, or merging multiple prediction sets in conformal inference. The central challenge in these problems is to aggregate evidence across multiple statistics or p-values while retaining rigorous finite-sample validity and high power. A large body of existing work addresses this challenge by first converting each statistic into a p-value and then combining these p-values using procedures that are valid under arbitrary dependence [e.g., 44, 56, 57, 72, 73]. While such p-value merging methods enjoy universal validity, their calibration is necessarily driven by worst-case dependence scenarios, which restricts the class of admissible merging functions and often leads to overly conservative procedures. An alternative *

Corresponding author: [email protected].

1

line of work considers max-type aggregation procedures [2, 62] with data-dependent calibration via Monte Carlo approximation. Although these methods can yield power improvements in practice, their guarantees are approximation-based, and uniform finite-sample type I error control remains unresolved. More recently, [26] propose a rank-transformed subsampling approach for aggregating exchangeable statistics or p-values that is asymptotically tight; however, their validity guarantees are asymptotic and require knowledge of the limiting null distribution of a single test statistic. In this work, we revisit aggregation through the exchangeability structure induced by groupinvariance schemes under the null hypothesis. Our main theoretical object is what we refer to as single-batch (SB) aggregation, a row-wise construction that applies a single collection of randomized datasets both to standardize individual statistics and to calibrate the aggregated evidence. This construction is closely related to nonparametric combination [49] and Westfall–Young-type permutation methods [75]. More broadly, the exchangeability structure underlying such aggregation methods has been used to establish finite-sample validity [47, 64]. Building on these foundations, we develop a finite-sample theory of power and dependence adaptivity for exchangeability-based aggregation. Our analysis shows that calibrating directly on the realized exchangeable array can yield less conservative thresholds than deterministic worst-case calibrations while preserving finitesample validity. We also develop sequential and two-batch extensions that broaden the scope of the framework.

1.1

Contributions

Our main contributions are detailed as follows. • We formalize single-batch aggregation and analyze its power. We show that its critical value strictly dominates any deterministic calibration valid under arbitrary p-value dependence (Theorem 2), and that it adapts to the dependence structure induced by the group-invariance scheme (Proposition 2 and Proposition 3). This can yield significant power improvements over classical p-value merging methods calibrated for worst-case dependence. • We develop a sequential alpha-spending extension of SB minimum aggregation (Section 4), maintaining non-asymptotic level control while allowing rejection as soon as early ordered statistics provide strong evidence. • We introduce a two-batch extension of the SB framework (Section 5) as a principled mechanism for separating standardization from calibration. This construction enables aggregation rules to be learned from a reference batch while preserving finite-sample validity on the testing batch, and offers substantial computational advantages in settings such as conformal prediction (Section 6.2). • We demonstrate the scope of the framework through applications to adaptive hypothesis testing and conformal prediction, where our methods recover and extend recent results with simpler arguments, sharper power guarantees, and improved empirical performance.

1.2

Related work

Several lines of work have addressed the problem of aggregating multiple tests under unknown dependence. 2

P-value merging under dependence. Classical p-value combination methods, such as Fisher’s method [16] and Stouffer’s method [69], are standard tools for combining evidence across multiple tests. Their usual calibrations, however, typically rely on independence assumptions among the input p-values and may fail to control type I error under arbitrary dependence. In parallel, a class of universally valid p-value combination methods that remain valid under arbitrary dependence has been established [e.g., 44, 56, 57] and more recently revisited and systematically characterized by [72, 73], among others. While these methods enjoy broad validity, their calibration is necessarily driven by worst-case dependence scenarios and can therefore be conservative. Recent work by [21] shows that p-value aggregation can be strictly improved under exchangeability of the input pvalues, by relaxing symmetry constraints and processing p-values sequentially. In contrast to these approaches, which operate at the level of p-values and are calibrated for broad dependence classes, we leverage the data-level exchangeable structure induced by group-invariance under the null. This allows our procedures to adapt to the specific dependence structure of the transformed statistics and can yield strict power improvements over methods calibrated at the p-value level. Max-type aggregation with Monte Carlo calibration. A different line of work focuses on what we refer to as max-type aggregation procedures with data-dependent calibration via Monte Carlo approximation [1, 2, 4, 18, 19, 60–63]. These methods construct a global test by rejecting whenever at least one marginal statistic exceeds an oracle threshold designed to tightly control the type I error rate. Since this oracle threshold either depends on the unknown joint distribution of the statistics or is difficult to compute in closed form, it is approximated in practice via Monte Carlo calibration. While this approach can yield non-conservative procedures, the resulting finite-sample behavior depends critically on how the Monte Carlo calibration is implemented. In particular, the calibration step uses a separate batch of transformed statistics to approximate the rejection probability but does not enforce joint exchangeability between the calibration batch and the original data. This can result in a type I error rate that exceeds the nominal level α for finite B, as illustrated in Section 2.4. As further discussed in Appendix B.2, the procedure can equivalently be viewed as aggregating p-values by taking their minimum and estimating a critical value via Monte Carlo approximation. The SB and TB constructions studied here avoid this finite-B approximation issue by calibrating directly across exchangeable testing rows (resulting in valid type I error control), while accommodating general merging functions beyond the minimum p-value. Nonparametric combination and permutation-based constructions. Related ideas also appear in earlier work on nonparametric combination (NPC) methods and permutation-based tests. The NPC framework, systematized by [49, Chapter 4.2], combines dependent permutation tests by applying a combining function to permutation distributions of individual statistics, and has been applied to replicated designs [68] as well as to testing global null hypotheses defined as intersections of multiple component hypotheses [9]. More recently, [64] consider goodness-of-fit tests that aggregate multiple standardized statistics computed from the same data, with validity ensured by exchangeability of replicas generated under the null. Closely related ideas appear in [47], who aggregate statistics across multiple function classes and calibrate the aggregate using a single collection of permutations. Thus, the exchangeability-based validity argument for SB aggregation follows the same principle as earlier permutation-combination work. Our contribution is instead to analyze this construction from the perspective of finite-sample power and dependence adaptivity, and to extend it through sequential spending and a two-batch construction that enables data-dependent aggregation while retaining non-asymptotic validity. Rank-transformed subsampling. [26] study rank-transformed subsampling methods for aggre3

gating exchangeable statistics or p-values, and establish asymptotic tightness under suitable regularity conditions. Their approach is complementary to ours in that it targets subsampling-based aggregation and asymptotic regimes, whereas we focus on aggregation under group-invariance and derive non-asymptotic validity and power guarantees.

1.3

Organization

The remainder of the paper is organized as follows. Section 2 introduces the setup and reviews existing approaches to aggregating multiple tests, including p-value merging methods and maxtype aggregation with Monte Carlo calibration. Section 3 presents the single-batch aggregation procedure, establishes its finite-sample validity, and develops a detailed power analysis, including adaptivity to dependence among p-values and several refined special cases. Section 4 presents the sequential extension of SB minimum aggregation. Section 5 introduces the two-batch aggregation procedure, establishes its finite-sample validity, and highlights settings in which separating standardization from calibration is essential. Section 6 presents examples on adaptive hypothesis testing and conformal prediction, which highlight the scope and practical implications of the proposed methods. Section 7 presents numerical experiments that corroborate our theoretical findings and demonstrate the empirical performance of the proposed methods before we conclude in Section 8. Additional technical results and omitted proofs are collected in the supplementary material.

1.4

Notation

For an integer n ≥ 1, let [n] := {1, 2, . . . , n} and [n]0 := {0, 1, 2, . . . , n}. For a real number x, we denote by ⌊x⌋ and ⌈x⌉ its floor and ceiling, respectively. A random variable X is said to be super-uniform if P(X ≤ t) ≤ t for all t ∈ [0, 1]. For q ∈ R, we denote by Quantileq {X0 , X1 , . . . , XB } the empirical q-quantile of X0 , X1 , . . . , XB , defined as  B 1 X Quantileq {X0 , X1 , . . . , XB } := inf t ∈ R : 1(Xi ≤ t) ≥ q , B+1 

i=0

with the conventions Quantileq = −∞ for q ≤ 0 and Quantileq = +∞ for q > 1. For a sequence d

of random variables (Yn ) and a random variable Y , we write Yn −→ Y to denote convergence in p distribution and Yn −→ Y to denote convergence in probability. For a given set A, we use |A| to denote its cardinality. For two real positive sequences (an ) and (bn ), we write an ≲ bn if there exists a constant C > 0 such that an ≤ Cbn . If both an ≲ bn and bn ≲ an hold, we write an ≍ bn .

2

Setup and existing aggregation methods

This section introduces the group-invariance setup and transformed arrays used throughout the paper. We then review the classical permutation test, which we use as an umbrella term for invariancebased testing procedures induced by group transformations. Although the term “randomization test” is also common in the literature [e.g., 39, page 832], we use the terminology “permutation test” throughout for brevity. Finally, we survey existing approaches for aggregating multiple tests, including methods that combine p-values via averaging or order statistics [e.g., 21, 72, 73], as well as max-type aggregation procedures based on Monte Carlo calibration [e.g., 2, 18, 19, 59, 62]. 4

2.1

Group-invariance and transformed arrays

We work within a general group-invariance framework. Let G be a finite group of transformations acting on the sample space, and let X denote the observed data. The classical group-invariance hypothesis, also referred to as the randomization hypothesis [39, Definition 17.2.1], is stated as follows. Definition 1 (Group-invariance hypothesis). Under the null hypothesis, the distribution of X is invariant under the action of G; that is, for every g ∈ G, the random variables g(X) and X have the same distribution. This framework encompasses a broad class of inference problems and underlies many distribution-free methods in statistics. In permutation testing, for example, G may consist of all permutations of the pooled sample in two-sample or k-sample problems, of sign-flip transformations in paired or symmetric designs, or of permutations that shuffle one variable while holding the other fixed in tests of independence [e.g., 17, 22, 39, 49]. More generally, conditional randomization tests arise when G is induced by resampling a subset of features conditional on the others [e.g., 6, 8]. Similar exchangeability arguments also underpin conformal inference, where they justify rank-based calibration of conformity scores [e.g., 3, 40, 71]. Let T 1 , . . . , T K be real-valued test statistics computed from the data X, and consider the problem of aggregating the evidence provided by these statistics into a single test of the group-invariance hypothesis. Let g0 := id and let g1 , . . . , gB ∈ G denote a collection of transformations. Unless stated otherwise, we assume that the transformations are chosen so that the transformed data g0 (X), g1 (X), . . . , gB (X) are exchangeable under the group-invariance hypothesis. This assumption includes, as special cases, the canonical Monte Carlo setting in which g1 , . . . , gB are i.i.d. draws from the uniform distribution on G conditional on X, as well as sampling without replacement from G \ {id}. Define the statistics computed from the transformed data by  Tbk := T k gb (X) , b = 0, 1, . . . , B, k = 1, . . . , K. (1) Under the group-invariance hypothesis, the collection of row vectors (Tb1 , . . . , TbK ) is exchangeable in the index b = 0, 1, . . . , B. This joint exchangeability is the basic structural fact behind rowwise permutation aggregation. Our analysis below asks how much power is gained by calibrating the aggregate on this realized exchangeable array, and how the construction can be extended to data-dependent aggregation rules.

2.2

Classical permutation tests

We begin by reviewing the classical permutation test for assessing the group-invariance hypothesis using a single test statistic T . The permutation test dates back to the seminal work of [17, 50] and has since been extensively studied; see, for example, [49] and [39, Chapter 17]. Given a test statistic T , designed such that larger values indicate stronger evidence against the null hypothesis, the permutation p-value corresponding to the observed statistic T0 is defined as B

p(T0 ) =

1 X 1(Tb ≥ T0 ), B+1 b=0

5

where T0 denotes the statistic computed from the original data, and T1 , . . . , TB denote the statistics computed from the transformed datasets as in (1). The permutation test rejects the null hypothesis when p(T0 ) ≤ α. The permutation test can equivalently be expressed in terms of the test statistic itself. In particular, rejecting the null hypothesis when p(T0 ) ≤ α is equivalent to rejecting when  T0 > Quantile1−α T0 , T1 , . . . , TB = T(⌈(1−α)(B+1)⌉) , where T(k) denotes the k-th smallest order statistic of T0 , T1 , . . . , TB . See e.g., Lemma S.4 in Appendix G. This formulation proves convenient in what follows, especially when we compare and calibrate aggregated statistics. Under exchangeability of T0 , T1 , . . . , TB , the permutation test has rejection probability at most α in finite samples; see, for example, [30, 52, 55], and Lemma S.4. As emphasized earlier, our primary interest lies in aggregating multiple test statistics T01 , . . . , T0K into a single test for evaluating the group-invariance hypothesis, while maintaining rigorous type I error control and favorable power properties. We next review existing approaches to this problem.

2.3

P-value merging under arbitrary dependence

In much of the existing literature on multiple testing and aggregation, individual test statistics are first converted into p-values, which are then combined to form a single global test. We review such p-value-based aggregation methods under arbitrary dependence. Let p1 , . . . , pK denote K individual p-values, each assumed to be super-uniform. A p-merging function is a measurable mapping f : [0, 1]K → [0, 1] such that, for any collection of input p-values, the aggregated random variable f (p1 , . . . , pK ) is itself super-uniform. Equivalently, a p-merging function combines multiple p-values into a single valid p-value while guaranteeing type I error control under arbitrary dependence structures among the inputs. Classical p-value merging procedures combine marginal p-values through deterministic rules that remain valid under arbitrary dependence, such as order-statistic merging and generalized-mean merging. Two prominent examples are the O-family [57] and the M-family [73]. These procedures are necessarily calibrated for worst-case dependence, which can lead to conservativeness and reduced power in realistic settings. Details on the O- and M-families, including their precise formulas, sharp constants, and admissibility properties, are recalled in Appendix B.1.

2.4

Max-type aggregation with Monte Carlo calibration

We review a class of max-type (MaxT) aggregation procedures, following [2] and [62], with earlier roots in [1, 4, 18, 19]. For a single statistic T k , the permutation test rejects the null hypothesis whenever  T0k > Quantile1−α T0k , T1k , . . . , TBk . To aggregate multiple statistics T01 , . . . , T0K , the procedures of [2, 62] construct a MaxT test that rejects the null hypothesis if at least one statistic exceeds its marginal permutation threshold after inflating the nominal significance level by a data-dependent factor ũα , namely, o n  max T0k − Quantile1−ũα T0k , . . . , TBk > 0. (2) k∈[K]

6

The correction factor ũα is intended to approximate the largest inflation of the nominal level that preserves type I error control. In practice, [2] and [62] estimate this quantity via Monte Carlo calibration     B n o  k 1 X k k ũα := sup u ∈ (0, 1) : 1 max TB+b − Quantile1−u T0 , . . . , TB >0 ≤α , B k∈[K]

(3)

b=1

where ũα is set to zero if the supremum is taken over an empty set.1 Operationally, this procedure k k to Monte Carlo-approximate the uses an additional batch of transformed statistics TB+1 , . . . , T2B population-level rejection probability, which may be viewed as an additional calibration step for type I error control. While the supremum in (3) is typically computed via bisection, we derive in Proposition S.1 a closed-form expression for ũα , eliminating the need for iterative search. Despite its practical appeal, the Monte Carlo calibration in (3) does not in general guarantee finite-sample level control. This can already be seen in the single-statistic case K = 1. Suppose that T0 , T1 , . . . , T2B are exchangeable and almost surely distinct, and let ũα be the Monte Carlo critical value obtained from (3) with K = 1. Then the resulting test satisfies  ⌊Bα⌋ + 1 ; P T0 > Quantile1−ũα {T0 , . . . , TB } = B+1 see Proposition S.8. This quantity can exceed α for finite B, although the discrepancy vanishes as B → ∞; the empirical level study in Figure 5 illustrates this finite-B size inflation. Thus, the Monte Carlo step is only an approximation, not a finite-sample rank calibration. The SB and TB procedures below avoid this issue by calibrating directly over exchangeable testing indices.

3

Single-batch aggregation

We begin with single-batch aggregation, which uses the same batch of transformations for both standardization and calibration. This section is organized as follows. We formulate the SB procedure in Section 3.1, establish its finite-sample validity in Section 3.2, and study its power properties in Section 3.3. In Section 3.4, we investigate its adaptivity to dependence among the test statistics. Section 3.5 provides a detailed treatment of the minimum merging function as a representative special case. Additional material on data-driven selection of merging functions within the SB framework is deferred to Appendix D.2.

3.1

Procedure

We formulate the SB aggregation procedure, which has its roots in [9, 47, 49, 64, 68, 75]. The central construction is to aggregate evidence across multiple test statistics while preserving exchangeability by applying the same set of transformations to all statistics simultaneously. While prior work has primarily focused on validity, a systematic analysis of power and adaptivity properties has remained largely unexplored. Moreover, some earlier formulations [e.g., 62] rely on Monte Carlo permutation p-values that do not fully preserve joint exchangeability across transformed copies, which can lead to slight finite-sample liberality as pointed out in Proposition S.8. Other approaches 1 Prior work allows non-uniform non-negative weights wk in Quantile1−uwk {T0k , . . . , TBk }. We take wk = 1 here for notational simplicity.

7

Algorithm 1 Single-Batch Aggregation (SB) Require: data X; statistics T 1 , . . . , T K ; transformations g0 , . . . , gB (g0 = id); merging function f ; level α. 1: Transformed statistics: For each b ∈ [B]0 and k ∈ [K], compute Tbk ← T k (gb (X)). 2: Standardization via p-values: For b ∈ [B]0 and k ∈ [K], compute the permutation p-value B

p Tbk



 1 X ← 1 Tik ≥ Tbk . B+1 i=0

3: Aggregation row-wise: For each b ∈ [B]0 , set

  fb ← f p Tb1 , . . . , p TbK . 4: Calibration: Compute the p-value B

pSB ←

 1 X 1 fb ≤ f0 . B+1 b=0

5: Decision rule: Reject if pSB ≤ α.

[e.g., 1, 2, 4, 18, 19] assume access to an oracle critical value, but such procedures are not practically implementable, requiring knowledge of the unknown joint null distribution. By contrast, the SB formulation enforces rank-based calibration across all transformations and is exactly implementable, requiring neither approximation nor oracle knowledge. Consider B + 1 transformations g0 , . . . , gB , with g0 being the identity. For each statistic T k and each transformation index b ∈ [B]0 , compute  Tbk := T k gb (X) . Using these values, define the permutation p-value B

p(Tbk ) :=

 1 X 1 Tik ≥ Tbk , B+1 i=0

b ∈ [B]0 , k ∈ [K],

(4)

which ranks each Tbk among {T0k , . . . , TBk }. Let f : [0, 1]K → R be a function that aggregates K permutation p-values into a single number. No structural assumptions on f are required for finite-sample validity of the SB procedure. For convenience of interpretation, however, we assume throughout that smaller values of f indicate stronger evidence against the null hypothesis, consistent with the interpretation of p-values. Typical examples include the minimum, the average, and the median of p-values. For each transformation index b, aggregate the p-values computed from the corresponding transformed dataset via  fb := f p(Tb1 ), . . . , p(TbK ) , b ∈ [B]0 . Crucially, the statistics Tb1 , . . . , TbK are all computed under the same transformation gb . Thus, the aggregation is performed row-wise across statistics sharing a common transformation, inducing a 8

coupling that preserves exchangeability. The SB aggregation test then compares the aggregated value f0 (corresponding to the original data) to its transformed counterparts: B

pSB :=

 1 X 1 fb ≤ f0 , B+1 b=0

and rejects the null hypothesis whenever pSB ≤ α. The SB aggregation procedure is summarized in Algorithm 1; a schematic illustration is deferred to Figure 4 in Appendix A.

3.2

Finite-sample validity

The finite-sample validity of the SB aggregation test follows directly from exchangeability. Under the group-invariance hypothesis, the row vectors (T01 , . . . , T0K ), . . . , (TB1 , . . . , TBK ) are exchangeable, since each row is obtained by applying the same transformation to the data. The column-wise p-value transformation preserves this exchangeability in the sense that permuting the rows of the statistics matrix induces the same permutation of the rows of the p-value matrix. Applying the merging function row-wise therefore yields aggregated values (f0 , . . . , fB ) that are also exchangeable. Consequently, the rank-based p-value pSB is super-uniform under the null, and the SB aggregation test controls the type I error rate at level α in finite samples. We formalize this statement in the following theorem. Theorem 1. Under the group-invariance hypothesis, the SB aggregation test satisfies  P pSB ≤ α ≤ α for all α ∈ (0, 1) and B ≥ 1. Moreover, if f0 , f1 , . . . , fB are distinct with probability one, then  ⌊(B + 1)α⌋ P pSB ≤ α = . B+1 We stress again that the above validity result holds for any choice of the merging function f . In particular, the p-value transformation in Algorithm 1 is not essential for finite-sample validity, and f may be applied directly to the collection of statistics {Tbk }k∈[K] computed from each transformed dataset [7, 54, 77]. Nevertheless, working with p-values provides a natural standardization across statistics, which reduces sensitivity to differences in scale. As alternative forms of standardization, one may also studentize each statistic prior to aggregation [e.g., 64], or transform each statistic using its null distribution function when the (asymptotic) null distribution is known [e.g., 26]. More generally, the validity result extends to procedures that break ties in the marginal statistics Tbk and/or in the merged statistics fb using auxiliary variables, provided that augmenting each row with these variables preserves exchangeability. This includes standard randomized constructions, such as lexicographical tie-breaking with i.i.d. uniform auxiliary variables. The same principle also accommodates deterministic, data-dependent choices of the auxiliary variables, enabling tiebreaking rules that exploit the structure of the aggregated statistics while preserving finite-sample validity. Concrete constructions and their power implications are deferred to Appendix B.6.

9

For later use, we record the following equivalent threshold representation of the SB decision rule, which rejects when  f0 < −Quantile1−α −f0 , −f1 , . . . , −fB := ûSB (5) α . To facilitate several subsequent arguments, we provide an equivalent characterization of ûSB α in terms of a supremum-based empirical quantile. Proposition 1. The SB threshold in (5) admits the equivalent representation 

ûSB α = sup

 B  1 X 1 fb ≤ u ≤ α . u∈R: B+1

(6)

b=0

This alternative characterization will be repeatedly invoked throughout the paper. In particular, it plays a central role in the power and adaptivity analysis.

3.3

Power dominance over deterministic calibration

We study the power properties of the SB aggregation test. Our main result (Theorem 2) shows that, for any merging function f , the SB procedure is at least as powerful as any test based on a deterministic threshold calibrated under arbitrary dependence among p-values. Moreover, the SB threshold adapts to the dependence structure of the p-values and can yield strict power improvements over worst-case dependence calibrations. We begin with a super-uniformity lemma that underpins the power analysis. A more general weighted version is given in Lemma S.8, which extends [28, Lemma A1] to arbitrary measures. Lemma 1. Let T0 , T1 , . . . , TB ∈ R and define B

pb :=

1 X 1(Ti ≥ Tb ), B+1 i=0

b ∈ [B]0 .

Then, for any α ∈ [0, 1] and any integer B ≥ 1, B

⌊(B + 1)α⌋ 1 X 1(pb ≤ α) ≤ ≤ α. B+1 B+1 b=0

Moreover, if T0 , . . . , TB are distinct, the first inequality holds with equality. By Lemma 1, conditional on the observed data X, the permutation p-value p(Tbk ) with b ∼ Unif([B]0 ) is super-uniform. In particular, for each k ∈ [K], B

 ⌊(B + 1)α⌋ 1 X 1 p(Tbk ) ≤ α ≤ ≤ α. B+1 B+1 b=0

This conditional super-uniformity underlies the finite-sample power comparison below.

10

Theorem 2. Fix α ∈ (0, 1). Let U = (U1 , . . . , UK ) be a random vector with an arbitrary joint K denote the class of all such distribution such that each coordinate Uk is super-uniform. Let Usup joint laws. Suppose that cα,K is a deterministic constant satisfying  sup Pν f (U1 , . . . , UK ) ≤ cα,K ≤ α. K ν∈Usup

Then it holds almost surely that cα,K < ûSB α . Consequently, the SB aggregation test defined in Algorithm 1 satisfies  P f0 ≥ ûSB ≤ P(f0 > cα,K ) . α Theorem 2 shows that the SB threshold strictly dominates any deterministic threshold ensuring type I error control under arbitrary dependence. Equivalently, the type II error of the SB test is no larger than that of any level-α test based on such a deterministic threshold. Despite the strength of the result, the argument is elementary. It uses only the fact that, conditional on the observed data, the permutation p-values p(Tb1 ), . . . , p(TbK ) are super-uniform when b is drawn uniformly from [B]0 . A deterministic threshold must be valid for every joint law with super-uniform margins, whereas the SB threshold is calibrated directly on the realized conditional empirical distribution. This is why SB obtains a larger critical value and hence higher power. It is worth emphasizing that Theorem 2 is entirely deterministic and imposes no assumptions on the transformations g1 , . . . , gB . Consequently, the result continues to hold for general weighted permutation schemes, including those considered in [52, Theorem 2]. As an immediate consequence, we obtain the following comparison with existing p-merging procedures calibrated under arbitrary dependence. Corollary 1. Let f be a classical p-merging rule calibrated under arbitrary dependence, such as an O-family or M-family rule; see Appendix B.1 for the definitions. Since such rules output valid p-values under arbitrary dependence, the corresponding deterministic calibration is cα,K = α. Consequently, the SB aggregation test is always at least as powerful as the corresponding worst-case calibrated test.

3.4

Adaptivity to dependence

We now illustrate how the SB procedure adapts to the unknown dependence structure among p-values. A key limitation of p-merging methods calibrated under arbitrary dependence is that their critical values are typically driven by worst-case dependence and can therefore be overly conservative outside special cases. In contrast, the SB procedure is data-dependent and can adapt to the dependence structure induced by the underlying data. We illustrate this adaptivity in two settings: the extreme case of perfect rank alignment and a more general asymptotic regime. Extreme case: perfect rank alignment. We begin by considering an extreme case in which no multiplicity adjustment is needed. The test statistics need not be identical across coordinates; it suffices that they induce the same ordering across transformations. Formally, suppose that for all k, k ′ ∈ [K] and all b, b′ ∈ [B]0 , ′ ′ Tbk ≤ Tbk′ ⇐⇒ Tbk ≤ Tbk′ . (7)

That is, the vectors (T0k , . . . , TBk ) are perfectly rank-aligned across coordinates, although their numerical values may differ (e.g., different scalings of the same statistic). Under (7), the permutation 11

p-values satisfy p(Tb1 ) = · · · = p(TbK )

for all b ∈ [B]0 ,

since each column induces the same ranking of the transformed statistics. Consider a monotone merging function f satisfying f (x, . . . , x) ≤ f (y, . . . , y)

⇐⇒

x ≤ y,

x, y ∈ R.

(8)

This condition is satisfied by many commonly used aggregation rules, including the minimum, the median, and the arithmetic mean. The next proposition shows that, under perfect rank alignment, the resulting SB aggregation test coincides with the usual permutation test based on a single statistic. Proposition 2. Consider the SB aggregation test in Algorithm 1. Assume that the merging function f satisfies the diagonal monotonicity condition (8). If the statistics {Tbk } satisfy the rank alignment condition (7), then B

pSB =

1 X 1(Tb ≥ T0 ), B+1 b=0

where Tb denotes any representative statistic (e.g., Tb1 ). Thus, whenever multiple test statistics convey the same ordering information across permutation samples, the SB procedure automatically detects this redundancy and reduces to a single-statistic permutation test. In this case, the effective number of tests is one, and no multiplicity correction is applied. Asymptotic adaptivity. We next develop more general adaptivity results using asymptotic arguments. Suppose that the limiting null distribution of the merging function f applied to permutation p-values were known. This distribution implicitly encodes the dependence structure among the pvalues and thus defines an oracle benchmark. In that case, one could reject the null hypothesis whenever f0 is less than its oracle α-quantile. We show below that, under suitable conditions, the SB threshold converges to this oracle α-quantile. Proposition 3. For each n, let X(n) denote the observed data set, and let fb,n , b ∈ [B]0 , be the SB aggregation scores computed from the corresponding permutation p-values. Suppose that, conditional on X(n) , the transformations g1 , . . . , gB are i.i.d. uniform on G. Fix α ∈ (0, 1), and consider any joint asymptotic regime in which B, n → ∞. Assume that d

′ (f1,n , f2,n ) −→ (V∞ , V∞ ), ′ are i.i.d. with distribution function F . Let Q⋆ := inf{u ∈ R : F (u) ≥ α} denote where V∞ and V∞ α the oracle lower α-quantile, and assume that it is well separated: for every ε > 0,

F (Q⋆α − ε) < α < F (Q⋆α + ε). p

⋆ Then the SB critical value satisfies ûSB α −→ Qα .

12

As a consequence, when both B and n are large, the SB threshold behaves as if the limiting null distribution of the merging function were known. This provides an asymptotic explanation for the adaptivity of the SB aggregation test to the unknown dependence structure among the test statistics or p-values. The joint convergence of (f1,n , f2,n ) to an i.i.d. limit is standard in the asymptotic theory of permutation distributions and is indeed a necessary condition for the permutation distribution to converge; see [13, Theorem 5.1]. A uniform version over classes of null distributions is deferred to Proposition S.4.

3.5

Minimum merging and Westfall–Young calibration

We now specialize the SB aggregation procedure to the minimum merging function, which recovers the Westfall–Young single-step method [75]; see also [43] for its asymptotic properties. Specifically, we consider  f p(Tb1 ), . . . , p(TbK ) = min p(Tbk ) =: fbmin , b ∈ [B]0 . (9) k∈[K]

Using the statistic formulation in (5), the SB minimum test rejects the null hypothesis whenever f0min < ûSB α,min , where  min min min ûSB . α,min := −Quantile1−α −f0 , −f1 , . . . , −fB The next proposition shows that the SB minimum test is bounded between the Bonferroni procedure and the unadjusted minimum test. Proposition 4. The SB minimum test satisfies, almost surely,       α k k SB k 1 min p(T0 ) ≤ ≤ 1 min p(T0 ) < ûα,min ≤ 1 min p(T0 ) ≤ α . K k∈[K] k∈[K] k∈[K] Moreover, ûSB α,min ≥ (⌊(B + 1)α/K⌋ + 1)/(B + 1). The additional term 1/(B + 1) in the second statement reflects the strict inequality in the SB decision rule: the SB minimum test rejects when mink∈[K] p(T0k ) < u rather than ≤ u as in the other procedures, which shifts the rejection threshold by one discretization step. The inequalities in Proposition 4 are tight with each bound attainable. In particular, when the coordinate-wise rejection events are disjoint, the Bonferroni correction can become tight (and is tight when the p-values are exactly uniform); in such cases the SB minimum test may coincide with the Bonferroni test (see Example S.1 for an explicit example). At the opposite extreme, when the rejection events coincide across coordinates, the SB minimum test reduces to the unadjusted single test (see Section 3.4). Together, these cases show that the SB minimum test interpolates between the Bonferroni and unadjusted procedures according to the dependence structure among the permutation p-values, while preserving level α. Additional material on data-driven aggregation over merging functions is deferred to Appendix D.2.

13

4

Sequential minimum aggregation

We next consider an alpha-spending version of SB minimum aggregation for ordered coordinates. In many applications, the coordinates are not merely an unordered collection of statistics, but have a natural order, such as increasing resolution, decreasing bandwidth, growing model complexity, or sequentially available data sources. In such settings, the alternative may already be visible in an early prefix of the sequence. Calibrating only the final minimum over all K coordinates can then be inefficient because later coordinates with little signal may still enter the calibration and make rejection harder. The general sequential aggregation principle is recorded in Appendix B.3. Here we specialize it to cumulative minima. The goal is to spend the global level α across a pre-specified sequence of prefix tests: at stage j, we inspect the cumulative minimum over the first j coordinates and spend only a portion αj of the total level. For b ∈ [B]0 and j ∈ [K], set mb,j := min1≤k≤j p(Tbk ). Let P α1 , . . . , αK ≥ 0 satisfy K j=1 αj ≤ α, and write qj := ⌊(B + 1)αj ⌋. The construction can be understood as a row-removal procedure. The set Sj−1 contains the permutation rows that have not yet been used for rejection at earlier stages. At stage j, among these surviving rows, we look at the prefix statistic mb,j and remove the rows whose values are among the most extreme according to the stage-j spending budget. If the observed row b = 0 is removed at any stage, the global null is rejected. Removing rows, rather than testing each prefix separately, keeps the eliminated sets disjoint and makes the finite-sample level calculation explicit. To define the procedure, initialize the survivor set as S0 = [B]0 . At stage j, given the current survivor set Sj−1 , let S

S

j−1 j−1 m(1),j ≤ · · · ≤ m(|S j−1 |),j

S

denote the order statistics of {mb,j : b ∈ Sj−1 }. Define the stage-j threshold cj := m(qj−1 , j +1),j eliminate Aj := {b ∈ Sj−1 : mb,j < cj }, and update Sj = Sj−1 \ Aj . The sequential minimum S aggregation test rejects H0 if 0 ∈ K j=1 Aj . The procedure is summarized in Algorithm 2, and the next proposition establishes its finite-sample validity. P Proposition 5. Fix α ∈ (0, 1), B ≥ 1, and α1 , . . . , αK ≥ 0 such that K j=1 αj ≤ α. Under the group-invariance null hypothesis, the sequential minimum aggregation test satisfies   K [ P 0∈ Aj ≤ j=1

K

 1 X (B + 1)αj ≤ α. B+1 j=1

Moreover, if for every j ∈ [K] the survivor values {mb,j : b ∈ Sj−1 } are almost surely distinct, then the first inequality is an equality. The proof is based on a simple exchangeability argument. The row-removal construction is equivariant with respect to permutations of the indices [B]0 : if the rows of the transformed array are permuted, then the survivor sets and eliminated sets are permuted in the same way. Under the null hypothesis, the row array is exchangeable, and therefore for each stage j, P(0 ∈ Aj ) = 14

E|Aj | . B+1

Algorithm 2 Sequential SB Minimum Aggregation (SeqSB) Require: data X; ordered statistics T 1 , . . . , T K ; transformations g0 , . . . , gB with g0 = id; spending sequence α1 , . . . , αK . 1: Set S0 ← [B]0 and mb,0 ← 1 for all b ∈ [B]0 . 2: for j = 1, . . . , K do 3: Compute Tbj ← T j (gb (X)) for all b ∈ [B]0 . 4: Compute permutation p-values p(Tbj ) as in (4) for all b ∈ Sj−1 . 5: Update mb,j ← min{mb,j−1 , p(Tbj )} for all b ∈ Sj−1 . 6: Set qj ← ⌊(B + 1)αj ⌋. 7: Let cj be the (qj + 1)-st order statistic of {mb,j : b ∈ Sj−1 }. 8: Set Aj ← {b ∈ Sj−1 : mb,j < cj }. 9: if 0 ∈ Aj then 10: Reject H0 and stop. 11: end if 12: Set Sj ← Sj−1 \ Aj . 13: end for 14: Do not reject H0 . The eliminated sets A1 , . . . , AK are disjoint by construction, and |Aj | ≤ qj . Hence   X K K [ P 0∈ Aj = P(0 ∈ Aj ) ≤ j=1

j=1

K

1 X qj . B+1 j=1

If the survivor values are distinct at every stage, then exactly qj rows are removed at stage j, giving equality in the first bound. The SB minimum rule is recovered by spending all level at the final coordinate, αK = α and αj = 0 for j < K. Thus, the sequential procedure can be viewed as a generalization of the SB minimum rule to ordered prefixes. It is not uniformly more powerful than the final SB minimum rule, since spending level early necessarily leaves less level for later coordinates. However, it can be advantageous when the signal is expected to appear in an early prefix. In that case, the stage-j threshold is calibrated only against the first j coordinates, rather than against the full minimum over all K coordinates, and hence may be less conservative. The procedure also has an operational advantage, as once the observed row is eliminated, the test can stop without computing later coordinates.

5

Two-batch aggregation and data-dependent rules

The SB procedure in Section 3 uses a single collection of transformed datasets both to (i) standardize each coordinate statistic via a permutation p-value map and (ii) calibrate the merged evidence through a permutation test. This section studies a TB aggregation procedure that leverages an extra batch for calibration while preserving finite-sample type I error control. Conceptually, TB aggregation proceeds as follows: a reference batch of transformations is used to construct a standardized p-value map for each coordinate statistic, and an independent testing batch is then used to run a permutation test on the merged evidence computed from these standardized p-values. 15

5.1

Holdout standardization

We begin with a formal description of the TB aggregation procedure. Let g0 = id and let g1 , . . . , g2B be additional transformations. As in the SB aggregation procedure, assume that g1 , . . . , g2B are chosen such that g0 (X), g1 (X), . . . , g2B (X) are exchangeable under the group-invariance hypothesis. We view {g0 , g1 , . . . , gB } as a testing batch, and {gB+1 , . . . , g2B } as a reference batch.2 Define the statistics  Tbk := T k gb (X) , b ∈ [B]0 , k ∈ [K]. For notational convenience, we re-index the reference batch by setting  g̃i := gB+i , i ∈ [B], T̃ik := T k g̃i (X) , i ∈ [B], k ∈ [K]. For each coordinate k ∈ [K] and each testing index b ∈ [B]0 , define the holdout permutation p-value by ranking the testing statistic Tbk against the reference batch {T̃1k , . . . , T̃Bk }: pHO Tbk



 B   X 1 k k := , 1+ 1 T̃i ≥ Tb B+1 i=1

b ∈ [B]0 , k ∈ [K].

Let f : [0, 1]K → R be a merging function that maps the K holdout p-values to a single realvalued score. As before, smaller values are interpreted as stronger evidence against the null hypothesis. For each b ∈ [B]0 , define  b ∈ [B]0 . f˜b := f pHO (Tb1 ), . . . , pHO (TbK ) , Finally, define the TB permutation p-value by comparing f˜0 to its testing-batch counterparts: pTB :=

B  1 X ˜ 1 fb ≤ f˜0 , B+1 b=0

and reject if pTB ≤ α.

(10)

In contrast to SB aggregation, where the same transformed statistics are used for both p-value construction and calibration, the TB procedure splits these roles across two independent batches conditional on X: a reference batch for standardization and a testing batch for calibration. The full procedure is summarized in Algorithm 3, and a schematic illustration is given in Figure 4 in Appendix A. As in Section 3.1, the decision rule (10) admits an equivalent quantile-based form:  f˜0 < −Quantile1−α −f˜0 , −f˜1 , . . . , −f˜B =: ûTB α . Equivalently, ûTB α can be written in the supremum-quantile form 

ûTB α = sup

 B  1 X ˜ u∈R: 1 fb ≤ u ≤ α . B+1

(11)

b=0

2 For notational convenience, we take the testing batch to have size B + 1 and the reference batch to have size B; however, all results extend directly to arbitrary batch sizes B1 and B2 .

16

Algorithm 3 Two-Batch Aggregation (TB) Require: data X; statistics T 1 , . . . , T K ; transformations g0 , . . . , g2B (g0 = id); merging function f ; level α. 1: Two batches: Define testing batch {g0 , . . . , gB } and reference batch {gB+1 , . . . , g2B }. 2: Transformed statistics: Compute Tbk ← T k (gb (X)) for b ∈ [B]0 , k ∈ [K]. 3: Transformed holdout statistics: Compute T̃ik ← T k (gB+i (X)) for i ∈ [B], k ∈ [K]. 4: Standardization via holdout p-values: For each b ∈ [B]0 , k ∈ [K], set pHO (Tbk ) ←

  B X  1 k k 1+ 1 T̃i ≥ Tb . B+1 i=1

5: Aggregation row-wise: For each b ∈ [B]0 , set

  f˜b ← f pHO Tb1 , . . . , pHO TbK . 6: Calibration: Compute the p-value

pTB ←

B  1 X ˜ 1 fb ≤ f˜0 . B+1 b=0

7: Decision rule: Reject if pTB ≤ α.

5.2

Finite-sample validity

Under the group-invariance hypothesis, the testing-batch row vectors (T01 , . . . , T0K ), . . . , (TB1 , . . . , TBK ) are conditionally exchangeable given the reference batch {(T̃i1 , . . . , T̃iK )}B i=1 . Since the holdout pvalue map and the merging function are applied row-wise, the merged values (f˜0 , . . . , f˜B ) remain conditionally exchangeable. Therefore, the rank-based p-value pTB in (10) is super-uniform under the null and controls the type I error rate at level α in finite samples as stated below. Theorem 3. Under the group-invariance hypothesis, the TB aggregation test controls the type I error rate at level α, that is,  P pTB ≤ α ≤ α for all α ∈ (0, 1) and B ≥ 1. Moreover, if f˜0 , f˜1 , . . . , f˜B are distinct with probability one, then  ⌊(B + 1)α⌋ . P pTB ≤ α = B+1 More generally, finite-sample validity still holds if the merging function is chosen based on the reference batch, so it may be data-dependent as long as it depends only on the reference information. However, this batch separation can lead to lower finite-sample power when B is small; a more detailed comparison with SB aggregation is deferred to Appendix E.3. The next subsection highlights settings where the separation is nevertheless useful. 17

5.3

Learning aggregation rules from the reference batch

A key advantage of TB aggregation is that the reference batch can be used to construct or choose the aggregation rule before the final calibration step. This enables inference tasks and aggregation strategies that are incompatible with the SB construction. We highlight two representative settings where the separation between reference and testing batches is essential. Merging conformal prediction sets. In conformal prediction, the score of the test point at its unknown true label is unavailable at calibration time. Because TB holdout p-values are computed from the reference batch alone, the calibration step can be precomputed and reused across candidate labels. This avoids the set-inversion bottleneck that would generally arise from an SB-style calibration incorporating each candidate test-point score into its own ranking set. We return to this connection and its algorithmic implications in Section 6.2 and Appendix B.5. Learning the aggregation map. Because the aggregation rule may be chosen using only the reference batch, TB aggregation allows data-dependent choices such as learned weights, projections, or dimension-reduced summaries of the coordinate-wise p-value vector. Conditional on the reference batch, the testing-batch rows remain exchangeable, so the final permutation calibration remains finite-sample valid. Such data-adaptive choices are not available under SB aggregation without additional correction, since learning the rule and calibrating the test on the same transformed rows would generally destroy the exchangeability underpinning the SB procedure. A PCA-based example is given in Appendix B.4.

6

Applications

This section presents notable theoretical and methodological applications of the preceding results. Section 6.1 and Appendix F consider general and kernel-based adaptive nonparametric testing with a focus on SB aggregation, while Section 6.2 examines both SB and TB aggregation in the context of conformal prediction, with particular emphasis on the computational advantages of TB aggregation. Additional material on SB aggregation of data-asymmetric tests is deferred to Appendix D.3.

6.1

Adaptive nonparametric testing

Many nonparametric testing problems require adaptation over unknown tuning parameters, such as smoothness or intrinsic dimension. We show that SB aggregation can replace Bonferroni aggregation in such adaptive constructions, preserving existing separation-rate guarantees while improving finitesample power. Setup and minimax separation radius. Let P be a model class and consider testing H0 : P ∈ P0

versus

H1 : P ∈ P1 (ρ),

where P0 ⊂ P is a class of null distributions under which the group-invariance hypothesis holds for a given transformation scheme. The alternative class is indexed by a separation parameter ρ > 0 and takes the generic form  P1 (ρ) := P ∈ P : d(P (0) , P (1) ) ≥ ρ , where (P (0) , P (1) ) denotes the pair of distributions whose equality characterizes the null, and d is a problem-specific distance or semi-metric. 18

For a test ∆ based on a sample of size n with error levels α ∈ (0, 1), β ∈ (0, 1 − α), define the uniform separation radius as o n ρ(∆; α, β, n) := inf ρ > 0 : sup PP (∆ = 1) ≤ α, sup PP (∆ = 0) ≤ β . P ∈P0

P ∈P1 (ρ)

Let Ωα denote the class of all level-α tests, and define the minimax benchmark ρ⋆ (α, β, n) := inf ρ(∆; α, β, n). ∆∈Ωα

We say that a test ∆ is minimax optimal if ρ(∆; α, β, n) ≲ ρ⋆ (α, β, n), where the inequality holds up to constants independent of (α, β, n) but may depend on other problem parameters. This formulation is standard in nonparametric testing; see, e.g., [31, 60]. Adaptive aggregation and transfer to SB calibration. A standard route to adaptivity aggregates a collection of tests {∆k }k∈[K] via Bonferroni correction:  ∆Bonf := 1

min p

(k)

k∈[K]

 ≤ α/K ,

where p(k) is the p-value associated with ∆k . This procedure is valid under arbitrary dependence and satisfies ρ(∆Bonf ; α, β, n) ≤ min ρ(∆k ; α/K, β, n). k∈[K]

In many minimax problems, the separation radius depends on α only through a slowly varying term, so Bonferroni aggregation over a tuning-parameter grid incurs only a logarithmic adaptivity cost. This strategy underlies adaptive permutation tests such as those in [5, 27, 33, 34, 46]. Under the same group-invariance scheme, Bonferroni aggregation can be replaced by SB calibration. By Theorem 1 and Proposition 4, the SB minimum test controls the type I error non-asymptotically and uniformly dominates the Bonferroni test. Consequently, existing Bonferroni-based adaptive separation bounds transfer directly to the SB-calibrated counterpart. More precisely, denoting by ∆SB the SB minimum test, we have ρ(∆SB ; α, β, n) ≤ ρ(∆Bonf ; α, β, n). This inequality is immediate from the finite-sample dominance of SB over Bonferroni calibration and requires no additional analysis beyond that of the existing adaptive constructions. In Appendix F, we instantiate this principle for adaptive two-sample testing and independence testing using kernel-based statistics, and show that the SB minimum test achieves the same adaptive separation rates as MaxT-based procedures [2, 62] while retaining finite-sample validity.

6.2

Conformal prediction

Conformal prediction [71] often yields multiple valid prediction sets arising from different nonconformity scores, feature representations, or training procedures. A natural question is how to merge such sets into a single prediction set that preserves finite-sample coverage while improving efficiency. In this section, we show how the SB and TB aggregation frameworks developed earlier can 19

be used to combine K conformal predictors. The key distinction between the two is computational: SB aggregation yields a fully self-calibrated conformal p-value but requires recalibration for each candidate label, whereas TB aggregation decouples calibration from label evaluation and enables efficient set inversion. Related approaches include validity-preserving selection of a single conformal predictor [76] and aggregation or selection rules valid under arbitrary dependence [20, 29]. Our approach instead uses permutation-based aggregation to exploit the underlying dependence among conformal p-values. Setup. To make the discussion concrete, we begin by recalling the standard split conformal setup [40, 48] with multiple scores. Let {(Xi , Yi )}ni=1 be a calibration sample and let (Xn+1 , Yn+1 ) be a new exchangeable observation. For each k ∈ [K], let sk : X × Y → R be a nonconformity score, and define Sik := sk (Xi , Yi ),

i ∈ [n],

and

k Sn+1 (y) := sk (Xn+1 , y),

y ∈ Y.

The associated split conformal p-value is pk (y) :=

1+

Pn

k k i=1 1{Si ≥ Sn+1 (y)}

n+1

,

(12)

which induces the prediction set Ck (Xn+1 ) := {y ∈ Y : pk (y) > α}.  By exchangeability, pk (Yn+1 ) is super-uniform, and hence P Yn+1 ∈ Ck (Xn+1 ) ≥ 1 − α. Our objective is to merge the collection {Ck }K k=1 into a single prediction set that retains finitesample coverage while exploiting dependence across scores to improve efficiency. To this end, we consider a merging function f : [0, 1]K → R, such as the minimum, median, or mean of the individual p-values. SB aggregation. For a candidate label y, SB aggregation constructs an aggregated conformal pvalue as follows. For each score k, form an augmented sample by appending the test-point score to the calibration scores and define ( Sik , i ∈ [n], Sik (y) := k Sn+1 (y), i = n + 1. Based on this augmented sample, define the column-wise permutation p-values Pik (y) :=

n+1 o 1 X n k 1 Sj (y) ≥ Sik (y) , n+1 j=1

i ∈ [n + 1], k ∈ [K],

(13)

k (y) = p (y). The standardized p-values are then aggregated across scores in a row-wise so that Pn+1 k manner using the merging function f ,  Mi (y) := f Pi1 (y), . . . , PiK (y) , i ∈ [n + 1].

The SB aggregated conformal p-value is obtained by calibrating the merged test-point value against its calibration counterparts, n+1

1 X pSB (y) := 1{Mi (y) ≤ Mn+1 (y)} , n+1 i=1

20

and the resulting merged conformal prediction set is CSB (Xn+1 ) := {y ∈ Y : pSB (y) > α}. When y = Yn+1 , the augmented score matrix is row-exchangeable. Since the permutationequivariant column-wise p-value map in (13) and the merging function f are applied row-wise, the merged values are exchangeable, so pSB (Yn+1 ) is super-uniform and  P Yn+1 ∈ CSB (Xn+1 ) ≥ 1 − α. The SB power comparison in Theorem 2 further implies that the SB merged conformal set is  uniformly no larger than the conformal set obtained by calibrating f p1 (y), . . . , pK (y) using any deterministic worst-case correction. The computational drawback is that pSB (y) must typically be evaluated over many candidate k (y) enters the ranking set in (13), so the column-wise p-values, labels y. The test-point score Sn+1 the merged values, and the final calibration all change with y. Thus, SB aggregation requires full recalibration for each candidate label. TB aggregation. We now introduce a TB construction whose key feature is that the calibration threshold can be computed once, independently of the candidate label y. Partition the calibration indices as [n] = Iref ∪ Iagg with |Iref | = n1 and |Iagg | = n2 (e.g., via a random split), and retain Xn+1 as the test feature. Using the reference batch, standardize each score via reference-based p-values. Specifically, for each k ∈ [K] and i ∈ Iagg , define P 1 + j∈Iref 1{Sjk ≥ Sik } k,ref := Pi . (14) n1 + 1 For a candidate label y, define analogously the test-point p-values P k (y)} 1 + j∈Iref 1{Sjk ≥ Sn+1 k,ref Pn+1 (y) := . n1 + 1

(15)

Aggregate across k using the merging function f to obtain   K,ref 1,ref (y) . Mi := f Pi1,ref , . . . , PiK,ref , i ∈ Iagg , (y), . . . , Pn+1 Mn+1 (y) := f Pn+1 Calibration compares the merged test-point value to the empirical distribution of the merged calibration values. Define  uTB α := − Quantile(1−α)(1+1/n2 ) −Mi : i ∈ Iagg , (16) CTB (Xn+1 ) := {y ∈ Y : Mn+1 (y) ≥ uTB α }. The factor (1 + 1/n2 ) accounts for calibrating with the n2 aggregation-batch values alone, rather than with the augmented set that also includes the test point. A key feature of the TB construction is that the resulting threshold uTB depends only on α {Mi }i∈Iagg and is thus independent of y. As a result, when many labels are evaluated, the calibration threshold is computed once and reused, and only the test-point quantity Mn+1 (y) needs to be updated for each candidate label y. This decoupling of calibration from label evaluation is the central computational advantage of TB aggregation over the SB construction, which requires full recalibration for each candidate label. 21

The finite-sample coverage guarantee follows from the same conditional exchangeability argument. Given the split, the reference-based p-value maps are fixed, so the aggregation-batch values and the test-point value at the true label are treated symmetrically. Thus,  P Yn+1 ∈ CTB (Xn+1 ) ≥ 1 − α. Example: TB intersection shortcut for residual scores. A particularly simple and useful specialization arises in regression with absolute residual scores. Suppose that sk (x, y) = |y − µ bk (x)|,

so that

Sik = |Yi − µ bk (Xi )|,

where µ bk : X → R is an arbitrary prediction function (e.g., a regression estimator) associated with score k and trained on separate data. Consider the minimum merging rule f (p1 , . . . , pK ) = min pk . k∈[K]

k ≤ · · · ≤ Rk Let {Rjk }j∈Iref denote the reference residuals Rjk := |Yj − µ bk (Xj )|, and write R(1) (n1 ) k := for their order statistics, with the convention R(n ∞. Then the TB set in (16) admits the 1 +1) following closed-form intersection representation.

Proposition 6. Suppose that f (p1 , . . . , pK ) = mink∈[K] pk . If uTB α = −∞ then CTB (Xn+1 ) = R,  TB  otherwise, letting ℓ := n1 + 2 − uα (n1 + 1) , it holds that CTB (Xn+1 ) =

K h \ k=1

i k k µ bk (Xn+1 ) − R(ℓ) ,µ bk (Xn+1 ) + R(ℓ) .

Thus, under minimum merging, TB aggregation reduces conformal set inversion to intersecting K residual-based intervals. Consequently, unlike the SB construction, Proposition 6 yields an explicit intersection form, so that conformal set inversion reduces to computing fixed residual quantiles rather than repeatedly recalibrating over candidate labels y.

7

Numerical experiments

In the main text, we evaluate the proposed methods in three representative settings: two-sample mean shift testing (Section 7.1), sequential two-sample nonparametric testing (Section 7.2), and conformal prediction (Section 7.3). Additional simulation results, including empirical level assessment, one-sample zero mean testing, and independence testing, are deferred to Appendix C. All experiments were conducted using a single NVIDIA A100 GPU (40GB VRAM). The code for reproducing the experiments is available at https://github.com/antoninschrab/sb-paper.

7.1

Two-sample mean shift testing

We evaluate the empirical power of the proposed tests using data drawn from a high-dimensional Laplace distribution with a targeted mean shift. Specifically, we generate two independent samples of sizes m and n in RD , where the features are initially drawn from a standard Laplace distribution. A signal is then introduced by replacing the values in the first d dimensions of the first sample 22

Power

1

0

1

2

4 8 16 32 64 128 Number of Shifted Dimensions d L1

19

39

L∞

L4 TB

256

SB

SB

59 79 99 199 499 Number of Permutations B L1 SB ties

L2

999

L3 TB

TB ties

Figure 1: Power estimation for two-sample d-dimensional mean shift detection. by ∆/2 and of the second sample by −∆/2, yielding a true mean difference of ∆ across the d shifted dimensions. The empirical power is averaged over 1000 independent repetitions at a nominal significance level of α = 0.05. For the first experiment varying signal sparsity in Figure 1, we fix the sample sizes to m = n = 100, the ambient dimension to D = 20000, and the number of permutations to B = 199. We vary the number of shifted dimensions d ∈ {1, 2, 4, 8, 16, 32, 64, 128, 256}, decaying the shift magnitude ∆ according to ∆ = 1.1/(1 + 0.0318(d − 1)0.826 ) to maintain comparable nontrivial power across regimes. For the second experiment, we fix m = n = 1000, D = 20000, d = 128, and ∆ = 0.1, while varying the number of permutations B from 19 to 999. For the group transformations, we use permutations of the pooled data to construct two permuted samples. We consider the classical permutation tests using the Lp norm of the difference in sample means as test statistics, as well as the SB and TB tests aggregating over the same three Lp norm statistics being considered. We implement the standard versions of SB and TB using tie-breaking completely at random, and their variants preserving ties (i.e., SB/TB ties). The left panel of Figure 1 demonstrates that the individual L1 , L4 , and L∞ tests achieve high power only against specific, distinct alternatives (i.e., Lp detects only dense signal for low p and only sparse signals for large p). In contrast, the combined SB and TB tests are highly adaptive, achieving high empirical power against all considered alternatives. This adaptivity only comes at a small cost in test power compared to the single best-performing test for any given regime. Furthermore, with B = 199 permutations, the SB and TB tests achieve exactly the same power. The right panel illustrates the effect of the permutation count B and tie-breaking strategies. Employing tie-breaking greatly increases the statistical power, especially for small values of B. For these small B values, the SB test outperforms the TB test (confirming the result of Proposition S.7); without tie-breaking, this difference in power is quite pronounced, whereas with tie-breaking, SB still outperforms TB, but the gap in power is much less significant. For large values of B, all adaptive variants (SB and TB, with and without tie-breaking) converge to the same power (confirming the result of Proposition S.6), again achieving an adaptive performance just below the best individual test. As such, we choose B = 199 for all our subsequent experiments since we observe that the power remains the same for larger values of B while the computational cost increases linearly with B. We refer to [15] for a method to significantly reduce the number of permutations required.

23

1

1

10

Power

Power

6 4

Decision Stage

8

2 0

0

0.2 SeqSB

0.4 0.6 Mixing Parameter  SB

WC

0.8

1

0

Oracle

1

1.01 SeqSB Power

1.02 1.03 Scale Parameter σ

1.04

1.05

SeqSB Average Decision Stage

Figure 2: Power estimation for kernel-based MMD two-sample nonparametric testing in a sequential setting.

7.2

Sequential two-sample nonparametric testing

We evaluate the empirical performance of the Sequential SB minimum test (SeqSB, see Algorithm 2) using simulated data evaluated across K = 10 sequential stages with equal spending αj = α/K for j ∈ [K], designed to introduce a controlled correlation structure along with a localized departure from the null. This experimental setting is closely related to the problem of change point detection. The sequential hypothesis testing framework proceeds as follows: we first receive data X, then at stage k = 1, . . . , K, we receive data Yk , and we test the global null that X and Yk are identically distributed for all k = 1, . . . , K. More specifically, we set m = n and first draw n = 1000 samples from N (0, I50 ) for X, and for each stage k = 1, . . . , K, we independently draw n = 1000 samples from N (0, I50 ) for Zk . For some scale parameter σ > 0 (σ = 1 corresponds to the null), we construct Y1 = Z1 , then Y2 = σf (Y1 , Z2 ), then Y3 = f (Y2 /σ, Z3 ), p and Yk = f (Yk−1 , Zk ) for stages k = 4, . . . , K. The transition function is defined as f (Y, Z) = 1 − (1 − ϵ)2 Y + (1 − ϵ)Z, ensuring that for any mixing parameter ϵ ∈ [0, 1], the distribution of Yk remains N (0, I50 ) for all k ̸= 2 and Y2 is i.i.d. N (0, σ 2 I50 ). If ϵ = 0, the samples Y1 , . . . , YK are mutually independent, and as ϵ increases, the correlation between the samples increases, with ϵ = 1 corresponding to perfect correlation across all stages. We set ϵ = 0.9 to introduce a strong correlation structure across the stages in the right panel of Figure 2. At each stage k = 1, . . . , K, we compute the Maximum Mean Discrepancy U-statistic (MMD, [24]) between X and Yk using a Laplace kernel with fixed bandwidth. Using the minimum merging function, we compare the proposed SeqSB test against the non-sequential SB test and to the Worst-Case (WC) baseline which corresponds to Bonferroni correction across the K stages. As a reference point, we also include the power of the Oracle MMD permutation test between X and Y2 , which leverages prior knowledge that the departure from the null occurs exactly (and only) at stage k = 2. The empirical power and the average decision stage (defined as the mean stage at which the sequential procedure terminates) are averaged over 1000 independent repetitions. All tests use B = 199 wild bootstrap MMD U-statistics, using the same Rademacher variables across all K stages to maintain the dependence structure across the test statistics. Note that we assume that the total number of stages K is known and fixed (see [67] for an anytime valid MMD-based test). The empirical results are summarized in Figure 2. The left panel illustrates the effect of the mixing parameter ϵ under a fixed departure scale of σ = 1.035. Notably, the SeqSB and SB tests 24

achieve identical empirical power, though the SeqSB test provides the critical operational advantage of early stopping. When ϵ = 0, the SB procedure coincides with the conservative WC Bonferroni baseline. However, as the dependence ϵ increases, the power of the SB tests rises, eventually achieving the exact same power as the Oracle single test as ϵ tends to 1. This behavior demonstrates that the SB methodology adaptively calibrates to the dependence structure across statistics, maintaining nominal level control without over-correcting for redundant information (Proposition 4). The right panel demonstrates the effect of varying the scale parameter σ from 1 to 1.05 under strong stage correlation (ϵ = 0.9). As the signal strength increases, the empirical power of the SeqSB test monotonically approaches 1. Concurrently, the average decision stage drops dramatically; for sufficiently strong departures from the null hypothesis, the SeqSB test is capable of rejecting the null as soon as signal is injected (e.g., at stage k = 2 out of 10 for σ = 1.05), underscoring its efficiency and responsiveness in sequential monitoring environments.

7.3

Conformal prediction

In Figure 3, we evaluate the proposed methods within a conformal prediction framework, aiming to construct valid prediction sets for a regression target. The objective is to ensure precise marginal coverage (defined as the probability that the true test response falls within the constructed set) while simultaneously minimizing the prediction set size (average interval width). Data is generated from a one-dimensional heteroscedastic regression model with covariates X ∼ Uniform(−2, 2) and responses Y = sin(3X) + ε, with noise ε ∼ N (0, (1 + |X|)). We evaluate K = 20 distinct predictors, each approximating the true function but injected with independent additive noise, as well as with noise shared across all K predictors. The first predictor has noise of much smaller scale. Formally, we define Y k = sin(3X)+sk (0.99εshared +0.01εk ) for independent noise variables εshared , ε1 , . . . , εK ∼ N (0, 1) and scaling parameter s1 = 0.01 and sk = 10 for k = 2, . . . , K. We use the absolute residuals as non-conformity scores and a target marginal coverage of 1 − α = 0.90. For SB we use n = 999, while for TB we split these samples into n1 = 500 for the reference batch and n2 = 499 for the aggregation batch, ensuring that ⌊(n + 1)α⌋/(n + 1) = ⌊(n2 + 1)α⌋/(n2 + 1) = α. We emphasize that our implementation using absolute residual scores is exact in the sense that it does not require a discretization of the response space, avoiding both the computational cost and the approximation errors inherent to grid-based evaluations. Validity metrics are averaged over 1000 independent repetitions, each considering 100 new test points (overall averaged over 100,000 test points). We systematically benchmark the Single-Batch (SB) and Two-Batch (TB) methods against Worst-Case (WC 1 batch and WC 2 batches) baselines using minimum, mean, and median merging functions, as detailed below. Finally, to assess computational scalability, end-to-end execution times are recorded strictly after Just-In-Time (JIT) compilation, ensuring a fair evaluation of runtime complexity as the number of test points varies (e.g., 500, 1000, 2000, 4000, 8000, 16000, 32000) with fixed n = 500 and K = 20. In this setting, we have K conformal prediction sets Cαk (Xn+1 ) := {y ∈ Y : pk (y) > α}, each constructed such that P(Yn+1 ∈ Cαk (Xn+1 )) ≥ 1 − α, for k ∈ [K]. The p-values pk are constructed either using the full data (i.e., single batch, (12)) or using a subsample of the data (i.e., two batches, (15)). The aim is then to combine C·1 (Xn+1 ), . . . , C·K (Xn+1 ) into a single set Cα (Xn+1 ) which still satisfies the marginal coverage guarantee P(Yn+1 ∈ Cα (Xn+1 )) ≥ 1 − α. The SB and TB methods presented in Section 6.2 provide valid constructions for this problem. Alternatively, one can also consider the Worst-Case (WC) baselines which construct CαWC (Xn+1 ) := {y ∈ Y : f (p1 (y), . . . , pK (y)) > α} for specifically designed p-merging functions (Section 2.3 and 25

1

Min

Mean

Target 1−α

Median SB

Runtime (s)

Efficiency

Coverage 0

101

40

20

0 TB

Min

Mean

Median

WC 1 batch

10−1 10−3 103 104 Number of test points

WC 2 batches

TB Min Prop 6

Figure 3: Marginal coverage, efficiency (average prediction set size) and runtimes (in seconds) for merging multiple conformal prediction sets. Appendix B.1) such as the minimum p-merging function f (p1 , . . . , pK ) = K min(p1 , . . . , pK ), the mean p-merging function f (p1 , . . . , pK ) = 2 mean(p1 , . . . , pK ) and the median p-merging function f (p1 , . . . , pK ) = 2 median(p1 , . . . , pK ), with scaling specifically chosen to ensure 1 − α marginal coverage under arbitrary dependence [73]. The empirical results of Figure 3 confirm the theoretical guarantees and highlight the computational superiority of the TB precomputation approach. As observed in the first analysis, both the SB and TB procedures tightly control the marginal coverage near the target level 1 − α for all merging functions. In contrast, the WC baselines are markedly conservative. For this experimental setting in which Y 1 is injected with less noise than the other predictors, the minimum merging methods yield significantly smaller prediction sets than the mean and median methods which are dominated by the noisy predictors. Furthermore, SB and TB achieve much smaller average prediction set sizes compared to WC, as the aggregation methods leverage the dependence structure across the K predictors. Crucially, while SB and TB exhibit identical statistical validity and efficiency, their execution times diverge significantly. The runtime plot demonstrates that the TB procedure generally executes considerably faster than SB, and that with the expression of Proposition 6 designed specially for the minimum merging function, TB is orders of magnitude faster than SB. This performance gap is directly explained by their time complexities: O(Kn log n + M Kn log(Kn)) for the general SB and TB variants3 , and O(Kn log n + M K) for the TB formulation in Proposition 6, where M represents the number of test points. This highlights the advantages of the TB approach for conformal prediction applications, while the SB procedure is superior for hypothesis testing applications. We refer the reader to Appendix B.5 for a detailed implementation of the TB and SB conformal prediction procedures, bypassing the need for discretization of the response space, and of their computational complexities. 3

This assumes that n1 and n2 are of the same order as n, and that the merging function admits an incremental update rule (e.g., the minimum, average, and median functions).

26

8

Discussion

This paper studies row-wise permutation aggregation for statistical evidence under exchangeability. Building on permutation-combination ideas, we characterize the finite-sample power and dependence adaptivity of SB aggregation and extend the framework through sequential spending and two-batch data-dependent aggregation. By operating at the level of transformed data rather than relying solely on worst-case super-uniformity, the proposed methods can exploit the underlying dependence across statistics while retaining finite-sample validity. The applications to adaptive testing and conformal prediction illustrate how this power theory and the TB extension can be used in practice. Several directions remain open. First, permutation p-values provide a canonical standardization, but for small B they coarsen evidence onto a sparse grid. Continuous standardizations, such as null-CDF or studentized standardizations, may improve power, and their gains remain to be characterized. Second, a general optimality theory for aggregation is missing. Aggregation adapts across heterogeneous alternatives but incurs a calibration cost; identifying when this cost is unavoidable, and proving minimax, oracle-adaptive, or lower-bound guarantees, are central questions. Third, the sequential theory could be extended beyond fixed alpha-spending. SeqSB gives finite-sample valid early stopping for ordered statistics, but leaves open how to choose or learn the spending sequence, handle streaming or adaptively ordered statistics, and characterize the tradeoff between early stopping, power, and calibration cost. Finally, it would be useful to extend the theory beyond exact exchangeability. The current guarantees assume that the transformed rows are exactly exchangeable, but in large or constrained transformation spaces one may only have approximate randomization. Characterizing how aggregation validity and power degrade under approximate exchangeability, and how to correct for this degradation, would substantially broaden the scope of the framework.

References [1] Albert, M. (2015). Tests of independence by bootstrap and permutation: an asymptotic and nonasymptotic study. Application to neurosciences. PhD thesis, Université Nice Sophia Antipolis. [2] Albert, M., Laurent, B., Marrel, A., and Meynaoui, A. (2022). Adaptive test of independence based on HSIC measures. The Annals of Statistics, 50(2):858–879. [3] Angelopoulos, A. N., Barber, R. F., and Bates, S. (2024). Theoretical Foundations of Conformal Prediction. arXiv preprint arXiv:2411.11824. [4] Baraud, Y., Huet, S., and Laurent, B. (2003). Adaptive tests of linear hypotheses by model selection. The Annals of Statistics, 31(1):225–251. [5] Berrett, T. B., Kontoyiannis, I., and Samworth, R. J. (2021). Optimal rates for independence testing via U-statistic permutation tests. The Annals of Statistics, 49(5):2457–2490. [6] Berrett, T. B., Wang, Y., Barber, R. F., and Samworth, R. J. (2020). The conditional permutation test for independence while controlling for confounders. Journal of the Royal Statistical Society Series B: Statistical Methodology, 82(1):175–197.

27

[7] Biggs, F., Schrab, A., and Gretton, A. (2023). MMD-FUSE: Learning and combining kernels for two-sample testing without data splitting. Advances in Neural Information Processing Systems, 36. [8] Candes, E., Fan, Y., Janson, L., and Lv, J. (2018). Panning for gold: ‘Model-X’ knockoffs for high dimensional controlled variable selection. Journal of the Royal Statistical Society Series B: Statistical Methodology, 80(3):551–577. [9] Caughey, D., Dafoe, A., and Seawright, J. (2017). Nonparametric combination (NPC): A framework for testing elaborate theories. The Journal of Politics, 79(2):688–701. [10] Cha, S., Lee, S., Schrab, A., and Kim, I. (2026). More Permutations Do Not Always Increase Power: Non-monotonicity in Monte Carlo Permutation Tests. arXiv preprint arXiv:2605.03886. [11] Chau, S. L., Schrab, A., Gretton, A., Sejdinovic, D., and Muandet, K. (2025). Credal twosample tests of epistemic uncertainty. In Proceedings of The 28th International Conference on Artificial Intelligence and Statistics, volume 258 of Proceedings of Machine Learning Research, pages 127–135. PMLR. [12] Choi, W. and Kim, I. (2023). Averaging p-values under exchangeability. Statistics & Probability Letters, 194:109748. [13] Chung, E. and Romano, J. P. (2013). Exact and asymptotically robust permutation tests. The Annals of Statistics, 41(2):484–507. [14] Cox, D. R. (1975). A note on data-splitting for the evaluation of significance levels. Biometrika, 62(2):441–444. [15] Domingo-Enrich, C., Dwivedi, R., and Mackey, L. (2025). Cheap permutation testing. arXiv preprint arXiv:2502.07672. [16] Fisher, R. A. (1925). Statistical Methods for Research Workers. Oliver and Boyd, Edinburgh. [17] Fisher, R. A. (1935). The Design of Experiments. Oliver and Boyd, Edinburgh. [18] Fromont, M. and Laurent, B. (2006). Adaptive goodness-of-fit tests in a density model. The Annals of Statistics, 34(2):680–720. [19] Fromont, M., Laurent, B., and Reynaud-Bouret, P. (2013). The two-sample problem for poisson processes: Adaptive tests with a nonasymptotic wild bootstrap approach. The Annals of Statistics, 41(3):1431–1461. [20] Gasparin, M. and Ramdas, A. (2024). Merging uncertainty sets via majority vote. arXiv preprint arXiv:2401.09379. [21] Gasparin, M., Wang, R., and Ramdas, A. (2025). Combining exchangeable p-values. Proceedings of the National Academy of Sciences, 122(11):e2410849122. [22] Good, P. (2005). Permutation, parametric and bootstrap tests of hypotheses. Springer. [23] Gretton, A. (2015). A simpler condition for consistency of a kernel independence test. arXiv preprint arXiv:1501.06103. 28

[24] Gretton, A., Borgwardt, K. M., Rasch, M. J., Schölkopf, B., and Smola, A. (2012). A kernel two-sample test. Journal of Machine Learning Research, 13(25):723–773. [25] Gretton, A., Herbrich, R., Smola, A., Bousquet, O., and Schölkopf, B. (2005). Kernel methods for measuring independence. Journal of Machine Learning Research, 6:2075–2129. [26] Guo, F. R. and Shah, R. D. (2025). Rank-transformed subsampling: inference for multiple data splitting and exchangeable p-values. Journal of the Royal Statistical Society Series B: Statistical Methodology, 87(1):256–286. [27] Hagrass, O., Sriperumbudur, B., and Li, B. (2024). Spectral regularized kernel two-sample tests. The Annals of Statistics, 52(3):1076–1101. [28] Harrison, M. T. (2012). Conservative hypothesis tests and confidence intervals using importance sampling. Biometrika, 99(1):57–69. [29] Hegazy, M., Aolaritei, L., Jordan, M. I., and Dieuleveut, A. (2025). Valid selection among conformal sets. In Advances in Neural Information Processing Systems, volume 38. [30] Hemerik, J. and Goeman, J. (2018). Exact testing with random permutations. Test, 27(4):811– 825. [31] Ingster, Y. and Suslina, I. A. (2012). Nonparametric goodness-of-fit testing under Gaussian models, volume 169. Springer Science & Business Media. [32] Janková, J., Shah, R. D., Bühlmann, P., and Samworth, R. J. (2020). Goodness-of-fit testing in high dimensional generalized linear models. Journal of the Royal Statistical Society Series B: Statistical Methodology, 82(3):773–795. [33] Kent, A., Berrett, T. B., and Yu, Y. (2026). Locally Differentially Private Two-Sample Testing. Biometrika, page asag034. [34] Kim, I., Balakrishnan, S., and Wasserman, L. (2022). Minimax optimality of permutation tests. The Annals of Statistics, 50(1):225–251. [35] Kim, I., Neykov, M., Balakrishnan, S., and Wasserman, L. (2024). Conditional independence testing for discrete distributions: Beyond χ2 - and G-tests. Electronic Journal of Statistics, 18(2):4767–4794. [36] Kim, I. and Ramdas, A. (2024). Bernoulli, 30(1):683–711.

Dimension-agnostic inference using cross U-statistics.

[37] Kim, I., Ramdas, A., Singh, A., and Wasserman, L. (2021). Classification accuracy as a proxy for two-sample testing. The Annals of Statistics, 49(1):411–434. [38] Kim, I. and Schrab, A. (2026). Differentially Private Permutation Tests. Journal of the American Statistical Association, pages 1–13. [39] Lehmann, E. and Romano, J. P. (2022). Testing Statistical Hypotheses. Springer Texts in Statistics. Springer, 4th edition.

29

[40] Lei, J., G’Sell, M., Rinaldo, A., Tibshirani, R. J., and Wasserman, L. (2018). Distribution-free predictive inference for regression. Journal of the American Statistical Association, 113(523):1094– 1111. [41] Liu, F., Xu, W., Lu, J., Zhang, G., Gretton, A., and Sutherland, D. J. (2020). Learning deep kernels for non-parametric two-sample tests. In International Conference on Machine Learning, pages 6316–6326. [42] Lundborg, A. R., Kim, I., Shah, R. D., and Samworth, R. J. (2024). The projected covariance measure for assumption-lean variable significance testing. The Annals of Statistics, 52(6):2851– 2878. [43] Meinshausen, N., Maathuis, M. H., and Bühlmann, P. (2011). Asymptotic optimality of the westfall–young permutation procedure for multiple testing under dependence. The Annals of Statistics, 39(6):3369–3391. [44] Meng, X.-L. (1994). Posterior predictive p-values. The Annals of Statistics, 22(3):1142–1160. [45] Moran, P. A. (1973). Dividing a sample into two parts a statistical dilemma. Sankhyā: The Indian Journal of Statistics, Series A, pages 329–333. [46] Mun, J., Kwak, S., and Kim, I. (2025). Minimax optimal two-sample testing under local differential privacy. Journal of Machine Learning Research, 26(252):1–79. [47] Paik, S., Celentano, M., Green, A., and Tibshirani, R. J. (2025). Integral Probability Metrics Meet Neural Networks: The Radon-Kolmogorov-Smirnov Test. Journal of Machine Learning Research, 26(86):1–57. [48] Papadopoulos, H. (2008). Inductive conformal prediction: Theory and application to neural networks. INTECH Open Access Publisher Rijeka. [49] Pesarin, F. and Salmaso, L. (2010). Permutation Tests for Complex Data: Theory, Applications and Software. Wiley Series in Probability and Statistics. John Wiley & Sons. [50] Pitman, E. J. (1937). Significance tests which may be applied to samples from any populations. Supplement to the Journal of the Royal Statistical Society, 4(1):119–130. [51] Pogodin, R., Schrab, A., Li, Y., Sutherland, D. J., and Gretton, A. (2024). Practical Kernel Tests of Conditional Independence. arXiv preprint arXiv:2402.13196. [52] Ramdas, A., Barber, R. F., Candès, E. J., and Tibshirani, R. J. (2023). Permutation tests using arbitrary permutation distributions. Sankhya A, 85(2):1156–1177. [53] Ramdas, A. and Wang, R. (2025). Hypothesis testing with E-values. Foundations and Trends® in Statistics, 1(1-2):1–390. [54] Ribero, M., Schrab, A., and Gretton, A. (2026). Regularized f -divergence kernel tests. In The 29th International Conference on Artificial Intelligence and Statistics. [55] Romano, J. P. and Wolf, M. (2005). Exact and approximate stepdown methods for multiple hypothesis testing. Journal of the American Statistical Association, 100(469):94–108. 30

[56] Rüschendorf, L. (1982). Random Variables with Maximum Sums. Advances in Applied Probability, 14(3):623–632. [57] Rüger, B. (1978). Das maximale Signifikanzniveau des Tests: „Lehne H0 ab, wenn k unter n gegebenen Tests zur Ablehnung führen.". Metrika, 25:171–178. [58] Schrab, A. (2025a). A practical introduction to kernel discrepancies: MMD, HSIC & KSD. arXiv preprint arXiv:2503.04820. [59] Schrab, A. (2025b). Optimal Kernel Hypothesis Testing. PhD thesis, UCL (University College London). [60] Schrab, A. (2025c). A unified view of optimal kernel hypothesis testing. arXiv preprint arXiv:2503.07084. [61] Schrab, A., Guedj, B., and Gretton, A. (2022a). KSD Aggregated Goodness-of-fit Test. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022. [62] Schrab, A., Kim, I., Albert, M., Laurent, B., Guedj, B., and Gretton, A. (2023). MMD aggregated two-sample test. Journal of Machine Learning Research, 24(194):1–81. [63] Schrab, A., Kim, I., Guedj, B., and Gretton, A. (2022b). Efficient aggregated kernel tests using incomplete U -statistics. Advances in Neural Information Processing Systems, 35:18793–18807. [64] Shah, R. D. and Bühlmann, P. (2018). Goodness-of-fit tests for high dimensional linear models. Journal of the Royal Statistical Society Series B: Statistical Methodology, 80(1):113–135. [65] Shekhar, S., Kim, I., and Ramdas, A. (2022). A permutation-free kernel two-sample test. Advances in Neural Information Processing Systems, 35:18168–18180. [66] Shekhar, S., Kim, I., and Ramdas, A. (2023). A permutation-free kernel independence test. Journal of Machine Learning Research, 24(369):1–68. [67] Shekhar, S. and Ramdas, A. (2024). Nonparametric two-sample testing by betting. IEEE Transactions on Information Theory, 70(2):1178–1203. [68] Solmi, F. and Onghena, P. (2014). Combining p-values in replicated single-case experiments with multivariate outcome. Neuropsychological Rehabilitation, 24(3-4):607–633. [69] Stouffer, S. A., Suchman, E. A., DeVinney, L. C., Star, S. A., and Williams Jr, R. M. (1949). The American Soldier: Adjustment during Army Life (Vol. 1). Princeton University Press, Princeton, NJ. [70] Tansey, W., Veitch, V., Zhang, H., Rabadan, R., and Blei, D. M. (2022). The holdout randomization test for feature selection in black box models. Journal of Computational and Graphical Statistics, 31(1):151–162. [71] Vovk, V., Gammerman, A., and Shafer, G. (2005). Algorithmic Learning in a Random World. Springer.

31

[72] Vovk, V., Wang, B., and Wang, R. (2022). Admissible ways of merging p-values under arbitrary dependence. The Annals of Statistics, 50(1):351–375. [73] Vovk, V. and Wang, R. (2020). Combining p-values via averaging. Biometrika, 107(4):791–808. [74] Vovk, V. and Wang, R. (2021). E-values: Calibration, combination, and applications. The Annals of Statistics, 49(3):1736–1754. [75] Westfall, P. H. and Young, S. S. (1993). Resampling-based multiple testing: Examples and methods for p-value adjustment. John Wiley & Sons. [76] Yang, Y. and Kuchibhotla, A. K. (2025). Selection and Aggregation of Conformal Prediction Sets. Journal of the American Statistical Association, 120(549):435–447. [77] Zhou, Z., Tian, X., Peng, L., Lei, C., Schrab, A., Sutherland, D. J., and Liu, F. (2025). Dual: Learning diverse kernels for aggregated two-sample and independence testing. In Advances in Neural Information Processing Systems, volume 38.

32

1 1(Tik ≥ Tbk) B+1 ∑ i=0

pkb =

T kb = T k(gbX)

Data X

T 10

T 20

T 30

···

T K0

p10

p20

p30

···

pK0

g1X

T 11

T 21

T 31

···

T K1

p11

p21

p31

···

pK1

T 12

T 22

3 2

1 2

2 2

3 2

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

p p p T · · · p for ··· T Supplementary material .. .. .. .. .. .. .. .. under .. Aggregation of Statistical Evidence Exchangeability . . . . . . . . . X

g2X

K 2

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

K 2

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

gBX

T 1B

f 10

f 30

X

T 101 f1

T 202 f1

T 303 f1

g1X

T 111 f2

T 212 f2

T 313 f2

g2X

T .12

T .22

T.32

.. . <latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

.. .. . <latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

f 1B

gBX

T 1B

.. .. . <latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

f 2B

T 2B

.. .. . <latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

P−value Matrix

pkb =

···

fM 0

p̄10

··· ··· <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

K TM f 10

··· ··· <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

··· <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

p̄30

p101 p̄1

p202 p̄1

T K1M f2

p111 p̄2

p212 p̄2

p303 · · · p̄1 · · ·

T.K2

p12

p22

.. .. . <latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

f 3B

··· <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

··· <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

.. . .. .1 <latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

X

X X

T 10

T kb = T k(gbX)

T 20

T 30

··· <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

T T 21

T T 31

···

gX g21X

T T 12

T T 22

T T 32

···

T.

T.

T.

gBX

T 1B

T 2B

T 3B

gBX

TB

TB

TB

f 10

f 20

f 30

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

p̄B

p̄B

T KB

p1B

p2B

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

p313 · · · p̄2 · · · <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

p32

.. · · · . .. . p̄3B · · · <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

··· <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

pK1M p̄ 2

f1 f¯2

pK2

f2.

.. . .. .M <latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

1 k k pkb =P−va ue1Ma (Ti ≥ xTb ) B+1 ∑ i=0 B 1 p = 1+ 1(T ≥ T )) B+1 ( ∑

.. <latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

.. <latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

p20

p30

TK T K1

1 pp110

2 pp210

3 pp310

TK T K2

p1 p121

p2 p221

K

1 .p2

2 .p2

p3 p321 ······

T KB

p1B

p2B

p3B

T KB

p1B

p2B

p3B

fM 0

p̄10

p̄20

T fM 2

p̄11

T.

..

..

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

=

p10

..

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

.. <latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

··· <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

··· ··· <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

3 .p2

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q

..

pK0

1 B 1 f¯ ≤ f¯ P−value ∑ ( bpSB 0) B+1 B b=0 1 1(f ≤ f0) B+1 ∑ b b=0

Reject H0 if pSB,F ≤ α Reject H0 if pSB ≤ α

.. .. . <latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

f¯B

p̄ B

pKB

P−value pSB,F

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

1 B 1(f im ≤ f bm) B+1 ∑ i=0

T K0

f¯0

f0 f¯1

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

p3B

fb = f(p1b, …, pKb)

pK0M p̄ 1

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

P−value Matrix

p̄m = b

T kb = T k(gbX) Tes ng Ba ch

T T 11

.. .. .

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

···

p̄ M 0

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

P−value Matrix Statistic Matrix(b) TB aggregation B

g1X X

g.2X

.. . .. .2

fM B

Aggregated Statistic Matrix

Merging p−values

1 B 1(Tik ≥ Tbk) B+1 ∑ i=0

p̄20

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

f bm = f m(p1b, …, pKb)

Transformed Data T ans o med Da a

pKB

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

T 3B

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

T kb = T k(gbX) f 20

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

p1B p2B p3B · · · T KB SB aggregation · · · (a)

T 3B

Statistic Matrix

Transformed Data

X

T 2B

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

fB Merging p−values f¯b = min(p̄1, …, p̄ M) b

b

Me g ng p−va ues = (p … pK)

K ppK10

P−va ue pTB

1 B 1( ≤ B+1 ∑ =

pK pK21 K .p2

..

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

Re ec H pTB ≤ α

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

gX gX

T f 11

T ans o med Da a

T f 31

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

··· <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

··· <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

T f 12

T f 22

T f 32

.. .

.. .

.. .

TB f 1B

TB f 2B

TB f 3B · · ·

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

gBX

T f 21

···

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

··· <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

pKB pKB

B

p̄30

···

p̄ M 0

f¯0

p̄21

p̄31

···

p̄ M 1

f¯1

p̄12

p̄22

p̄32

···

p̄ M 2

f¯2

.. .

.. .

.. .

.. .

.. .

.. .

T KB fM B

p̄1B

p̄2B

p̄3B

p̄ M B

f¯B

K

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

··· ···

TK fM 1

Re e ence Ba ch Aggregated Statistic Matrix T = T (g X) f bm = f m(p1b, …, p M b)

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

)

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

··· <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

P−value Matrix

1 B p̄m = 1(f im ≥ f bm) b B+1 ∑ i=0

P−value pSB,F

1 B 1(f¯b ≤ f¯0) B+1 ∑ b=0

Reject H0 if pSB,F ≤ α

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

Merging p−values f¯b = min(p̄1b, …, p̄ M b)

Figure 4: Schematic illustrations of (a) the SB aggregation procedure of Algorithm 1 and (b) the K and merging function f . P−value Matrix TB aggregation procedure ofTesting Algorithm 3, using statistics T 1 , . . . , TMerging p−values Batch Transformed Data

X

pkb =

T kb = T k(gbX)

p20

p30

···

pK0

f0

T K1

p11

p21

p31

···

pK1

f1

T K2

p12

p22

p32

···

pK2

f2

T 10

T 20

T 30

···

T K0

g1X

T 11

T 21

T 31

···

g2X

T 12

T 22

T 32

···

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

33

fb = f(p1b, …, pKb)

p10

X

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

B 1 1+ 1(T̃ ki > Tbk)) ( ∑ B+1 i=1

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

P−value pTB

1 B 1(fb ≤ f0) B+1 ∑ b=0

Reject H if

The supplementary material contains schematic illustrations (Appendix A), background and methodological details (Appendix B), additional simulation results (Appendix C), extensions of the aggregation framework (Appendix D), power refinements (Appendix E), adaptive testing details (Appendix F), technical lemmas and auxiliary tools (Appendix G), and proofs (Appendices H and I).

A

Schematic illustrations

This section presents schematic illustrations of the SB and TB aggregation procedures, which can be found in Figure 4.

B

Background and methodological details

This section collects background material, methodological reformulations, and auxiliary examples used to situate the proposed procedures.

B.1

Classical p-value merging families

This subsection recalls the two classical p-merging families referenced in Section 2.3. Let p1 , . . . , pK denote K super-uniform p-values. O-family. Rüger’s O-family [57] is based on scaled order statistics and is defined as   K f (p1 , . . . , pK ) = min p ,1 , k (k) where p(k) denotes the k-th smallest p-value among (p1 , . . . , pK ). This family includes several classical procedures as special cases, such as the Bonferroni method (k = 1), the median rule (k = ⌈K/2⌉), and the maximum p-value rule (k = K). All members of the O-family are precise p-merging functions, in the sense that their super-uniformity property is tight. M-family. The M-family [73] is parameterized by r ∈ [−∞, ∞] and is defined as f (p1 , . . . , pK ) = min{ar,K Mr,K (p1 , . . . , pK ), 1} , where Mr,K (p1 , . . . , pK ) denotes the generalized mean K

Mr,K (p1 , . . . , pK ) =

1 X r pk K

!1/r .

k=1

This family encompasses a wide range of commonly used aggregation rules, including the minimum (r = −∞), maximum (r = ∞), harmonic (r = −1), geometric (r = 0), and arithmetic (r = 1) means. The scaling constant ar,K is chosen to ensure that f is a precise p-merging function; see, for example, [73, Table 1]. A notable special case is the arithmetic mean (r = 1), for which the optimal scaling constant is a1,K = 2, a result that dates back to classical work [44, 56]. More generally, the asymptotically sharp scaling constant for the M-family is (r + 1)1/r for r > 0, e for r = 0, and r 1+1/r for r < −1; see, for example, [73, Table 1]. r+1 K 34

Recent work by [72] further investigates the structure of admissible p-merging functions, namely those that cannot be uniformly improved while preserving universal validity under arbitrary dependence. Their results provide a complete characterization of admissibility within the O- and M-families. These procedures require deterministic worst-case calibration under arbitrary dependence. Under the group-invariance hypothesis, our main result (Corollary 1) shows that SB aggregation uniformly dominates deterministically calibrated p-merging rules, including the O- and M-families.

B.2

MaxT p-value formulation and comparison with SB and TB

This subsection provides additional details on the MaxT aggregation procedure reviewed in Section 2.4, including a closed-form calibration threshold, a p-value formulation, and a comparison with the SB and TB procedures. Closed-form MaxT threshold. The estimated critical value ũα in (3) can be written in closed form. This avoids the bisection procedure employed in prior work [2, 62] and reveals the p-value structure of the MaxT calibration. Proposition S.1. For each b ∈ [B], define B

1 X k ≤ Tjk ). 1(TB+b k∈[K] B + 1

ub := min

j=0

Then, for α ∈ (0, 1), the estimated critical value ũα defined in (3) admits the representation ũα = u(⌊Bα⌋+1) , where u(1) ≤ u(2) ≤ · · · ≤ u(B) are the order statistics of {u1 , u2 , . . . , uB }. Proof. See Appendix I.9. The quantity ub is the minimum, over coordinates k, of the permutation p-value that the calk would receive when ranked against the testing batch T0k , . . . , TBk . Thus, ũα ibration statistic TB+b is the empirical α-quantile of these calibration-batch minimum p-values. This gives the following equivalent p-value view of the MaxT decision rule. P-value formulation of MaxT. Let B

p(T0k ) :=

1 X 1(Tbk ≥ T0k ) B+1 b=0

denote the permutation p-value associated with T0k . Using the equivalent formulations explained in Section 2.2, the decision rule in (2) can be written as min p(T0k ) ≤ ũα .

k∈[K]

Hence the MaxT procedure can be interpreted as a minimum p-value test with a Monte Carlocalibrated correction factor. This viewpoint makes the comparison with SB and TB minimum 35

aggregation transparent: all three procedures aggregate through the minimum p-value, but they differ in how the calibration set is constructed. Comparison with SB minimum aggregation. By Proposition 1 and Lemma S.4, the SB minimum threshold admits the following equivalent representation. Write  k k QSB k (u) := Quantile1−u T0 , . . . , TB . Then  ûSB = sup u ∈ (0, 1) : α,min

   B n o 1 X 1 max Tbk − QSB (u) > 0 ≤ α . k B+1 k∈[K] b=0

This representation makes explicit the key distinction between SB minimum aggregation and the existing MaxT aggregation reviewed in Section 2.4: the SB threshold ûSB α,min relies on a single k k batch of transformed statistics T0 , . . . , TB both to approximate the conditional rejection probability and to compute the relevant quantiles. In contrast, MaxT procedures employ an additional k solely for calibration. This modification resolves the type I error control issue k , . . . , T2B batch TB+1 identified in Section 2.4, while remaining data-dependent and adaptive to the dependence structure among the p-values. Comparison with TB minimum aggregation. For the minimum merge f (p1 , . . . , pK ) = mink∈[K] pk , define  k k k QTB b,k (u) := Quantile1−u Tb , T̃1 , . . . , T̃B . The TB critical value can be written explicitly as  u ∈ (0, 1) : = sup ûTB α

   B n o 1 X 1 max Tbk − QTB (u) > 0 ≤ α . b,k B+1 k∈[K] b=0

The difference from the MaxT critical value in (3) lies in how the calibration is constructed. In the TB procedure, each testing statistic Tbk is ranked against a reference batch that explicitly includes Tbk itself, and the randomization average in the definition of ûTB α is taken over b = 0, . . . , B. In contrast, the MaxT calibration does not include Tbk in the ranking set and averages only over the calibration indices b = 1, . . . , B. Although seemingly minor, these differences are structural: by including Tbk in the ranking set and averaging over b = 0, . . . , B, the TB construction is always well-defined, restores exact conditional exchangeability and achieves finite-sample type I error control.

B.3

General sequential aggregation principle

This subsection records the general sequential aggregation principle underlying the SeqSB construction in Section 4. Let I1 , . . . , IJ ⊆ [K] be prespecified index sets, not necessarily nested, and for each j ∈ [J] let hj : [0, 1]|Ij | → R be a measurable function, where smaller values indicate stronger evidence against the null. For b ∈ [B]0 and j ∈ [J], define the stage-j score  zb,j := hj (p(Tbk ))k∈Ij . (17)

36

P Let α1 , . . . , αJ ∈ [0, 1] be stage-wise significance budgets with Jj=1 αj ≤ α. Set S0 := [B]0 , and define recursively, for j ∈ [J],   X 1 cj := sup u ∈ R : 1(zb,j ≤ u) ≤ αj , B+1 (18) b∈Sj−1

Aj := { b ∈ Sj−1 : zb,j < cj },

Sj := Sj−1 \ Aj .

The sequential aggregation test rejects whenever 0 ∈

SJ

j=1 Aj .

Proposition S.2. Fix α ∈ (0, 1), B ≥ 1, and stage-wise budgets α1 , . . . , αJ ∈ [0, 1] satisfying PJ α ≤ α. Let (Aj )Jj=1 denote the sequence of eliminated sets produced by the recursive procej j=1 dure (18). Under the group-invariance hypothesis,   J [ P 0∈ Aj ≤ j=1

J

1 X ⌊(B + 1)αj ⌋ ≤ α. B+1 j=1

Moreover, when {zb,j : b ∈ Sj−1 } are almost surely distinct for every j ∈ [J], the first inequality is tight. Proof. See Appendix I.10.

B.4

PCA-based aggregation using the reference batch

This subsection gives a concrete example of a data-dependent aggregation rule that is enabled by the TB split. The idea is to learn a low-dimensional representation of the coordinate-wise p-value vector using only the reference batch, and then use this learned representation when aggregating the testing batch. For each testing index b ∈ [B]0 , let  Pb := pHO (Tb1 ), . . . , pHO (TbK ) , and let P̃1 , . . . , P̃B denote the corresponding reference-batch vectors, defined by  P̃b :=

 B B   1 X 1 X 1 1 K K 1 T̃j ≥ T̃b , . . . , 1 T̃j ≥ T̃b , B B j=1

j=1

b ∈ [B].

Using only the reference batch, let Σ̂ denote the empirical covariance matrix of {P̃i }B i=1 , and let K v̂ ∈ R be a unit leading eigenvector of Σ̂, selected according to a fixed deterministic tie-breaking and sign convention. Define the aggregation map fPC (p1 , . . . , pK ) := ⟨v̂, (p1 , . . . , pK )⟩. The merged testing-batch statistics are then given by f˜b = fPC (Pb ) for b ∈ [B]0 . Intuitively, this construction retains only the dominant direction of variation of the p-value vector under the null, thereby removing redundancy across coordinates. This is useful when the number of coordinates K is large and the coordinate-wise p-values exhibit redundancy or an approximately low-dimensional dependence structure. In such regimes, fixed aggregation rules such as uniform 37

averaging can overweight redundant directions. Because the projection vector v̂ is measurable with respect to the reference batch, it is fixed after conditioning on the reference information, and the testing-batch rows remain exchangeable. As a result, the corresponding TB permutation p-value remains finite-sample valid. Unlike standard sample-splitting schemes that partition the data sample X, the split here is performed only over the transformations, so all observations are used for both learning the aggregation map and performing the test.

B.5

SB and TB conformal prediction algorithms

Recall from Section 6.2 that the SB conformal prediction set is defined as CSB (Xn+1 ) := {y ∈ Y : 1 Pn+1 pSB (y) > α} for the SB p-value pSB (y) := n+1 1{M i (y) ≤ Mn+1 (y)} , with merged values i=1  1 Pn+1 1 k K k Mi (y) := f Pi (y), . . . , Pi (y) of rank-transformed values Pik (y) := n+1 j=1 1{Sj (y) ≥ Si (y)} k (y) := s (X with residual scores Sik := sk (Xi , Yi ) for i ∈ [n] and Sn+1 n+1 , y). When using the k k k absolute residual scores Si := |Yi − µ bk (Xi )|, i ∈ [n] and Sn+1 (y) := |y − µ bk (Xn+1 )|, we note that the indicator 1(Sik ≥ |y − µ bk (Xn+1 )|) changes its truth value only when the test-point score equals a calibration residual, i.e., |y − µ bk (Xn+1 )| = Sik . Between any two such points, none of the indicators change, meaning all p-values, merged scores Mi , and the final aggregated p-value are completely constant. As such, pSB is a step function of y with breakpoints at {b µk (Xn+1 )±Sik : k ∈ [K], i ∈ [n]}, and the exact continuous prediction set can be perfectly recovered by evaluating the functions at exactly one representative point in each open cell and at the boundary breakpoints themselves. This allows for exact efficient computation of the SB conformal prediction set, entirely bypassing the need for discretization of Y = R, as detailed below. First, sort the set {b µk (Xn+1 ) ± Sik : k ∈ [K], i ∈ [n]}, removing duplicates, to obtain ordered breakpoints ỹ1 < . . . < ỹL . We need to evaluate pSB at these breakpoints to determine whether each of the singleton sets {ỹℓ }, ℓ ∈ [L] should be included in the SB conformal prediction set CSB (Xn+1 ). We also need to evaluate pSB at reference points within the intervals (ỹℓ , ỹℓ+1 ), ℓ ∈ [L]0 , with ỹ0 := −∞ and ỹL+1 := ∞ to determine whether each of these should be included in CSB (Xn+1 ). Since the conformal p-value function is piecewise constant, evaluating the merged scores at exactly one representative point perfectly determines the acceptance or rejection of the entirety of the interval. Furthermore, stepping onto a breakpoint, and subsequently stepping off of it into the next open cell, triggers incremental updates for the exact same subset of coordinate-index pairs. To this end, we construct the evaluation points as   ỹ1 − 1 if ℓ = 0,    ỹ if ℓ = 2j − 1 for j ∈ {1, . . . , L}, j yℓ = (19)  (ỹj + ỹj+1 )/2 if ℓ = 2j for j ∈ {1, . . . , L − 1},    ỹ + 1 if ℓ = 2L, L with corresponding intervals   (−∞, ỹ1 )   {ỹ } j Iℓ =  (ỹ j , ỹj+1 )    (ỹ , ∞) L

if ℓ = 0, if ℓ = 2j − 1 for j ∈ {1, . . . , L}, if ℓ = 2j for j ∈ {1, . . . , L − 1}, if ℓ = 2L,

38

(20)

and corresponding k-indices   [K]    K(ỹ ) j Kℓ =  K(ỹj )    K(ỹ ) L

if ℓ = 0, if ℓ = 2j − 1 for j ∈ {1, . . . , L}, if ℓ = 2j for j ∈ {1, . . . , L − 1}, if ℓ = 2L,

(21)

where K(y) := {k ∈ [K] : ∃i ∈ [n] s.t. y = µ bk (Xn+1 ) ± Sik }, and letting J k (y) := {i ∈ [n] : y = k µ bk (Xn+1 ) ± Si }, for k ∈ Kℓ we define   [n] if ℓ = 0,    J k (ỹ ) if ℓ = 2j − 1 for j ∈ {1, . . . , L}, j Jℓk = (22) k  J (ỹj ) if ℓ = 2j for j ∈ {1, . . . , L − 1},    J k (ỹ ) if ℓ = 2L. L Here, yℓ is the representative evaluation point, Iℓ is its associated interval or singleton, and Kℓ , Jℓk are the subsets of coordinates and indices that require incremental updates at step ℓ. For each evaluation point yℓ for ℓ ∈ [2L]0 , we need to evaluate pSB (yℓ ), which requires computing the p-values Pik and the merged values Mi for all i ∈ [n + 1] and k ∈ [K]. To evaluate this sequence efficiently without incurring the full O(nK) cost at every step, we update the p-values and merged scores on the fly. At the leftmost evaluation point y0 , we initialize the p-values Pik (y0 ) and merged scores Mi (y0 ) from scratch. As we sweep left to right for subsequent steps ℓ > 0, only the specific coordinates Kℓ and indices Jℓk for k ∈ Kℓ triggered by the breakpoint undergo a change, computable in at most O(log n) time. If the test score crosses a calibration residual, the respective indicator sum either increases or decreases by 1. We then re-evaluate only the affected merged scores Mi and dynamically adjust their position in the sorted dynamic data structure M (e.g., an order-statistic tree, balanced binary search tree, or Fenwick tree). Merging functions such as the minimum and median can be updated incrementally in time O(log K) (time O(1) for the mean), without requiring the full O(K) cost of applying the merging function. However, for a generic merging function, this cost is inevitable. The exact SB conformal prediction procedure with absolute residuals is presented in Algorithm 4, along with a detailed analysis of the computational complexity for each step. The TB conformal prediction procedure can be exactly and efficiently evaluated in a similar manner, with a much simpler implementation of the update for each evaluation point, as presented in Algorithm 5. Finally, we detail in Algorithm 6 the exact implementation of the expression in Proposition 6 for the minimum merging function, which drastically reduces the computational complexity of the algorithm. While the methods are presented in Section 6.2 for the case of a single test point Xn+1 , in practice we m }M consider a batch of M test points {Xn+1 m=1 simultaneously, as conformal prediction is typically deployed to generate valid prediction sets for an entire batch of unlabelled observations, with M often taking very large values in real-world applications. This practical reality of evaluating large test batches highlights a critical computational dichotomy between the approaches. For the SB procedure (Algorithm 4), the necessity of dynamically updating the calibration merged scores Mi and Mn+1 across all evaluation points imposes a total computational runtime that scales as O(M Kn log(Kn)) for minimum, mean and median merging functions. The generic TB inversion (Algorithm 5) relaxes this burden by relying on a static reference threshold, which avoids the need to update Mi but still yields a comparable test-time cost of 39

O(M Kn1 log(Kn1 )) assuming n1 and n2 are of the same order. However, the true computational advantage emerges under the minimum merging rule with the implementation of Proposition 6 (Algorithm 6). By entirely bypassing the grid-free sweep in favor of directly intersecting fixed residual quantiles, the per-test-point complexity collapses from O(Kn1 log(Kn1 )) to merely O(K). In modern deployment regimes where M ≫ n, this reduction to a strictly O(Kn1 log(Kn1 ) + M K) bottleneck isolates the test-time cost from the calibration size, transforming an otherwise prohibitive search into a highly scalable procedure. Algorithm 4 Exact SB Conformal Prediction for General Merging Function m }M ; prediction functions {b Require: Calibration data {(Xi , Yi )}ni=1 ; test points {Xn+1 µk } K m=1 k=1 ; merging function f ; level α. 1: Compute Sik ← |Yi − µ bk (Xi )| for k ∈ [K] and i ∈ [n]. ▷ O(Kn) 2: Sort {Sik : i ∈ [n]} for each k ∈ [K]. ▷ O(Kn log n) 3: for m = 1, . . . , M do m ) ← ∅. 4: Initialize CSB (Xn+1 m )±S k : k ∈ [K], i ∈ [n]}, remove duplicates, get ỹ < · · · < ỹ . ▷ O(Kn log(Kn)) 5: Sort {b µk (Xn+1 1 L i as in (19), (20), (21) and (22). ▷ O(Kn) 6: Construct the evaluation tuples (yℓ , Iℓ , Kℓ , Jℓk )2L ℓ=0 k m )| for k ∈ [K]. 7: Initialize Sn+1 ← |y0 − µ bk (Xn+1 ▷ O(K) k k k 8: Initialize pi ← Pi (y0 ) for k ∈ [K] and i ∈ [n + 1] using sorted {Sj : j ∈ [n]}. ▷ O(Kn log n)  9: Initialize Mi ← f p1i , . . . , pK fori ∈ [n + 1]. ▷ O(Kn) i 10: Initialize M ← sort M1 , . . . , M ▷ O(n log n) Pn using a dynamicdata structure. 1 11: Compute pSB (y0 ) ← n+1 ▷ O(log n) 1 + ni=1 1 Mi ≤ Mn+1 via binary search on M. m m 12: if pSB (y0 ) > α then CSB (Xn+1 ) ← CSB (Xn+1 ) ∪ I0 . 13: for ℓ = 1, . . . , 2L do 14: for k ∈ Kℓ do k m )|. 15: Update Sn+1 ← |yℓ − µ bk (Xn+1 k k 16: Update pn+1 ← Pn+1 (yℓ ) using sorted {Sjk : j ∈ [n]}. ▷ O(log n)  1 K 17: Update Mn+1 ← f pn+1 , . . . , pn+1 . ▷ O(log K) for min/mean/med, O(K) for generic 18: for i ∈ Jℓk do k 19: Update pki ← Pik (yℓ ) by updating only 1(Sn+1 ≥ Sik ). ▷ O(1) 20: Remove Mi from sorted M. ▷ O(log n) 21: Update Mi ← f p1i , . . . , pK ▷ O(log K) for min/mean/med, O(K) for generic i . 22: Insert Mi into M preserving the sorted order. ▷ O(log n) 23: end for 24: end for  P 1 25: Compute pSB (yℓ ) ← n+1 1 + ni=1 1 Mi ≤ Mn+1 via binary search on M. ▷ O(log n) m ) ← C (X m ) ∪ I . 26: if pSB (yℓ ) > α then CSB (Xn+1 SB ℓ n+1 27: end for 28: end for m ) : m ∈ [M ]}. 29: Return: {CSB (Xn+1 30: Complexity: O(Kn log n + M Kn log(Kn)) for min, mean, and median f , 31: O(Kn log n + M Kn(log(Kn) + K)) for generic f .

40

Algorithm 5 Exact TB Conformal Prediction for General Merging Function Require: Calibration data {(Xi , Yi )}ni=1 ; disjoint index batches Iref ∪ Iagg = [n] with |Iref | = n1 m }M ; prediction functions {b and |Iagg | = n2 ; test points {Xn+1 µk }K m=1 k=1 ; merging function f ; level α. 1: Compute Sjk ← |Yj − µ bk (Xj )| for k ∈ [K] and j ∈ Iref . ▷ O(Kn1 ) 2: Sort {Sjk : j ∈ Iref } for each k ∈ [K]. ▷ O(Kn1 log n1 ) k,ref

← Pik,ref for k ∈ [K] and i ∈ Iagg using sorted {Sjk : j ∈ Iref }. ▷ O(Kn2 log n1 )  1,ref 4: Compute Mi ← f pi , . . . , pK,ref for i∈ Iagg . ▷ O(Kn2 ) i 5: Compute uTB ← −Quantile −M : i ∈ I . ▷ O(n i agg 2 log n2 ) (1−α)(1+1/n2 ) α 6: for m = 1, . . . , M do m ) ← ∅. 7: Initialize CTB (Xn+1 m 8: Sort {b µk (Xn+1 )±Sjk : k ∈ [K], j ∈ Iref }, remove duplicates, get ỹ1 < . . . < ỹL ▷ O(Kn1 log(Kn1 )) 9: Construct (yℓ , Iℓ , Kℓ )2L ℓ=0 as in (19), (20) and (21) using Iref instead of [n] for Kℓ . ▷ O(Kn1 ) k m )| for k ∈ [K]. 10: Initialize Sn+1 ← |y0 − µ bk (Xn+1 ▷ O(K) k,ref k,ref k ▷ O(K log n1 ) 11: Initialize pn+1 ← Pn+1 (y0 ) for k ∈ [K] using sorted {Sj : j ∈ Iref }. 1,ref K,ref  12: Initialize Mn+1 ← f pn+1 , . . . , pn+1 . ▷ O(K) TB m m 13: if Mn+1 ≥ uα then CTB (Xn+1 ) ← CTB (Xn+1 ) ∪ I0 . 14: for ℓ = 1, . . . , 2L do 15: for k ∈ Kℓ do m )|. k ← |yℓ − µ bk (Xn+1 16: Update Sn+1 k,ref k ▷ O(log n1 ) 17: Update pk,ref n+1 ← Pn+1 (yℓ ) using sorted {Sj : j ∈ Iref }. 18: end for K,ref  19: Update Mn+1 ← f p1,ref n+1 , . . . , pn+1 . ▷ O(log K) for min/mean/med, O(K) for generic m )←C m 20: if Mn+1 ≥ uTB then CTB (Xn+1 TB (Xn+1 ) ∪ Iℓ . α 21: end for 22: end for m ) : m ∈ [M ]}. 23: Return: {CTB (Xn+1 24: Complexity: O(K(n1 +n2 ) log n1 +n2 log n2 +M Kn1 log(Kn1 )) for min, mean, and median f , 25: O(K(n1 +n2 ) log n1 +n2 log n2 +M Kn1 (log(Kn1 ) + K)) for generic f . 3: Compute pi

41

Algorithm 6 Efficient Exact TB Conformal Prediction for Minimum Merging (Proposition 6) Require: Calibration data {(Xi , Yi )}ni=1 ; disjoint index batches Iref ∪ Iagg = [n] with |Iref | = n1 m }M ; prediction functions {b and |Iagg | = n2 ; test points {Xn+1 µk }K m=1 k=1 ; level α. k 1: Compute Sj ← |Yj − µ bk (Xj )| for k ∈ [K] and j ∈ Iref . ▷ O(Kn1 ) k k k 2: Sort {Sj : j ∈ Iref } to get S(1) ≤ · · · ≤ S(n ) for each k ∈ [K]. ▷ O(Kn1 log n1 ) 1 k 3: Set S(n +1) ← ∞ for each k ∈ [K]. ▷ O(K) 1 k,ref

4: Compute pi

← Pik,ref for k ∈ [K] and i ∈ Iagg using sorted {Sjk : j ∈ Iref }. ▷ O(Kn2 log n1 ) k,ref

5: Compute Mi ← mink∈[K] pi

for i ∈ Iagg .  −Mi : i ∈ Iagg . TB 7: if uα = −∞ then ℓ ← n1 + 1 else ℓ ← n1 + 2 − ⌈uTB α (n1 + 1)⌉. 8: for m = 1, . . . , M do h i TK m )← m ) − Sk , µ m ) + Sk . 9: CTB (Xn+1 µ b (X b (X k k n+1 n+1 k=1 (ℓ) (ℓ) 10: end for m ) : m ∈ [M ]}. 11: Return: {CTB (Xn+1 12: Complexity: O(K(n1 + n2 ) log n1 + n2 log n2 + M K). 6: Compute uTB α ← −Quantile(1−α)(1+1/n2 )

B.6

▷ O(Kn2 ) ▷ O(n2 log n2 ) ▷ O(1) ▷ O(K)

Powerful tie-breaking strategies

When the underlying data distributions are discrete, identical evaluations across permutations occur with positive probability. While such ties may arise naturally at the level of individual marginal statistics Tbk , this degeneracy is severely compounded when these evaluations are aggregated via merging functions, particularly those with a low-cardinality image space such as the minimum merging function. Because such merging functions project multi-dimensional vectors onto highly constrained supports, structural ties on the final merged test statistics occur with substantially higher frequency. The standard p-value transformation, evaluated conservatively without tie-breaking, assigns the maximum of the tied ranks to all identical evaluations, inducing a conservative bias that artificially limits statistical power. Resolving these ties is therefore practically necessary. As briefly explained in Section 3.2, tie-breaking can be implemented by augmenting each transformed row with auxiliary variables Ub1 , . . . , UbL . For example, the marginal tie-broken p-values can be written as pU (Tbk ) :=

B o 1 X n k 1 1 (Ti , Ui , . . . , UiL ) ≥lex (Tbk , Ub1 , . . . , UbL ) , B+1 i=0

where ≥lex denotes the usual lexicographical order. The final merged statistic can be ranked analogously by pSB,U :=

B 1 X  1 (fi , −Ui1 , . . . , −UiL ) ≤lex (f0 , −U01 , . . . , −U0L ) . B+1 i=0

Provided that the augmented sequence of tuples (Tb1 , . . . , TbK , fb , Ub1 , . . . , UbL )B b=0 remains exchangeable under the null hypothesis, these lexicographical transformations preserve finite-sample exchangeability and hence level control. While ties among the statistics T0k , . . . , TBk are typically rare due to their often continuous support, ties among the merged values f0 , . . . , fB occur far more 42

frequently because the merging function often projects these values onto a highly-constrained lowcardinality support (e.g., minimum merging function). Importantly, when breaking fb -ties, the lexicographically tie-broken p-value is never larger than its conservative counterpart obtained without tie-breaking. Consequently, any power guarantee established for the conservative procedure immediately carries over to the tie-broken procedure. We formalize specific strategies for choosing the auxiliary variables within the SB testing framework, which readily generalize to TB and data-driven SB. • No tie-breaking (conservative baseline). Setting all auxiliary variables equal, i.e., Ub1 = · · · = UbL = c, recovers the standard conservative ranking in which tied evaluations receive the maximum tied rank. This choice is computationally trivial and guarantees a valid test. However, failing to separate the transformed or merged statistics yields an upwardly biased p-value, resulting in an overly conservative procedure with suboptimal power. • Uniformly random tie-breaking. Ties may be broken by appending an independent and identically distributed random sequence, e.g., Ub1 ∼ Uniform(0, 1), to each permutation and evaluating the lexicographical ordering of (Tbk , Ub1 ) for each k, and of (fb , −Ub1 ). Although this approach typically improves empirical power relative to the conservative no-tie-breaking baseline, it does not exploit the information contained in the test statistics themselves. This motivates the following data-dependent strategies, which construct the auxiliary variables to leverage this structural information. • Data-dependent tie-breaking. Rather than breaking ties at random, the auxiliary variables can instead be constructed systematically to favor permutations that exhibit stronger evidence against the null. The exchangeability framework developed here permits such data-dependent tie-breaking rules while preserving finite-sample validity. We propose two concrete constructions below, although the underlying principle is not restricted to these examples. (i) Data-dependent tie-breaking via studentised statistics. To implement this idea, we define the primary auxiliary sequence as the logsumexp, which is a smooth approximation to the maximum of the studentized statistics, i.e., Ub1 := log

K X k=1

  exp (Tbk − T̄bk )/σbk .

(23)

Here, T̄bk and σbk denote centering and scaling parameters, which may, for example, be taken as the empirical mean and standard deviation of {T0k , . . . , TBk } \ {Tbk } in a leave-one-out fashion [64, Eq. (4)]. This preserves the exchangeability of the augmented sequence and hence finite-sample validity. This auxiliary variable systematically breaks ties in favor of permutations exhibiting stronger studentized evidence against the null. This can improve power relative to purely random tie-breaking. 43

(ii) Data-dependent tie-breaking via p-value aggregation. Alternatively, ties can first be structurally resolved by deriving auxiliary sequences from the conservative (notie-breaking) marginal p-values p̌1b , . . . , p̌K b . For instance, we may compute Ub1 := −

K X

log(p̌kb ),

k=1

and

Ub2 := −

K X

p̌kb

k=1

relying on Fisher’s and Edgington’s methods. The primary variable Ub1 breaks all ties for which the product of the marginal p-values differs, while the secondary variable Ub2 subsequently breaks all ties for which their sum differs. Consequently, any potential ties surviving both aggregations must necessarily have equal sums and products of marginal p-values (e.g., exact same multiset occurring in a different dimensional ordering). These remaining symmetric ties can then potentially be broken by evaluating Ub3 as a softmax merging of studentized statistics as in (23). Nevertheless, inherent ties may still manifest if, for instance, the resampling procedure samples the exact same group transformation multiple times. Since identical transformations yield indistinguishable evaluations across all deterministic constructions, these remaining degeneracies can only be resolved by appending a terminal, independent uniform random variable, e.g., UbL ∼ Uniform(0, 1), thereby preserving finite-sample validity while breaking the remaining ties.

C

Additional simulation results

This section reports simulation results that complement the experiments in Section 7.

C.1

Type I error control

To estimate the type I error of the various tests in Figure 5 we use the two-sample testing framework with m = n = 50 i.i.d. samples drawn from the uniform distribution on (0, 1), aggregating over statistics formed by the Lp norms of the difference in means, and using permutations of the pooled samples. Nonetheless, we stress that these level results hold more generally for any distribution, sample size and statistic in any framework that tests the group-invariance null hypothesis (Definition 1), since they rely only on the exchangeability of the last B + 1 values. For illustration purposes, only B = 10 permutations are used in this setting, and uniformly random tie-breaking is implemented. While the level achieved by (Data-Driven) SB/TB holds independently of the value of K, the results for SeqSB depend on K, and the level presented for MaxT only holds for K = 1. The estimated levels are averaged over 20,000 repetitions, and the theoretical levels are: ⌊(B + 1)α⌋/(B + 1) for SB, TB and their data-driven variants (Theorems 1 and 3), K⌊(B + 1)α/K⌋/(B + 1) for SeqSB with αj = α/K, j ∈ [K] (Proposition 5), and (⌊Bα⌋ + 1)/(B + 1) for MaxT (Proposition S.8). As seen in Figure 5, the estimated levels of all implemented tests match their theoretical levels, confirming the validity of the theoretical analysis, as well as the correctness of our implementation. Figure 5 highlights the fact that MaxT does not control the type I error at the desired level α for most values of α for a fixed number B of transformations, even in the simplest case with K = 1. Therefore, the MaxT procedure is not a valid test, and hence is not considered in our power 44

1 SB/TB

SeqSB

MaxT

Type I Error

Data-Driven SB/TB

0

0

1

0

1 0 Significance Level α

1

0

Nominal Level α

Estimated Level

Theoretical Level

1

Figure 5: Type I error estimation for the SB/TB, Data-Driven SB/TB, SeqSB and MaxT tests. experiments in the rest of this section to ensure fair comparisons across tests. We note that for K > 1 the MaxT test exhibits a similar behavior (failing to control the type I error at level α) but deviates from the theoretical level derived for K = 1. All other tests have type I error bounded above by α as desired. Due to the discrete nature of the permutation test, the test level is not always exactly α, illustrating the importance of choosing B appropriately; see [10] for a formal study of the choice of B in Monte Carlo permutation tests. In particular, the (Data-Driven) SB/TB tests of Appendix D.2 achieve exact level α whenever (B + 1)α is an integer; in the following experiments we use B = 199 transformations and level α = 0.05. The same parameter choice ensures that SeqSB with αj = α/K and K = 10 also achieves exact level α. The classical worst-case tests of Section 2.3, which are constructed to be valid under arbitrary dependence, are in general much more conservative, with much lower type I error, than the permutation-calibrated tests presented here.

C.2

One-sample zero mean testing

To evaluate the proposed methods under varying degrees of dependence across statistics, we generate n = 10, 000 samples from a d-dimensional multivariate normal distribution with ambient dimension d = 40. The covariance matrix is constructed such that all dimensions share a pairwise feature correlation parameter ρ, representing the off-diagonal entries, while the diagonal variances are set to 1. We consider two distinct alternative hypotheses: a sparse setting in the left panel of Figure 6 where only the first dimension contains a signal (a mean shift of µ1 = 0.03, with µj = 0 for j ̸= 1), and a dense setting for the right panel where the signal is distributed across all dimensions (µj = 0.015 for all j). For each regime, we assess the performance of the worst-case (WC) and singlestep (SB) procedures using the minimum, median, and mean aggregation functions over the absolute one-sample t-statistics computed for each dimension. Additionally, we evaluate the Data-Driven SB test, defined in Appendix D.2, which adapts across the three merging functions. As a benchmark in the sparse regime, we include an Oracle test, which computes the classical permutation one-sample t-test strictly on the first dimension, assuming prior knowledge of the true signal location. For the group transformations, we multiply each data point by a random sign (i.e., Rademacher random variable) to simulate the case of zero mean, leveraging the symmetric property of the multivariate normal distribution. All empirical power calculations are averaged over 1000 independent repetitions at a nominal significance level of α = 0.05. The empirical results, illustrated in Figure 6, reveal several key dynamics of the testing proce45

Power

1

0

0.0

0.2

0.4 0.6 Correlation ρ

SB Data-Driven SB Min SB Median SB Mean

0.8

1.0

Oracle WC Min WC Median WC Mean

0.0

0.2

0.4 0.6 Correlation ρ

SB Data-Driven SB Min SB Median SB Mean

0.8

1.0

WC Min WC Median WC Mean

Figure 6: Power estimation for one-sample zero mean testing of an equicorrelated multivariate Gaussian. dures. First, the worst-case (WC) tests are highly conservative, achieving near-zero empirical power across almost all settings. In the sparse signal regime (left panel), among the three single-step methods, only the SB Min test demonstrates strong performance. Because the mean shift is isolated to a single dimension, the minimum p-value effectively targets the only active statistic; taking the mean or median heavily dilutes this localized signal with the remaining d − 1 noise dimensions. Conversely, in the dense signal regime (right panel), where the mean shift is applied across all dimensions, the SB Mean and SB Median tests successfully aggregate the distributed signal and consequently outperform the SB Min test. Crucially, across both regimes, the SB Data-Driven test proves to be highly adaptive, consistently achieving statistical power that approximates the single best-performing aggregation method. Finally, the correlation ρ significantly, yet oppositely, impacts test power depending on the signal structure: increasing ρ in the sparse setting leads to higher power, whereas increasing ρ in the dense setting leads to lower power. In the dense case, strong correlation reduces the effective number of independent signals available for aggregation. However, in the sparse case, highly correlated features effectively constrain the background noise, allowing the isolated signal to stand out. Notably, at ρ = 1, the SB Min and SB Data-Driven tests achieve the exact same empirical power as the single Oracle test, demonstrating that, unlike their WC counterparts, these adaptive procedures do not need to over-correct for perfectly redundant data (as explained in the discussion following Proposition 4).

C.3

Independence nonparametric testing

We evaluate the proposed methods for independence testing using the Hilbert-Schmidt Independence Criterion (HSIC, [25]) in Figure 7. The data is generated from a Dirichlet distribution with a uniform concentration parameter vector 1d . Specifically, we generate a sample of size N in ambient dimension d, and partition each observation into two views, X ∈ Rd−c and Y ∈ Rc , to evaluate the dependence between them. In our first experiment, we vary the sample size N from 500 to 5000 while fixing the dimensions to d = 500 and c = 200. In the second experiment, we fix N = 500, d = 200, and c = 80, and examine the impact of the bandwidth grid size. The HSIC V-statistic,

46

Power

1

0

1000

2000 3000 Sample Size N

SB Data-Driven SB Min SB Median SB Mean

4000

5000

TB Data-Driven TB Min TB Median TB Mean

1

100 200 300 Number of Bandwidths K SB Data-Driven SB Min SB Median SB Mean

400

WC Min WC Median WC Mean

Figure 7: Power estimation for kernel-based HSIC independence nonparametric testing. with Gaussian kernels, is computed over a multi-scale grid where the total number of bandwidths evaluated, K, corresponds to the product of the number of candidate bandwidths for X and Y (i.e., K = KX × KY ). Empirical power is averaged over 1000 independent repetitions, all tests are calibrated using B = 199 permutations, and the nominal significance level is set to α = 0.05. The paired data is transformed by permuting the data within one sample to break the dependence structure. We compare the Single-Step (SB) methods against the Two-Batch (TB) and WorstCase (WC) methods using three standard merging functions: the minimum, mean, and median of the individual test statistics. Furthermore, we assess the performance of Data-Driven SB/TB tests, defined analogously to Appendix D.2, which are adaptive to the choice of these three merging functions. The experimental results of Figure 7 highlight the significant advantages of the proposed SB testing procedure. As shown in the left panel, the SB method consistently outperforms the TB method, with the power differential becoming particularly pronounced at larger sample sizes. The Data-Driven SB and TB procedures are observed to be adaptive, especially for larger sample sizes. Furthermore, the right panel demonstrates that the SB procedure vastly outperforms the classical WC approach; this holds strictly true across all considered merging functions (minimum, mean, and median). When comparing the merging strategies themselves, the minimum function strictly achieves the highest power in the experiment varying the sample size (left panel). In the experiment varying the number of paired kernel bandwidths (right panel), each of the three merging functions dominates in a distinct regime depending on the grid density. In this setting, SB Data-Driven is seen to match the power of the best-performing merging function across all regimes. The right panel also reveals critical dynamics regarding the bandwidth selection process: the initial, narrow bandwidth grid is poorly calibrated, yielding near-zero power. As we increase the total number of bandwidths K, the grid broadens to encompass better-calibrated bandwidths, resulting in a sharp, immediate increase in empirical power. However, continuing to increase K inevitably introduces poorly calibrated, noisy bandwidths into the collection, which begins to penalize the overall power. Crucially, as K becomes increasingly large, the power of the SB procedure stabilizes and reaches a robust plateau, whereas the power of the conservative WC procedure deteriorates completely to zero. 47

Power

1

0

0.0

0.2

0.4 0.6 Perturbation Scale

MMDAgg-SB MMDAgg-TB

0.8

1.0

MMDAgg-Bisection

0.0

0.2

0.4 0.6 Perturbation Scale

HSICAgg-SB HSICAgg-TB

0.8

1.0

HSICAgg-Bisection

Figure 8: Power estimation for aggregated kernel-based MMDAgg and HSICAgg nonparametric testing.

C.4

Improved MMDAgg and HSICAgg optimal tests

In the experimental setting of Figure 8, we implement the SB versions of the MMDAgg and HSICAgg tests [2, 62] based on unbiased U-statistics. These implementations are exact: they correspond strictly to the SB tests analyzed in Appendix F, requiring no further approximations or empirical parameter tuning. In particular, the collections of bandwidths for Gaussian kernels are implemented exactly as prescribed by theory (28, 29).4 These bandwidth collections are designed for the Sobolev smoothness assumption, which is satisfied by the data-generating process in this experiment as we use specially-constructed perturbed uniform distributions. For two-sample testing, the aim is to detect the difference between the perturbed and unperturbed one-dimensional uniform distributions, while for independence testing the goal is to detect the dependence between the two dimensions of a joint perturbed uniform distribution (with uniform marginals). We refer the reader to [62, Eq. 17] and [2, Eq. 4.3] for specific expressions of the perturbed uniform distributions. We control the signal strength by varying the perturbation scale from 0 (the null hypothesis) to 1 (the theoretical maximum for non-negative densities), while maintaining a sample size of 500, a test level of α = 0.05 and a permutation number of B = 199 across both experiments. We refer to [15] for a method to significantly reduce the computational cost of permutations while maintaining high power. In Figure 8, we compare the empirical power (averaged over 1000 independent repetitions) of the SB tests and of their TB variants against the original MMDAgg and HSICAgg tests, which are defined via their statistic view with rejection event (2) and Monte Carlo calibration threshold (3) where the supremum is being further approximated via a bisection method. As we show in Proposition S.1, the threshold (3) can actually be computed exactly, bypassing the estimator error induced by the bisection approximation. This corresponds to the MaxT variants of the tests (Section 2.4). However, as shown in Proposition S.8, these do not guarantee finite-sample level control across all settings, and are hence excluded from our experiments, while the more conservative bisection-based MaxT variants are included for comparison with [2, 62]. The proposed TB variants (with minimum merging function) of MMDAgg and HSICAgg resolve this lack of finite-sample validity while 4

The isotropic Sobolev assumption implies uniform smoothness across all dimensions; therefore, we use the same bandwidth for both kernels in the HSIC computation. We evaluate the more general setting of kernel-specific bandwidths using a two-dimensional grid in Appendix C.3.

48

remaining closely related to their MaxT counterparts (see Appendix B.2), benefit from a p-value formulation and achieve higher empirical power. The computationally more efficient SB variants are observed to be statistically more powerful, validating our theory (see Appendix E.3). Indeed, for both the MMDAgg and HSICAgg comparisons of Figure 8, the SB method strictly outperforms the TB approach, which in turn outperforms the standard MaxT-Bisection method. While this power advantage is significant when using B = 199 permutations, it is expected that all three tests perform comparably when B is sufficiently large. In practice, the implementations of [2, 62] use many more permutations (e.g., B = 2000) to stabilize the power. Consequently, the superiority of the SB procedure can be interpreted dually: it achieves strictly higher power for a fixed small number of permutations, or equivalently, it drastically reduces the computational cost required to match the power of the Bisection method. We encourage practitioners to use the newly proposed MMDAgg-SB and HSICAgg-SB versions of the tests, whose minimax optimality and validity are established in Appendix F.

D

Extensions of the aggregation framework

This section collects extensions of the basic SB and TB aggregation framework beyond the core p-value aggregation procedures.

D.1

Exchangeable e-values

Although our primary focus is on p-value aggregation, the SB aggregation framework extends naturally to e-values [74], which quantify evidence via non-negative random variables E satisfying E[E] ≤ 1 under the null. While simple averaging of e-values preserves validity, it can be conservative when the expectation of the average is strictly smaller than one. Under the group-invariance hypothesis, however, additional structure is available: the row-wise merged quantities (f0 , . . . , fB ) are exchangeable in the index b. Exploiting this exchangeability, for any measurable score map ψ : R → [0, ∞), the self-normalized statistic (B + 1)ψ(f0 ) ESB := PB , ψ(f ) b b=0

ESB = 1 if

B X

ψ(fb ) = 0,

b=0

is an e-value satisfying E[ESB ] = 1 under the null. Thus, exchangeability yields an exactly calibrated e-value that avoids the potential conservativeness of generic averaging schemes. This construction corresponds to the soft-rank e-value [53, Chapter 1.7].

D.2

Data-driven aggregation of merging functions

This subsection extends the aggregation framework from combining statistics to combining the merging rules themselves. Building on this perspective, we develop a data-driven aggregation scheme that exploits redundancy among merging rules to retain the power of the best-performing rule while preserving finite-sample validity. A related strategy has been explored by [26] in the context of subsampling-based aggregation; however, a formal theoretical power analysis has been lacking. While we formulate this data-driven scheme within the single-batch (SB) framework for clarity, the underlying methodology readily extends to the two-batch (TB) procedure.

49

X

X

T 10

T 20

T 30

···

T K0

p10

p20

p30

···

pK0

f0

g1X

T 11

T 21

T 31

···

T K1

p11

p21

p31

···

pK1

f1

g2X

T 12

T 22

T 32

···

T K2

p12

p22

p32

···

pK2

f2

.. .

.. .

.. .

.. .

.. .

.. .

.. .

.. .

.. .

.. .

gBX

T 1B

T 2B

T 3B

T KB

p1B

p2B

p3B

pKB

fB

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

··· <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

1 B pkb = 1(Tik ≥ Tbk) B+1 ∑ i=0

T kb = T k(gbX)

X

T 10

T 20

T 30

···

T K0

p10

p20

p30

···

pK0

g1X

T 11

T 21

T 31

···

T K1

p11

p21

p31

···

pK1

g2X

T 12

T 22

T 32

···

T K2

p12

p22

p32

···

pK2

.. .

.. .

.. .

.. .

.. .

.. .

.. .

.. .

gBX

T 1B

T 2B

T 3B

···

T KB

p1B

p2B

p3B

f 10

f 20

f 30

···

fM 0

p̄10

p̄20

f 11

f 21

f 31

···

fM 1

p̄11

f 12

f 22

f 32

···

fM 2

.. .

.. .

.. .

f 1B

f 2B

f 3B

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

.. .

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

···

pKB

p̄30

···

p̄ M 0

f¯0

p̄21

p̄31

···

p̄ M 1

f¯1

p̄12

p̄22

p̄32

···

p̄ M 2

f¯2

.. .

.. .

.. .

.. .

.. .

.. .

fM B

p̄1B

p̄2B

p̄3B

p̄ M B

f¯B

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

···

Reject H0 if pSB ≤ α

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

···

P−value pSB

1 B 1(fb ≤ f0) B+1 ∑ b=0

P−value Matrix

Statistic Matrix

Transformed Data

X

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

··· <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

P−value Matrix

Aggregated Statistic Matrix

1 B 1(f¯b ≤ f¯0) B+1 ∑ b=0

Reject H0 if pSB,F ≤ α

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

Merging p−values f¯b = min(p̄1, …, p̄ M)

1 B p̄m = 1(f im ≤ f bm) b B+1 ∑ i=0

f bm = f m(p1b, …, pKb)

P−value pSB,F

b

b

Figure 9: Schematic illustration of the Data-driven SB aggregation procedure of Appendix D.2 P−value Matrix 1 Merging TestingTBatch 1 , . . . , T K and merging Transformed B (Algorithm 7), using statistics f , . . . , f M . p−values 1 functions Data

X

T kb = T k(gbX)

pkb =

1+ 1(T̃ ki ≥ Tbk)) B+1 ( ∑ i=1

f˜b = f(p1b, …, pKb)

X

T 10

T 20

T 30

···

T K0

p10

p20

p30

···

pK0

f˜0

g1X

T 11

T 21

T 31

···

T K1

p11

p21

p31

···

pK1

f˜1

g2X

T 12

T 22

T 32

···

T K2

p12

p22

p32

···

pK2

f˜2

.. .

.. .

.. .

.. .

.. .

.. 50 .

.. .

.. .

.. .

.. .

gBX

T 1B

T 2B

T 3B

T KB

p1B

p2B

p3B

pKB

f˜B

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

··· <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

··· <latexit sha1_base64="JXzBjyZgmsI8ags2tosNzBng8sY=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9Wik0PTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66LqX1Zr97VK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/ALFFjzg=</latexit>

<latexit sha1_base64="qDf6pzroMMJvMcIkHWdR/tYPv+o=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSRS1GPRi8cKthbaUDabTbt2sxt2J4VS+h+8eFDEq//Hm//GbZuDtj4YeLw3w8y8MBXcoOd9O4W19Y3NreJ2aWd3b/+gfHjUMirTlDWpEkq3Q2KY4JI1kaNg7VQzkoSCPYbD25n/OGLacCUfcJyyICF9yWNOCVqp1R1FCk2vXPGq3hzuKvFzUoEcjV75qxspmiVMIhXEmI7vpRhMiEZOBZuWuplhKaFD0mcdSyVJmAkm82un7plVIjdW2pZEd67+npiQxJhxEtrOhODALHsz8T+vk2F8HUy4TDNkki4WxZlwUbmz192Ia0ZRjC0hVHN7q0sHRBOKNqCSDcFffnmVtC6q/mW1dl+r1G/yOIpwAqdwDj5cQR3uoAFNoPAEz/AKb45yXpx352PRWnDymWP4A+fzB85dj0s=</latexit>

P−value pTB

1 B 1(f˜b ≤ f˜0) B+1 ∑ b=0

Reject H0 if pTB ≤ α

Algorithm 7 Data-Driven Single-Batch Aggregation Require: data X; statistics T 1 , . . . , T K ; transformations g0 , . . . , gB (g0 = id); merging functions f 1 , . . . , f M ; level α. 1: Transformed statistics: For each b ∈ [B]0 and k ∈ [K], compute Tbk ← T k (gb (X)). 2: Standardization via p-values: For each b ∈ [B]0 and k ∈ [K], compute the permutation p-value B

p Tbk



 1 X ← 1 Tik ≥ Tbk . B+1 i=0

3: Multiple aggregation row-wise: For each b ∈ [B]0 and m ∈ [M ], set

  fbm ← f m p Tb1 , . . . , p TbK . 4: Standardization via p-values: For each b ∈ [B]0 and m ∈ [M ], compute the permutation

p-value

B

 p≤ fbm ←

 1 X 1 fim ≤ fbm . B+1 i=0

5: Aggregation row-wise: For each b ∈ [B]0 , set

 ← min p≤ fbm . pmin b m∈[M ]

6: Calibration: Compute the p-value

pSB,F ←

B  1 X ≤ pmin . 1 pmin 0 b B+1 b=0

7: Decision rule: Reject if pSB,F ≤ α.

Procedure. Let F = {f 1 , . . . , f M } denote a collection of merging functions, such as the minimum, median, mean, or other order-statistic–based rules. For each m ∈ [M ], apply the SB aggregation procedure of Section 3.1 to obtain merged statistics f0m , f1m , . . . , fBm . Importantly, for each b ∈ [B]0 , the values fb1 , . . . , fbM are obtained by applying the merging functions in F to the same collection of permutation p-values computed from the b-th transformed dataset. With a slight abuse of notation, define the associated lower-tail permutation p-values B

p≤ (fbm ) :=

 1 X 1 fim ≤ fbm , B+1 i=0

b ∈ [B]0 .

For each transformation index b, aggregate across merging functions by taking the minimum p-value := min p≤ (fbm ), pmin b m∈[M ]

51

b ∈ [B]0 .

The data-driven aggregation test is then defined by B  1 X pSB,F := 1 pmin ≤ pmin , 0 b B+1 b=0

and rejects the null hypothesis whenever pSB,F ≤ α. Finite-sample validity follows directly from the exchangeability of (pmin b )b∈[B]0 under the group-invariance hypothesis. This procedure is fully presented in Algorithm 7 and is illustrated in Figure 9. To study the power of the proposed data-driven aggregation procedure, we first formalize the notion that, in many applications, only a small number of merging functions are effectively distinct. This captures redundancy across merging rules through the ranks they induce on the same transformed rows. Effective multiplicity. For ε ≥ 0, define the (random) effective multiplicity n o Neff (ε) = min |S| : S ⊆ [M ], min p≤ (fbm ) ≤ (1 + ε) min p≤ (fbm ) ∀b ∈ [B]0 . m∈S

m∈[M ]

Thus, Neff (ε) is the smallest number of representative merging functions whose minimum p-value approximates, up to a multiplicative factor 1 + ε, the minimum over all M merging functions. We assume that there exist deterministic constants N ≤ M and η ∈ [0, 1] such that (24)

P(Neff (ε) ≤ N ) ≥ 1 − η.

Condition (24) is satisfied whenever the collection {p≤ (fbm )}m∈[M ] exhibits redundancy, for instance when many merging functions induce nearly identical ranks on the same transformation. In the extreme case where all merging functions induce identical ranks, one has Neff (0) = 1 almost surely. The following proposition quantifies the power of the data-driven aggregation procedure under this effective multiplicity condition. Proposition S.3. Fix α ∈ (0, 1) and ε ≥ 0. Suppose that the effective multiplicity condition (24) holds for some N ≤ M and η ∈ [0, 1]. Then the type II error of the data-driven aggregation procedure satisfies    m P pSB,F > α ≤ η + P min p≤ (f0 ) > aε,N ≤ η + min P(p≤ (f0m ) > aε,N ) m∈[M ]

m∈[M ]

for aε,N := α/(N (1 + ε)). Proof. See Appendix I.3. Proposition S.3 shows that the power of the data-driven aggregation procedure is comparable to that of the best merging function in the collection, up to a multiplicative factor (1 + ε)N in the significance level and an additive error η. In particular, the procedure behaves as if only N effectively distinct merging functions were considered. As a special case, taking ε = 0, η = 0, and N = M yields     α α m P pSB,F > α ≤ P min p≤ (f0 ) > ≤ min P p≤ (f0m ) > , M M m∈[M ] m∈[M ] 52

which coincides with the classical Bonferroni-type bound obtained by calibrating each merging function at level α/M . Existing sufficient conditions under which a single permutation p-value is powerful, such as those studied in [34, 38], can be invoked to further instantiate and interpret the power bound above. Uniform power improvement over single merging rules. Proposition S.3 quantifies the power of the data-driven aggregation test in terms of redundancy among the candidate merging functions, as captured by the effective multiplicity Neff (ε). This perspective is particularly informative when many rules are nearly equivalent (so that Neff (ε) ≪ M ). At the same time, even when the candidate rules are all effectively distinct (so that Neff (0) = M ), the permutation-calibrated aggregation can strictly outperform every fixed choice of a single merging function under heterogeneous (e.g., mixture) alternatives, where different rules are optimal in different sub-regimes. We delve into this phenomenon in more detail in Appendix E.2.

D.3

SB aggregation of data-asymmetric tests

We next specialize the preceding SB framework to settings in which the test statistic is asymmetric in the data. Such asymmetry arises when a portion of the sample is used to construct a data-dependent object that is subsequently evaluated on held-out data, so that the resulting test depends on a particular ordering or partition of the sample. Examples include both early data-splitting procedures [14, 45] and more recent nonparametric tests involving learned or data-adaptive components [e.g., 11, 32, 35–37, 41, 42, 51, 65, 66, 70]. In such settings, it is natural to repeat the test over multiple random (or predetermined) splits and to aggregate the resulting p-values to improve power and stability. Let ∆1 , . . . , ∆K denote tests obtained by applying a fixed asymmetric testing procedure to K independent random orderings or partitions of the data, and let p(k) be the p-value associated with ∆k . Even when P the vector (p(1) , . . . , p(K) ) is exchangeable under the null hypothesis, the arithmetic K (k) is not, in general, super-uniform [12, 21], and therefore cannot be used −1 mean p̄ := K k=1 p directly as a valid p-value. While a classical result of [56] shows that the rescaled statistic 2p̄ is super-uniform, this worst-case correction is often conservative and may substantially reduce power in practice. Assume now that the group-invariance hypothesis holds. For each transformation index b ∈ [B]0 (k) and repetition k ∈ [K], let pb denote the permutation p-value obtained by applying the same P (k) transformation to all K repetitions. Define the row-wise averages p̄b := K −1 K k=1 pb for b ∈ [B]0 , P and define the SB average p-value by pSB,avg := (B + 1)−1 B b=0 1(p̄b ≤ p̄0 ). By Theorem 1, this pvalue is valid under the group-invariance hypothesis. Moreover, Theorem 2 and Proposition 3 imply that it is structure-adaptive and uniformly dominates the worst-case corrected average 2p̄0 in terms of power. It is also worth noting that exploiting the additional finite-population structure inherent in the SB construction, the worst-case correction factor can be sharpened from 2 to 2(B +1)/(B +2); see Corollary S.2. Taken together, these results establish SB aggregation as a principled approach for aggregating asymmetric tests under the group-invariance hypothesis. Consistency transfer from individual tests. Crucially, the uniform power dominance established in Theorem 2 has an immediate implication for power consistency. Since Theorem 2 shows that the type II error of the SB procedure is uniformly upper bounded by that of any deterministic worst-case calibrated test, any consistency guarantee established for the latter automatically transfers to the SB procedure as shown below. 53

n Corollary S.1. Let K = Kn be an arbitrary (possibly diverging) sequence, and let p10,n , . . . , pK 0,n be exchangeable permutation p-values indexed by n. Assume that, for some (and hence also for all) k ∈ [Kn ],  sup PP pk0,n > α → 0 for every α ∈ (0, 1).

P ∈P

Then the SB average aggregation test satisfies  sup PP pSB,avg > α → 0 P ∈P

for every α ∈ (0, 1),

without any restriction on the growth rate of Kn . Proof. See Appendix I.6. The converse does not hold in general: when the individual tests are (nearly) independent, the SB average test can be consistent even if none of the individual tests are.

E

Power refinements

This section collects refinements of the power and adaptivity analysis for the proposed aggregation procedures.

E.1

Uniform asymptotic adaptation

Here, we present a version of the asymptotic adaptivity result in Proposition 3 which holds uniformly over a class of null distributions. Proposition S.4. Let P be a class of null distributions. Under the setup of Proposition 3, assume that for each P ∈ P, the pointwise convergence conditions hold with F replaced by FP , and that sup sup PP (f1,n ≤ t, f2,n ≤ s) − FP (t)FP (s) → 0.

P ∈P t,s∈R

Write Q⋆α,P := inf{u ∈ R : FP (u) ≥ α}, and assume the uniform margin condition: for every ε > 0,  inf min α − FP (Q⋆α,P − ε), FP (Q⋆α,P + ε) − α > 0.

P ∈P

Then, for every ε > 0,  ⋆ sup PP |ûSB α − Qα,P | > ε → 0.

P ∈P

Proof. See Appendix I.2.

54

E.2

Uniform power improvement over single merging rules

In this subsection, we show that SB aggregation of multiple merging rules (Appendix D.2) can achieve uniform power improvement over each individual rule. The key mechanism is that aggregation can succeed whenever at least one candidate rule provides sufficiently strong evidence, while incurring only a limited calibration penalty. We start with a simple lemma. Lemma S.1. Suppose that the permutation p-values are defined by B

p≤ (fbm ) =

1 X 1{fim ≤ fbm }, B+1 i=0

b ∈ [B]0 , m ∈ [M ].

Suppose that the minimum permutation p-value across merging rules attains the smallest possible grid value 1/(B + 1), that is, 1 . (25) pmin = min p≤ (f0m ) = 0 B+1 m∈[M ] In this case, the SB aggregation p-value admits the representation pSB,F =

W , B+1

W ≤ M,

where = 1/(B + 1)} W := {b ∈ [B]0 : pmin b denotes the number of permutation indices attaining the minimal grid value. As a result, whenever M ≤ ⌊(B + 1)α⌋, the aggregated test necessarily rejects. Proof. See Appendix I.11. Lemma S.1 shows that if any candidate merging rule attains the smallest possible permutation p-value, namely 1/(B + 1), on the observed data, then the second-stage permutation calibration incurs a penalty of at most M/(B + 1). In particular, if B is large enough so that M/(B + 1) ≤ α, then this event alone guarantees rejection by the aggregated test. The next proposition demonstrates that, under a mixture alternative in which different components favor different merging rules, the aggregated test can achieve power one, even though no single merging rule is uniformly powerful across all components. Proposition S.5. Fix α ∈ (0, 1) and suppose that M ≤ ⌊(B + 1)α⌋. Consider an alternative distribution of the data that is a finite mixture P =

J X

πj Pj ,

where

j=1

πj > 0,

J X

πj = 1.

j=1

such that for each component Pj there exists an index m(j) ∈ [M ] satisfying   1 m(j)  PPj p≤ f0 = = 1, B+1 55

(26)

and moreover no single merging function succeeds on all components, in the sense that for every fixed m ∈ [M ] there exists some j ∈ [J] with πj > 0 such that   PPj p≤ f0m ≤ α < 1. (27) Then the aggregated test has power one under the mixture:  PP pSB,F ≤ α = 1, whereas every fixed-rule SB test has strictly smaller mixture power:   PP p≤ f0m ≤ α < 1 for all m ∈ [M ]. Proof. See Appendix I.12.

E.3

Power properties of TB aggregation

The analysis of the TB aggregation procedure differs from that of SB aggregation since the testing and reference batches are distinct, and hence the deterministic super-uniformity argument of Lemma 1, which underpins the power analysis of SB aggregation, no longer applies. Nevertheless, as the number of transformations B increases, we show that the SB and TB thresholds become asymptotically close. This follows from the fact that, under i.i.d. uniform sampling of transformations, both procedures converge to the idealized procedure that enumerates all transformations in G. Although threshold convergence alone does not imply identical power, it plays a crucial asymptotic role. Under mild regularity conditions, such as continuity of the limiting distribution at the threshold, it ensures that the power guarantees established for SB aggregation carry over to TB aggregation as B → ∞. The following proposition formalizes the asymptotic threshold equivalence and provides an exponential-in-B control. Proposition S.6. Assume that the SB and TB procedures use the same continuous merging function TB f : [0, 1]K → R, and that the number of coordinates K is fixed. Let ûSB α and ûα denote the SB and TB thresholds defined in (6) and (11), respectively. Assume that, conditional on the observed data X, the transformations g1 , . . . , g2B are i.i.d. uniform draws from G. Then for every ε > 0, there exist constants C(ε, f, K) > 0 and c(ε, f ) > 0 such that  SB −c(ε,f )B P |ûTB , for all sufficiently large B. α − ûα | > ε ≤ C(ε, f, K) e p

SB In particular, ûTB α − ûα −→ 0 at an exponential rate in B uniformly over all data-generating distributions for X.

Proof. See Appendix I.4. Although the SB and TB thresholds become asymptotically equivalent as B → ∞, their finite-B power may differ. In particular, TB can be less powerful than SB even though it uses an additional batch of transformations. The loss arises from the fact that TB holdout p-values are constructed without including the observed statistic T0k in the ranking set, whereas the SB permutation pvalues do include T0k . When B is small, this distinction can be decisive: under a strong signal, the inclusion of T0k systematically inflates the SB calibration p-values away from the smallest grid 56

point 1/(B + 1), thereby increasing the SB critical value and facilitating rejection. By contrast, TB calibration p-values can still hit the smallest grid point with non-vanishing probability, creating ties that obstruct rejection under a strict comparison. We now formalize this mechanism via an explicit construction, showing that, for fixed and small B, SB aggregation can achieve asymptotic power one whereas TB aggregation cannot. Proposition S.7. Fix an integer B ≥ 1 and let α ∈ (1/(B + 1), 1). There exist a sequence of alternatives and a merging function for which, with B fixed, the SB aggregation test achieves asymptotic power one, whereas the TB aggregation test has asymptotic power strictly bounded away from one. More precisely, under a strong-signal regime and some regularity conditions specified in Appendix I.5, P(pTB ≤ α) → P(W ≤ ⌊(B + 1)α⌋ − 1)

as n → ∞,

where W follows a negative hypergeometric distribution with probability mass function  2B−k−1 P(W = k) =

B−1  2B B

,

k = 0, 1, . . . , B.

In particular, this limit is strictly smaller than one for all B ≥ 1 and α ∈ (1/(B + 1), 1). By contrast, under the same conditions, P(pSB ≤ α) → 1

as n → ∞.

Proof. See Appendix I.5. For illustration, when 1/(B +1) < α < 2/(B +1), we have ⌊(B +1)α⌋ = 1 and hence the limiting TB power reduces to P(W = 0) = 1/2, which holds for all B ≥ 1. This shows that, for fixed B, the TB construction can suffer a non-vanishing power loss relative to SB aggregation, despite using an additional batch of transformations.

E.4

Worst-case bound for quasi-arithmetic mean aggregation

In this subsection, we derive finite-sample worst-case bounds for quasi-arithmetic mean aggregation under the SB construction, and quantify how the resulting calibration constants improve upon existing super-uniform bounds by exploiting the permutation structure. (k)

Lemma S.2. Let {pb : b ∈ [B]0 , k ∈ [K]} be permutation p-values constructed as in (4). Let ϕ : (0, 1] → R be a strictly increasing and continuous function. Define the row-wise quasi-arithmetic mean ! K 1 X  (k)  −1 p̄ϕ,b := ϕ ϕ pb , b ∈ [B]0 . K k=1

For each ℓ ∈ {1, . . . , B + 1}, define the deterministic thresholds     ℓ X 1 j . tϕ,B (ℓ) := ϕ−1  ϕ ℓ B+1 j=1

57

Define the calibration function Gϕ,B : (0, 1] → [0, 1] by n o 1 max ℓ ∈ [B + 1]0 : tϕ,B (ℓ) ≤ t , Gϕ,B (t) := B+1 with the convention tϕ,B (0) := 0. Then, for any t ∈ (0, 1], it holds deterministically that B

1 X 1(p̄ϕ,b ≤ t) ≤ Gϕ,B (t). B+1 b=0

Consequently, under the group-invariance hypothesis, P(p̄ϕ,0 ≤ t) ≤ Gϕ,B (t),

t ∈ (0, 1].

In particular, Gϕ,B (p̄ϕ,0 ) is a valid p-value. Proof. See Appendix I.13. Lemma S.2 recovers the power-mean aggregation scheme by choosing ϕ(x) = xr (r > 0). In this case, it is useful to obtain an explicit linear upper bound on the stepwise calibration function Gϕ,B , as such bounds directly yield simple and interpretable worst-case corrections for the aggregated p-values. Accordingly, for r > 0 we define cr,B :=

max

 1≤ℓ≤B+1 1 Pℓ

j=1

jr

1/r .

This constant is chosen so that for all t ∈ (0, 1].

Gϕ,B (t) ≤ cr,B t

Indeed, recall that the calibration function Gϕ,B is defined by n o 1 Gϕ,B (t) = max ℓ ∈ [B + 1]0 : tϕ,B (ℓ) ≤ t . B+1 Thus, if tϕ,B (ℓ) ≤ t for some ℓ, then ℓ/(B + 1) ≤ Gϕ,B (t), and for the maximal admissible index ℓ⋆ (t) we have Gϕ,B (t) = ℓ⋆ (t)/(B + 1). Consequently, any inequality of the form ℓ ≤Ct B+1

whenever tϕ,B (ℓ) ≤ t

immediately implies the linear bound Gϕ,B (t) ≤ C t. For ϕ(x) = xr with r > 0, we have 

tϕ,B (ℓ) =

ℓ 1X

1  B+1 ℓ

1/r jr

,

j=1

and the smallest constant C for which the above inequality holds uniformly over ℓ ∈ {1, . . . , B + 1} is precisely cr,B . 58

It is instructive to contrast the constant cr,B with its worst-case counterpart in [73]. When only super-uniformity of the aggregated p-value is assumed, without any additional structural information, the optimal worst-case linear bound for the power mean is given by the universal constant cr = (r + 1)1/r , which is sharp in the class of all super-uniform random variables. In the present setting, however, the aggregated quantities arise from permutation p-values and therefore inherit a finite-population structure that is not captured by super-uniformity alone. This additional structure allows the worst-case correction factor to be strictly improved at finite B. The constant cr,B quantifies the optimal linear bound that exploits this permutation structure. The following lemma shows that cr,B is strictly smaller than cr for any finite B, while converging to cr as B → ∞. Lemma S.3. Fix r > 0 and recall the constant cr,B defined above. Let cr := (r + 1)1/r . Then: (i) For every B ≥ 1, cr,B < cr . (ii) The sequence B 7→ cr,B is nondecreasing and as B → ∞.

cr,B ↑ cr Proof. See Appendix I.14.

As a direct consequence of Lemma S.2 and Lemma S.3, we obtain the following worst-case linear bound for arithmetic-mean aggregation of permutation p-values. (k)

Corollary S.2. Let {pb : b ∈ [B]0 , k ∈ [K]} be permutation p-values constructed as in (4). For each b ∈ [B]0 , define the row-wise average K

1 X (k) p̄b := pb . K k=1

Then, for any t ∈ [0, 1], it holds deterministically that   B 2(B + 1) 1 X 1(p̄b ≤ t) ≤ min 1, t . B+1 B+2 b=0

Pℓ

j ℓ+1 j=1 B+1 = 2(B+1) . 2ℓ = max1≤ℓ≤B+1 ℓ+1 = Hence the associated linear bound constant is c1,B = max1≤ℓ≤B+1 1 Pℓℓ j j=1 ℓ 2(B+1) B+2 , which yields the claim.

Proof. Apply Lemma S.2 with ϕ(x) = x, so that p̄ϕ,b = p̄b and tϕ,B (ℓ) = 1ℓ

This improves upon the classical result of [56] stating that twice the arithmetic mean of pvalues is itself a valid p-value. Here, we show that, in the permutation setting, the scaled quantity p̄b 2 (B + 1)/(B + 2) is valid. See Appendix D.3 for details.

59

F

Adaptive testing details

This section details how SB aggregation yields finite-sample valid adaptive kernel tests. We instantiate SB aggregation to adaptive kernel two-sample MMD and independence HSIC testing5 by aggregating permutation p-values over a dyadic bandwidth grid. As discussed in Section 2.4, the existing MaxT-based adaptive MMD and HSIC procedures of [2] and [62] either rely on oracle critical values that are not directly implementable in practice or use Monte Carlo calibration that does not guarantee finite-sample control of the type I error. Consequently, the corresponding adaptive separation rates are not established under rigorous finite-sample validity. In contrast, the SB-based adaptive MMD and HSIC tests proposed here are provably finite-sample valid and straightforward to implement. Moreover, since the SB minimum test uniformly dominates Bonferroni calibration in type II error, existing adaptive separation bounds for Bonferroni/MaxT procedures [2, 62, 63] transfer directly, yielding the same adaptive rates. Throughout this section, we assume that g1 , . . . , gB are drawn independently and uniformly from the collection of all permutations when implementing SB aggregation. We can also consider other groups of transformations, such as pairwise permutations, which correspond to a wild bootstrap [e.g., 62, 63] and can also be used to establish adaptivity. For simplicity, however, we focus on uniform permutations.

F.1

Adaptive MMD test i.i.d.

We start with the adaptive two-sample testing problem considered by [62]. Let (X1 , . . . , Xm ) ∼ P i.i.d.

and (Y1 , . . . , Yn ) ∼ Q be independent samples taking values in Rd , and consider testing H0 : P = Q

versus

H1 : P ̸= Q.

Let fP and fQ denote the densities of P and Q with respect to the Lebesgue measure. Separation radius. To quantify testing difficulty, we adopt the notion of a uniform separation radius. Let ∆ be a level-α test based on (X1 , . . . , Xm , Y1 , . . . , Yn ). For a function class C and constants ρ, M > 0, define the alternative class n o FρM (C) := (fP , fQ ) : max(∥fP ∥∞ , ∥fQ ∥∞ ) ≤ M, fP − fQ ∈ C, ∥fP − fQ ∥2 ≥ ρ , where ∥ · ∥2 denotes the L2 norm. For β ∈ (0, 1), the (uniform) separation radius of ∆ over C is defined as n o ρ(∆; C, M, α, β) := inf ρ > 0 : sup P(fP ,fQ ) (∆ = 0) ≤ β . (fP ,fQ )∈FρM (C)

Throughout this subsection, we write ρ(∆) for ρ(∆; C, M, α, β). For simplicity, we regard α, β ∈ (0, 1) as fixed constants but one can make the dependence on α and β explicit in the separation radius, as done in [60]. Sobolev smoothness class. Following [62], we model smooth alternatives using a Sobolev ball. For s > 0 and R > 0, define Z n o ˆ(ξ)|2 dξ ≤ (2π)d R2 , Sds (R) := f ∈ L1 (Rd ) ∩ L2 (Rd ) : ∥ξ∥2s | f 2 Rd

5

We simply refer the reader to [58] for introductory details on MMD [24] and HSIC [25].

60

where fˆ denotes the Fourier transform of f . In what follows, we take C = Sds (R).

Quadratic-time MMD statistic. To test equality of distributions, we employ the quadratic-time satisfying Gi ∈ L1 (R) ∩ L2 (R) and RMMD statistic [24]. Let G1 , . . . , Gd be one-dimensional kernels d Gi = 1. For a bandwidth vector λ = (λ1 , . . . , λd ) ∈ (0, ∞) , define the product kernel kλ (x, y) :=

d Y 1



λi

i=1

Gi

x i − yi λi

 ,

which we assume to be characteristic on Rd . The squared population MMD is MMD2λ (P, Q) := ∥µP − µQ ∥2Hk , λ

which admits an unbiased U-statistic estimator given by 2

\ λ (Xm , Yn ) := MMD

1 m(m − 1)

+

1 n(n − 1)

X

kλ (Xi , Xi′ )

1≤i̸=i′ ≤m m

X 1≤j̸=j ′ ≤n

kλ (Yj , Yj ′ ) −

n

2 XX kλ (Xi , Yj ). mn i=1 j=1

Bandwidth collection. Adaptivity with respect to the unknown smoothness parameter s is achieved by aggregating tests across a suitably chosen collection of bandwidths. Following [62, Corollary 10], we consider the dyadic grid mo  n l m+n , (28) Λ := (2−k , . . . , 2−k ) ∈ (0, ∞)d : k = 1, . . . , d2 log2 log log(m+n) and let K := |Λ|. This dyadic grid balances the bias–variance tradeoff across smoothness levels and ensures that the optimal bandwidth is approximated up to logarithmic factors. For each k ∈ [K], with associated bandwidth λ(k) ∈ Λ, define the statistic 2

\ λ(k) (Xm , Yn ). T k (Xm , Yn ) := MMD Permutation tests and SB minimum aggregation. To obtain finite-sample valid inference, we employ permutation tests. Let (U1 , . . . , Um+n ) denote the pooled sample obtained by concatenating (X1 , . . . , Xm ) and (Y1 , . . . , Yn ). For each permutation g of {1, . . . , m + n}, define the permuted g g g g n k samples Xm = (Ug(i) )m i=1 and Yn = (Ug(m+j) )j=1 , and let T (Xm , Yn ) be the corresponding statistic. Let g1 , . . . , gB be i.i.d. uniform permutations and set g0 to be the identity. For b ∈ [B]0 and k ∈ [K], write gb Tbk := T k (Xm , Yngb ).

We form permutation p-values by ranking each Tbk within its permutation distribution, B

 p Tbk :=

 1 X 1 Tik ≥ Tbk , B+1 i=0

61

and aggregate them row-wise using the minimum merging function  fbmin := min p Tbk , b ∈ [B]0 . k∈[K]

The SB minimum-aggregated p-value is then defined as pSB,MMD :=

B  1 X 1 fbmin ≤ f0min . B+1 b=0

Under H0 , the vectors (Tb1 , . . . , TbK ), b ∈ [B]0 , are row-wise exchangeable. Since the permutation p-value map and the merging function are applied identically to each row, the aggregated values (f0min , . . . , fBmin ) are exchangeable, and hence pSB,MMD is super-uniform (Theorem 1). The resulting SB minimum test rejects H0 when pSB,MMD ≤ α.

Adaptive separation rate. We now analyze the testing power of the SB minimum test. By Theorem 2, the SB minimum test, denoted by ∆SB,MMD , enjoys uniform dominance in type II error over the Bonferroni-corrected MMD procedure constructed on the same bandwidth grid. For each k ∈ [K] with λ(k) ∈ Λ, let ∆λ(k) denote the level-(α/K) permutation test based on the statistic 2 \ λ(k) , and define the Bonferroni-aggregated test as MMD 



∆Bonf,MMD := 1 min pλ(k) ≤ α/K . k∈[K]

Under the conditions of [62, Theorem 6] and for a sufficiently large B, the permutation MMD test ∆λ with a fixed bandwidth λ satisfies the separation bound 2

ρ(∆λ ) ≲

d X

λ2s i +

i=1

log(1/α) √ . (m + n) λ1 · · · λd

Applying this bound at level α/K and invoking the uniform power dominance of SB aggregation yields ( d ) X (k) log(K/α) 2 2 2s q ρ(∆SB,MMD ) ≤ ρ(∆Bonf,MMD ) ≲ min (λi ) + . k∈[K] (k) (k) i=1 (m + n) λ1 · · · λd ∗

Balancing the upper bounds in terms of λ(k) , we choose λ(k ) = (2−k , . . . , 2−k ) with  m l 2 m+n ∗ k = log2 . 4s + d log log(m + n) Substituting this choice into the above bound yields 2

ρ(∆SB,MMD ) ≲



log log(m + n) m+n

 4s

4s+d

,

which coincides with the adaptive separation rate established in [62, Corollary 10]. 62

F.2

Adaptive HSIC test

We consider adaptive independence testing via the Hilbert–Schmidt Independence Criterion (HSIC), building on the framework of [2] and its permutation extension in [63]. Let {(Xi , Yi )}ni=1 be i.i.d. observations in Rdx × Rdy with joint distribution PXY and marginals PX and PY . We test versus

H0 : PXY = PX ⊗ PY

H1 : PXY ̸= PX ⊗ PY .

HSIC statistic. Let k and ℓ be characteristic kernels on Rdx and Rdy , respectively, and define the  product kernel κ (x, y), (x′ , y ′ ) := k(x, x′ ) ℓ(y, y ′ ). The population HSIC [25] is given by  HSICk,ℓ (PXY ) = MMD2κ PXY , PX ⊗ PY , and vanishes if and only if PXY = PX ⊗ PY . An unbiased estimator takes the form of a fourth-order U-statistic [2, 34, 63]: \k,ℓ (Zn ) = HSIC

1 n(n − 1)(n − 2)(n − 3)

X

hHSIC k,ℓ (Zi , Zj , Zr , Zs ),

(i,j,r,s)∈i4n

where Zi = (Xi , Yi ) and i4n denotes the set of all 4-tuples of distinct indices from {1, . . . , n}. For a kernel s on a Euclidean space, define hMMD (a1 , a2 ; a3 , a4 ) := s(a1 , a2 ) − s(a1 , a4 ) − s(a2 , a3 ) + s(a3 , a4 ), s and set hHSIC k,ℓ (z1 , z2 , z3 , z4 ) =

1 MMD (y1 , y2 ; y3 , y4 ), (x1 , x2 ; x3 , x4 ) hMMD h ℓ 4 k

with zi = (xi , yi ). Bandwidth collection. We consider translation-invariant product kernels indexed by bandwidth parameters, analogously to the adaptive MMD construction. Let R G1 , . . R. , Gdx and L1 , . . . , Ldy be one-dimensional kernels satisfying Gi , Lj ∈ L1 (R) ∩ L2 (R) and Gi = Lj = 1. For bandwidth vectors λ ∈ (0, ∞)dx and µ ∈ (0, ∞)dy , define ′

kλ (x, x ) :=

dx Y 1 i=1

λi

 Gi

xi − x′i λi



ℓµ (y, y ) :=

,

dy Y 1 j=1

µj

 Lj

yj − yj′ µj

 .

We assume that kλ and ℓµ are characteristic, ensuring that the associated HSIC detects independence.6 The Gaussian kernels employed in [2] correspond to a special case of this family, and the analysis of [63, Theorem 3] applies to the general product-kernel setting. Consider the dyadic bandwidth collection formed by jointly varying the X- and Y -bandwidths. Specifically, we define Λ as the collection of concatenated bandwidth pairs (λ, µ) ∈ (0, ∞)dx +dy of the form  mo n l 2 n log , (29) Λ := (2−k , . . . , 2−k ) ∈ (0, ∞)dx +dy : k = 1, . . . , dx +d 2 log log n y 6

Following [23], it is sufficient to assume that G1 , . . . , Gdx and L1 , . . . , Ldy are characteristic.

63

which coincides with the collection used in [63, Theorem 3]. For each k ∈ [K] with K = |Λ|, let (λ(k) , µ(k) ) denote the k-th bandwidth pair in Λ, and define \k ,ℓ (Zn ). T k (Zn ) := HSIC (k) (k) λ

µ

Permutation tests and SB aggregation. Under H0 , the joint distribution is invariant under permutations of the Y -coordinates relative to the X-coordinates. Let g1 , . . . , gB be i.i.d. uniform permutations of {1, . . . , n}, with g0 the identity. Define n  \k ,ℓ Xi , Ygb (i) i=1 , b ∈ [B]0 . Tbk := HSIC (k) (k) λ

µ

The associated permutation p-values are B

 p Tbk :=

 1 X 1 Tik ≥ Tbk . B+1 i=0

Aggregating row-wise using the minimum merging function yields  fbmin := min p Tbk , k∈[K]

pSB,HSIC :=

B  1 X 1 fbmin ≤ f0min . B+1 b=0

The resulting test ∆SB,HSIC rejects H0 whenever pSB,HSIC ≤ α. By row-wise exchangeability under H0 , pSB,HSIC is super-uniform (Theorem 1), and hence ∆SB,HSIC is a valid level-α test. Adaptive separation rate. Recall that ρ(∆λ,µ ) denotes the uniform separation radius of the level α HSIC test with bandwidth (λ, µ), defined with respect to the alternative class FρM Sdsx +dy (R) . Let fXY denote the joint density of (X, Y ) and let fX ⊗ fY denote the product of the marginal densities. In this formulation, separation from the null corresponds to the L2 distance between fXY and fX ⊗ fY , with Sobolev smoothness imposed on this density difference through Sdsx +dy (R). Under the conditions of [63, Theorem 3] and for a sufficiently large B, the HSIC permutation test with fixed bandwidth (λ, µ) satisfies ρ(∆λ,µ )2 ≲

dx X

λ2s i +

i=1

dy X

log(1/α) log(1/α)3/2 p µ2s + + . i 3/2 λ · · · λ µ · · · µ n µ · · · µ n λ · · · λ 1 1 d d 1 1 d d x y x y i=1

By the uniform power dominance of SB aggregation (Theorem 2), the SB-aggregated HSIC test satisfies ( d dy x X X (k) 2s (k) 2 ρ(∆SB,HSIC ) ≲ min (λi ) + (µi )2s k∈[K]

i=1

i=1

log(K/α)

log(K/α)3/2

)

+ q + . (k) (k) (k) (k) (k) (k) (k) (k) n3/2 λ1 · · · λdx µ1 · · · µdy n λ1 · · · λdx µ1 · · · µdy Since the dyadic grid satisfies K ≍ log n, the additional Bonferroni factor contributes only a log log n ∗ ∗ term. Balancing the bias and variance terms and assuming 4s ≥ dx + dy ,7 choosing λ(k ) = µ(k ) = 7

The condition 4s ≥ dx + dy ensures that the term involving n−3/2 does not dominate at the optimal bandwidth.

64

(2−k , . . . , 2−k ) with k∗ =

l

m  2 log2 log nlog n 4s + dx + dy

yields 2

ρ(∆SB,HSIC ) ≲



log log n n



4s 4s+dx +dy

,

which matches the adaptive minimax separation rate for independence testing up to the standard iterated logarithmic factor.

G

Technical lemmas and auxiliary tools

This section presents several useful facts on empirical quantiles and p-values that are used throughout the proofs. We begin by collecting standard facts relating empirical quantiles, rank-based p-values, and their equivalence under exchangeability. Here and throughout, Z(1) ≤ · · · ≤ Z(n) denote the order statistics, with the convention Z(0) = −∞ and Z(n+1) = +∞. Lemma S.4 (Lemmas S.14 and S.16, [38]). Let α ∈ [0, 1]. Let Z1 , . . . , Zn ∈ R. Then, for each i ∈ [n], n

 1X 1(Zj ≥ Zi ) ≤ α ⇐⇒ Zi > Quantile1−α Z1 , . . . , Zn . n j=1

Furthermore when Z1 , . . . , Zn are exchangeable, it holds that     X n  1 1(Zj ≥ Zi ) ≤ α = P Zi > Quantile1−α Z1 , . . . , Zn ≤ α. P n j=1

Lemma S.5 (Fact 2.10, [3]). Let α ∈ [0, 1]. For any Z1 , . . . , Zn ∈ R,  Quantileα Z1 , . . . , Zn = Z(⌈αn⌉) . Lemma S.6 (Lemma 3.4, [3]). Let α ∈ [0, 1]. For any Z1 , . . . , Zn , Zn+1 ∈ R,   Zn+1 ≤ Quantileα Z1 , . . . , Zn , Zn+1 ⇐⇒ Zn+1 ≤ Quantileα(1+1/n) Z1 , . . . , Zn . Lemma S.7. For any Z1 , . . . , Zn ∈ R and α ∈ [0, 1], the following holds:     n n 1X 1X sup t : 1(Zi ≤ t) ≤ α = sup t : 1(Zi < t) ≤ α . n n i=1

i=1

Moreover, both sides are equal to Z(⌊nα⌋+1) where Z(1) ≤ Z(2) ≤ · · · ≤ Z(n) are the order statistics of Z1 , . . . , Zn . Proof. See Appendix I.7. 65

The following lemma establishes a new general super-uniformity property for weighted rank functionals on arbitrary measure spaces. Lemma S.8. Let (Ω, F, µ) be a measure R space. Let t : Ω → R := R ∪ {−∞, +∞} be measurable and w : Ω → [0, ∞) be measurable with Ω w dµ < ∞. Then for every α ≥ 0, Z  Z ′ ′ ′ w(ω) 1 w(ω ) 1{t(ω ) ≥ t(ω)} dµ(ω ) ≤ α dµ(ω) ≤ α. Ω

Proof. See Appendix I.8. The above lemma immediately implies the following corollary, which corresponds to [28, Lemma A1]. Corollary S.3. Let t0 , . . . , tn ∈ [−∞, ∞], w0 , . . . , wn ∈ [0, ∞), and α ≥ 0. Then ( n ) n X X wk 1 wi 1{ti ≥ tk } ≤ α ≤ α. i=0

k=0

Proof. Apply Lemma S.8 with Ω = {0, 1, . . . , n}, µ the counting measure, t(k) = tk , and w(k) = wk .

H

Proofs of the results in the main text

This section collects the proofs of the results stated in the main text.

H.1

Proof of Theorem 1

Under the group-invariance hypothesis, the row vectors (T01 , . . . , T0K ), . . . , (TB1 , . . . , TBK ) are exchangeable. Since the p-value map (T0k , . . . , TBk ) 7→ (p(T0k ), . . . , p(TBk )) and the merging map (p(Tb1 ), . . . , p(TbK )) 7→ fb are applied identically for each index b ∈ [B]0 , the aggregated values (f0 , . . . , fB ) are exchangeable as well. Hence the rank p-value B

1 X pSB = 1(fb ≤ f0 ) B+1 b=0

is super-uniform by Lemma S.4, which yields P(pSB ≤ α) ≤ α. If, moreover, f0 , . . . , fB are distinct a.s., then pSB is uniform on the grid {1/(B +1), . . . , 1}, and therefore P(pSB ≤ α) = ⌊(B +1)α⌋/(B + 1).

H.2

Proof of Proposition 1

We start from the definition of the SB threshold, ûSB α = −Quantile1−α {−f0 , . . . , −fB }, and expand 1 PB the empirical quantile using our convention Quantileq {x0 , . . . , xB } = inf{t : B+1 b=0 1(xb ≤ t) ≥ q}. Then ûSB α = − inf

n t:

B

o 1 X 1(−fb ≤ t) ≥ 1 − α B+1 b=0

66

n = − inf t : n = sup u : n = sup u : n = sup u :

B

o 1 X 1(−fb > t) ≤ α B+1

1 B+1 1 B+1 1 B+1

b=0 B X b=0 B X b=0 B X b=0

1(−fb > −u) ≤ α 1(fb < u) ≤ α

o

o

o 1(fb ≤ u) ≤ α ,

where the second line uses the identity 1(x ≤ t) = 1 − 1(x > t), the third line is the change of variables u = −t, and the last line follows from Lemma S.7 (which shows that replacing ≤ by < inside the supremum does not change its value). This completes the proof.

H.3

Proof of Lemma 1

We begin by noting that the second inequality is an immediate consequence of a general weighted result due to [28, Lemma A1] recalled in Corollary S.3. Specifically, for arbitrary T0 , . . . , TB ∈ [−∞, ∞] and nonnegative weights w0 , . . . , wB , that result implies X  B B X wi 1 wj 1(Tj ≥ Ti ) ≤ α ≤ α. i=0

j=0

Specializing to uniform weights w0 = · · · = wB = 1/(B +1) yields the second claim of Lemma 1. We therefore concentrate on establishing the first inequality, for which we provide a direct and sharper argument in the PB P uniform-weights setting. Let Rj = B j=0 1(Rj ≤ k) ≤ k for all i=0 1(Ti ≥ Tj ) so that pj = Rj /(B + 1). We claim that k ∈ [B]0 . To see this, write T(0) ≤ T(1) ≤ T(2) ≤ . . . ≤ T(B) for the order statistics and observe B X j=0

X  X X  B B B B X 1(Rj ≤ k) = 1 1(Ti ≥ Tj ) ≤ k = 1 1(Ti ≥ T(j) ) ≤ k . j=0

i=0

j=0

i=0

Now note that for each j ∈ [B]0 , B X i=0

1(Ti ≥ T(j) ) ≥ B − j + 1.

Therefore B X j=0

1(Rj ≤ k) ≤

B X j=0



1 B−j+1≤k =

B X j=0



1 B−k+1≤j =

B X

1 = k,

j=B−k+1

as desired. To complete the proof, observe that B X j=0

1(pj ≤ α) =

B X j=0

1(Rj ≤ (B + 1)α) = 67

B X j=0

1(Rj ≤ ⌊(B + 1)α⌋) ≤ ⌊(B + 1)α⌋.

Consequently, we prove the first claim that B

1 X ⌊(B + 1)α⌋ 1(pj ≤ α) ≤ ≤ α. B+1 B+1 j=0

When T0 , . . . , TB are distinct, we have B X j=0

1(Rj ≤ ⌊(B + 1)α⌋) = ⌊(B + 1)α⌋,

which proves the second claim. This completes the proof of Lemma 1.

H.4

Proof of Theorem 2

Fix the observed data X and transformations g0 , . . . , gB with g0 being the identity. For each b ∈ [B]0 , define  fb = f p(Tb1 ), . . . , p(TbK ) . By Lemma 1, the p-values p(Tb1 ), . . . , p(TbK ) are super-uniform conditional on X, with randomness arising only through b ∼ Unif([B]0 ). Hence, by the defining property of cα,K , B

1 X 1(fb ≤ cα,K ) ≤ α. B+1

(30)

b=0

1 PB Let F (u) := B+1 b=0 1(fb ≤ u) denote the empirical distribution function of {f0 , . . . , fB }. Since F is right-continuous and non-decreasing from 0 to 1, the alternative characterization of ûSB α in SB almost ) > α, while (30) gives F (c ) ≤ α. Therefore c < û Proposition 1 implies F (ûSB α,K α,K α α surely, which implies the claimed bound on the type II error.

H.5

Proof of Corollary 1

By Theorem 2, the SB aggregation test calibrated at level α uniformly dominates any deterministic test obtained by worst-case calibration of the same merging function. For the O- and M-families, the prior work [e.g., 73] shows that the corresponding merging functions produce super-uniform aggregated p-values under arbitrary dependence, so that the worst-case calibration constant satisfies cα,K = α. Combining these facts yields the stated uniform power dominance.

H.6

Proof of Proposition 2

Recall that the merging function f satisfies the diagonal monotonicity condition (8), namely f (x, . . . , x) ≤ f (y, . . . , y)

⇐⇒

x ≤ y.

Under the rank alignment condition (7), the vectors (T0k , . . . , TBk ) induce the same ordering for all k ∈ [K]. Consequently, the permutation p-values satisfy p(Tb1 ) = · · · = p(TbK ) =: pb , 68

b ∈ [B]0 .

Therefore, for each b ∈ [B]0 ,  f p(Tb1 ), . . . , p(TbK ) = f (pb , . . . , pb ), and by diagonal monotonicity, for any b, b′ ∈ [B]0 , f (pb , . . . , pb ) ≤ f (pb′ , . . . , pb′ )

⇐⇒

pb ≤ pb′ .

Using the p-value representation of the SB test, we obtain B   1 X  pSB = 1 f p(Tb1 ), . . . , p(TbK ) ≤ f p(T01 ), . . . , p(T0K ) B+1

=

1 B+1

b=0 B X b=0

1(pb ≤ p0 ).

Finally, recall that by definition of the permutation p-values, B

pb =

1 X 1(Tik ≥ Tbk ) B+1 i=0

for any fixed k ∈ [K]. Since the ordering of Tik does not depend on k under (7), the right-hand side is the same for all k. Hence pb ≤ p0 if and only if Tbk ≥ T0k for any k ∈ [K]. In particular, fixing an ⋆ arbitrary representative coordinate k ⋆ ∈ [K] and writing Tb := Tbk for all b ∈ [B]0 , we obtain the equivalent condition Tb ≥ T0 . Combining the above identities yields B

pSB =

1 X 1(Tb ≥ T0 ), B+1 b=0

which coincides with the usual permutation test based on a single statistic. This completes the proof.

H.7

Proof of Proposition 3

The argument follows the standard analysis of empirical quantiles for a permutation distribution (e.g., [39, Theorem 17.2.3]). The only difference is that, instead of enumerating all transformations in a finite group, we work with statistics computed from i.i.d. random transformations. We therefore verify directly that the Monte Carlo empirical distribution converges to the limiting null distribution, and then apply a standard quantile-continuity argument. Proof. For each (B, n), define the empirical distribution functions B

F̂B,n (t) :=

1 X 1(fb,n ≤ t), B+1

B

F̃B,n (t) :=

b=0

1 X 1(fb,n ≤ t), B b=1

Since F̂B,n (t) =

B 1 1(f0,n ≤ t), F̃B,n (t) + B+1 B+1 69

t ∈ R.

we have the uniform bound sup F̂B,n (t) − F̃B,n (t) ≤ t∈R

1 . B+1

Thus it suffices to establish convergence of F̃B,n , and then transfer it to F̂B,n and to the corresponding empirical quantiles. Step 1: Pointwise convergence of the empirical CDF. Fix a distribution P and a continuity point t of F∞,P (t) := PP (V∞,P ≤ t). Consider B

F̃B,n (t) =

1 X 1(fb,n ≤ t). B b=1

Since (f1,n , . . . , fB,n ) are exchangeable (under both the null and the alternative),   EP F̃B,n (t) = PP (f1,n ≤ t) =: pn,P (t). Moreover, B B      1 X 1 X  EP 1(fb,n ≤ t) + 2 EP 1(fb,n ≤ t)1(fb′ ,n ≤ t) EP F̃B,n (t)2 = 2 B B ′ b=1

=

b,b =1 b̸=b′

B−1 1 pn,P (t) + qn,P (t), B B

where qn,P (t) := PP (f1,n ≤ t, f2,n ≤ t). Therefore,  B−1   1 VarP F̃B,n (t) = pn,P (t) − pn,P (t)2 + qn,P (t) − pn,P (t)2 . B B d

Recall that t is a continuity point of F∞,P . By the assumed joint convergence (f1,n , f2,n ) −→ ′ ′ (V∞,P , V∞,P ) with V∞,P , V∞,P i.i.d., the Portmanteau theorem applied to the continuity sets (−∞, t] 2 and (−∞, t] yields qn,P (t) → F∞,P (t)2 ,

pn,P (t) → F∞,P (t),

so qn,P (t) − pn,P (t)2 → 0. Hence VarP (F̃B,n (t)) → 0 as B, n → ∞. Chebyshev’s inequality gives p

F̃B,n (t) − pn,P (t) −→ 0,

and thus

p

F̃B,n (t) −→ F∞,P (t).

Step 2: Pointwise quantile convergence. Recall that the SB threshold is n o ûSB Q⋆α,P := inf{u ∈ R : F∞,P (u) ≥ α}. α = sup u ∈ R : F̂B,n (u) ≤ α , Fix P . By the assumption, for every ε > 0, F∞,P (Q⋆α,P − ε) < α and F∞,P (Q⋆α,P + ε) > α. 70

Using Step 1 (pointwise) and the O(B −1 ) bound between F̂B,n and F̃B,n , p

p

F̂B,n (Q⋆α,P − ε) −→ F∞,P (Q⋆α,P − ε),

F̂B,n (Q⋆α,P + ε) −→ F∞,P (Q⋆α,P + ε),

where Q⋆α,P − ε and Q⋆α,P + ε are assumed to be continuity points of F∞,P without loss of generality. Consequently,     PP F̂B,n (Q⋆α,P + ε) > α → 1. PP F̂B,n (Q⋆α,P − ε) ≤ α → 1, p

⋆ SB ⋆ On the intersection of these events we have Q⋆α,P − ε ≤ ûSB α ≤ Qα,P + ε, hence ûα −→ Qα,P .

H.8

Proof of Proposition 4

For each k ∈ [K] and b ∈ [B]0 , define B

pkb := p(Tbk ) =

1 X 1(Tik ≥ Tbk ), B+1

fb := fbmin = min pkb .

i=0

k∈[K]

Recall from Proposition 1 that the SB critical value for the minimum merge satisfies n ûSB = sup u∈R: α,min

B o 1 X 1(fb ≤ u) ≤ α . B+1

(31)

b=0

All statements below are deterministic conditional on the realized matrix {Tbk }b,k , hence they hold almost surely. Step 1: Bonferroni rejection implies SB rejection. For the minimum merge f min (u1 , . . . , uK ) = mink∈[K] uk , the Bonferroni bound implies that the deterministic worst-case calibration constant satisfies cα,K = α/K. Hence, by Theorem 2, the SB critical value obeys ûSB α,min > cα,K = α/K k k SB almost surely. Therefore, mink p(T0 ) ≤ α/K implies mink p(T0 ) < ûα,min , proving the claim. Step 2: SB rejection implies unadjusted minimum rejection. We prove the contrapositive. Suppose that f0 = mink p(T0k ) > α, and let k⋆ ∈ [K] be an index attaining the minimum, i.e., p(T0k⋆ ) = f0 . Consider the set of transformations whose k⋆ -statistic is at least as large as the observed one, A := {b ∈ [B]0 : Tbk⋆ ≥ T0k⋆ }. By definition of the permutation p-value, |A| = (B + 1) p(T0k⋆ ) = (B + 1)f0 . Moreover, for any b ∈ A, the statistic Tbk⋆ is more extreme than T0k⋆ , so its permutation p-value in coordinate k⋆ cannot be larger: p(Tbk⋆ ) ≤ p(T0k⋆ ) = f0 . Since fb = mink p(Tbk ) ≤ p(Tbk⋆ ), we obtain fb ≤ f0 for every b ∈ A. Consequently, B

1 X |A| 1(fb ≤ f0 ) ≥ = f0 > α. B+1 B+1 b=0

71

Thus u = f0 is not feasible in (31), which implies ûSB α,min ≤ f0 . In particular, under f0 > α the event SB f0 < ûα,min is impossible, proving     k SB k 1 min p(T0 ) < ûα,min ≤ 1 min p(T0 ) ≤ α . k

k

Step 3: Lower bound on ûSB α,min . By Theorem 2, we have ûSB α,min >

α . K

Since each permutation p-value lies on the grid   1 2 , ,...,1 , B+1 B+1 the merged values fbmin also lie on this grid. By the order-statistic representation of the SB threshold, ûSB α,min belongs to the same grid. Therefore, the smallest grid value strictly larger than α/K is ⌊(B + 1)α/K⌋ + 1 , B+1 which implies ûSB α,min ≥

⌊(B + 1)α/K⌋ + 1 . B+1

Example S.1. We construct a finite-sample example in the standard randomization setup (statistics computed from transformed data) such that     α k k SB 1 min p(T0 ) ≤ = 1 min p(T0 ) < ûα,min almost surely. (32) K k∈[K] k∈[K] Step 1: Parameters. Fix integers B ≥ 1, K ≥ 2, and α ∈ (0, 1), and define   (B + 1)α . m := K Assume (B + 1)α/K ∈ / Z so that m α m+1 < < . B+1 K B+1 Step 2: Randomization setup. Let the data vector X have size B + 1 and define X = (U0 , . . . , UB ),

i.i.d.

U0 , . . . , UB ∼ Unif(0, 1).

Thus the sample size in this construction is n = B + 1. Let π : {0, . . . , B} → {0, . . . , B} be the cyclic permutation π(j) := j + 1 (mod B + 1). For each b ∈ {0, . . . , B}, define gb as the coordinate permutation induced by the b-fold composition π b (i.e., π applied b times), namely (gb (X))j := Xπb (j)

for all j ∈ {0, . . . , B}. 72

In particular, (gb (X))0 = Ub , and therefore {(gb (X))0 : b = 0, . . . , B} = {U0 , . . . , UB }. Step 3: Coordinate statistics. Let R1 , . . . , RK ⊂ {1, . . . , B +1} be disjoint sets of size m. Define the rank of the first coordinate by r(x) := 1 +

B X j=0

1{xj < x0 }.

For each k ∈ [K], define T k (x) := 1{r(x) ∈ Rk },

Tbk := T k (gb (X)).

Since the Ub are continuous, all ranks are distinct almost surely. Step 4: Permutation p-values. For each k and b, define the permutation p-value B

p(Tbk ) =

1 X 1(Tik ≥ Tbk ). B+1 i=0

Because exactly m ranks belong to Rk , we have almost surely  m  , r(gb (X)) ∈ Rk , k p(Tb ) = B + 1 1, otherwise. Step 5: Minimum merged values. Define fbmin := min p(Tbk ). k∈[K]

Then almost surely  m S  , r(gb (X)) ∈ K k=1 Rk , min fb = B + 1 1, otherwise. Since the sets Rk are disjoint,   B 1 X m Km min 1 fb ≤ = ≤ α, B+1 B+1 B+1 b=0

whereas B

 1 X 1 fbmin ≤ 1 = 1 > α. B+1 b=0

73

Step 6: SB threshold and equivalence. By definition, n u:

ûSB α,min = sup

B

o 1 X 1(fbmin ≤ u) ≤ α = 1. B+1 b=0

Therefore, k min p(T0k ) < ûSB α,min ⇐⇒ min p(T0 ) = k

k

m . B+1

m < α/K < 1, we conclude that Using B+1

min p(T0k ) ≤ k

α K

min p(T0k ) < ûSB α,min ,

⇐⇒

k

which proves (32).

H.9

Proof of Proposition 5

This is the special case of Proposition S.2 obtained by taking J = K, Ij = [j], and hj (x1 , . . . , xj ) = min1≤k≤j xk .

H.10

Proof of Theorem 3

Condition on the reference batch {(T̃i1 , . . . , T̃iK )}B i=1 . Under the group-invariance hypothesis, the testing-batch row vectors (T01 , . . . , T0K ), . . . , (TB1 , . . . , TBK ) are exchangeable given the reference batch. Since the holdout p-value map pHO (·) and the merging function f are applied row-wise and identically across b ∈ [B]0 , the merged values (f˜0 , . . . , f˜B ) are conditionally exchangeable as well. Therefore the rank p-value B

1 X ˜ pTB = 1(fb ≤ f˜0 ) B+1 b=0

is conditionally super-uniform by Lemma S.4, implying P(pTB ≤ α) ≤ α after removing the conditioning. If f˜0 , . . . , f˜B are distinct a.s., then pTB is uniform on the grid {1/(B +1), . . . , 1} conditional on the reference batch, and hence P(pTB ≤ α) = ⌊(B + 1)α⌋/(B + 1).

H.11

Proof of Proposition 6

Fix the split with |Iref | = n1 and write u := uTB α . Under the absolute residual score sk (x, y) = k |y − µ bk (x)|, the reference residuals are Rj := |Yj − µ bk (Xj )| for j ∈ Iref , and for any candidate label y ∈ R the reference-based p-value (15) becomes n o P 1 + j∈Iref 1 Rjk ≥ |y − µ bk (Xn+1 )| k,ref Pn+1 (y) = . n1 + 1 74

With the minimum merge f (u1 , . . . , uK ) = mink∈[K] uk , we have k,ref Mn+1 (y) = min Pn+1 (y). k∈[K]

By definition of the TB set in (16), y ∈ CTB (Xn+1 )

⇐⇒

Mn+1 (y) ≥ u

⇐⇒

k,ref Pn+1 (y) ≥ u for all k ∈ [K].

TB If uTB α = −∞, we trivially have CTB (Xn+1 ) = Y = R, so assume uα is finite for the remainder of k,ref the proof. Now fix k ∈ [K] and set r := |y − µ bk (Xn+1 )|. The condition Pn+1 (y) ≥ u is equivalent to

X

1+

j∈Iref

1{Rjk ≥ r} ≥ u(n1 + 1).

Since the left-hand side is integer-valued, this is equivalent to X 1{Rjk ≥ r} ≥ ⌈u(n1 + 1)⌉ − 1. j∈Iref

Let so that

m := ⌈u(n1 + 1)⌉ − 1,

ℓ := n1 − m + 1 = n1 + 2 − ⌈u(n1 + 1)⌉ ,

which matches the definition in the proposition. The inequality above therefore says that at least m reference residuals are ≥ r, i.e., at most n1 − m = ℓ − 1 reference residuals are < r. Equivalently, r is no larger than the ℓ-th smallest reference residual: X k 1{Rjk ≥ r} ≥ m ⇐⇒ r ≤ R(ℓ) , j∈Iref

k with the convention R(n = ∞ covering the case m = 0. Thus, y ∈ CTB (Xn+1 ) is equivalent to 1 +1) the following condition holding for all k ∈ [K]: k,ref (y) ≥ u Pn+1

⇐⇒ ⇐⇒

k |y − µ bk (Xn+1 )| ≤ R(ℓ) h i k k y∈ µ bk (Xn+1 ) − R(ℓ) , µ bk (Xn+1 ) + R(ℓ) .

Intersecting over k = 1, . . . , K yields CTB (Xn+1 ) =

K h i \ k k µ bk (Xn+1 ) − R(ℓ) , µ bk (Xn+1 ) + R(ℓ) ,

k=1

as claimed.

I

Proofs of the results in the supplementary material

This section collects the proofs of the results stated in the supplementary material. 75

I.1

Finite-sample limitation of Monte Carlo MaxT calibration

Consider the simplified setting K = 1. For notational convenience, write Tb = Tbk for b = 0, 1, . . . , 2B and k = 1. In this case, the critical value ũα in (3) reduces to ( ) B   1 X  ũα = sup u ∈ (0, 1) : 1 TB+b > Quantile1−u T0 , T1 , . . . , TB ≤α . B b=1

The resulting MaxT test (2) becomes equivalent to rejecting the null hypothesis when  T0 > Quantile1−ũα T0 , T1 , . . . , TB .

(33)

The same intuition applies when the coordinate-wise statistics are perfectly rank-aligned, so that the coordinate-wise rejection events coincide. Proposition S.8. Suppose that T0 , T1 , . . . , T2B are exchangeable and take distinct values with probability one. Fix α ∈ (0, 1). Then the type I error probability of the test in (33) satisfies  ⌊Bα⌋ + 1 . P T0 > Quantile1−ũα {T0 , . . . , TB } = B+1 Proof. Let S := {T0 , . . . , TB } denote the calibration set and H := {TB+1 , . . . , T2B } the holdout set. Then B

B

b=1

b=1

X 1 X 1{TB+b > Quantile1−u (S)} ≤ ⌊Bα⌋. 1{TB+b > Quantile1−u (S)} ≤ α ⇐⇒ B Set k = ⌊Bα⌋. The event above is equivalent to stating that at most k holdouts exceed the (1 − u)quantile of S. This is the same as stating that at least B − k holdouts are less than or equal to this quantile. This is equivalent to H(B−k) ≤ Quantile1−u (S), where H(B−k) denotes the (B − k)-th order statistic of the holdout set H. The definition of ũα in (3) (with K = 1 and wk = 1) is thus  ũα = sup u ∈ (0, 1) : H(B−k) ≤ Quantile1−u (S) . Let S(1) ≤ S(2) ≤ . . . ≤ S(B+1) denote the order statistics of S. By Lemma S.5, we have Quantile1−u (S) = S(⌈(1−u)(B+1)⌉) and define S(0) = −∞ by convention. We partition the sample space into B + 2 disjoint events: E1 = {H(B−k) ≤ S(1) },

Er = {S(r−1) < H(B−k) ≤ S(r) }, EB+2 = {S(B+1) < H(B−k) }.

We now compute ũα on each event:

76

for r ∈ {2, . . . , B + 1},

• On E1 , H(B−k) ≤ S(1) . For any u ∈ (0, 1), we have ⌈(1 − u)(B + 1)⌉ ≥ 1, so Quantile1−u (S) = S(⌈(1−u)(B+1)⌉) ≥ S(1) . Thus, H(B−k) ≤ Quantile1−u (S) holds for all u ∈ (0, 1). The supremum is ũα = 1. This gives Quantile1−ũα (S) = Quantile0 (S) = S(⌈0·(B+1)⌉) = S(0) = −∞. • On Er for r ∈ {2, . . . , B + 1}, we have S(r−1) < H(B−k) ≤ S(r) . The condition H(B−k) ≤ Quantile1−u (S) = S(⌈(1−u)(B+1)⌉) holds if and only if r ≤ ⌈(1 − u)(B + 1)⌉. This inequality r−1 is equivalent to r − 1 < (1 − u)(B + 1), or u < 1 − B+1 . Therefore, the supremum is ũα = r−1 B+2−r 1 − B+1 = B+1 . This gives Quantile1−ũα (S) = Quantile r−1 (S) = S(⌈ r−1 (B+1)⌉) = S(r−1) . B+1

B+1

• On EB+2 , we have S(B+1) < H(B−k) . Since Quantile1−u (S) ≤ S(B+1) for all u ∈ (0, 1), the condition H(B−k) ≤ Quantile1−u (S) never holds. The set of valid u is empty. By convention for this type of test, we define the supremum over an empty set (as a subset of [0, 1]) to be 0. Thus, ũα = 0. This gives Quantile1−ũα (S) = Quantile1 (S) = S(⌈1·(B+1)⌉) = S(B+1) . Using these calculations, we can express the type I error probability as X  B+2  P T0 > Quantile1−ũα (S) = P {T0 > Quantile1−ũα (S)} ∩ Er =

=

r=1 B+2 X

r=1 B+2 X r=1

P T0 > S(r−1) ∩ Er



 P T0 > S(r−1) | Er P(Er ).

Let Ar = {T0 > S(r−1) }. We claim that Ar and Er are independent; the proof is deferred to Lemma S.9. Using this independence, we have X  B+2  P T0 > Quantile1−ũα (S) = P T0 > S(r−1) P(Er ). r=1

As shown above, we have the identity  B+2−r . P T0 > S(r−1) = B+1 Substituting this in, the type I error is B+2

1 X (B + 2 − r)P(Er ). B+1 r=1

PB Note that Er occurs if and only if V := b=0 1{Tb < H(B−k) } = r − 1 under the distinctness assumption. By re-indexing the sum using v = r − 1, we have B+2

B+1

r=1

v=0

1 X 1 X (B + 2 − r)P(Er ) = (B + 1 − v)P(V = v) B+1 B+1 77

1 = B+1

B+1 X

B+1 X

v=0

v=0

(B + 1)P(V = v) −

! vP(V = v)

B + 1 − E[V ] 1 ((B + 1) · 1 − E[V ]) = . B+1 B+1 PB P By linearity of expectation, E[V ] = B b=0 P(Tb < H(B−k) ). By exb=0 E[1{Tb < H(B−k) }] = changeability of T0 , . . . , TB , all these probabilities are equal: =

E[V ] = (B + 1)P(T0 < H(B−k) ). To compute P(T0 < H(B−k) ), consider the set U = {T0 } ∪ H = {T0 , TB+1 , . . . , T2B }. This set has B + 1 exchangeable elements. Let R0 be the rank of T0 within U . By exchangeability, R0 ∼ Uniform{1, . . . , B + 1}. The event T0 < H(B−k) means that T0 is smaller than at least k + 1 elements of H (namely, H(B−k) , . . . , H(B) ). This implies that the rank R0 of T0 in U can be at most (B + 1) − (k + 1) = B − k. Thus, P(T0 < H(B−k) ) = P(R0 ≤ B − k) = B−k B+1 . Substituting this into the expression for E[V ]: E[V ] = (B + 1) ·

B−k = B − k. B+1

Finally, the type I error is  B + 1 − E[V ] B + 1 − (B − k) k+1 P T0 > Quantile1−ũα (S) = = = . B+1 B+1 B+1 Substituting k = ⌊Bα⌋, the type I error is ⌊Bα⌋+1 B+1 . Lemma S.9. For each r ∈ {1, . . . , B + 2}, the events Ar := {T0 > S(r−1) } and Er are independent. Proof. Let GS = σ(S(1) , . . . , S(B+1) ) be the sigma-field generated by the order statistics of S, and let GH = σ(H(1) , . . . , H(B) ) be the sigma-field generated by the order statistics of H. Let J ∈ {1, . . . , B + 1} denote the rank of T0 within S. By exchangeability and distinctness, conditional on GS , J is uniformly distributed on {1, . . . , B + 1}. Moreover, block-exchangeability (i.e., invariance under permutations within each block {0, . . . , B} and {B + 1, . . . , 2B}) implies that P(J = j | GS , GH ) = P(J = j | GS ) =

1 , B+1

j = 1, . . . , B + 1.

Hence J is conditionally independent of GH given GS . Recall that for r ∈ {1, . . . , B + 2}, Ar = {T0 > S(r−1) } = {J ≥ r},

Er = {S(r−1) < H(B−k) ≤ S(r) }.

Using iterated conditioning, P(Ar ∩ Er | GS , GH ) = P(Ar | GS , GH ) 1(Er )

= P(J ≥ r | GS , GH ) 1(Er )

= P(J ≥ r | GS ) 1(Er ). 78

Taking expectations first with respect to GH conditional on GS yields P(Ar ∩ Er | GS ) = P(J ≥ r | GS ) P(Er | GS ). Finally, taking expectations over GS and using P(J ≥ r | GS ) = (B + 2 − r)/(B + 1), we obtain P(Ar ∩ Er ) =

B+2−r P(Er ) = P(Ar ) P(Er ), B+1

which proves independence.

I.2

Proof of Proposition S.4

We use the notation from Appendix H.7. For each (B, n), define B

F̂B,n (t) :=

1 X 1(fb,n ≤ t), B+1

B

F̃B,n (t) :=

b=0

1 X 1(fb,n ≤ t). B b=1

As before, supt |F̂B,n (t) − F̃B,n (t)| ≤ 1/(B + 1). Proof. Step 1’: Uniform convergence of the empirical CDF. Assume the uniform marginal and joint approximation conditions used in the uniform statement: sup sup PP (f1,n ≤ t) − FP (t) → 0,

P ∈P t∈R

sup sup PP (f1,n ≤ t, f2,n ≤ t) − FP (t)2 → 0,

P ∈P t∈R

where FP (t) = PP (V∞,P ≤ t). These can be deduced from the more general pointwise convergence condition in the statement of Proposition S.4. Fix η > 0 and write Xb,n (t) = 1{fb,n ≤ t}. For the P Monte Carlo CDF, F̃B,n (t) = B1 B b=1 Xb,n (t), exchangeability yields   B−1 1 CovP X1,n (t), X2,n (t) VarP F̃B,n (t) = VarP (X1,n (t)) + B B 1 ≤ + CovP (X1,n (t), X2,n (t)) , 4B where the 1/4 factor arises from the fact that X1,n (t) is a Bernoulli random variable, whose variance is at most 1/4. Moreover, CovP (X1,n (t), X2,n (t)) = PP (f1,n ≤ t, f2,n ≤ t) − PP (f1,n ≤ t)2

≤ PP (f1,n ≤ t, f2,n ≤ t) − FP (t)2 + 2 PP (f1,n ≤ t) − FP (t) .

Taking supP ∈P supt∈R , the right-hand side tends to 0 by assumption. Hence  sup sup VarP F̃B,n (t) → 0. P ∈P t∈R

Fix η > 0. By Chebyshev’s inequality,  sup sup PP |F̃B,n (t) − PP (f1,n ≤ t)| > η → 0.

P ∈P t∈R

79

Combining with uniform marginal convergence gives  sup sup PP |F̃B,n (t) − FP (t)| > 2η → 0.

P ∈P t∈R

Finally, since supt∈R |F̂B,n (t) − F̃B,n (t)| ≤ 1/(B + 1), the same conclusion holds with F̂B,n in place of F̃B,n . Step 2’: Uniform quantile convergence. For the uniform statement, use the assumed uniform margin condition: for each ε > 0,  δ(ε) := inf min α − FP (Q⋆α,P − ε), FP (Q⋆α,P + ε) − α > 0. P ∈P

Then, for any P ∈ P, ⋆ ⋆ {ûSB α > Qα,P + ε} ⊆ {F̂B,n (Qα,P + ε) ≤ α} n o ⊆ |F̂B,n (Q⋆α,P + ε) − FP (Q⋆α,P + ε)| ≥ δ(ε) ,

and similarly ⋆ ⋆ {ûSB α < Qα,P − ε} ⊆ {F̂B,n (Qα,P − ε) > α} n o ⊆ |F̂B,n (Q⋆α,P − ε) − FP (Q⋆α,P − ε)| ≥ δ(ε) .

Therefore,    ⋆ ⋆ ⋆ sup PP |ûSB − Q | > ε ≤ sup P | F̂ (Q + ε) − F (Q + ε)| ≥ δ(ε) P B,n P α α,P α,P α,P P ∈P P ∈P   + sup PP |F̂B,n (Q⋆α,P − ε) − FP (Q⋆α,P − ε)| ≥ δ(ε) P ∈P

 ≤ 2 sup sup PP |F̂B,n (t) − FP (t)| ≥ δ(ε) → 0 P ∈P t∈R

by Step 1’. This proves uniform consistency of ûSB α .

I.3

Proof of Proposition S.3

Write  fb := pmin = min p≤ fbm , b m∈[M ]

b ∈ [B]0 ,

so that the data-driven procedure is exactly the SB test applied to the statistic sequence (fb )b∈[B]0 , i.e., B

pSB,F =

1 X 1(fb ≤ f0 ). B+1 b=0

Let ûSB α denote the corresponding SB threshold. By Proposition 1, n u∈R:

ûSB α = sup

B o 1 X 1(fb ≤ u) ≤ α . B+1 b=0

80

(34)

Since the SB test rejects when f0 < ûSB α as in (5), we have   P pSB,F > α = P f0 ≥ ûSB α . Let A := {Neff (ε) ≤ N } so that P(Ac ) ≤ η by (24). On the event A, by the definition of Neff (ε) there exist indices m1 , . . . , mN ∈ [M ] such that for all b ∈ [B]0 ,  m  min p≤ fb j ≤ (1 + ε) min p≤ fbm = (1 + ε)fb , j∈[N ]

m∈[M ]

or equivalently, fb ≥

1 gb , 1+ε

m  where gb := min p≤ fb j . j∈[N ]

(35)

Define the SB threshold for the sequence (gb )b∈[B]0 by n := ûSB,(g) sup u∈R: α

B o 1 X 1(gb ≤ u) ≤ α . B+1 b=0

From (35), on the event A each fb is bounded below by gb /(1 + ε). Consequently, for all b ∈ [B]0 and all u,  1(fb ≤ u) ≤ 1 gb ≤ (1 + ε)u . Therefore, the empirical distribution function of (fb )b∈[B]0 is dominated by that of (gb )b∈[B]0 after rescaling the argument by the factor (1 + ε). Invoking the characterization of the SB threshold in (34), we obtain the lower bound ûSB α ≥

1 ûSB,(g) 1+ε α

on A.

Next, by Theorem 2 and the fact that the Bonferroni threshold cα,N = α/N satisfies the condition of that theorem for the minimum merging function with N coordinates, we conclude that the corresponding SB threshold satisfies ûαSB,(g) >

α . N

Finally, P pSB,F > α



 α ,A (1 + ε)N   α ≤ η + P f0 > (1 + ε)N    α m = η + P min p≤ f0 > (1 + ε)N m∈[M ]    α ≤ η + min P p≤ f0m > , (1 + ε)N m∈[M ]

c = P(f0 ≥ ûSB α ) ≤ P(A ) + P

 f0 >

where the last inequality uses {minm Xm > t} = ∩m {Xm > t}. This proves the claim. 81

I.4

Proof of Proposition S.6

Fix K and work conditionally on the observed data X. Throughout, we assume that, conditional on X, the transformations g1 , . . . , g2B are i.i.d. uniform draws from G. For any function F , we denote its left limit at t by F (t−) := lims↑t F (s). Step 1 (Holdout vs. within-testing-batch p-values). For each k ∈ [K], define the empirical CDFs 1 k Fbref (t) := B

B X i=1

B

1(T̃ik ≤ t),

k Fbtest (t) :=

1 X 1(Tjk ≤ t). B+1 j=0

For each b ∈ [B]0 , B

pHO (Tbk ) =

o X 1 n B bk 1+ F (T k −), 1(T̃ik ≥ Tbk ) = 1 − B+1 B + 1 ref b i=1

while the SB (within-batch) permutation p-values satisfy B

p(Tbk ) =

1 X k (Tbk −). 1(Tjk ≥ Tbk ) = 1 − Fbtest B+1 j=0

Hence, k k (t−) − Fbtest (t−) + max pHO (Tbk ) − p(Tbk ) ≤ sup Fbref

b∈[B]0

t∈R

1 . B+1

(36)

Step 2 (Uniform DKW control). Let 1 k Fbtest,MC (t) := B

B X j=1

1(Tjk ≤ t).

Then k k sup Fbtest (t−) − Fbtest,MC (t−) ≤ t∈R

1 . B+1

k and F bk Conditional on X, both Fbref test,MC are empirical CDFs of independent i.i.d. samples from the same randomization distribution. By the Dvoretzky–Kiefer–Wolfowitz (DKW) inequality applied k and F bk to Fbref test,MC , and a union bound, for every η > 0,

 k k P sup Fbref (t−) − Fbtest (t−) > 2η + t∈R

1 B+1

 X

2

≤ 4e−2Bη .

A union bound over k ∈ [K] (with K fixed) yields, for every η > 0,   2 2 P εB > 2η + ≤ 4Ke−2Bη , εB := max max pHO (Tbk ) − p(Tbk ) . B+1 k∈[K] b∈[B]0

82

In particular, for any δ > 0, there exist constants c > 0 and B0 < ∞ such that for all B ≥ B0 , 2

P(εB > δ) ≤ 4K e−cBδ . Step 3 (From p-values to merged values via uniform continuity). Define the TB merged values  f˜b := f pHO (Tb1 ), . . . , pHO (TbK ) , b ∈ [B]0 , and the SB merged values  fb := f p(Tb1 ), . . . , p(TbK ) ,

b ∈ [B]0 .

Since f is continuous on the compact set [0, 1]K , it is uniformly continuous. Thus, for every ε > 0 there exists δf (ε) > 0 such that whenever ∥x − y∥∞ ≤ δf (ε), one has |f (x) − f (y)| ≤ ε. Recalling the definition of εB in Step 2, we therefore have the implication εB ≤ δf (ε)

=⇒

max |f˜b − fb | ≤ ε.

b∈[B]0

Consequently, with ∆B := maxb∈[B]0 |f˜b − fb |,  P(∆B > ε) ≤ P εB > δf (ε) .

(37)

Combining (37) with the exponential tail bound for εB from Step 2 shows that P(∆B > ε) decays exponentially fast in B (with constants depending only on f and ε). Step 4 (From merged values to thresholds). Let m := ⌊(B + 1)α⌋ + 1 and write osm (·) for the m-th order statistic. Since order statistics are 1-Lipschitz with respect to the sup-norm, SB ˜ ˜ |ûTB α − ûα | = osm (f0 , . . . , fB ) − osm (f0 , . . . , fB ) ≤ ∆B .

Therefore, for every ε > 0,  SB P |ûTB α − ûα | > ε ≤ P(∆B > ε), and the right-hand side decays exponentially fast in B by Step 3. This completes the proof.

I.5

Proof of Proposition S.7

We start with an explicit construction underlying Proposition S.7 in Appendix I.5.1, followed by the assumptions in Appendix I.5.2. The proof of the limiting rejection probabilities is given in Appendix I.5.3. Throughout we use the minimum p-value merging rule, which makes the finite-B discretization effect most transparent. For notational convenience in the comparative analysis across SB and TB, we redefine the TB procedure by treating the transformations for b = 1, . . . , B as the reference batch and the subsequent B transformations for b = 0, B + 1, . . . , 2B as the testing batch. Since the 2B sampled transformations are independent and identically distributed conditional on the data, this relabeling is distributionally equivalent to the convention used in Algorithm 3. 83

I.5.1

Setup

Fix integers B ≥ 1 and K ≥ 1. For each n, let X(n) denote the observed data and let g0 = id and g1 , . . . , g2B be Monte Carlo transformations sampled from the randomization mechanism. For each coordinate k ∈ [K], define  k := T k gb (X(n) ) , Tb,n b ∈ [2B]0 . SB p-values (first batch). For b ∈ [B]0 and k ∈ [K], define the usual permutation p-values B

pkb,n :=

 1 X k k 1 Ti,n ≥ Tb,n . B+1

(38)

i=0

Let the merged SB values be k p̃SB b,n := min pb,n , k∈[K]

b ∈ [B]0 ,

and define the SB aggregation p-value (under the minimum merger) as B

 1 X SB 1 p̃SB pSB,n := b,n ≤ p̃0,n . B+1 b=0

TB holdout p-values (second batch, first batch as reference). For i ∈ [B] and k ∈ [K], k , . . . , T k }, define the holdout p-values by ranking the second batch against the reference set {T1,n B,n pHO,k B+i,n :=

  B X  1 k k 1+ 1 Tℓ,n ≥ TB+i,n . B+1

(39)

ℓ=1

= pk0,n by by the same formula with B + i replaced by 0. Note that pHO,k We also define pHO,k 0,n 0,n construction. Let the merged TB testing-batch values be HO,k p̃TB 0,n := min p0,n , k∈[K]

HO,k p̃TB B+i,n := min pB+i,n , k∈[K]

i ∈ [B],

and define the TB aggregation p-value as the rank p-value of the observed merged value among the B + 1 testing-batch merged values: ( ) B X  1 TB pTB,n := 1+ 1 p̃TB . B+i,n ≤ p̃0,n B+1 i=1

I.5.2

Assumptions

We impose three simplifying conditions.

84

(A1) Perfect rank alignment across coordinates. For all k, k ′ ∈ [K], all b, b′ ∈ [2B]0 , and all n ∈ N, k Tb,n ≤ Tbk′ ,n

⇐⇒

almost surely.

k Tb,n ≤ Tbk′ ,n

In particular, this implies that all permutation and holdout p-values are invariant in k, and hence the minimum merger reduces to a single coordinate: 1 p̃SB b,n = pb,n ,

b ∈ [B]0 ,

HO,1 1 p̃TB 0,n = p0,n = p0,n ,

HO,1 p̃TB B+i,n = pB+i,n ,

i ∈ [B].

(A2) Exchangeability and conditioning on full order statistics. For each fixed n and k, k ,...,Tk conditional on the data used to compute the statistic, the Monte Carlo draws T1,n 2B,n are i.i.d. from the randomization distribution and are distinct with probability one (e.g., after jittering). Let  k k Gn := σ T(1),n , . . . , T(2B),n denote the σ-field generated by the full collection of order statistics. Conditional on Gn , the k ,...,Tk ranks of T1,n 2B,n form a uniformly random permutation of {1, . . . , 2B}. (A3) Strong-signal regime. Assume that   1 1 P T0,n > max Tb,n → 1 1≤b≤B

I.5.3

as n → ∞.

Proof of the statement

We work under the setup and assumptions described in Appendices I.5.1 and I.5.2. We show that, for every α ∈ (0, 1),   1 , P(pSB,n ≤ α) → 1 α ≥ B+1 and P(pTB,n ≤ α) → P(W ≤ ⌊(B + 1)α⌋ − 1) ,

n → ∞,

where W is supported on {0, . . . , B} with 2B−k−1 B−1  2B B



P(W = k) =

,

k = 0, . . . , B.

Define the event n o 1 1 En := T0,n > max Tb,n . 1≤b≤B

By (A3), we have P(En ) → 1. Moreover, by (A1), it suffices to restrict attention to the single coordinate k = 1.

85

1 < T 1 for all b ∈ [B]. Therefore, Step 1 (SB behavior on En ). On the event En , we have Tb,n 0,n by (38),

p10,n =

1 . B+1

For each b ∈ [B], the corresponding SB permutation p-value satisfies ( ) B X 1 2 1 1 1 1 1(T0,n ≥ Tb,n )+ p1b,n = 1(Tℓ,n ≥ Tb,n ) ≥ B+1 B+1

on En ,

ℓ=1

1 ≥ T 1 ) = 1 and the summation term is at least 1 (it includes ℓ = b). Consequently, on since 1(T0,n b,n SB B En , the merged SB value p̃SB 0,n is the unique minimum among {p̃b,n }b=0 , and hence

pSB,n =

1 B+1

on En .

It follows that  P(pSB,n ≤ α) = P(En ) 1

 1 ≤ α + P(pSB,n ≤ α, Enc ), B+1

and therefore 

1 P(pSB,n ≤ α) → 1 α ≥ B+1

 .

HO,1 Step 2 (TB behavior on En ). On the event En , we have p̃TB 0,n = p0,n = 1/(B + 1). Define

 B X Wn := 1 pHO,1 B+i,n = i=1

1 B+1



  B X 1 1 = 1 TB+i,n > max Tℓ,n . 1≤ℓ≤B

i=1

1 1 1 ,...,T1 Let Gn denote the σ-field generated by the order statistics (T(1),n , . . . , T(2B),n ) of {T1,n 2B,n }. 1 1 Conditional on Gn , the ranks of T1,n , . . . , T2B,n form a uniformly random permutation of {1, . . . , 2B} 1 , . . . , T 1 } is a by (A2). Consequently, the set of ranks occupied by the reference batch {T1,n B,n uniformly random subset of size B from {1, . . . , 2B}. Let  1 1 Rn := max ranks of T1,n , . . . , TB,n .

Then Rn ∈ {B, B + 1, . . . , 2B} and, conditional on Gn ,  r−1 P(Rn = r | Gn ) =

B−1  , 2B B

r = B, . . . , 2B.

Moreover, since no reference statistic can exceed its own maximum rank Rn , all statistics with rank exceeding Rn must belong to the testing batch. It therefore follows deterministically that Wn = 2B − Rn . 86

Hence, conditional on Gn , the distribution of Wn is given by  2B−k−1 P(Wn = k | Gn ) =

B−1  2B B

k = 0, 1, . . . , B,

,

which does not depend on the realized values of the order statistics. On the event En , the TB aggregation p-value satisfies pTB,n =

1 + Wn . B+1

Writing m := ⌊(B + 1)α⌋, we therefore have {pTB,n ≤ α} = {Wn ≤ m − 1} on En . By the law of total probability, P(pTB,n ≤ α) = P(pTB,n ≤ α, En ) + P(pTB,n ≤ α, Enc ). The second term is bounded by P(Enc ) = o(1). For the first term, note that P(Wn ≤ m − 1, En ) = P(Wn ≤ m − 1) − P(Wn ≤ m − 1, Enc ), and since P(Wn ≤ m − 1, Enc ) ≤ P(Enc ) = o(1), we obtain P(Wn ≤ m − 1, En ) = P(W ≤ m − 1) + o(1), where W follows a negative hypergeometric random variable with probability mass function  2B−k−1 P(W = k) =

B−1  2B B

k = 0, . . . , B,

,

corresponding to a population of size 2B with B successes, and a stopping rule of r = 1 failure. Combining the above displays yields P(pTB,n ≤ α) = P(W ≤ ⌊(B + 1)α⌋ − 1) + o(1), which establishes the stated limit. Finally, note that P(W = B) =

1 2B B

 > 0.

Since ⌊(B + 1)α⌋ ≤ B for all α ∈ (0, 1), the limiting TB rejection probability is strictly less than one. This completes the proof.

87

I.6

Proof of Corollary S.1

We begin by establishing pointwise consistency. Define K

p̄0,n :=

n 1 X pk0,n Kn

k=1

as the average of the individual permutation p-values. By Theorem 2, the SB procedure calibrated with any deterministic merging function dominates the corresponding worst-case calibrated test. In particular, since K

f (u1 , . . . , uKn ) :=

n 2 X uk Kn

k=1

is a valid merging function [44, 56], implying    P pSB,avg > α = P p̄0,n ≥ ûαSB,avg ≤ P p̄0,n > α/2 , it suffices to show that p̄0,n → 0 in probability under the alternative. By Markov’s inequality and the layer-cake representation, ! Kn 1 X k P(p̄0,n ≥ α/2) = P p0,n ≥ α/2 Kn k=1 Z  2 2 1 1 ≤ E[p0,n ] = P p10,n > t dt, α α 0 R1 where the last equality follows from the identity E[X] = 0 P(X > t) dt for nonnegative random variables bounded by 1. By assumption, P(p10,n > t) → 0 for every t ∈ (0, 1), and since 0 ≤ p10,n ≤ 1, the dominated convergence theorem yields P(p̄0,n ≥ α/2) → 0. This establishes pointwise consistency of the SB average aggregation test. We now turn to uniform consistency. Assume that the individual permutation p-values are uniformly consistent over P, i.e., sup PP (p10,n > t) → 0

P ∈P

for every t ∈ (0, 1).

Fix α ∈ (0, 1). Applying Markov’s inequality and the layer-cake representation uniformly over P ∈ P gives Z 2 1 sup PP (p̄0,n > α/2) ≤ sup PP (p10,n > t) dt. α P ∈P 0 P ∈P The integrand converges pointwise to zero and is uniformly bounded by 1. Therefore, the dominated convergence theorem implies sup PP (p̄0,n > α/2) → 0,

P ∈P

which establishes uniform consistency over P and completes the proof. 88

I.7

Proof of Lemma S.7

For simplicity, write     n n 1X 1X A := sup t : 1(Zi ≤ t) ≤ α and B := sup t : 1(Zi < t) ≤ α . n n i=1

i=1

1 Pn

1 Pn

Since Fn− (t) := n i=1 1(Zi < t) ≤ Fn (t) := n i=1 1(Zi ≤ t), it is clear that {t : Fn (t) ≤ α} ⊆ {t : Fn− (t) ≤ α} and thereby A ≤ B. For the reverse direction, i.e., A ≥ B, assume that A < B for contradiction. Since B is the supremum, there is a sequence sk ↗ B with Fn− (sk ) ≤ α for all k. Choose k sufficiently large so that sk > A and sk is not equal to any of Z1 , . . . , Zn . Then we have Fn (sk ) = Fn− (sk ) ≤ α. Therefore Fn (sk ) ≤ α contradicts the definition of A since sk > A. This contradiction shows that A ≥ B and hence A = B. Finally, let k = ⌊nα⌋ + 1. For any t < Z(k) , at most k − 1 observations are less than or equal P to t, so n1 ni=1 1(Zi ≤ t) ≤ k−1 n ≤ α. This implies that t is in the set whose supremum defines A. Since this holds for all t < Z(k) , we have A ≥ Z(k) . Conversely, for any t ≥ Z(k) , at least k P observations are less than or equal to t, so n1 ni=1 1(Zi ≤ t) ≥ nk > α. This implies that t is not in the set whose supremum defines A. Therefore, we must have A ≤ Z(k) . Combining both directions gives A = Z(k) .

I.8

Proof of Lemma S.8

Let B(R) denote the Borel σ-algebra on R = R∪{−∞, +∞}. Define a finite measure ν on (R, B(R)) as the pushforward of w dµ by t: Z := ν(B) w(ω) 1{t(ω) ∈ B} dµ(ω), B ∈ B(R). Ω

For any x ∈ R, ν([x, ∞]) = Hence, Z

Z w(ω) 1

Z Ω

w(ω) 1{t(ω) ≥ x} dµ(ω).

w(ω ′ ) 1{t(ω ′ ) ≥ t(ω)} dµ(ω ′ ) ≤ α



 dµ(ω) = ν {x ∈ R : ν([x, ∞]) ≤ α} .

Define G(x) := ν([x, ∞]). Then G is non-increasing, since [y, ∞] ⊆ [x, ∞] whenever x < y, which implies G(y) ≤ G(x). Let S := {x ∈ R : G(x) ≤ α},

xα := inf S,

with the convention inf ∅ = +∞. Because G is non-increasing, the set S is upward closed: if x ∈ S and y ≥ x, then G(y) ≤ G(x) ≤ α, hence y ∈ S. Consequently, S is of the form [xα , ∞] or (xα , ∞]. 89

If S = ∅, then ν(S) = 0 ≤ α, and the claim holds. Hence assume S ̸= ∅.

Case 1: xα ∈ S. In this case, S ⊆ [xα , ∞], and therefore

ν(S) ≤ ν([xα , ∞]) = G(xα ) ≤ α. Case 2: xα ∈ / S. Then S ⊆ (xα , ∞]. By the definition of xα = inf S and the assumption S ̸= ∅, there exists a sequence (xm )m≥1 ⊂ S such that xm ↓ xα . By monotonicity of G, ν([xm , ∞]) = G(xm ) ≤ α

for all m ≥ 1.

Moreover, the sets [xm , ∞] form an increasing sequence and satisfy [ [xm , ∞] = (xα , ∞]. m≥1

By continuity of measures for increasing sequences, ν((xα , ∞]) = lim ν([xm , ∞]) ≤ α. m→∞

Since S ⊆ (xα , ∞], we conclude that ν(S) ≤ α.

In both cases, ν(S) ≤ α, which proves the claim.

I.9

Proof of Proposition S.1

Observe that the maximum of events can be reduced to a union of events as o n k max TB+b − Quantile1−u {T0k , T1k , . . . , TBk } > 0 k∈[K] o [ n k > Quantile1−u {T0k , T1k , . . . , TBk } . ⇐⇒ TB+b k∈[K]

k be the j-th order statistic of {T k , T k , . . . , T k } and then Lemma S.5 yields that Let T(j) 0 1 B k Quantile1−u {T0k , T1k , . . . , TBk } = T(⌈(1−u)(B+1)⌉) .

Therefore, we have k k TB+b > T(⌈(1−u)(B+1)⌉)

⇐⇒ ⌈(1 − u)(B + 1)⌉ ≤

B X

k 1(TB+b > Tjk ).

j=0

Since ⌈x⌉ ≤ m is equivalent to x ≤ m for integers m, we have ⌈(1 − u)(B + 1)⌉ ≤

B X j=0

k 1(TB+b > Tjk ) ⇐⇒ (1 − u)(B + 1) ≤

B X

k 1(TB+b > Tjk ),

j=0

which can be rearranged to ukb :=

B+1−

PB

k k j=0 1(TB+b > Tj )

(B + 1)

B

=

X 1 k 1(TB+b ≤ Tjk ) ≤ u. (B + 1) j=0

90

Recalling ub = mink∈[K] ukb , we have n o k max TB+b − Quantile1−u {T0k , T1k , . . . , TBk } > 0 ⇐⇒ ub ≤ u.

k∈[K]

Therefore, the definition of ũα in (3) can be rewritten as   B  1 X 1 ub ≤ u ≤ α = u(⌊Bα⌋+1) ũα = sup u > 0 : B b=1

where the last equality follows from Lemma S.7. If ub = 0 for all b ∈ [B], then the supremum is taken over an empty set and ũα is set to zero by convention. It then indeed holds that ũα = 0 = u(⌊Bα⌋+1) . Hence the claim follows.

I.10

Proof of Proposition S.2

S We prove the type I error bound P(0 ∈ Jj=1 Aj ) ≤ α by establishing exchangeability of the index 0 with each b ∈ [B] via a row-permutation equivariance argument, then derive tightness under almost surely distinct scores. S Step 1: Reduction. The sets A1 , . . . , AJ are pairwise disjoint, since Aj ⊆ Sj−1 = [B]0 \ j−1 ℓ=1 Aℓ for each j ∈ [J], so   X J J [ P 0∈ Aj = P(0 ∈ Aj ). j=1

j=1

It therefore suffices to show P(0 ∈ Aj ) ≤ ⌊(B + 1)αj ⌋/(B + 1) for each j ∈ [J] separately.

Step 2: Setup. For b ∈ [B]0 and k ∈ [K], define Rbk := p(Tbk ),

Rb := (Rb1 , . . . , RbK ),

R := (Rb )b∈[B]0 .

For any permutation π of [B]0 , let πR denote the matrix obtained by permuting the rows of R, i.e., (πR)b := Rπ(b) . Under the group-invariance hypothesis, the rows R0 , . . . , RB are exchangeable, so d

πR = R for every permutation π of [B]0 . Step 3: Equivariance of survival sets. The key property of the stage scores is their rowpermutation equivariance: for every permutation π of [B]0 , every b ∈ [B]0 , and every ℓ ∈ [J],   k zb,ℓ (πR) = hℓ ((πR)kb )k∈Iℓ = hℓ (Rπ(b) )k∈Iℓ = zπ(b),ℓ (R), (40) which follows immediately from the definition (17) and (πR)b = Rπ(b) . We use (40) to establish equivariance of the survival sets Sℓ by induction on ℓ. Fix j ∈ [J]. Set S0 (R) := [B]0 , and define recursively, for ℓ ∈ [j], n cℓ (R) := sup u ∈ R :

1 B+1

X b∈Sℓ−1 (R)

91

o  1 zb,ℓ (R) ≤ u ≤ αℓ ,

and  Aℓ (R) := b ∈ Sℓ−1 (R) : zb,ℓ (R) < cℓ (R) ,

Sℓ (R) := Sℓ−1 (R) \ Aℓ (R).

We establish by induction on ℓ that  Sℓ (πR) = π −1 Sℓ (R)

(41)

for every permutation π of [B]0 and every ℓ ∈ {0} ∪ [j]. The base case ℓ = 0 holds trivially since S0 (πR) = S0 (R) = [B]0 . For the inductive step, assume (41) holds for ℓ − 1, where ℓ ∈ [j]. By (40), zb,ℓ (πR) = zπ(b),ℓ (R) for all b ∈ [B]0 . Using the inductive hypothesis Sℓ−1 (πR) = π −1 (Sℓ−1 (R)), we obtain, for every u ∈ R, X X   1 zπ(b),ℓ (R) ≤ u 1 zb,ℓ (πR) ≤ u = b∈π −1 (Sℓ−1 (R))

b∈Sℓ−1 (πR)

X

=

b∈Sℓ−1 (R)

 1 zb,ℓ (R) ≤ u ,

where the last equality uses that b 7→ π(b) is a bijection from π −1 (Sℓ−1 (R)) onto Sℓ−1 (R). Hence cℓ (πR) = cℓ (R). Consequently, b ∈ Aℓ (πR) ⇐⇒ b ∈ Sℓ−1 (πR) and zb,ℓ (πR) < cℓ (πR)

⇐⇒ π(b) ∈ Sℓ−1 (R) and zπ(b),ℓ (R) < cℓ (R) ⇐⇒ π(b) ∈ Aℓ (R)

 ⇐⇒ b ∈ π −1 Aℓ (R) ,

so Aℓ (πR) = π −1 (Aℓ (R)), and therefore    Sℓ (πR) = Sℓ−1 (πR) \ Aℓ (πR) = π −1 Sℓ−1 (R) \ π −1 Aℓ (R) = π −1 Sℓ (R) , which completes the induction. Step 4: Type I error bound. Fix b ∈ [B]0 , and let π be any permutation of [B]0 with π(0) = b. Since π(0) = b and the induction gives Aj (πR) = π −1 (Aj (R)),  b ∈ Aj (R) ⇐⇒ 0 ∈ π −1 Aj (R) ⇐⇒ 0 ∈ Aj (πR). d

Since πR = R, we have P(b ∈ Aj ) = P(0 ∈ Aj ) for all b ∈ [B]0 , and therefore   B |Aj | 1 X P(0 ∈ Aj ) = P(b ∈ Aj ) = E . B+1 B+1 b=0

Since zb,j takes finitely many values as b ranges over Sj−1 , there exists ϵ > 0 such that (cj − ϵ, cj ) contains no score zb,j with b ∈ Sj−1 . Hence X X |Aj | 1 1 = 1(zb,j < cj ) = 1(zb,j ≤ cj − ϵ) ≤ αj B+1 B+1 B+1 b∈Sj−1

b∈Sj−1

92

almost surely,

Since |Aj | is integer-valued, the bound |Aj |/(B + 1) ≤ αj a.s. implies P(0 ∈ Aj ) ≤ ⌊(B + 1)αj ⌋/(B + 1). Summing over j = 1, . . . , J and applying subadditivity of the floor function gives   X J J [ P 0∈ Aj = P(0 ∈ Aj ) j=1

j=1

J

 1 X (B + 1)αj B+1 j=1   P (B + 1) Jj=1 αj ⌊(B + 1)α⌋ ≤ ≤ α. ≤ B+1 B+1

(42)

Step 5: Tightness under almost surely distinct scores. Suppose now that for every j ∈ [J], the stage-j scores {zb,j : b ∈ Sj−1 } are almost surely distinct. We claim that   J [ P 0∈ Aj = j=1

J

1 X ⌊(B + 1)αj ⌋. B+1 j=1

Set qj := ⌊(B + 1)αj ⌋ for j ∈ [J]. It suffices to show |Aj | = qj almost surely for each j ∈ [J]. Indeed, if this holds, then by the exchangeability argument above, P(0 ∈ Aj ) =

qj E|Aj | = , B+1 B+1

j ∈ [J],

and since A1 , . . . , AJ are disjoint by construction,   X J J [ P 0∈ Aj = P(0 ∈ Aj ) = j=1

j=1

J

1 X qj . B+1 j=1

We record for later use that J X ℓ=1

J k j X αℓ ≤ ⌊(B + 1)α⌋ ≤ B. qℓ ≤ (B + 1)

(43)

ℓ=1

For j = 1, we have S0 = [B]0 , so |S0 | = B + 1 ≥ q1 + 1. Let z(1),1 < · · · < z(B+1),1 denote the ordered stage-1 scores in S0 . Since the scores are almost surely distinct and q1 = ⌊(B + 1)α1 ⌋, for every u ∈ R, 1 #{b ∈ S0 : zb,1 ≤ u} ≤ α1 B+1

⇐⇒

#{b ∈ S0 : zb,1 ≤ u} ≤ q1

⇐⇒

u < z(q1 +1),1 .

Hence the feasible set in the definition of c1 is exactly (−∞, z(q1 +1),1 ), so c1 = z(q1 +1),1 and A1 = {b ∈ S0 : zb,1 < z(q1 +1),1 }. Therefore |A1 | = q1 almost surely. For the inductive step, let j ∈ {2, . . . , J} and assume |Aℓ | = qℓ almost surely for all ℓ < j. Then |Sj−1 | = B + 1 −

j−1 X ℓ=1

|Aℓ | = B + 1 − 93

j−1 X ℓ=1

qℓ ≥ qj + 1,

by (43). Let z(1),j < · · · < z(|Sj−1 |),j denote the ordered stage-j scores in Sj−1 . Repeating the argument for j = 1 with Sj−1 in place of S0 , 1 #{b ∈ Sj−1 : zb,j ≤ u} ≤ αj ⇐⇒ u < z(qj +1),j . B+1 Hence cj = z(qj +1),j and Aj = {b ∈ Sj−1 : zb,j < z(qj +1),j }, so |Aj | = qj almost surely. This completes the induction and therefore the proof of tightness under almost surely distinct scores.

I.11

Proof of Lemma S.1

Recall that B

p≤ (fbm ) =

1 X 1{fim ≤ fbm } B+1 i=0

∈ {1/(B + 1), 2/(B + 1), . . . , 1},

b ∈ [B]0 , m ∈ [M ].

also takes values in the same grid, it On the event (25), we have pmin = 1/(B + 1). Since pmin 0 b follows that for all b ∈ [B]0 , min = 1/(B + 1)}. ≤ pmin {pmin 0 } = {pb b

Therefore,

pSB,F =

B 1 X  min 1  W 1 pb = = , B+1 B+1 B+1 b=0

= 1/(B + 1)} . pmin b

where W = {b ∈ [B]0 : It remains to show that W ≤ M . Fix m ∈ [M ] and define the (possibly empty) set of strict minimizers o n Sm := b ∈ [B]0 : fbm < fim for all i ̸= b . By definition, Sm has cardinality at most one. Moreover, for the usual permutation p-value based on the ordering of {fim }i∈[B]0 ,

1 ⇐⇒ b ∈ Sm , B+1 because p≤ (fbm ) = 1/(B + 1) holds if and only if exactly one value in the multiset {fim }i∈[B]0 is ≤ fbm , i.e., fbm is a strict minimum. Consequently, n 1 o pmin = =⇒ ∃ m ∈ [M ] such that b ∈ Sm . b B+1 Hence p≤ (fbm ) =

W ≤

M [ m=1

If M ≤ ⌊(B + 1)α⌋, then pSB,F =

Sm ≤

M X m=1

|Sm | ≤ M.

⌊(B + 1)α⌋ W ≤ ≤ α, B+1 B+1

which concludes the proof. 94

I.12

Proof of Proposition S.5

We prove the two claims separately. m(j) 

Claim 1. Fix j ∈ [J]. By assumption (26), p≤ f0

1 = B+1 almost surely under Pj . Since

m(j)

1 almost surely under Pj as pmin = minm∈[M ] p≤ (f0m ) ≤ p≤ (f0 ), it follows that pmin = B+1 0 0 well. Applying Lemma S.1 with M ≤ ⌊(B + 1)α⌋, we conclude that pSB,F ≤ α Palmost surely under Pj . Since j ∈ [J] was arbitrary, the same holds under the mixture P = j πj Pj , giving PP (pSB,F ≤ α) = 1.

Claim 2. Fix m ∈ [M ]. By assumption, there exist j ∈ [J] with πj > 0 and ε > 0 such that   PPj p≤ f0m ≤ α ≤ 1 − ε. Therefore, J   X   PP p≤ f0m ≤ α = πℓ PPℓ p≤ f0m ≤ α ℓ=1

≤ πj (1 − ε) +

I.13

X ℓ̸=j

πℓ = 1 − πj ε < 1.

Proof of Lemma S.2

Fix t ∈ (0, 1]. Let S := {b ∈ [B]0 : p̄ϕ,b ≤ t} and ℓ := |S|. The case ℓ = 0 is trivial, so assume ℓ ≥ 1. Upper bound. By definition of p̄ϕ,b and monotonicity of ϕ, p̄ϕ,b ≤ t implies K1 Summing over b ∈ S yields K XX (k)  ϕ pb ≤ ℓKϕ(t).

(k) k=1 ϕ(pb ) ≤ ϕ(t).

PK

(44)

b∈S k=1 (k)

(k)

(k)

(k)

Lower bound. Fix k ∈ [K] and let p(1) ≤ · · · ≤ p(B+1) denote the order statistics of (p0 , . . . , pB ). By Lemma 1, for α = (j − 1)/(B + 1) we have   B 1 X j−1 j−1 (k) ≤ , 1 pb ≤ B+1 B+1 B+1 b=0

(k)

(k)

so at most j − 1 values are ≤ (j − 1)/(B + 1). Hence p(j) > (j − 1)/(B + 1), and since each pb on the grid {1/(B + 1), . . . , 1}, it follows that (k)

p(j) ≥

j , B+1

(45)

j ∈ [B + 1].

Since ϕ is increasing and |S| = ℓ, X b∈S

(k)  ϕ pb ≥

ℓ X j=1

(k)  ϕ p(j) ≥

95

 ℓ X ϕ j=1

j B+1

lies

 ,

where the last inequality uses (45). Summing over k = 1, . . . , K yields K XX b∈S k=1

(k)  ϕ pb ≥ K

 ℓ X ϕ j=1

j B+1



(46)

.

Combine. Combining (44) and (46) gives ℓ

1X ϕ(t) ≥ ϕ ℓ j=1



j B+1

 ,

hence t ≥ tϕ,B (ℓ), by monotonicity of ϕ and the definition of tϕ,B (ℓ). Therefore ℓ ≤ max{ℓ′ : tϕ,B (ℓ′ ) ≤ t}, which proves the deterministic inequality B

1 X 1(p̄ϕ,b ≤ t) ≤ Gϕ,B (t). B+1 b=0

(1)

(K)

For the probabilistic statement, under the group-invariance hypothesis the rows (pb , . . . , pb ), b ∈ [B]0 , are exchangeable. Since p̄ϕ,b is a measurable function of the bth row, the vector (p̄ϕ,0 , . . . , p̄ϕ,B ) is exchangeable as well. Thus " # B 1 X P(p̄ϕ,0 ≤ t) = E 1(p̄ϕ,b ≤ t) , B+1 b=0

and taking expectations of the deterministic bound yields P(p̄ϕ,0 ≤ t) ≤ Gϕ,B (t). Finally, we prove the validity. By the deterministic inequality above, B

1 X 1(p̄ϕ,b ≤ p̄ϕ,0 ) ≤ Gϕ,B (p̄ϕ,0 ) B+1 b=0

holds almost surely. Under the group-invariance hypothesis, the row-wise statistics (p̄ϕ,0 , . . . , p̄ϕ,B ) are exchangeable, and therefore ! B 1 X P(Gϕ,B (p̄ϕ,0 ) ≤ α) ≤ P 1(p̄ϕ,b ≤ p̄ϕ,0 ) ≤ α ≤ α, B+1 b=0

where the last inequality follows from Lemma S.4. This shows that Gϕ,B (p̄ϕ,0 ) is super-uniform and hence a valid p-value.

I.14

Proof of Lemma S.3

Define gr (ℓ) := 

ℓ 1 Pℓ ℓ

r j=1 j

96

1/r ,

ℓ ∈ N,

so that cr,B = max1≤ℓ≤B+1 gr (ℓ). Step 1: upper bound by cr . Since x 7→ xr is increasing on [0, ∞), ℓ X

j

Z ℓ

r

>

xr dx =

0

j=1

ℓr+1 , r+1

ℓ ≥ 1.

Therefore ℓ

ℓr 1X r j > ℓ r+1

gr (ℓ) < (r + 1)1/r = cr .

=⇒

j=1

Taking the maximum over ℓ ≤ B + 1 gives cr,B < cr , proving (i).

Step 2: monotonicity and convergence. Monotonicity in B is immediate because the maximization set {1, . . . , B + 1} increases with B. To identify the limit, note that ℓ

1X ℓ j=1

 r Z 1 j 1 → xr dx = ℓ r+1 0

(ℓ → ∞),

i.e., a standard Riemann-sum limit. Hence  gr (ℓ) = 

ℓ   1X j r

j=1

−1/r 

→ (r + 1)1/r = cr .

Given ε > 0, choose ℓ large with gr (ℓ) > cr − ε. For all B such that B + 1 ≥ ℓ, cr,B ≥ gr (ℓ) > cr − ε. Together with cr,B < cr , this yields cr,B ↑ cr .

97

Record · ID 381764 · SHA-256 0fd453142f8e64d4
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.