Risk-Limiting Audits for Parliamentary Majorities⋆ Jack Freestone1[0009−0008−2983−6676] , Dennis Leung2[0000−0002−5924−8161] , and Damjan Vukcevic3[0000−0001−7780−9586] 1
arXiv:2607.21082v1 [stat.AP] 23 Jul 2026
School of Mathematical and Physical Sciences, Macquarie University, Australia 2 School of Mathematics and Statistics, University of Melbourne, Australia 3 Department of Econometrics and Business Statistics, Monash University, Australia [email protected]
Abstract. Existing methods for risk-limiting audits typically focus on certifying individual contests. In parliamentary elections, however, the politically relevant outcome is often whether a party has won enough seats to form government, not whether every reported seat outcome is correct. Extending on the work of Mohanty et al. [9], we formulate the certification of a parliamentary majority as a partial conjunction testing problem: it is enough to verify that the reported winning party truly won at least a majority of its reported seats. Building on the SHANGRLA auditing framework, we construct a sequential audit statistic for the majority outcome by combining seat-level statistics. We then propose adaptive sampling strategies that allocate auditing effort across seats, including variants that learn to avoid spending excessive effort on seats that appear unlikely to have been truly won. Using simulations based on synthetic and real data, from the 2014 Indian Lok Sabha election, we show that auditing the parliamentary majority can substantially reduce the number of ballots inspected (by almost a thousand-fold) compared to certifying every reported winning seat.
1
Introduction
A risk-limiting audit (RLA) is a post-election procedure that sequentially samples ballots, accumulating statistical evidence until either the reported outcome can be certified4 or the audit escalates to a full hand count [8]. Its guarantee is that if the reported outcome is wrong, the audit certifies it with probability at most α (the risk limit), a pre-specified value. If the reported outcome is correct, ⋆
Authors listed alphabetically. Accepted for E-Vote-ID 2026. This work was supported by the Australian Research Council (OPTIMA ITTC IC200100009). 4 We use ‘certify’ (and ‘certification’) throughout in a purely statistical sense: to verify the result of the election with risk limit α. This is distinct from certification as a legal act, whose meaning vary across jurisdictions.
2
Freestone J, Leung D, Vukcevic D
certification typically needs only a small fraction of the ballots. The SHANGRLA framework of Stark [13] reduces the auditing of any electoral system (a social choice function) to testing a finite collection of assertions about population means; the ALPHA test supermartingale [14] provides a flexible sequential test for each. These tools, however, were developed primarily for single-contest elections. In a parliamentary system, government is determined not by any single race but by whether a party controls a majority of seats, a joint outcome whose auditing has received comparatively little attention. As first observed by Mohanty et al. [9], certifying a parliamentary majority does not require certifying every individual seat. We formalise this as a partial conjunction hypothesis test [1]: a procedure that rejects at least some (but not necessarily all) null hypotheses from a large collection. This pools evidence across seats, certifying the majority even when some seats have not individually accumulated enough evidence, and can substantially reduce the number of ballots that must be hand-inspected. We develop a suite of adaptive sampling schemes that dynamically allocate auditing effort across seats. These include greedy strategies that concentrate sampling on weaker seats (where stronger evidence is required) and filtering strategies that redirect effort away from seats where the reported outcome seems to be false. We evaluate our schemes via simulation across a range of parliamentary configurations, including scenarios based on the 2014 Indian Lok Sabha election, and show they require observing orders of magnitude fewer ballots than audits that aim to certify all reported winning seats.
2
RLA for a parliamentary majority by sequentially testing a partial conjunction hypothesis
Let S denote the set of all seats in a parliament, each corresponding to a separate contest. As a concrete example, the Lok Sabha (Lower House) of the Indian Parliament has |S| = 543 seats corresponding to 543 constituencies [9]. Whichever party controls the majority, that is, r ≡ ⌊|S|/2⌋ + 1 of the |S| seats, can form the government for the next term. Suppose W ⊂ S is the subset of seats reportedly won by the winning party after a given election (necessarily |W| ⩾ r). To confirm that the new government can be rightfully formed, we need to verify that at least r of the seats in W have been truly won by the winning party, which can be formulated as testing the null hypothesis Hr/W : fewer than r seats in W have been truly won by the winning party.
Risk-Limiting Audits for Parliamentary Majorities
3
This is the complementary null of the assertion that the winning party won at least r seats. If Hr/W is rejected, the majority outcome is considered certified. To conduct an audit with risk limit α, we can sample the ballots cast for the seats in W sequentially, until their revealed evidence suggests that Hr/W can be rejected at level α. We conceptualise such sampling as happening iteratively in ‘rounds’, indexed by a time variable t. Before any sampling has taken place we have t = 0, and at each subsequent time t ∈ {1, 2, . . . }, the auditor has the option to draw one cast ballot for each seat s ∈ W without replacement.5 Let ( Ds,t+1 ≡
1,
if a ballot for seat s is drawn at time t + 1,
0,
otherwise,
be the auditor’s sampling decision for time t + 1 for seat s. The decision Ds,t+1 can therefore only depend on the information from the ballots sampled up to time t; we denote this information as Ft . To evaluate the evidence against a given null hypothesis, H, we need to define a suitable test statistic, Ut . We will calculate this iteratively based on the information revealed in Ft , giving rise to a sequence of values, U ≡ (Ut )t⩾0 , referred to as a (test) process. We want processes that satisify the following: 1. We say that U is anytime-valid for testing H if Pr sup Ut ⩾ 1/α ⩽ α for all α ∈ (0, 1) when H is true.
(1)
t⩾0
This allows us to reject H, with risk limit α, once U exceeds 1/α. 2. If H is false, we want to reject it efficiently. We seek a process U that will grow rapidly in response to new observations, so it can quickly exceed 1/α. In the remainder of this section we construct a process (Er/W in Theorem 1) that is anytime-valid for testing Hr/W . Then in Section 3 we explore various sampling schemes that guide the drawing decisions Ds,t across time, aiming to minimize the total number of ballots sampled before Hr/W can be rejected. 2.1
A test process for certifying a single seat
As a starting point, we define for each seat s ∈ W the null hypothesis Hs : the winning party has not truly won the seat s, 5
This choice is primarily for clear exposition. In practice, a real audit would aim to have very few rounds, and would draw a much larger sample of ballots in each round. The methods we develop here are applicable to any such sampling scheme.
4
Freestone J, Leung D, Vukcevic D
and let Ns be the total number of ballots cast for seat s, which we label bs,1 , . . . , bs,Ns . We assume the social choice function for this seat can be decomposed into a finite set of Js assertions using the SHANGRLA framework s such that, by [13], i.e., there exist Js non-negative assorter functions {As,j }Jj=1 letting xs,j,i ≡ As,j (bs,i ), Hs is false if and only if N
x̄s,j ≡
s 1 X xs,j,i > 1/2 Ns i=1
for all j = 1, . . . , Js .
(2)
We can therefore write Hs =
Js [
Hs,j ,
(3)
j=1
where Hs,j : x̄s,j ⩽ 1/2 is the null hypothesis that the jth assertion is not true. As an example, consider a seat contested by just two candidates, Alice (the reported winner) and Bob. Here there is a single assertion (Js = 1), whose assorter As,1 scores a ballot for Alice as 1, a ballot for Bob as 0, and any other (e.g. invalid) ballot as 1/2. The assertion mean x̄s,1 is then the average of these scores across all Ns ballots, which exceeds 1/2 if and only if Alice received more votes than Bob. s Suppose (Bs,i )N i=1 represents the full sequence of ballots that would be drawn randomly without replacement for seat s, which is necessarily a random permuNs Ns s tation of {bs,i }N i=1 . Let (Xs,j,i )i=1 = (As,j (Bs,i ))i=1 represent the corresponding Ns (random) sequence of values from {xs,j,i }i=1 . Since a ballot for seat s isn’t necessarily drawn in every round, one should keep in mind that Bs,i is only the ith ballot drawn for seat s and its sampling will take place at a time t ⩾ i. Let ns (t) denote the number of ballots drawn for seat s up to and inclusive of time t. Define the process Ms,j ≡ (Ms,j,t )t⩾0 by letting Ms,j,0 ≡ 1 and ns (t)
Ms,j,t ≡
Y
(1 + λs,j,i (Xs,j,i − µs,j,i ))
for all t = 1, 2, . . . ,
(4)
i=1 P N /2− i−1 X
s,j,k k=1 is the hypothesised mean of remaining assorter where µs,j,i ≡ s Ns −i+1 values just before the ith ballot is drawn assuming x̄s,j = 1/2 (the boundary value for Hs,j to hold), and λs,j,i ∈ (0, µ−1 s,j,i ) is a parameter depending only on FTs (i)−1 , where Ts (i) ≡ min{t : ns (t) = i} is the first time that i ballots have been sampled for seat s. Each λs,j,i controls how aggressively the process responds to the value Xs,j,i − µs,j,i , and can be set based on the information available from all the revealed ballots before the ith ballot is drawn for seat s. Larger values of λs,j,i produce faster growth of Ms,j,t when Hs,j is false.
Risk-Limiting Audits for Parliamentary Majorities
5
Note that Ms,j is an example of a betting process [18], introduced for election auditing by the ALPHA process of Stark [14]. Unlike the original ALPHA, Ms,j has the variation that its value may not update at every time step because sampling doesn’t necessarily occur for seat s at all t. Alternatively, we can write Ms,j,t = Ms,j,t−1 · Ds,t · 1 + λs,j,ns (t) (Xs,j,ns (t) − µs,j,ns (t) ) + 1 − Ds,t (5) which gives Ms,j,t = Ms,j,t−1 when Ds,t = 0 and elucidates how the betting process evolves from t − 1 to t. Importantly, Ms,j is a test supermartingale under Hs,j . Specifically, it is nonnegative (Ms,j,t ⩾ 0 for all t ⩾ 0), has initial value 1 (Ms,j,0 = 1), and E[Ms,j,t | Ft−1 ] ⩽ Ms,j,t−1
(6)
when Hs,j is true. We establish these facts in Appendix A. Ville’s inequality [15] then implies that Ms,j is anytime-valid for testing Hs,j . To test the seat as whole, we combine these together into a seat-level process Es,t ≡ min Ms,j,t . 1⩽j⩽Js
(7)
By virtue of the ‘min’ construction, this is anytime-valid for testing Hs (see Theorem 1, below). However, our goal is not test each seat on its own, but to use these processes as a building block for testing more complex hypotheses. 2.2
A test process for certifying a parliamentary majority
Let C ⊂ W be a subset of the reported winning seats, and consider \ H1/C ≡ Hs ,
(8)
s∈C
the intersection null hypothesis that Hs is true for all seats in C. Note that H1/C is true when no (< 1) seats in C are truly won by the winning party, and hence its notational similarity to Hr/W . Now define Y E1/C,t ≡ Es,t . (9) s∈C
This provides a stepping-stone to a process for testing Hr/W , because we can write that hypothesis as the following union [ Hr/W = H1/C . (10) C⊂W: |C|=|W|−r+1
Our test process for a parliamentary majority can now be defined as follows: |W|−r+1
Er/W,t ≡
Y s=1
E(s),t = min{E1/C,t : |C| = |W| − r + 1}
(11)
6
Freestone J, Leung D, Vukcevic D
where E(1),t ⩽ · · · ⩽ E(|W|),t are the order statistics of {Es,t }s∈W at time t. In other words, Er/W,t is the product of the smallest |W| − r + 1 values among {Es,t }s∈W , and the second equality in (11) follows naturally. It is easy to check that all of Es,t , E1/C,t and Er/W,t are computable from Ft . The following theorem proven in Appendix B states that they indeed define processes that are anytime-valid for testing Hs , H1/C and Hr/W . Theorem 1. The processes Es ≡ (Es,t )t⩾0 , E1/C ≡ (E1/C,t )t⩾0 and Er/W ≡ (Er/W,t )t⩾0 are anytime-valid for testing Hs , H1/C and Hr/W , respectively. Intuitively, using Er/W to test Hr/W makes sense: In light of the union representation (10), Hr/W is false if and only if all intersection nulls H1/C with |C| = |W| − r + 1 are false; the ‘min’ representation in (11) says that Er/W,t is large enough to reject Hr/W only when, for all C with cardinality |W| − r + 1, the value of E1/C,t is large enough to reject the corresponding hypothesis H1/C . 2.3
Related works
In the statistics literature, Hr/W is a partial conjunction hypothesis, testing whether at least r of the seat-level nulls {Hs }s∈W are false. The notion is due to Benjamini and Heller [1] in the context of replicability analysis; Bogomolov and Heller [2] provide a recent survey. Our construction sits in the e-value and e-process framework [11, 16]: each of Es,t , E1/C,t and Er/W,t is an e-process for its null, and combining the seat-level ones into Er/W,t is a sequential Benjamini– Heller-type construction [6, 7]. As shown in Appendix B, it stays valid even when sampling induces cross-seat dependence, extending Vovk and Wang [17]. Mohanty et al. [9, Section 4] first proposed certifying the overall winning party rather than every seat, for a more efficient audit, and sketched a test of Hr/W , but did not implement nor evaluate it. Like us, they built an anytimevalid Es,t for each Hs , but then transformed it into a p-value Ps,t ≡ 1/Es,t and combined these across each subset C ⊂ W by Fisher’s method [5]. This is the stratified union-intersection idea of SUITE [10], applied across seats rather than across strata of a single contest. Working directly with the e-processes instead allows us to use arbitrary sampling schemes Ds,t with easy theoretical justification, including the adaptive schemes in Section 3 which induce dependence across s, and is uniformly more efficient than the p-value approach; see Appendix C. A final pair of RLA methods also test a union-of-intersections null, but within a single contest. AWAIRE [3] audits a single instant-runoff election. Its component tests share the same population of ballots, so they are not conditionally independent and must be combined by averaging rather than as a product. In our setting, since each seat is a separate contest, their e-processes are conditionally independent and can be multiplied. Spertus et al. [12] stratify a single
Risk-Limiting Audits for Parliamentary Majorities (a) Non-adaptive All seats sampled
(b) Greedy
(c) Filtered
(d) Greedy Filtered
Active set: dt = 2
Active set: |W|−r+1 = 3
Active set: dt = 2
active set
active set
1/α
20
7
active set
1/α
1/α
1/α
Es, t
15
10
5
0
s₁
(false)
s₂
s₃
s₄
s₅
s₁
(false)
s₂
s₃
s₄
s₅
s₁
(false)
s₂
s₃
s₄
s₅
0.04 0.71 0.89 0.97 0.99 pos terior prob. πs
Sampled this round
Not sampled (above cutoff)
s₁
(false)
s₂
s₃
s₄
s₅
0.04 0.71 pos terior prob. πs
Excluded (πs ≤ τ)
Fig. 1. Illustration of the four sampling schemes at a single round, with |W| = 5 seats and r = 3. Bars show the current values of Es,t ; the dashed red line is 1/α. Blue bars indicate seats sampled in the next round; grey bars indicate seats above the active-set cutoff; orange bars indicate seats excluded by the posterior filter (πs ⩽ τ ). Seat s1 is falsely reported. (a) Non-adaptive: every seat is sampled. (b) Greedy: only the dt weakest seats are sampled. (c) Filtered: only seats among the |W| − r + 1 weakest with πs > τ are sampled. (d) Greedy Filtered: combines the dt active-set window of the Greedy scheme with the πs > τ filter of the Filtered scheme.
contest and, like us, combine by products. They pool evidence about one contestwide margin across strata. In contrast, we pool evidence about many separate contests, of which only r need to be verified.
3
Adaptive sampling schemes
Given how Er/W,t is defined in (11), the ballot drawing decisions Ds,t should ideally be made so that the product of the smallest |W| − r + 1 values among {Es,t }s∈W can exceed 1/α with minimal sampling. We now describe several schemes for adaptively determining Ds,t . These are illustrated in Figure 1.
Non-adaptive. This scheme sets Ds,t = 1 for all s ∈ W and all t. The auditor makes no attempt to reduce the number of ballots to be sampled.
Greedy. Since the winning party can be immediately certified at time t as long as the product of the smallest |W| − r + 1 values among {Es,t }s∈W exceeds 1/α, a natural heuristic is to ‘greedily’ sample ballots for the seats with the lowest seat-level processes, in the hope that the additional ballots will increase these
8
Freestone J, Leung D, Vukcevic D
values. Let dt ∈ {1, . . . , |W|} be a cutoff sample size. The Greedy scheme sets ( 1, if Es,t ⩽ E(dt ),t , Ds,t+1 = 0, if Es,t > E(dt ),t , where E(1),t ⩽ · · · ⩽ E(|W|),t are the order statistics of {Es,t }s∈W at time t. The question is how to choose dt . Define |W|−r+i n o Y δt = max i : i ∈ {1, . . . , r} and E(s),t < 1/α ,
(12)
s=i
with the convention that δt = 0 if the maximum is taken over an empty set (meaning the winning party can already be certified at time t). Let ( |W|, if δt = r, dt = (13) δt + a, if δt < r, for a constant a ∈ {0, 1, . . . , |W| − r}. When δt = r, even the product of the largest |W| − r + 1 elements among {Es,t }s∈W is below 1/α, so we sample all |W| seats. When δt < r, the order statistics E(1),t , . . . , E(δt ),t are precisely those for which the product of any consecutive set of |W|−r+1 containing them would not exceed 1/α, and at a minimum we should gather more evidence for the δt seats corresponding to these processes. The parameter a controls how aggressively the scheme minimises per-round sampling: a = 0 is the most aggressive, while larger values of a hedge by sampling additional seats beyond the strict minimum. Since sampling is without replacement, it is possible for every one of the dt selected weakest seats to have already exhausted its ballot supply (i.e., all Ns ballots have been drawn) without achieving certification (Er/W,t ⩾ 1/α). In this case, the round would issue no new samples and the audit would stall. To avoid this, we apply the following fallback: if all dt weakest seats are exhausted, we incrementally increase dt by one—extending the selection to the next-weakest seat—and repeat until either at least one selected seat has remaining ballots or dt = |W|. This guarantees that in each round the sampling effort is directed at the weakest seats whose ballot supply has not yet been exhausted. Filtered. The Greedy scheme concentrates all sampling effort on the seats with the weakest seat-level processes. While intuitive, this can fail when some of the weakest seats correspond to falsely reported outcomes, i.e., seats where the winning party did not actually prevail. For such a seat s, the process Es won’t grow, which will lead to Er/W stalling or growing very slowly. We define a Bayesian inference approach that aims to filter out such seats. For each seat s ∈ W, we monitor the observed ballots and update a posterior
Risk-Limiting Audits for Parliamentary Majorities
9
distribution on the true ballot-type proportions. After each round, we use this to calculate a posterior probability, πs , that we use to exclude unpromising seats from further sampling. Our approach uses a canonical Dirichlet–multinomial conjugate setup. We omit most of the details (they are standard) and just sketch out the main steps. For seat s with Ls ballot types, set a prior on the ballot-type probabilities ps = (ps,1 , . . . , ps,Ls ) ∼ Dirichlet(αs,0 ), with hyperparameters αs,0 ≡ p̂s,0 κs,0 based on an initial value p̂s,0 and concentration parameter κs,0 . After observing ns (t) ballots at time t, use conjugacy6 to give the posterior ps | Ft ∼ Dirichlet(αs,ns (t) ). Let θs,j be the true assorter mean for assertion j. The posterior on ps induces a posterior on θs,j via its assorter function. The joint posterior on θs,j across all j could be used to calculate the posterior probability that seat s was truly won (θs,j > 1/2 for all j). For convenience, and for a fixed ϵ ⩾ 0, we instead define πs ≡ min Pr(θs,j > 1/2 + ϵ | Ft ), 1⩽j⩽Js
(14)
so that πs is small when at least one assertion has substantial posterior evidence that its margin does not exceed ϵ. We investigate only ϵ = 0, which deprioritises seats that appear falsely reported. A positive ϵ would additionally deprioritise seats won by too small a margin to certify feasibly. (Determining an efficient, adaptive choice of ϵ is left for future work.) We calculate each Pr(θs,j > 1/2 + ϵ | Ft ) via a normal approximation to the posterior, or an exact calculation in the case of Ls = 2 (two ballot types). The Filtered scheme restricts attention to the |W| − r + 1 seats with the smallest current seat-level processes Es,t since only these seats enter the product defining Er/W,t in (11). Among this active set, it samples those whose posterior probability exceeds a threshold τ :
Ds,t+1 =
1,
if Es,t ⩽ E(|W|−r+1),t and πs > τ,
0,
otherwise.
(15)
Choosing τ small (we set τ = 0.01) makes the filter liberal: a seat is excluded only when there is substantial evidence that it was not truly won. If no seat in the active set satisfies both πs > τ and has remaining ballots, the πs > τ condition is dropped and Ds,t+1 = 1 is set for every seat in the active set that still has ballots to draw. 6
The conjugacy relies on the sampled ballots being treated as multinomial, an approximation since sampling is without replacement. This model is used only to choose where the filter samples next, which is a predictable decision, so by Theorem 1 the risk limit holds regardless of how accurate the model is.
10
Freestone J, Leung D, Vukcevic D
Greedy Filtered. The Greedy scheme is efficient in benign cases (no falsely reported seats), while the Filtered scheme is more robust when the audit encounters false seats. We combine them together by using both the adaptive window dt (the ‘active set’) and the πs > τ selection filter. To ensure we always draw at least one ballot, we use the following fallback rules: (1) The active set is expanded one seat at a time—adding the next-weakest seat—first to ensure it contains a seat with remaining ballots, and, if necessary, further (up to a maximum size of |W| − r + 1) to ensure it contains a seat that also satisfies πs > τ . (2) Similar to the fallback of the Filter method, if the above rule selects no seats, the πs > τ filter is dropped and the scheme sets Ds,t+1 = 1 for every seat in that window with remaining ballots.
4
Results
We evaluate the sampling schemes on plurality elections audited by ballotpolling7 without replacement. Throughout, we assume there are no invalid ballots8 and set a risk limit of α = 0.05. Section 4.1 shows results from controlled two-candidate simulations (with further results in Appendix D), while in Section 4.2 we simulate audits based on the 2014 Indian parliamentary election. Code and data to reproduce our results are available at: https://github.com/f reejstone/RLA parliamentary majorities code.
Audit strategies. We compare six strategies for certifying the parliamentary majority (testing Hr/W ): Non-adaptive; Greedy with a = 0 or a = 3 in (13); Filtered with τ = 0.01; and Greedy Filtered with a = 0 or a = 3, and τ = 0.01. For comparison, we include two further strategies to act as benchmarks: (1) Top-r seats is a two-phase rule that first samples from R, the r winning seats with the largest reported seat margins, switching to the remaining |W| − r winning seats if the first r seats are exhausted. During the first phase, it effectively attempts to certify all seats in R, i.e., to reject Hr/R . (2) All seats aims to certify all reported winning seats, rather than just a majority. It uses the same sampling rule as Non-adaptive, but sets r = |W|, giving it the stricter target of rejecting H|W|/W . 7
Ballot-polling audits require only the paper trail and the reported outcome, so can be used for any election with a paper trail. This is the simplest scenario and the only one we investigate here. More efficient audits are possible if electronic records of the ballots are available; only some jurisdictions would have these. 8 This is purely for simplicity. The methods can cope with invalid ballots: the assorters would score these as 1/2 (see Section 2.1).
Risk-Limiting Audits for Parliamentary Majorities
11
Tuning parameters. For testing each assertion, we use Ms,j from (4) and the truncated-shrinkage strategy of ALPHA [14] to set λs,j,i . This uses tuning parameters d and η0 , which we set to d = 200 and η0 = 0.51 as per Ek et al. [4]. For the Dirichlet prior used in the filtering strategies, we set the hyperparameters such that the prior mean for each assertion was 0.51 and κs,0 = 200 (for plurality voting this involves solving a set of linear equations). This was to mimic the role of η0 and d in the ALPHA truncated-shrinkage estimator. 4.1
Simulated plurality contests
We first evaluated the audit strategies using synthetic two-candidate plurality contests. Consider a parliament with |S| = 100 seats, so a majority requires r = ⌊|S|/2⌋ + 1 = 51 seats. We fixed the number of reported winning seats at |W| = 60 and each seat to have Ns = 5000 ballots. For each seat s, we set a target proportion of ballots for the reported winner, ptarget,s . The reported winning candidate receives ⌊Ns × ptarget,s ⌋ ballots and the remaining ballots are assigned to the opposing candidate. We allowed some reported winning seats to be false. If nfalse > 0, then nfalse of the seats in W were generated with ptarget,s = 0.48, so that the reported winner did not truly win that seat. The remaining |W| − nfalse seats are correctly reported; we considered two scenarios for setting ptarget,s for these seats: 1. (Homogeneous margins) All truly won seats share the same target vote share ptarget,s = ptarget . We varied ptarget ∈ {0.52, 0.55, 0.60}. 2. (Heterogeneous margins) The target vote share for each truly won seat was drawn independently as ptarget,s ∼ Beta(p̄target κhet , (1 − p̄target )κhet ) , and then truncated to lie in [0.51, 0.999], ensuring that the reported winner truly wins the seat. We fixed κhet = 30, corresponding to moderate heterogeneity, and varied p̄target ∈ {0.52, 0.55, 0.60}. We also varied nfalse ∈ {0, 3, 5}, giving 3 × 3 = 9 configurations for each of the two scenarios. We replicated each one 100 times. Figure 2 shows our results, in terms of the number of ballots sampled until certification. All six strategies for certifying the parliamentary majority were markedly more efficient than the two benchmarks. All seats necessarily uses the most ballots: certifying H|W|/W requires every reported seat to individually exceed the threshold. Top-r seats is handicapped in the same way: during its first phase it effectively certifies Hr/R , which is again driven by the minimum process. In contrast, certifying Hr/W depends on the product of the bottom |W| − r + 1 seat-level processes, so stronger seats can carry the minimum. Among the
12
Freestone J, Leung D, Vukcevic D 0 incorrectly reported
3 incorrectly reported
5 incorrectly reported
300,000
Homogeneous (W = 60)
30,000
10,000
3,000
Heterogeneous (W = 60, kappa = 30)
Total ballots sampled (mean +/− 2 SD, log scale)
100,000
300,000
100,000
30,000
10,000
p = 0.52
p = 0.55
p = 0.60
p = 0.52
p = 0.55
p = 0.60
p = 0.52
p = 0.55
p = 0.60
True winning share All seats
Non−adaptive
Greedy (a=3)
Greedy Filtered
Reported top−r seats
Greedy
Filtered
Greedy Filtered (a=3)
Method
Fig. 2. Simulated two-candidate plurality contests. The total number of ballots sampled to certification: each point shows the mean across 100 replicates and error bars denote ±2 standard deviations. Columns correspond to the number of falsely reported seats, nfalse . Rows correspond to different ways to set the margins: top row shows Scenario 1, bottom row shows Scenario 2. Results for the All seats method are only reported for nfalse = 0, since for nfalse > 0 it tests H|W|/W , which is a true null and requires a full recount with probability at least 1 − α.
majority schemes, Non-adaptive was least efficient, since it samples every seat regardless of how informative each is. The ‘filtered’ variants outperform their pure-greedy counterparts once nfalse > 0, as expected, by redirecting effort away from seemingly false seats. We provide some further results in Appendix D, in which we also vary the number of reported winning seats |W| ∈ {51, 60, 80} and the heterogeneity parameter κhet ∈ {10, 30, 100}. Our broad conclusions remained the same.
4.2
Indian general election
To evaluate the audit framework on a realistic parliament-scale election, we simulated audits using data from the 2014 Indian general election of the Lok Sabha (Lower House). We used candidate-level vote totals for all 543 constituencies
Risk-Limiting Audits for Parliamentary Majorities
13
from a community-scraped ECI dataset9 , originally sourced from the Election Commission of India’s results portal. The Bharatiya Janata Party (BJP) won |W| = 282 of the 543 seats, forming a single-party majority. Certifying a BJP majority requires verifying that at least r = ⌊543/2⌋ + 1 = 272 of these 282 reported winning seats were truly won. Across the 282 BJP seats, the number of candidates ranged from Ls = 5 to Ls = 43 (median 15), and the number of ballots per seat ranged from approx. 87,000 to 1,512,000 (median 1,010,000), summing to a total of about 282 million ballots. The winner’s vote share ranged from 0.264 to 0.758 (median 0.495). Because each Ls -candidate plurality contest has Js = Ls − 1 SHANGRLA assertions, the minimum assorter mean across these head-to-head comparisons is the bottleneck for certifying a seat. In the BJP seats, this minimum assorter mean ranged from 0.5002 to 0.7812 (median 0.5831). We generated nfalse falsely reported seats by swapping the vote totals of the reported winner and the runner-up. The BJP remains the reported winner, but the runner-up is the true winner in the underlying ballot population. We apply this perturbation to the nfalse most marginal BJP seats, where marginality is measured by the two-candidate margin between the BJP candidate and the runner-up. We vary nfalse ∈ {0, 3, 5}. We replicated each of these three configurations 100 times. Figure 3 reports the total number of ballots sampled before certification. We see qualitative behaviour similar to Section 4.1. All adaptive schemes dominated the benchmarks by one to three orders of magnitude: the All seats baseline required roughly 250 million ballots, and the Reported top-r seats benchmark required around 5 million across all nfalse , whereas the others typically required only a few hundred thousand. Within the parliamentary-majority schemes, the Greedy (a = 0) scheme degraded sharply as we increased nfalse (going from about 400k to 3.8M ballots), demonstrating the price of pouring effort into false seats. Widening the active set to a = 3 restored robustness at nfalse = 0, where it is the most efficient scheme overall. However, once nfalse > 0, the ‘filtered’ schemes took over, showing consistently superior performance when nfalse = 5.
5
Discussion
We have developed a risk-limiting audit for parliamentary majorities. We formalised this task as a partial conjunction hypothesis test and constructed a sequential parliament-level audit statistic by combining seat-level statistics. The resulting procedure controls the risk of certifying an incorrect majority outcome, while allowing the auditor to adaptively allocate sampling effort across seats. 9
https://github.com/datameet/india-election-data/tree/master/parliament-electio ns/election2014, accessed on 2026-04-19.
14
Freestone J, Leung D, Vukcevic D 0 incorrectly reported
3 incorrectly reported
5 incorrectly reported
Total ballots sampled (log scale)
100,000,000
10,000,000
1,000,000
)
d re
=3 (a
lte
d
Fi re e
dy
Fi
lte
re
)
ed er Fi lt
dy re e G G
dy
=3 (a y
ed re
ive pt
re e G
da
rs p−
−a on
to or ep R
G
ts
ts ea
ea Al ls te
d
Fi dy
re e
N
)
d re
=3 (a
lte
d
Fi
lte
re
)
ed er Fi lt
dy re e G G
dy
ive
=3 (a y
ed re
pt
re e G
da
rs p−
−a on
to or ep R
G
ts
ts ea
ea Al ls te
d
Fi dy
re e
N
)
d re
=3 (a
lte
d
Fi dy
lte
re
)
ed er Fi lt re e G G
dy
ive
=3 (a y
ed re G
pt
re e
da
N
on
−a
p− to
d te or ep R
G
ea rs
Al ls
ea
ts
ts
100,000
Fig. 3. India 2014. Box plots of the total number of ballots sampled to certification across 100 replicates, with one panel per number of falsely reported seats nfalse . Results for the All seats method are omitted for nfalse > 0, like in Figure 2.
Although our experiments focused on ballot-polling audits of plurality elections, the methodology is not limited to this setting. Because it is built on SHANGRLA assertions, it inherits that framework’s generality, covering a wide variety of social choice functions, other audit styles such as comparison audits, and settings with missing ballots. It can stretch further still, targeting any hypothesis of the form Hr/W for any value of r, not just majorities. In practice, each constituency needs the same infrastructure as any singlecontest RLA. The extra burden imposed by the majority audit is cross-seat coordination: a central authority must aggregate data from the sampled ballots, update Er/W,t , and issue the next round’s instructions. This overhead favours a few large-batch rounds rather than the single-ballot rounds we used for exposition (n.b., the risk limit holds for any batch size). Designing this coordination to run efficiently in practice is an important problem in its own right, one we leave for future work. The burden is nonetheless far lighter than auditing every seat: the total sample is orders of magnitude smaller, and the adaptive schemes concentrate effort on a few seats, allowing most constituencies to stop early. In our experiments we only considered ballot-polling audits. When the relevant election infrastructure is available, more sophisticated and efficient audit designs are possible, such as ballot-level comparison audits. The sample size required for such audits typically grows linearly with respect to the reciprocal of the margin, whereas for ballot-polling it grows quadratically [8]. Thus, combining our method with a comparison audit should lead to even smaller sample sizes, in the order of thousands of ballots for the Indian election example.
Risk-Limiting Audits for Parliamentary Majorities
15
Our audit certifies that a reported winning party won a majority of seats in its own right. This is the pertinent question when a single party, or a predeclared alliance, wins an outright majority, as in our Indian example. When no party/alliance wins such a declared majority, our method is not relevant: if a governing coalition is negotiated only after the election, the majority to certify is not known while the ballots remain available, and such elections are better served by auditing the individual seats. However, because our audit is built from ordinary seat-level assertions, it can be run in tandem with a full seat-by-seat audit: the majority can be certified ‘midstream’, as soon as the parliamentlevel statistic Er/W,t crosses 1/α, while the seat-by-seat audit continues towards certifying every individual seat or recounting those that appear miscounted. Several directions remain for future work: (1) It would be useful to compare our sampling strategies in diverse scenarios, such as other election and audit types. (2) The efficiency of the audit depends on the choice of sampling scheme and ALPHA tuning parameters; future work could study alternative choices. This includes investigating the benefit of setting them based on reported vote totals/margins, in scenarios where these are available. (3) Determining optimal ways to set sample sizes per seat for each round would be useful for practice. (4) The process Er/W is anytime-valid for testing Hr/W , but might not be optimal. Similarly, the process Es,t is one option for summarizing seat-level information, but might not be optimal. Investigating if there are more efficient alternatives would be of interest.
References [1] Benjamini, Y., Heller, R.: Screening for partial conjunction hypotheses. Biometrics 64(4), 1215–1222 (2008) [2] Bogomolov, M., Heller, R.: Replicability across multiple studies. Statistical Science 38(4), 602–620 (2023), Preprint: arXiv:2210.00522 [3] Ek, A., Stark, P.B., Stuckey, P.J., Vukcevic, D.: Adaptively weighted audits of instant-runoff voting elections: AWAIRE. In: E-Vote-ID 2023. LNCS, vol. 14230, pp. 35–51. Springer (2023), Preprint: arXiv:2307.10972 [4] Ek, A., Stark, P.B., Stuckey, P.J., Vukcevic, D.: Efficient weighting schemes for auditing instant-runoff voting elections. In: Financial Cryptography and Data Security. FC 2024. LNCS, vol. 14746, pp. 18–32. Springer (2025), Preprint: arXiv:2403.15400 [5] Fisher, R.A.: Statistical methods for research workers. Hafner Publishing Co., New York, fourteenth edn. (1973) [6] Gablenz, P., Sabatti, C.: Catch me if you can: signal localization with knockoff e-values. Journal of the Royal Statistical Society Series B: Statistical Methodology 87(1), 56–73 (2025), Preprint: arXiv:2306.09976
16
Freestone J, Leung D, Vukcevic D
[7] Hoang, A.T., Dickhaus, T.: Combining independent p-values in replicability analysis: A comparative study. Journal of Statistical Computation and Simulation 92(10), 2184–2204 (2022), Preprint: arXiv:2104.13081 [8] Lindeman, M., Stark, P.B.: A gentle introduction to risk-limiting audits. IEEE Security & Privacy 10(5), 42–49 (2012) [9] Mohanty, V., Culnane, C., Stark, P.B., Teague, V.: Auditing Indian elections. In: E-Vote-ID 2019. LNCS, vol. 11759, pp. 150–165. Springer (2019), Preprint: arXiv:1809.04235 [10] Ottoboni, K., Stark, P.B., Lindeman, M., McBurnett, N.: Risk-limiting audits by stratified union-intersection tests of elections (SUITE). In: EVote-ID 2018. LNCS, vol. 11143, pp. 174–188. Springer (2018), Preprint: arXiv:1809.04235 [11] Ramdas, A., Grünwald, P., Vovk, V., Shafer, G.: Game-theoretic statistics and safe anytime-valid inference. Statistical Science 38(4), 576–601 (2023), Preprint: arXiv:2210.01948 [12] Spertus, J.V., Sridhar, M., Stark, P.B.: Sequential stratified inference for the mean. arXiv:2409.06680 (2026) [13] Stark, P.B.: Sets of half-average nulls generate risk-limiting audits: SHANGRLA. In: Financial Cryptography and Data Security. FC 2020. LNCS, vol. 12063, pp. 319–336. Springer (2020), Preprint: arXiv:1911.10035 [14] Stark, P.B.: ALPHA: Audit that learns from previously hand-audited ballots. The Annals of Applied Statistics 17(1), 641–679 (2023), Preprint: arXiv:2201.02707 [15] Ville, J.: Etude critique de la notion de collectif, Monographies des Probabilites, vol. 3. Gauthier-Villars Paris (1939) [16] Vovk, V., Wang, R.: E-values: Calibration, combination and applications. The Annals of Statistics 49(3), 1736–1754 (2021), Preprint: arXiv:1912.06116 [17] Vovk, V., Wang, R.: Merging sequential e-values via martingales. Electronic Journal of Statistics 18(1), 1185–1205 (2024), Preprint: arXiv:2007.06382 [18] Waudby-Smith, I., Ramdas, A.: Estimating means of bounded random variables by betting. Journal of the Royal Statistical Society Series B: Statistical Methodology 86(1), 1–27 (2024), Preprint: arXiv:2010.09686
A
Proof that Ms,j is a test supermartingale under Hs,j
Non-negativity holds since Xs,j,i ⩾ 0 and λs,j,i ⩽ µ−1 s,j,i in (4), and Ms,j,0 = 1 by definition. To prove (6), from (5) we can see that, when Hs,j is true, E[Ms,j,t | Ft−1 ] = Ms,j,t−1 Ds,t · 1 + λs,j,ns (t) (E[Xs,j,ns (t) | Ft−1 ] − µs,j,ns (t) ) + 1 − Ds,t ⩽ Ms,j,t−1 ,
Risk-Limiting Audits for Parliamentary Majorities
17
where the equality holds because the quantities Ms,j,t−1 , Ds,t , λs,j,ns (t) and µs,j,ns (t) only depend on Ft−1 ,10 and the inequality holds because E[Xs,j,ns (t) | Ft−1 ] ⩽ µs,j,ns (t) under Hs,j : here µs,j,ns (t) is the mean of the remaining assorter values under x̄s,j = 1/2, the boundary (hence maximal) value for Hs,j .
B
Proof of Theorem 1
Es is anytime-valid for Hs : Suppose Hs is true. By the union (3), some js ∈ {1, . . . , Js } has Hs,js true, and (7) gives Es,t ⩽ Ms,js ,t for all t ⩾ 0. It thus suffices to show that Pr(supt⩾0 Ms,js ,t ⩾ 1/α) ⩽ α under Hs,js . This holds because Ms,js is a test supermartingale (Appendix A) and Ville’s inequality [15]. T E1/C is anytime-valid for H1/C : Suppose H1/C = s∈C Hs is true, so Hs holds for all s ∈ C. By the union (3), each s ∈ C has an assertion js ∈ {1, . . . , Js } Q with Hs,js true; use these to define M1/C,t ≡ s∈C Ms,js ,t . From (7), E1/C,t ⩽ M1/C,t for all t ⩾ 0. By this bound, it suffices to show that M1/C is a test supermartingale and apply Ville’s inequality, as follows. From the expression in (5), we have that E[M1/C,t | Ft−1 ] " # Y = M1/C,t−1 E Ds,t · 1 + λs,js ,ns (t) (Xs,js ,ns (t) − µs,js ,ns (t) ) + 1 − Ds,t | Ft−1 s∈C (a)
Y
= M1/C,t−1
Ds,t · 1 + λs,js ,ns (t) (E[Xs,js ,ns (t) | Ft−1 ] − µs,js ,ns (t) ) + 1 − Ds,t
s∈C (b)
⩽ M1/C,t−1 , where (a) holds because, conditional on Ft−1 , the factors Ds,t , λs,js ,ns (t) and µs,js ,ns (t) are constants and the Xs,js ,ns (t) are independent across s ∈ C; (b) follows as in Appendix A. Q Since Ms,js ,t ⩾ 0 and M1/C,0 = s∈C Ms,js ,0 = 1, M1/C is a test supermartingale, and Ville’s inequality gives anytime-validity for H1/C . Er/W is anytime-valid for Hr/W : Suppose Hr/W is true. By (10), there is a subset C0 ⊂ W with |C0 | = |W| − r + 1 and H1/C0 true. From above, E1/C0 is anytime-valid for testing H1/C0 , so Pr(supt E1/C0 ,t ⩾ 1/α) ⩽ α. The ‘min’ representation (11) gives Er/W,t ⩽ E1/C0 ,t for all t, so Pr(supt Er/W,t ⩾ 1/α) ⩽ Pr(supt E1/C0 ,t ⩾ 1/α) ⩽ α, proving the claim. 10
Here λs,j,ns (t) depends on FTs (ns (t))−1 , and FTs (ns (t))−1 ⊂ Ft−1 since Ts (ns (t)) ⩽ t. Likewise µs,j,ns (t) depends only on the first ns (t) − 1 ballots for seat s, all revealed by time t − 1.
18
C
Freestone J, Leung D, Vukcevic D
Comparison with the Fisher combination approach of Mohanty et al. [9]
We first show that our direct way of combining the seat-level e-processes via products is substantially more efficient than the Fisher combination proposal of Mohanty et al. [9, Section 4], holding the sampling scheme fixed and common to both methods. Fix a time t and let k ≡ |W| − r + 1. Certifying Hr/W requires rejecting the intersection null H1/C for every C ⊂ W with |C| = k (10). For a given C, our Q process uses E1/C,t = s∈C Es,t and rejects when it reaches 1/α (Theorem 1). Mohanty et al. [9] instead form seat-level p-values Ps,t ≡ 1/Es,t and combine them by Fisher’s method [5] to form the statistic X Y −2 ln Ps,t = 2 ln Es,t . s∈C
s∈C
This Fisher combined statistic is referenced to a standard χ22k distribution under Q the null H1/C . That is, their test rejects H1/C when s∈C Es,t ⩾ exp(χ22k,1−α /2), where χ22k,1−α is the 1 − α quantile of the χ22k distribution. In both cases, the critical subset C is the one with the smallest product, namely the one with the k smallest seat-level processes, so the majority decision reduces to a threshold Qk on s=1 E(s),t . Our method uses 1/α, theirs uses exp(χ22k,1−α /2). The two thresholds coincide only when k = 1, since χ22,1−α = 2 ln(1/α). For k ⩾ 2, the Fisher threshold is strictly larger, so our test rejects whenever theirs does and is uniformly at least as powerful. At α = 0.05, our threshold is 1/α = 20, while theirs is ≈115 at k = 2 and ≈2.3 × 107 at k = 11 (ten spare seats). Because a seat-level process grows geometrically in the number of ballots drawn, this gap inflates the required sample by a factor of about χ22k,1−α /(2 ln(1/α)), roughly 5.7 at k = 11, which is a substantial difference in efficiency. Finally, we point out that Mohanty et al. [9] justified the theoretical validity of their Fisher combination proposal incorrectly. They mistakenly treated their p-value processes Ps,t as independent across seats, and hence the typical quantile χ22k,1−α under Fisher’s combining method can be used as the rejection threshold. However, this proof argument only works under the non-adaptive sampling scheme in Section 3, because all the other sampling schemes in Section 3 necessarily inject dependence between the e-processes (and therefore the transformed p-value processes) across seats. In contrast, our exposition above actually proves the anytime validity of their method correctly, by demonstrating that theirs is simply uniformly more conservative than our direct e-process-based approach.
D
Further simulations
Figures 4–7 show extra simulations results beyond those presented in Section 4.1.
Risk-Limiting Audits for Parliamentary Majorities
19
Homogeneous Margins −− W = 51, 60, 80 S = 100 total seats, N = 5000 ballots/seat, alpha = 0.05
0 incorrectly reported
3 incorrectly reported
5 incorrectly reported
100,000
W = 51
30,000
100,000
W = 60
Total ballots sampled (mean +/− 2 SD, log scale)
10,000
300,000
30,000
10,000
3,000
100,000
W = 80
10,000
1,000 p = 0.52
p = 0.55
p = 0.60
p = 0.52
p = 0.55
p = 0.60
p = 0.52
p = 0.55
p = 0.60
True winning share All seats
Non−adaptive
Greedy (a=3)
Greedy Filtered
Reported top−r seats
Greedy
Filtered
Greedy Filtered (a=3)
Method
Fig. 4. Extended simulation results for the two-candidate plurality contests of Scenario 1 from Section 4.1. The plot matches the first row of Figure 2, but additionally varies |W| ∈ {51, 60, 80}.
20
Freestone J, Leung D, Vukcevic D
Heterogeneous Margins −− W = 51, 60, 80, kappa = 10 S = 100, N = 5000 ballots/seat, alpha = 0.05 0 incorrect
3 incorrect
5 incorrect
200,000
W = 51
100,000
100,000
W = 60
Total ballots sampled (mean +/− 2 SD, log scale)
50,000 300,000
30,000
100,000
W = 80
10,000
1,000 mean p = 0.52
mean p = 0.55
mean p = 0.60
mean p = 0.52
mean p = 0.55
mean p = 0.60
mean p = 0.52
mean p = 0.55
mean p = 0.60
Mean winning share All seats
Non−adaptive
Greedy (a=3)
Greedy Filtered
Reported top−r seats
Greedy
Filtered
Greedy Filtered (a=3)
Method
Fig. 5. Extended simulation results for the two-candidate plurality contests of Scenario 2 from Section 4.1. The plot matches the second row of Figure 2, but additionally varies |W| ∈ {51, 60, 80} and sets κhet = 10.
Risk-Limiting Audits for Parliamentary Majorities
21
Heterogeneous Margins −− W = 51, 60, 80, kappa = 30 S = 100, N = 5000 ballots/seat, alpha = 0.05 0 incorrect
3 incorrect
5 incorrect
200,000
W = 51
100,000
300,000
100,000
W = 60
Total ballots sampled (mean +/− 2 SD, log scale)
50,000
30,000
10,000
300,000
100,000
W = 80
30,000
10,000
3,000
mean p = 0.52
mean p = 0.55
mean p = 0.60
mean p = 0.52
mean p = 0.55
mean p = 0.60
mean p = 0.52
mean p = 0.55
mean p = 0.60
Mean winning share All seats
Non−adaptive
Greedy (a=3)
Greedy Filtered
Reported top−r seats
Greedy
Filtered
Greedy Filtered (a=3)
Method
Fig. 6. Extended simulation results for the two-candidate plurality contests of Scenario 2 from Section 4.1. The plot matches the second row of Figure 2, but additionally varies |W| ∈ {51, 60, 80} and sets κhet = 30.
22
Freestone J, Leung D, Vukcevic D
Heterogeneous Margins −− W = 51, 60, 80, kappa = 100 S = 100, N = 5000 ballots/seat, alpha = 0.05 0 incorrect
3 incorrect
5 incorrect
300,000
100,000
W = 51
300,000
100,000
W = 60
Total ballots sampled (mean +/− 2 SD, log scale)
30,000
30,000
10,000
100,000
W = 80
10,000
1,000 mean p = 0.52
mean p = 0.55
mean p = 0.60
mean p = 0.52
mean p = 0.55
mean p = 0.60
mean p = 0.52
mean p = 0.55
mean p = 0.60
Mean winning share All seats
Non−adaptive
Greedy (a=3)
Greedy Filtered
Reported top−r seats
Greedy
Filtered
Greedy Filtered (a=3)
Method
Fig. 7. Extended simulation results for the two-candidate plurality contests of Scenario 2 from Section 4.1. The plot matches the second row of Figure 2, but additionally varies |W| ∈ {51, 60, 80} and sets κhet = 100.