ALIVE: Warnings Before Exclusion in Budgeted Multi-Source Learning Xiyang Zhang1,2 , Hongzhi Wang1∗ , Yuanhe Tian2 1
Harbin Institute of Technology 2 Zhongguancun Academy [email protected], [email protected], [email protected]
arXiv:2607.29400v1 [cs.LG] 31 Jul 2026
Abstract A routing decision can be revised at the next transaction, but a latched source exclusion persists across later decisions. We ask what evidence should authorize these unequalpersistence actions when finite-population auditing and learning share a budget. ALIVE (Action-Layered Intervention via Evidence) is an auditable control layer: one randomized without-replacement prefix supplies cached evidence, heuristic warnings drive non-latching floor-bounded routing, and only two fresh simultaneous certificate separations may latch an exclusion request subject to capacity-feasible activation. Conditional on fixed support and labels under an ideal uniform audit permutation, any predictable controller preserving this interface inherits an anytime familywise bound of δ on acting against a source that fails the pre-fixed absolute or relative strict-majority-disagreement predicate. With a published known-size, all-strict-majority PPR engine, median evidence count fell from 304 to 96 identities in e40 and from 171 to 62 in e60, while both engines used 48 in e80. In the matched CIFAR controller, the persistent-action layer added +0.1935 accuracy-AUBC percentage points over routing-only in all ten paired seed clusters. The +0.1954-point full-system contrast against CBR was also positive but did not meet the predeclared multiplicity-adjusted criterion (conditional Holm-adjusted sign-flip reference value =.097656). On a fixed natural panel, exploratory PPR used a median closure prefix of 95 rather than 105 for exploratory Serfling/FPC, but still exposed 88.0% of the panel and had no downstream task. Together these results map a restraint–power–cost–utility boundary: the action contract controls a defined persistent decision, while net value depends on evidence margin, audit cost, and budget regime.
Introduction In a crowd-labeled or repeated-label learner, reducing an annotator’s share for the next batch is reversible; excluding that annotator from future optimization is not. The latter changes subsequent data availability and source opportunity, even when a capacity check occasionally suppresses the exclusion. A noisy score should therefore not silently escalate into a persistent action. We study the evidence-to-action interface needed to separate these decisions when auditing and learning draw from one budget. ∗
Corresponding author.
This distinction exposes three questions. Statistical validity asks whether evidence supports a pre-specified action predicate. Semantic validity asks whether that predicate identifies error or harm. Decision validity asks whether acting yields net downstream value after audit cost. ALIVE controls the first question for a finite-population majoritydisagreement target and evaluates the latter two separately; it does not treat disagreement as ground truth. Our setting is the shared-identity, multi-annotator case of multi-source learning, not disjoint domain streams or federated clients. ALIVE (Action-Layered Intervention via Evidence) makes this separation operational. It consumes one randomized, without-replacement prefix of the complete all-source identity intersection. A heuristic warning may alter non-latching, floor-bounded routing and request more audit. A persistent request requires simultaneous source-by-prefix confidence separation on two fresh evidence advances, and activation additionally requires that remaining sources can fill the optimizer batch. Evidence acquired during a decision is not action-defining until the next decision. Conditional on fixed support and potential labels under an ideal uniform audit permutation, any predictable controller that preserves this prefix and certificate interface inherits an anytime familywise bound of δ on persistent action against a source that fails either the pre-fixed absolute-threshold or relative strict-majority-disagreement condition. Warning quality is outside this statement; changing routing can change evidence delay but not the declared false-action target.
Contributions. Action-layered formulation: we separate heuristic warnings, latched requests, and capacity-feasible active exclusions as actions with different persistence under one budget. Interface-preserving evidence contract: we specify the randomized-prefix, frozen-cache, latch, and timing invariants under which predictable controllers inherit familywise false-persistent-action control. We also derive a schedule-conditional audit-to-latch bound and an analytical direct-contrast strengthening of the evaluated coordinatewise certificate. Cross-layer evaluation: we measure null restraint, evidence delay, panel exposure, matched persistentaction value, and end-to-end operating regimes instead of treating certificate validity as a proxy for learning benefit.
Related Work Source filtering and aggregation. Closest to a persistent source action are confidence-based worker eviction, disagreement-based reputation against adversaries, and label-consistency filtering before aggregation (Joglekar, Garcia-Molina, and Parameswaran 2013; Jagabathula, Subramanian, and Venkataraman 2014; Li et al. 2024). Those methods target worker error, adversarial aggregation, or label purification. ALIVE instead gives anytime familywise control of false persistent actions for a declared finite-population disagreement target and separates a non-latching warning from a persistent request. Active learning with imperfect labelers jointly chooses instances and oracles, estimates expertise, or trades quality against cost (Donmez and Carbonell 2008; Donmez, Carbonell, and Schneider 2009; Zheng, Scott, and Deng 2010; Yan et al. 2011, 2012; Huang et al. 2017; Chakraborty 2020; Gao and Saar-Tsechansky 2020). Latent-label and repeated-label models aggregate answers or estimate reliability (Dawid and Skene 1979; Raykar et al. 2010; Sheng, Provost, and Ipeirotis 2008; Karger, Oh, and Shah 2014). ALIVE does not estimate gold labels or source accuracy; it uses stable shared identities to bind a specified evidence event to actions of different duration. Sequential finite-population evidence. Samplingwithout-replacement concentration is classical (Hoeffding 1963; Serfling 1974); dedicated confidence sequences and gambling constructions can be tighter (Waudby-Smith and Ramdas 2020; Ryu and Wornell 2024). The base controller uses a Hoeffding comparison with explicit alpha spending. We derive a strict-majority direct-contrast certificate that pathwise contains this base certificate, and instantiate a prior–posterior ratio (PPR) alternative with a uniform prior when the complete population size is known and every identity has a strict majority. These are separate level-δ engines. On a common source–prefix path, the nested direct/base certificate union equals the direct certificate; a union with non-nested PPR requires a pre-fixed error split. The concentration tools are established; our contribution is the action interface, its inheritance result, and the measurement of delay and downstream cost. Risk-sensitive actions. Conservative exploration and selective prediction limit risky decisions (Wu et al. 2016; Garcelon et al. 2020; Chow 1970; El-Yaniv and Wiener 2010). Their constraints apply to rewards or predictions, whereas ALIVE controls a finite-population source-action predicate and allows capacity to suppress activation without clearing the evidence latch. Budget-sensitive data selection. Coresets, gradient matching, uncertainty, and class balancing choose training examples (Sener and Savarese 2018; Killamsetty et al. 2021; Settles 2009; Bengar et al. 2022); budget-dependent strategy selection motivates a strong cheap incumbent (Hacohen, Dekel, and Weinshall 2022; Hacohen and Weinshall 2023; Zhang et al. 2023). We therefore hold the selector fixed in the matched action ablation. Compute-aware selection charges selection and optimization (Yin and Rush 2025; Wan, Zhang,
and Jin 2025); our shared ledger has the same accounting aim but is not a hardware or monetary cost model. Robust losses act after acquisition (Natarajan et al. 2013; Patrini et al. 2017; Han et al. 2018; Li, Socher, and Hoi 2020); ALIVE changes source opportunity and eligibility rather than correcting labels or estimating a noise transition.
Setting and Shared Ledger We study shared-identity multi-annotator learning with S ≥ 3 sources and zero-based decisions t = 0, 1, . . .. At decision t, the learner chooses integer source counts nt = (n1t , . . . , nSt ), acquires candidate window Wt , selects training batch Bt ⊆ Wt , and updates parameters. An observation contains features, a source label Ys (i), an opaque source key s, and, when available, shared identity i. No action-defining branch receives an environment name, corruption rate, cleansource flag, validation outcome, or test metric. Fixed labels, S ≥ 3, and nonempty common support are guarantee preconditions rather than facts inferred at runtime: the current implementation does not verify label stability, and an empty intersection disables certification instead of aborting the incumbent learner. Every method obeys the declared accounting contract X acquire + cscore + cmaintain + ctrain ct ≤ B. (1) t t t t
Complete audit groups occupy candidate slots and are charged once. Entropy ranking pays one forward pass, CBR no ranking pass, and Full-Consensus every source query and tally. Each method stops before exceeding B. Thus “same budget” means one frozen reproducibility ledger, not equal labels, gradient steps, seconds, FLOPs, energy, or money.
Evidence-Gated Source Actions Randomized shared-identity prefix Let U be the fixed intersection of all sources’ audit-eligible supports. Before audit labels are read, the controller draws one domain-separated pseudorandom permutation π of U , idealized as uniform in the theory. Each request consumes the next unused identities and admits only complete all-source groups. Adaptive request sizes may change prefix length, but cannot select, revisit, or replace an identity according to its labels. Evidence acquired at decision t changes the controller only at t + 1. One-decision execution and audit. At decision t, predicates are evaluated only from the prefix cached through t−1; the resulting states fix audit share, routing, integer allocation, exclusion request, and eligibility. The next complete groups and ordinary candidates are then acquired. New groups may enter the audit data structure, but the frozen cache prevents them from changing action t. After scoring, capacity fallback is recorded, the eligible batch is selected and trained, and all transitions and charges are logged; new evidence can first act at t + 1. A strict replay auditor checks four invariants before reading utility: admitted identities are exactly one prefix with no duplicates or partial groups; Provisional sources remain
eligible; first latch requires two certificate-positive fresh advances and can never clear; and recomputed event charges reproduce an in-budget ledger. These are trajectory facts, not inferences from accuracy or a method name. For identity i, let m(i) be its unique strict cross-source majority label; identities without a strict majority abstain. If a1 , . . . , an are the first n comparable identities induced by π, define n X Zs (i) = I{Ys (i) ̸= m(i)}, pbs,n = n−1 Zs (ak ). k=1
(2) The focal source participates in its majority. This is not a leave-one-source- out or truth estimator. Before labels are read, the action designer fixes τ ∈ (0, 1). ALIVE uses τ = 1/S as a symmetric exact-null policy reference, not an estimated or universal corruption rate or a consequence of majority voting.
Warning, certificate, and actions After every source has at least eight comparable observations, source s has the empirical warning X 1 Ws,n = I pbs,n > τ, pbs,n > pbj,n . (3) S−1 j̸=s
This warning is heuristic. For δ = 0.05, set ( r ) log(2Sn(n + 1)/δ) rn = min 1, , 2n Ls,n = max{0, pbs,n − rn },
Us,n = min{1, pbs,n + rn }. (4) A fresh evidence block supports separation when 1 X Uj,n . (5) Cs,n = I Ls,n > τ, Ls,n > S−1 j̸=s
A persistent request first latches only after Cs,n = 1 at two consecutive fresh-evidence advances and the estimated remaining horizon is at least two decisions. These are nested prefixes, not independent replications. Figure 1 summarizes the state machine. Warnings change the next floor-bounded opportunity without latching or changing eligibility; they do not undo past acquisition or training. Two certificate-positive advances latch a request, while each transaction activates ineligibility only if remaining-source capacity fills the batch. The frozen 12.5/25% audit shares, 15/40% routing floor/cap, two confirmations, and horizon guard are heuristics rather than theorem consequences. Audit floors preserve observability but do not recover a source; a utility certificate, recovery rule, or valueof-information schedule is not part of the tested method. ALIVE–CBE (class-balanced entropy) composes the controller with its namesake P selector: a charged forward pass computes ht (x) = − k pθt (k | x) log pθt (k | x), a largestremainder quota covers observed classes when feasible, and the highest entropy eligible candidates are selected within class. This selector is an incumbent, not part of the novelty claim.
Matched controller variants ALIVE–CBR (class-balanced random) changes only withinclass ranking to a seeded random permutation and retains the controller, class quota, timing, and ledger. Its horizon still uses the counterfactual CBE step cost so selector choice cannot move certification, while realized ranking cost is zero. This is the matched strong-incumbent composition, not a new selector. The supplement specifies Full-Consensus and evidence- or clock-timed selector switches as diagnostic boundaries and reports their complete outcomes.
Evidence-to-Action Guarantee Condition on common support U and the complete potentiallabel table Y = {Ys (i) : s ∈ [S], i ∈ U }. The comparable population is ( ) S X A = i ∈ U : max I{Ys (i) = y} > S/2 , y
s=1
M = |A|. The theorem idealizes the pseudorandom audit permutation as uniform and takes probability over a hypothetical re-randomization conditional on the same (U, Y ). Domainseparating the audit seed from source-noise and optimizer seeds prevents shared generator state, but is not itself a probabilistic guarantee. For M ≥ 1, define the fixed finitepopulation target X ps = M −1 I{Ys (i) ̸= m(i)}. i∈A
Let X N = s : ps ≤ (S − 1)−1 pj or ps ≤ τ . j̸=s
Write AN , LN , and CN for any active exclusion, latched request, and certificate-positive prefix, respectively, among sources in N . The proof separates the evidence engine from the controller through AN ⊆ LN ⊆ CN ⊆ Eδc .
(6)
The action contract supplies the first two inclusions. The base result below supplies the last one with alpha-spent Hoeffding intervals; the PPR specialization later gives a separate level-δ engine under stronger population conditions. Theorem 1 (randomized-prefix coverage). If π | (U, Y ) is a uniform permutation, then the intervals in Equation (4) satisfy Pr{ps ∈ [Ls,n , Us,n ], s ∈ [S], n ∈ [M ] | U, Y } ≥ 1 − δ. π (7) Denote the simultaneous-coverage event inside the braces by Eδ .
ALIVE EVIDENCE-GATED SOURCE ACTIONS
cached evidence fixes the state used at transaction t
CACHED PREFIX
WARNING + CS
CLEAR
PROVISIONAL
CERTIFIED
shared-ID order complete to t-1
simultaneous [L,U] separation event C
base route all eligible
reversible route floor; all eligible
latched after 2 C exclude if capacity
TRANSACTION AT t
new evidence is committed only after the action
1 ROUTE + ALLOCATE
2 ACQUIRE GROUP
3 SCORE + UPDATE
4 COMMIT EVIDENCE
state-fixed policy
next shared ID
capacity; train
first visible at t+1
Certificate = peer-disagreement separation; not truth, harm, or downstream safety. ONE PROXY LEDGER
AUDIT | ACQUIRE | ROUTE | SCORE | TRAIN | MAINTAIN
Figure 1: One-decision evidence-to-action transaction. State, routing, allocation, exclusion request, and scoring eligibility are fixed from evidence cached through t − 1. New groups may update the audit data structure during the transaction, but the frozen cache lets them affect actions only at t + 1. Two certificate-positive advances latch a request; optimizer ineligibility is activated only when remaining-source capacity can fill the batch. Proof sketch. Restricting a uniform permutation of U to the fixed subset A gives a uniform permutation of A. For fixed (s, n), the first n binary disagreements are sampled without replacement from a fixed Bernoulli population. Hoeffding’s 2 comparison gives Prπ {|b ps,n − ps | ≥ ϵ | U, Y } ≤ 2e−2nϵ . When the unclipped radius is below one, substitution bounds failure by δ/[Sn(n + 1)]; otherwise the clipped interval is [0, 1]. Union bounding over sources and prefixes uses P −1 = 1. Simultaneity in n permits the prefix n≥1 [n(n+1)] length and decision time to depend on the observed prefix without another optional-stopping correction. □ Theorem 1 is an exact-real statement; the base controller uses ordinary binary floating point without outward rounding and is not bit-level verified. In the synthetic PPR study, membership in the conservative weak-boundary PPR set is exact-integer, but outward-padded floating hull endpoints enter the inherited predicate. Natural PPR instead evaluates its action-defining count inequalities exactly in integers. Corollary 1 (false action-predicate certification). On Eδ the two null branches are explicit. P If ps ≤ τ , then Ls,n ≤ ps ≤ τ ; if ps ≤ (S − 1)−1 j̸=s pj , then Ls,n ≤ ps ≤ P P (S − 1)−1 j̸=s pj ≤ (S − 1)−1 j̸=s Uj,n . Either branch contradicts one certificate inequality. Therefore Pr{∃s ∈ N ever certified | U, Y } ≤ δ. (8) π
The two-confirmation and horizon filters can only remove latch and action events relative to certificate-positive prefixes; the confirmations use nested prefixes and do not change δ to δ 2 . If M = 0, no disagreement mean or certification event is defined and the implementation retains [0, 1]. Action-inheritance corollary (false persistent actions). Let a predictable controller choose warning actions, routing, and the number of next prefix groups from evidence available before each decision. Provided it never skips, repeats, or replaces an identity according to labels, and every
hard action on source s requires a latch whose first entry followed a certificate-positive prefix, the event of ever acting is contained in the event of ever certifying. Equation (8) and Equation (6) therefore bound by δ the probability of any hard action on s ∈ N . Capacity fallback can only remove active-exclusion events; it cannot create one without a latch. The guarantee is consequently modular to non-latching policy changes that preserve the randomized prefix. It neither covers warning actions nor licenses a label-dependent audit rule. Corollary 2 (sufficient detection at a covered P prefix). For ∆abs = ps − τ and ∆rel = ps − (S − 1)−1 j̸=s pj , simultaneous coverage implies that Equation (5) holds whenever both gaps are positive and rn < min{∆abs /2, ∆rel /4}. Indeed, Ls,n ≥ ps − 2rn and Uj,n ≤ pj + 2rn ; the supplement gives the complete argument. Write gs = min{∆abs /2, ∆rel /4} and n⋆s = min{n ∈ [M ] : rn < gs }. With A = U , every positive transaction advance followed by a fresh evaluation, predictable per-transaction prefix increments at most b, a following fresh evaluation, and first-entry horizon at least two, Corollary 3 gives nlatch,s ≤ n⋆s + 2b, using at most S(n⋆s + 2b) cached action-defining audit labels; capacity may still suppress activation. This exact-real, schedule-conditional bound excludes all other ledger costs and does not ensure utility. If n⋆s does not exist by M , the theorem supplies error control but no detection-power claim. The audit requirement can be substantial: for S = 4 and δ = 0.05, rn is 0.149 at n = 384, 0.097 at n = 1000, and 0.034 at n = 10,000. The Hoeffding-spending radius also remains nonzero at a full census, so this is a transparent conservative bound rather than a tight delay characterization. Under the nominal constructions, the sufficient boundary first occurs at n = 1,782, 731, and 196 of M = 10,000
common identities for e40, e60, and e80, respectively; e20 has ∆abs < 0, so no prefix can satisfy it. Direct-contrast alternative. The supplement derives a strict-majority direct Hoeffding certificate that pathwise contains the evaluated coordinatewise certificate, including under clipping. It remains unevaluated, its Bluebirds census bound is still below zero. It is a separate level-δ procedure and is not unioned with either reported certificate. Known-size PPR specialization under strict majorities. Suppose all N = |U | = M identities are known in advance to have strict majorities. For Ks = N ps , the observed count xs,n , hypergeometric likelihood ℓs,n (k), uniform prior π0 (k) = 1/(N + 1), and posterior πs,n , define Cs,n = {k : π0 (k)/πs,n (k) < S/δ}. Here and below, the ratio uses the extended-real convention π0 (k)/0 = +∞. The posterior-ratio process is a nonnegative test supermartingale; Ville’s inequality (Waudby-Smith and Ramdas 2020) and a source union give Prπ {Ks ∈ Cs,n , ∀s, n | U, Y } ≥ 1 − δ. On feasible counts, the ratio is 1/[(n + 1)ℓs,n (k)]; the hull of Cs,n /N supplies anytime intervals and is exact at census. This is an alternative level-δ procedure: combining it with a non-nested engine requires a pre-fixed error split. Replacing unknown M by |U | when any identity lacks a strict majority is invalid. Synthetic and Bluebirds replays used exact-integer membership in the conservative weak-boundary PPR set at N = M = 10,000 and 108, respectively, with a strict majority at every identity. The fixed-panel replay is exploratory rather than a new confirmatory family. Exploratory finite-population correction (FPC). The base radius does not vanish at a census. After the base certificate closed no natural-panel outlier path, we separately evaluated a frozen, non-confirmatory Serfling/census sensitivity. At an ordinary comparable prefix (before support exhaustion), it replaces rn by n−1 , |U | ( r ) ρn,U log(2Sn(n + 1)/δ) FPC rn = min 1, . 2n ρn,U = 1 −
If the unknown comparable population has size M ≤ |U |, the Serfling factor 1 − (n − 1)/M is no larger, so using |U | is conservative (Serfling 1974; Bardenet and Maillard 2015); the same alpha-spending argument applies. Once every identity in U is exhausted, all comparable values are known and the sensitivity closes with Ls = Us = ps , making a false population-outlier closure impossible. Ordinary requests still need two fresh prefixes and horizon; a census closure is statistical evidence for a future capacity decision, not an activation. This separately labeled sensitivity was not used in the CIFAR trajectories and is not confirmatory evidence. For external truth y ⋆ (i), let qs and w be, respectively, the source and strict-majority error rates restricted to A. Then |ps − qs | ≤ w. Thus translation to source accuracy requires
an additional bound on wrong majorities. The theorem says nothing about warnings, utility, generalization, or whether exclusion is beneficial. In the ordered-risk experiments, at least three construction-clean sources make every retained majority correct, so w = 0 and ps = qs inside that simulator. In the rotating exact null, all four values satisfy ps = 1/4, so Equation (8) applies to any source. These consequences depend on the constructions; they do not turn majority disagreement into a general truth or deployment guarantee. The evaluated bound is deliberately loose and is not silently replaced. The FPC diagnostic above tests one classical correction under a separate label; other general, without-replacement confidence sequences that incorporate identities without a strict majority remain untested.
Experimental Protocol Data and environments. CIFAR-100 is used only for controller development; the reported utility study uses a subsequent CIFAR-10 validation cache (Krizhevsky 2009). Both are encoded once by an ImageNet-pretrained ResNet-18 (He et al. 2016; Deng et al. 2009), with split seed 2701. Each run seed draws without replacement a 10,000-example pool from the fixed 40,000-example training cache; paired methods share that pool and all runs use the same fixed 10,000example validation cache. A one-hidden- layer MLP of width 256 is trained with AdamW (Loshchilov and Hutter 2019) (learning rate 10−3 , weight decay 10−4 ). Candidate, maximum batch, and minimum batch sizes are 512, 256, and 32. The ordered-risk environments contain three constructionclean sources. In e20/e40/e80, one fourth source has a fixed 20%, 40%, or 80% set of identity-level label changes; repeated access to an identity returns the same label. In e60, two of five sources have independently fixed 60% change sets. The earlier development stage also contains an e80 lowoverlap diagnostic, reported in the supplement. These are programmatic stress tests rather than natural annotator data. The exact null assigns one wrong label per identity and rotates that role evenly across four sources, producing a unique 3–1 majority and exactly ps = 1/4 for every source. Budgets are 5, 10, 15, and 20% of a fixed anchor. The main ALIVE–CBE/CBR comparisons use paired seeds 40– 49; the known-size PPR study, which has a strict majority at every identity, uses disjoint seeds 60–69. Paired methods share pools, budgets, and environment definitions. Methods, grids, endpoints, and consumers were fixed stagewise before their corresponding outcomes; later variants remain validation-adaptive, not independent confirmation. The fixed validation cache supports conditional comparisons, and the split-specific official test cache remains unopened. Detailed protocols, complete results, study sequence, and provenance controls are in the supplement. Natural complete-panel audit. Bluebirds contains 108 binary tasks fully labeled by 39 workers (4,212 labels) (Welinder et al. 2010). The base audit replays 10,000 fixedseed pseudorandom permutations at (δ, τ ) = (0.05, 1/S), with two fresh confirmations and an empirical compara-
tor. Ground truth enters only offline majority-vote and prespecified Dawid–Skene comparisons (Dawid and Skene 1979). FPC and exact-PPR replays were designed after their parent natural-panel outcomes and frozen before their own outcomes. They are exploratory same-panel timing analyses; the panel has no features, routing, capacity, or training utility.
Env. Warning Latched request Active Median tact
Comparisons and inference. Strict consumers require complete finite grids, finite metrics, no overrun, and exact reconstruction of prefix, state, routing, eligibility, fallback, and the proxy ledger. ALIVE–CBR holds CBR and controller timing fixed: routing-only is the matched action ablation, while standalone CBR is the predeclared full-system reference. Empirical hard action is an uncalibrated diagnostic, not a same-FWER baseline. Other controller compositions and aggregation boundaries are reported in the supplement. Accuracy is primary and macro-F1 diagnostic; normalized area under the budget curve (AUBC) integrates the four budgets. Effects are averaged across environments within each of ten paired seed clusters, not across 40 environment signenumeration reference values require unestablished joint sign-exchangeability (Ernst 2004); Holm adjusts each prefixed family (Holm 1979). Paired intervals are descriptive, and non-rejection is not equivalence.
Table 1: ALIVE–CBR action funnel. Entries before the median count runs in which the event occurred at least once (40 per environment). The implementation records certification and latch together after the required two certificate-positive advances; every latched request activated, with zero capacityblocked transaction steps. Across 160 runs, the controller spent 796 transaction-steps in the non-latching Provisional state. The 160 latched source paths remained active for a median 87 transactions (range 24–151). Complete-group audit occupied 1,021,396/7,884,800 (13.0%) acquired candidate slots. tact is the zero-based first active-exclusion decision; “—” denotes no activation.
Results Contract conformance and null restraint. On 40 repeated ALIVE–CBE exact-null cells, there were zero certificates, requests, and active exclusions; the uncalibrated empirical controller acted in all 40. These cells share one construction and are a mechanism diagnostic, not 40 independent confirmations of the nominal FWER. Across the audited ALIVE–CBE and ALIVE–CBR grids, strict replay found no certificate reopening, eligibility violation, prefix skip or duplicate, label-selected identity, or budget overrun. Every action used only evidence cached through the preceding decision. In the ordered-risk ALIVE–CBR grid, all 40 cells in each of e40 and e80 latched and activated exactly the designated source, and all 40 e60 cells did so for both designated sources. All 40 e20 cells abstained from latching and active exclusion, although 16/40 issued at least one non-latching warning. Because the e20 designated source has ps = 0.20 < τ = 0.25, this is the pre-fixed target boundary rather than a lowrisk power claim. Evidence-engine efficiency. The disjoint-seed PPR study completed 320 utility and 40 null runs with all integrity and action-timing checks satisfied. On the 120 high-risk curves, PPR separated no later than Hoeffding on every curve and earlier on 100. Median PPR versus Hoeffding evidence counts were 96 versus 304 identities in e40, 62 versus 171 in e60, and 48 versus 48 in e80; all 40 null curves remained action-free. Thus the tighter published engine materially reduces delay in some margin–support regimes without claiming universal improvement. Earlier evidence did not guarantee uniform downstream value. PPR full minus routing-only averaged +0.2235 accuracy-AUBC pp, while e40 was −0.0555 pp; the predeclared all-environment utility criterion therefore remained
e20 e40 e60 e80
16/40 40/40 40/40 40/40
0/40 40/40 40/40 40/40
0/40 40/40 40/40 40/40
— 12.5 7 3
unmet. The mechanism gain and downstream robustness are distinct findings. Matched value of persistent action. Within the frozen ALIVE–CBR controller, adding the persistent-action layer to routing-only increased accuracy AUBC by +0.1935 pp, with descriptive 95% interval [+0.1172, +0.2697], positive differences in all ten seed clusters, and conditional Holm reference p = 0.005859. This matched contrast estimates the persistent-action increment while holding audit, routing, selector, timing, and pools fixed; it does not establish superiority to a separate learning system. A post-outcome descriptive decomposition clarifies the shared-budget mechanism. Audit-only minus CBR was −0.2274 pp, and routing-only minus audit-only was +0.2293 pp: non-latching routing approximately repaid the audit opportunity cost before persistent action supplied the +0.1935point increment. These contrasts are outside Holm and do not revise the predeclared endpoint. Full-system operating regimes. ALIVE–CBR minus standalone CBR was +0.1954 pp, with descriptive 95% interval [+0.0033, +0.3874] and positive differences in 7/10 seed clusters. The point estimate is positive, but the contrast did not meet the predeclared multiplicity-adjusted robustness criterion (conditional Holm reference p = 0.097656). The aggregate separates into two regimes. Accuracy AUBC was negative in e20 and e40 (−0.142 and −0.242 pp) but positive in e60 and e80 (+0.621 and +0.545 pp). At 5% budget, e20/e40 lost 0.523/0.693 pp. The persistentaction layer can add value inside the matched controller, yet audit cost and operating regime determine whether the full system improves on CBR. Complete CBE, predecessor, selector-switch, and aggregation results remain in the supplement. Natural-panel audit cost. The fixed Bluebirds population has 16 relative-disagreement outliers and 23 null workers. Base ALIVE closed none of 160,000 outlier-worker–replay paths; the empirical immediate rule closed every outlier
Reader question
Evidence-bound answer
Statistical role
Does exact-null evidence trigger ALIVE–CBE: 0/40; empirical immediate: 40/40 unsupported active exclusion?
Repeated mechanism diagnostic; not independent FWER replication How does PPR evidence delay com- Median separation prefix, PPR / Hoeffding: e40 96/304, e60 Pre-known N = M ; every pare with Hoeffding? 62/171, e80 48/48 identity has a strict majority; no later on 120/120 curves What is the matched persistent- ALIVE–CBR − routing-only: +0.1935 [+0.1172, +0.2697], Met predeclared conditional action increment? 10/10 seed wins reference criterion, p = .005859 Does the full system outperform ALIVE–CBR − CBR: +0.1954 [+0.0033, +0.3874], 7/10 seed Positive mean; system robuststandalone CBR? wins ness criterion unmet (conditional ref. p = .097656)
Table 2: Evidence for the main CIFAR-10 validation questions. The null mechanism diagnostic uses ALIVE–CBE; the disjointseed delay comparison uses separate level-0.05 PPR and Hoeffding engines for the same action target. Each identity count is one complete all-source audit group, not total label-query or monetary cost. Utility contrasts use ALIVE–CBR. Effects are method minus comparator in accuracy AUBC (pp); brackets are descriptive paired-t 95% intervals, and wins count positive differences over ten paired seed clusters. The audit-only decomposition and complete secondary results are in the supplement. The exact-null cells are repeated diagnostics; the official test remained unopened.
Env. e20 e40 e60 e80
5%
10%
15%
20% AUBC
−0.523 −0.230 +0.090 −0.051 −0.142 −0.693 −0.226 −0.111 −0.085 −0.242 +0.323 +0.525 +0.787 +0.778 +0.621 +0.210 +0.514 +0.683 +0.666 +0.545
Table 3: Environment–budget mean paired accuracy differences (ALIVE–CBR minus CBR, pp; seeds 40–49). AUBC is the normalized trapezoidal integral across budgets; positive values favor ALIVE–CBR.
path but requested at least one null worker in every replay. Majority-vote truth accuracy was 0.759 versus 0.889 for Dawid–Skene, underscoring that disagreement and truth are different targets. The exploratory Serfling procedure closed 96,117/160,000 paths before census and 63,883 at census, with median prefix 105. Separately level-0.05 PPR closed 128,917 before census and 31,083 at census, with median prefix 95/108 = 88.0%, or 3,705/4,212 complete-panel labels exposed. Both had zero null closures; PPR was no later on all paired paths and earlier on 80.6%. These are panel-exposure counts rather than query savings, and the replay has no downstream utility endpoint.
Discussion and Limitations The three validity questions introduced in Section resolve differently. The randomized-prefix theorem addresses statistical validity under its finite-population assumptions. Semantic validity does not follow: the focal source participates in the majority, a correct minority expert can be an outlier, and correlated majority error can be invisible. Decision validity also remains empirical. The matched increment is positive here, but the full-system regimes show why certificate validity alone cannot answer it. Four limitations bound the evidence. First, the guarantee assumes fixed support and labels, an ideal uniform permuta-
tion of the complete intersection, and at least one comparable identity; partial overlap, drift, changing sources, and calldependent labels are outside scope. The base construction abstains on identities without a strict majority and excludes them from its estimand, whereas the evaluated PPR specialization requires known N = M and a strict majority at every identity. Second, τ = 1/S and the zero-margin relative predicate are policy choices, not a utility-derived exclusion rule; no recovery or utility certificate is implemented. Third, downstream evidence uses synthetic label changes, a reused CIFAR validation cache, ten seed clusters, and no published same-target, same-FWER action baseline. Bluebirds is one small complete panel without routing or training. Finally, the ledger is a reproducible proxy rather than wall-clock, FLOPs, energy, or monetary cost. Independent natural downstream evaluation, resource sensitivity, partial-overlap certificates, and drift-aware recovery remain open.
Conclusion ALIVE treats persistent source action as a controlled interface rather than a side effect of a warning score. Its randomized prefix, evidence cache, latch, and capacity check let predictable routers inherit a defined familywise falseaction bound. In the frozen study, a published PPR engine sharply reduced evidence delay in e40/e60, and persistent action added a positive matched increment over routing-only. The full-system endpoint and natural-panel cost show that net utility still depends on margin, audit cost, and budget regime. The resulting contract provides an auditable basis for designing and evaluating persistent source actions.
Use of Generative AI GPT-5.6 Sol was used for language polishing. All AI-assisted revisions were reviewed and verified by the authors, who take full responsibility for the content of this paper.
References Bardenet, R.; and Maillard, O.-A. 2015. Concentration Inequalities for Sampling without Replacement. Bernoulli, 21(3): 1361–1385. Bengar, J. Z.; van de Weijer, J.; Fuentes, L. L.; and Raducanu, B. 2022. Class-Balanced Active Learning for Image Classification. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 3707–3716. Chakraborty, S. 2020. Asking the Right Questions to the Right Users: Active Learning with Imperfect Oracles. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, 3365–3372. Chow, C. K. 1970. On Optimum Recognition Error and Reject Tradeoff. IEEE Transactions on Information Theory, 16(1): 41–46. Dawid, A. P.; and Skene, A. M. 1979. Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm. Journal of the Royal Statistical Society: Series C (Applied Statistics), 28(1): 20–28. Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and FeiFei, L. 2009. ImageNet: A Large-Scale Hierarchical Image Database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 248–255. Donmez, P.; and Carbonell, J. G. 2008. Proactive Learning: Cost-Sensitive Active Learning with Multiple Imperfect Oracles. In Proceedings of the 17th ACM Conference on Information and Knowledge Management, 619–628. Donmez, P.; Carbonell, J. G.; and Schneider, J. G. 2009. Efficiently Learning the Accuracy of Labeling Sources for Selective Sampling. In Proceedings of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 259–268. El-Yaniv, R.; and Wiener, Y. 2010. On the Foundations of Noise-Free Selective Classification. Journal of Machine Learning Research, 11: 1605–1641. Ernst, M. D. 2004. Permutation Methods: A Basis for Exact Inference. Statistical Science, 19(4): 676–685. Gao, R.; and Saar-Tsechansky, M. 2020. Cost-Accuracy Aware Adaptive Labeling for Active Learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, 2569–2576. Garcelon, E.; Ghavamzadeh, M.; Lazaric, A.; and Pirotta, M. 2020. Conservative Exploration in Reinforcement Learning. In Proceedings of the Twenty-Third International Conference on Artificial Intelligence and Statistics, volume 108 of Proceedings of Machine Learning Research, 1431–1441. Hacohen, G.; Dekel, A.; and Weinshall, D. 2022. Active Learning on a Budget: Opposite Strategies Suit High and Low Budgets. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, 8175–8195. PMLR. Hacohen, G.; and Weinshall, D. 2023. How to Select Which Active Learning Strategy is Best Suited for Your Specific Problem and Budget. In Advances in Neural Information Processing Systems, volume 36, 13395–13407.
Han, B.; Yao, Q.; Yu, X.; Niu, G.; Xu, M.; Hu, W.; Tsang, I.; and Sugiyama, M. 2018. Co-Teaching: Robust Training of Deep Neural Networks with Extremely Noisy Labels. In Advances in Neural Information Processing Systems, volume 31, 8527–8537. He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 770–778. Hoeffding, W. 1963. Probability Inequalities for Sums of Bounded Random Variables. Journal of the American Statistical Association, 58(301): 13–30. Holm, S. 1979. A Simple Sequentially Rejective Multiple Test Procedure. Scandinavian Journal of Statistics, 6(2): 65–70. Huang, S.-J.; Chen, J.-L.; Mu, X.; and Zhou, Z.-H. 2017. Cost-Effective Active Learning from Diverse Labelers. In Proceedings of the 26th International Joint Conference on Artificial Intelligence, 1879–1885. Jagabathula, S.; Subramanian, L.; and Venkataraman, A. 2014. Reputation-Based Worker Filtering in Crowdsourcing. In Advances in Neural Information Processing Systems, volume 27, 2492–2500. Curran Associates, Inc. Joglekar, M.; Garcia-Molina, H.; and Parameswaran, A. 2013. Evaluating the Crowd with Confidence. In Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 686–694. ACM. Karger, D. R.; Oh, S.; and Shah, D. 2014. Budget-Optimal Task Allocation for Reliable Crowdsourcing Systems. Operations Research, 62(1): 1–24. Killamsetty, K.; Sivasubramanian, D.; Ramakrishnan, G.; De, A.; and Iyer, R. 2021. Grad-Match: Gradient Matching Based Data Subset Selection for Efficient Deep Model Training. In Proceedings of the 38th International Conference on Machine Learning, 5464–5474. Krizhevsky, A. 2009. Learning Multiple Layers of Features from Tiny Images. Technical report, University of Toronto. Li, J.; Jiang, L.; Li, C.; and Zhang, W. 2024. Label Consistency-Based Worker Filtering for Crowdsourcing. In Proceedings of the 40th Conference on Uncertainty in Artificial Intelligence, volume 244 of Proceedings of Machine Learning Research, 2226–2237. PMLR. Li, J.; Socher, R.; and Hoi, S. C. H. 2020. DivideMix: Learning with Noisy Labels as Semi-Supervised Learning. In International Conference on Learning Representations. Loshchilov, I.; and Hutter, F. 2019. Decoupled Weight Decay Regularization. In International Conference on Learning Representations. Natarajan, N.; Dhillon, I. S.; Ravikumar, P. K.; and Tewari, A. 2013. Learning with Noisy Labels. In Advances in Neural Information Processing Systems, volume 26, 1196–1204. Patrini, G.; Rozza, A.; Menon, A. K.; Nock, R.; and Qu, L. 2017. Making Deep Neural Networks Robust to Label Noise: A Loss Correction Approach. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1944–1952.
Raykar, V. C.; Yu, S.; Zhao, L. H.; Valadez, G. H.; Florin, C.; Bogoni, L.; and Moy, L. 2010. Learning from Crowds. Journal of Machine Learning Research, 11: 1297–1322. Ryu, J. J.; and Wornell, G. W. 2024. Gambling-Based Confidence Sequences for Bounded Random Vectors. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, 42856–42869. PMLR. Sener, O.; and Savarese, S. 2018. Active Learning for Convolutional Neural Networks: A Core-Set Approach. In International Conference on Learning Representations. Serfling, R. J. 1974. Probability Inequalities for the Sum in Sampling without Replacement. The Annals of Statistics, 2(1): 39–48. Settles, B. 2009. Active Learning Literature Survey. Technical Report 1648, University of Wisconsin–Madison. Sheng, V. S.; Provost, F.; and Ipeirotis, P. G. 2008. Get Another Label? Improving Data Quality and Data Mining Using Multiple, Noisy Labelers. In Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 614–622. Wan, W.; Zhang, W.; and Jin, C. 2025. Computational Budget Should Be Considered in Data Selection. In Advances in Neural Information Processing Systems, volume 38. Waudby-Smith, I.; and Ramdas, A. 2020. Confidence Sequences for Sampling without Replacement. In Larochelle, H.; Ranzato, M.; Hadsell, R.; Balcan, M.-F.; and Lin, H.-T., eds., Advances in Neural Information Processing Systems, volume 33, 20204–20214. Curran Associates, Inc. Welinder, P.; Branson, S.; Belongie, S.; and Perona, P. 2010. The Multidimensional Wisdom of Crowds. In Advances in Neural Information Processing Systems, volume 23, 2424– 2432. Wu, Y.; Shariff, R.; Lattimore, T.; and Szepesvári, C. 2016. Conservative Bandits. In Proceedings of the 33rd International Conference on Machine Learning, volume 48 of Proceedings of Machine Learning Research, 1254–1262. Yan, Y.; Rosales, R.; Fung, G.; and Dy, J. G. 2011. Active Learning from Crowds. In Proceedings of the 28th International Conference on Machine Learning, 1161–1168. Yan, Y.; Rosales, R.; Fung, G.; Farooq, F.; Rao, B.; and Dy, J. 2012. Active Learning from Multiple Knowledge Sources. In Proceedings of the Fifteenth International Conference on Artificial Intelligence and Statistics, volume 22 of Proceedings of Machine Learning Research, 1350–1357. Yin, J. O.; and Rush, A. M. 2025. Compute-Constrained Data Selection. In International Conference on Learning Representations. Zhang, J.; Shao, S.; Verma, S.; and Nowak, R. 2023. Algorithm Selection for Deep Active Learning with Imbalanced Datasets. In Advances in Neural Information Processing Systems, volume 36, 9614–9647. Zheng, Y.; Scott, S.; and Deng, K. 2010. Active Learning from Multiple Noisy Labelers with Varied Costs. In 2010 IEEE International Conference on Data Mining, 639–648.
Supplementary Material for “ALIVE: Warnings Before Exclusion in Budgeted Multi-Source Learning” Xiyang Zhang1,2 , Hongzhi Wang1∗ , Yuanhe Tian2 1
Harbin Institute of Technology 2 Zhongguancun Academy [email protected], [email protected], [email protected]
Scope and Reader Map ALIVE separates empirical warnings, latched exclusion requests, and capacity-feasible active exclusions under one shared learning budget. This supplement is organized by scientific function. Section gives the method variants, exact transaction, and accounting rules; Section gives the finitepopulation guarantees and alternative evidence engines; and Section reports the complete-panel stress test. Sections and specify the experimental estimands, decision criteria, and complete results. Resource limitations, study sequence, reproducibility, and non-claims follow in Sections – . The statistical model requires S ≥ 3, stable record keys and labels, and a nonempty all-source identity intersection. These are guarantee preconditions, not runtime-verified facts: empty support disables certification and retains the incumbent. The method does not cover disjoint, pairwise-only, or dynamically changing source sets. No action-defining branch consumes construction risk flags, environment IDs, validation outcomes, or test metrics. The statistical statement is deliberately narrow. It controls false certification of a source that fails either conjunct of the pre-fixed absolute-and-relative strict-majority-disagreement predicate in a fixed finite population under an ideal uniform audit permutation. It does not certify truth, corruption, harm, downstream safety, accuracy, utility, regret, calibration, or generalization. The non-latching Provisional state has no confidence guarantee. Its routing state can clear, but past acquisition and training are not undone. This separation is the design principle: empirical warnings may alter non-latching routing, whereas only simultaneous confidence separation may latch an exclusion request; active ineligibility remains capacity-feasible per transaction.
Evidence predicates and state machine For source s, let pbs,n be its disagreement rate over the first n comparable audit identities, and let [Ls,n , Us,n ] be Equation (4). Before audit labels are read, the action designer fixes τ ∈ (0, 1); the reported experiments use τ = 1/S. This defines the action target, not an estimated corruption rate. After every source has at least eight comparable observations, define 1 X Ws,n = I pbs,n > τ, pbs,n > pbj,n , (1) S−1 j̸=s 1 X Cs,n = I Ls,n > Uj,n , Ls,n > τ . (2) S−1 j̸=s
The peer comparator in Equation (2) is the mean of peer upper bounds, not their maximum. It excludes the focal source, whereas the strict majority defining disagreement includes it. Equation (1) is heuristic and has no confidence interpretation. Confirmation streaks advance or reset only when the comparable prefix grows. A source enters Certified after two consecutive fresh Cs,n = 1 advances and only when the estimated remaining decision horizon is at least two; certification is then latched for the stationary run. A not-yetcertified source with Ws,n = 1 is Provisional; otherwise it is Clear. The aggregate state is certified if any source is latched, provisional if none is latched but at least one warning is active, and clear otherwise. Nested confirmations are not independent trials and do not change δ to δ 2 .
Exact per-decision procedure For ALIVE–CBE and ALIVE–CBR, one transaction is:
Complete Algorithms and Accounting Setting and method families Let the source set be [S], the candidate-window size at zerobased decision t be Nt , and the requested optimizer batch size be Kt . Controller evidence used at decision t contains only complete audit groups collected through t−1. The fixed TS common audit support is U = s=1 Us . ∗
Corresponding author.
1. Read the cached comparable prefix and compute pbs,n , Ws,n , [Ls,n , Us,n ], and Cs,n . Advance/reset only fresh confirmation streaks; retain all latches. This freezes every action-defining state for the transaction. 2. Request audit fraction 0.25 exactly when at least one source is provisional and none is certified; otherwise request 0.125. Compute the floor/ceiling-feasible integer source allocation from the cached state. 3. Acquire the candidate window, placing next-unused audit identities into complete all-source groups and filling
Method
Routing evidence
Exclusion evidence
Incumbent
ALIVE–CBE ALIVE–CBE routing-only Empirical quarantine (CBE) ALIVE–CBR ALIVE–CBR routing-only ALIVE–PPR
empirical warning empirical warning empirical warning empirical warning empirical warning empirical warning
class-balanced entropy (CBE) CBE CBE class-balanced random (CBR) CBR CBR
ALIVE–PPR routing-only
CBR
never
uniform CBR
Empirical quarantine (CBR) Adaptive-Switch
empirical warning; identical shadow PPR certificate shadow warning audit; action disabled empirical warning empirical warning
certificate, latched never warning, non-latching certificate, latched never exact uniform-PPR-hull certificate, latched never
warning, non-latching certificate, latched
Fixed-Switch CBR baseline Full-Consensus
empirical warning none none
CBR CBR until the first certificate decision; CBE from the next decision onward CBR for t < 8; CBE for t ≥ 8 CBR strict-majority-label CBE
Audit-only diagnostic
certificate, latched never never
Table 1: Reader-facing method variants. Ablations share their parent method’s audit, allocator, and incumbent. The audit-only control is an outcome-informed exploratory diagnostic; Full-Consensus is a separate aggregation boundary.
State
Entry
Action
Clear
no warning or certificate
Provisional
active empirical warning
Certified
two consecutive fresh certificate advances and estimated remaining horizon ≥ 2
uniform acquisition; 12.5% audit; all candidates eligible non-latching monitored routing; 25% audit if no source is certified; candidates remain eligible latched exclusion request and monitoring; 12.5% audit; capacity-feasible activation
Table 2: Aggregate controller semantics. Capacity fallback suppresses one exclusion transaction but does not clear a certificate. residual slots by ordinary source sampling. Completed groups are appended to the audit data structure, but the frozen cache prevents them from changing any action in this transaction; they can first affect decision t + 1. 4. Score all acquired candidates. ALIVE–CBE performs one charged model forward and uses predictive entropy. ALIVE–CBR draws one deterministic random permutation from the run generator, maps it to strictly descending scores, and charges no ranking forward. 5. Use the cached state to form optimizer eligibility. ALIVE requests exclusion only for certified sources, the empirical-hard-action ablation uses current warnings, and routing-only never excludes. Activate a request only if the remaining candidates can fill Kt ; otherwise log capacity fallback and admit all candidates without clearing a latch. 6. Compute class quotas, restrict to eligible candidates, se-
lect within class by descending score, train one update, and charge every acquisition, scoring, maintenance, and selected training example. 7. Log the audit-permutation prefix, candidate and selected indices, source IDs, states, warning/certificate values, requested/realized audit, integer allocations, eligibility, fallback, ranking phase, and ledger events. The controller’s first-entry horizon uses the counterfactual ALIVE–CBE per-step denominator including an entropy-forward charge even for ALIVE–CBR and preswitch Adaptive-Switch. This keeps the incumbent from changing the certificate-entry rule. The actual ledger nevertheless charges only executed work.
Exact-integer PPR procedure ALIVE–PPR retains the ALIVE–CBR incumbent, warning and routing policy, uniform without-replacement audit prefix, two fresh-evidence confirmations, horizon guard, monitoring, capacity fallback, ledger, and training protocol. It replaces only the action-defining confidence endpoints by conservative weak-boundary enlargements of the PPR hulls in Theorem 3. Preflight requires a known N = M = 10,000, all-source complete support, and a strict majority at every identity. Given δ = a/b, membership of count k is computed by the exact integer comparison N k N −k a ≤ bS(n + 1) . n xs,n n − xs,n Retaining equality is weakly more conservative than the theorem’s strict set. Hull endpoints are divided by N and padded outward before the inherited strict floating-point certificate predicate. The full method uses the resulting latched certificate for hard-action eligibility; routing-only computes and latches
the identical certificate but sets the hard-action policy to never. Paired paths must therefore be identical until the first certificate and may diverge afterward only through eligibility, selection, and downstream model quantities. The reported fresh-seed PPR replication made no controller or statistical change relative to the initial PPR evaluation; it corrected only the analyzer–comparator interface described in Section .
Class quotas and CBR ranking Let Et be eligible candidate indices, Ct their observed classes, nc class capacity, and K = min(Kt , |Et |). If K ≤ 0 or no class is present, selection is empty. If K < |Ct |, the K largest observed classes receive one slot. Otherwise each observed class receives one slot; the residual is distributed proportionally to nc − 1, floored, then completed by largest fractional remainder without exceeding capacity. ALIVE–CBE ranks within class by entropy. ALIVE–CBR draws randperm(|Wt |), maps it to strictly descending scores 1, 1 − 1/|Wt |, . . . , 1/|Wt |, restricts to Et , and applies the same quota. The fixed class and tie order is part of replay.
Capped-simplex source allocation Without an actioned source, integer allocation is uniform and remainder slots follow the fixed source order. Otherwise, each provisional or certified source receives a feasible lower monitoring target f = 0.15, non-actioned sources receive the residual, and no continuous share exceeds c = max(0.40, 1/S). Residual mass is redistributed through the capped simplex; therefore 15% is neither an equality nor an upper bound for an actioned source. Shares are multiplied by Nt , floored, and completed by largest remainder. A repair enforces kmin = min(0.15, 1/S)Nt − 10−12 by transfers from sources above kmin ; if Skmin > Nt , the prerepair feasible allocation is retained. For Nt = 512, S = 4, three monitored sources receive (103, 102, 102) and the other source 205, up to the fixed source order; all four monitored sources receive 128 each. Largest-remainder integerization can exceed the continuous ceiling by one candidate.
Audit construction and quarantine On first use, a separate audit generator permutes U once. With requested fraction αt , Rt = ⌊αt Nt ⌋, Gt = ⌊Rt /S⌋ , gt = min{Gt , min nst , |U | − cursort }. s
Exactly gt identities are queried from every source, giving Sgt audit slots. Shortfall separately records complete-group rounding, allocation clipping, support exhaustion, and finalbudget shrinkage. Audit slots occupy the ordinary source allocation and are charged once. Only a unique strict majority creates a comparable observation; identities without a strict majority abstain. The cursor cannot skip, revisit, or replace identities. Exhaustion freezes the prefix rather than substituting source-specific records. For the confidence-gated methods, let Qt contain candidates from certified sources. Exclusion is requested when
Qt ̸= ∅ and is active only if |Wt \ Qt | ≥ Kt . Otherwise all candidates are eligible and capacity fallback is logged. A suppressed request is not an active quarantine or false isolation. The 12.5/25% audit shares and 15/40% monitoring floor/cap are fixed heuristics. After a latch, the monitoring floor and 12.5% audit maintain observability of additional source divergence and capacity state; they do not recover the latched source. Their continuing ledger charge is a declared utility cost and a plausible contributor to the lack of a CBR win. An independently validated recovery/epoch reset or value-of-information audit schedule is future work and was not tested.
Adaptive and fixed ranking schedules Adaptive-Switch inherits the ALIVE–CBE audit, state, routing, horizon, monitoring, and quarantine without change. Before certification it uses CBR and has zero ranking cheapevaluation cost. On the first decision with a nonempty certified-source latch, quarantine may act immediately, but ranking is still CBR. That transaction irreversibly arms entropy. At exactly the next decision, CBE becomes active and the actual candidate-forward cost is reserved before selection and charged once. The switch never rolls back and has no environment branch. Fixed-Switch preserves the same controller and quarantine, but fixes ranking independently of controller state: CBR for t = 0, . . . , 7 and CBE for every t ≥ 8. Decision 8 is reserved before selection. Its clock was fixed from the pooled zero-based median first-certificate decision 7 over 120 ALIVE–CBR high-risk curves, plus the one-decision delay used by Adaptive-Switch. No Adaptive-Switch outcome contributed.
Full-consensus boundary Full-Consensus is not a state-controller row. Each decision samples exactly 512 identity positions with replacement from the all-source intersection, retains duplicates, and materializes the aligned proposal from every source. It fails closed unless aligned features are exactly equal. All S labels are queried; a target is retained only when one label count is strictly greater than S/2, with no fallback or resampling for an identity without a strict majority. Proposal features are ranked by entropy, the same quota routine selects at most 256 examples, and training uses logical source consensus. No feedback, warning, audit, routing, certificate, allocation, or quarantine path executes. Per full step, Full-Consensus charges S × 512 source labels, 512 entropy scores, S × 512 majority-tally events, and the realized training count. Nominal full-step totals are 343.04 units for four sources and 358.40 for five. The last step is shrunk before selection and then charged at realized count. The synthetic construction has three clean sources, so strict majority equals simulator truth; this is a strong clean-target boundary, not a natural-worker truth-inference result.
Ledger semantics Strict auditors reconstruct every ledger event and require realized cost not to exceed the declared budget. Acquisition/audit
Stage
Unit cost
Count
source acquisition cheap evaluation fine evaluation training maintenance
0.02
acquired candidate/sourcelabel slots actually scored candidates
0.20 1.00 0.01
refresh
configured
0.05
fine-scored candidates selected examples declared maintenance/tally events realized refresh action
Table 3: Proxy-accounting coefficients used throughout the experiments. These units are not FLOPs, energy, currency, or universal latency.
slots are not double charged. ALIVE–CBE, post-switch Adaptive-Switch, Fixed-Switch after decision 7, and FullConsensus charge entropy forwards; ALIVE–CBR, CBR, and pre-switch Adaptive-Switch do not. Trained-example exposure counts selected indices, including repeat exposures.
Finite-Population Audit Guarantee Probability space and estimand Let U be the finite common audit support and Y = {Ys (i) : s ∈ [S], i ∈ U } the complete potential-label table. All results condition on realized (U, Y ). The mathematical randomization is π | (U, Y ) ∼ Unif(SU ). (3) No independence across sources or identities and no i.i.d. superpopulation model are assumed. The audit-seed lookup is generated separately from source-noise and training generators. Domain separation prevents accidental generator-state sharing; it does not prove Equation (3). Once a seed is fixed, the realized permutation and all decisions are deterministic. The nominal probability refers to ideal rerandomization conditional on the same (U, Y ), not residual randomness in a run or a frequency over ten seeds. Define ( ) S X A = i ∈ U : max I{Ys (i) = y} > S/2 , y
s=1
M = |A|. For i ∈ A, the strict-majority label m(i) is unique. If M = 0, the comparable-prefix set is empty and the implementation retains [0, 1] intervals without advancing certification. Assume M ≥ 1, and define X Zs (i) = I{Ys (i) ̸= m(i)}, ps = M −1 Zs (i). i∈A
Lemma 1 (restriction of a random permutation). Conditional on (U, Y ), the relative order induced by π on fixed subset A is uniform over the M ! permutations of A.
Proof. Fix an ordering of A. For each choice of the M positions occupied by A, exactly (|U | − M )! permutations of the remaining identities induce that ordering. Summing | over the |U position choices gives the same count for every M ordering of A. □
Simultaneous source-by-prefix coverage Let a1 , . . . , aM be the order induced on A. For s ∈ [S], n ∈ [M ], and δ ∈ (0, 1), define n X pbs,n = n−1 Zs (ak ), k=1
r
log(2Sn(n + 1)/δ) , rn = min{1, ren }, 2n Ls,n = max{0, pbs,n − rn }, Us,n = min{1, pbs,n + rn }. (4) Theorem 1 (randomized audit coverage). Under Equation (3), Pr{∀s ∈ [S], ∀n ∈ [M ] : ps ∈ [Ls,n , Us,n ] | U, Y } ≥ 1−δ. ren =
π
Proof. By Lemma 1, for fixed s, n, the values Zs (a1 ), . . . , Zs (an ) form a simple random sample without replacement from the fixed Bernoulli population {Zs (i) : i ∈ A}. Hoeffding’s comparison for sampling without replacement (Hoeffding 1963) gives Pr{|b ps,n − ps | ≥ ϵ | U, Y } ≤ 2 exp(−2nϵ2 ). π
If ren < 1, substitution yields 2Sn(n + 1) δ 2 exp − log . = δ Sn(n + 1) If ren ≥ 1, then rn = 1 and clipping produces [0, 1], with failure probability zero. Let F denote any source–prefix failure. A union bound gives Pr(F | U, Y ) ≤ π
M S X X
δ Sn(n + 1) s=1 n=1
≤ δ, because
P
−1
n≥1 [n(n + 1)]
= 1. □
Adaptive-prefix consequence. The implementation consumes one cached permutation and adds only complete allsource groups. It may use past evidence to choose a later audit request, stop at a budget boundary, or stop at support exhaustion. After majority abstentions, observed comparable identities still form an initial segment of (a1 , . . . , aM ). The theorem is simultaneous in n, so evaluating at these adaptive times requires no additional optional-stopping correction. Label-dependent identity selection, revisits, or calldependent labels would violate the premise. This is Hoeffding’s known comparison plus explicit alpha spending, not a new confidence-sequence construction. Tighter finite-population corrections and dedicated withoutreplacement confidence sequences (Serfling 1974; WaudbySmith and Ramdas 2020) are not used by the evaluated CIFAR controller; the separately labeled natural diagnostic below applies the classical Serfling factor.
Hard-action corollary Let U −s,n =
1 X Uj,n . S−1 j̸=s
Corollary 1 (false action-predicate certification). multaneous coverage, 1 X Ls,n > U −s,n =⇒ ps > pj , S−1
On si-
j̸=s
and Ls,n > τ implies ps > τ . Thus, for 1 X N = s : ps ≤ pj or ps ≤ τ , S−1
Corollary 3 (schedule-conditional audit-to-latch label bound). Assume A = U , and let 0 = n0 < n1 < · · · be the cached comparable-prefix sizes at fresh evaluations. Every positive transaction audit advance produces the next fresh evaluation, and each such predictable increment is at most b ∈ N. For the gaps in Corollary 2, define gs = min{∆abs /2, ∆rel /4}, n⋆s = min{n ∈ [M ] : rn < gs }. Suppose gs > 0, n⋆s exists, the schedule reaches a first nk⋆ ≥ n⋆s and has a following fresh evaluation nk⋆ +1 , and the firstentry horizon guard at that second evaluation is at least two. On the simultaneous-coverage event, source s latches by nlatch,s ≤ n⋆s + 2b.
j̸=s
Pr{∃s ∈ N ever entering Certified | U, Y } ≤ δ. π
Proof.
On coverage, ps ≥ Ls,n >
1 X 1 X Uj,n ≥ pj . S−1 S−1 j̸=s
j̸=s
For the first branch of N , certification is contained in the complement of simultaneous coverage. For the second branch, Ls,n > τ contradicts ps ≥ Ls,n and ps ≤ τ . Taking their union proves the claim. □ Any predictable warning, routing, or adaptive audit-size policy preserves this bound when it consumes only the next groups of the same prefix and every hard action requires a latch whose first entry followed a certificate-positive prefix. Event-wise, ever acting is then contained in ever certifying. The bound does not cover warning quality or label-dependent identity selection. Corollary 2 (sufficient detection at a covered prefix). Suppose source s has positive gaps 1 X ∆abs = ps − τ, ∆rel = ps − pj . S−1 j̸=s
On simultaneous coverage, the certificate predicate in Equation (2) holds at prefix n if ∆abs ∆rel rn < min , . 2 4 Proof. Coverage gives pbs,n ≥ ps − rn and pbj,n ≤ pj + rn . Clipping can only strengthen the resulting one-sided inequalities, so 1 S−1
X j̸=s
Ls,n ≥ ps − 2rn , 1 X Uj,n ≤ pj + 2rn . S−1 j̸=s
The absolute inequality follows from 2rn < ∆abs ; subtracting the peer bound from the focal lower bound leaves more than ∆rel − 4rn > 0. Both certificate conjuncts therefore hold. □
At most S(n⋆s + 2b) cached action-defining audit labels are used through the latch. Including the new audit groups acquired after state is frozen in the latch transaction gives at most S(n⋆s + 3b) audit labels by the transaction’s end. If |Wt \ Qt | ≥ Kt , activation occurs in that same transaction; otherwise the request remains latched but this result gives no activation-delay bound. The label counts exclude ordinary candidate acquisition, scoring, maintenance, and training; they are neither wall-clock nor total-budget bounds. Proof. Put c = 2S/δ. For positive real n, the derivative of log(cn(n + 1))/n has the sign of 1+
n − log(cn(n + 1)), n+1
which is negative because S ≥ 3, δ ∈ (0, 1), and n ≥ 1. Hence rn decreases. By minimality of k ⋆ and the increment bound, nk⋆ ≤ n⋆s + b,
nk⋆ +1 ≤ n⋆s + 2b.
Both radii are below gs , so Corollary 2 makes both fresh evaluations certificate-positive on simultaneous coverage. The second advance completes the two-fresh streak, and the assumed horizon guard latches the request. Because A = U , each prefix identity contributes one label from every source, proving the cached count. State is frozen before the transaction’s new acquisition, which contributes at most one further Sb block and proves the post-transaction count. The capacity statement follows from the activation predicate. Finally, activation is contained in latch, which is contained in certification, so Corollary 1 still bounds false activation by δ. □ The preceding results prove the main-paper detection statement and its schedule consequence. They are sufficient conditions, not necessary sample-complexity bounds. Two fresh-prefix confirmations and the horizon guard remain required before a latch. They only remove certification events, so they neither weaken Corollary 1 nor square its error level. Because a latched exclusion request requires certification, active exclusion of a member of N is also contained in the same event. The corollaries say nothing about utility, safety, generalization, or whether exclusion improves learning.
Derived alternative confidence constructions The direct construction remains unevaluated. The PPR construction was preselected for the synthetic ALIVE–PPR trajectories reported below. Each is a complete alternative levelδ certificate. By the pathwise dominance proved below, on any common source–prefix path the base certificate event is contained in the direct certificate event. Their pointwise union therefore equals the direct certificate and incurs no additional multiplicity. A union with a non-nested engine, such as PPR, requires a pre-fixed error split. Theorem 2 (strict-majority direct-contrast bounds). Put h = S − 1, q = ⌊h/2⌋, and X Ds (i) = Zs (i) − h−1 Zj (i),
lower
ds = M −1 dbs,n = n−1
Ds (i) = ps − h−1
i∈A n X
Ds (ak ),
X
j̸=s
cj = Uj,n − pbj,n = min{rn , 1 − pbj,n }. The evaluated relative inequality has a nonnegative right side, so it implies Ls,n = pbs,n − rn and X dbs,n > rn + h−1 cj . j̸=s
For rn ∈ [0, 1], cj ≥ rn (1 − pbj,n ). Moreover, at most q peers dissent on each identity, so p̄−s,n ≤ q/h. Thus X h−1 cj ≥ rn (1 − p̄−s,n ) ≥ rn (1 − q/h) ≥ rn q/h,
j̸=s
X
Proof. For rn = 1, the evaluated strict relative inequality is impossible because Ls,n = 0 while peer upper bounds are nonnegative. Suppose rn < 1, and write X p̄−s,n = h−1 pbj,n ,
pj ,
j̸=s
j̸=s
because q/h ≤ 1/2. Therefore dbs,n > rn (1 + q/h) = Rrn . The absolute conjunct Ls,n > τ is unchanged. □ For the Bluebirds population, h = 38, q = 19, and Rr108 = 0.4174, whereas the largest population relative margin is 0.3826. Thus even the dominating direct certificate fails its relative conjunct at census; it does not resolve the evaluated base method’s zero-power finding.
R = 1 + q/h,
k=1
Gs,n = dbs,n − Rrn . Under Equation (3), Pr Ls,n ≤ ps , Gs,n ≤ ds , π
∀s ∈ [S], ∀n ∈ [M ] | U, Y
≥ 1 − δ.
Consequently, the strict alternative certificate D Cs,n = I{Ls,n > τ, Gs,n > 0}
(5)
controls at level δ the probability of ever certifying any s ∈ N. Proof. A unique strict majority leaves at most q dissenting sources on each comparable identity. If source s agrees, then Ds (i) ≥ −q/h; if it dissents, then 0 ≤ Ds (i) ≤ 1. Hence every Ds (i) ∈ [−q/h, 1], an interval of width R. For ren < 1, one-sided without-replacement Hoeffding bounds give
Theorem 3 (known-size uniform-PPR sets with strict majorities). Assume every identity in U is known in advance to have a strict majority, so N = |U | =P M is known. For n source s, let Ks = N ps and xs,n = j=1 Zs (aj ). For k ∈ {0, . . . , N }, define the hypergeometric likelihood N −k k ℓs,n (k) =
xs,n
n−xs,n N n
,
with infeasible binomial coefficients equal to zero. Under the uniform prior π0 (k) = 1/(N + 1), let π0 (k)ℓs,n (k) πs,n (k) = PN , u=0 π0 (u)ℓs,n (u) Cs,n = {k : π0 (k)/πs,n (k) < S/δ}.
(6)
δ , π 2Sn(n + 1) δ Pr{dbs,n − ds ≥ Rrn | U, Y } ≤ . π 2Sn(n + 1)
Here and below, the ratio uses the extended-real convention π0 (k)/0 = +∞. Then Pr Ks ∈ Cs,n , ∀s ∈ [S],
For ren ≥ 1, rn = 1 makes both lower bounds pathwise valid: Ls,n = 0 ≤ ps and Gs,n ≤ 1 − R = −q/h ≤ ds . Summing the over sources and prefixes uses P two failure probabilities −1 D [n(n + 1)] = 1. On the resulting event, Cs,n =1 n≥1 strictly implies ps > τ and ds > 0, which excludes both branches of N . □
When Cs,n ̸= ∅, its hull divided by N is an anytime-valid interval for ps ; an empty set triggers no certificate. Replacing [Ls,n , Us,n ] in the original strict certificate by these hulls therefore controls the same false-action family.
Pr{b ps,n − ps ≥ rn | U, Y } ≤
Proposition (pathwise dominance). At the same (δ, τ ), every prefix satisfying the evaluated coordinatewise certificate in Equation (2) also satisfies Equation (5), including under clipping.
π
∀n ∈ {0, . . . , N } | U, Y
≥ 1 − δ.
Proof. Fix s and its true count Ks = k. Let Pk denote the law of the count history under that count and let Q = PN u=0 π0 (u)Pu be the mixture law. Under the count filtration Fns = σ(xs,0 , . . . , xs,n ), the ratio PN π0 (u)ℓs,n (u) π0 (k) = u=0 Es,n (k) = πs,n (k) ℓs,n (k)
equals Q(hn )/Pk (hn ) on every history hn in the support of Pk . It is a nonnegative test supermartingale initialized at one. Indeed, if Ek (hn ) is the set of one-step extensions having positive Pk -probability, then P hn+1 ∈Ek (hn ) Q(hn+1 ) EPk [Es,n+1 (k) | hn ] = Pk (hn ) Q(hn ) ≤ = Es,n (k). Pk (hn ) The inequality allows mixture mass on extensions that are impossible under the true count k; this support leakage is why equality need not hold. At census, Es,N (k) = π0 (k) = 1/(N +1), consistent with a supermartingale rather than a martingale in general. Ville’s inequality (WaudbySmith and Ramdas 2020) under the random-permutation law gives Prπ {∃n : Es,n (k) ≥ S/δ | U, Y } ≤ δ/S. The strict set boundary in Equation (6) makes exclusion of the true count exactly such a crossing; union bounding over S sources proves the claim and requires no independence across sources. Vandermonde’s identity gives N X u N −u u=0
x
n−x
=
N +1 , n+1
so for the uniform prior π0 (k) 1 = πs,n (k) (n + 1)ℓs,n (k) on feasible support. If δ = a/b in lowest terms, exact membership is N k ∈ Cs,n ⇐⇒ a < bS(n + 1) n k N −k × . xs,n n − xs,n The inequality is strict: equality is excluded by the definition of Cs,n . At n = N , only k = xs,N is feasible, hence the hull is the exact singleton {ps }. □ Posterior ratios may be evaluated descriptively in log space, with zero likelihood mapped to infinity. The actiondefining ALIVE–PPR membership test instead uses the exact-integer non-strict variant of the membership inequality above, conservatively retaining equality. Hull endpoints are converted to binary floating point with 64-ULP outward padding before the inherited certificate predicate; this narrows but does not formally eliminate program-level arithmetic risk. This specialization applies to Bluebirds because all 108 tasks are known, observed, and have a strict majority, and to the mechanically checked synthetic ALIVE–PPR panels with N = M = 10,000. It does not apply when any identity lacks a strict majority, making M unknown: substituting |U | for M changes both the support and likelihood and is not a conservative shortcut. Synthetic downstream PPR trajectories and the later exploratory Bluebirds mechanism replay are reported below under separate claim boundaries.
Majority identifiability Let y ⋆ (i) be an external true label on A, and define X qs = M −1 I{Ys (i) ̸= y ⋆ (i)}, i∈A
w=M
−1
X
I{m(i) ̸= y ⋆ (i)}.
i∈A
Let ds = ps −
1 X pj , S−1
d⋆s = qs −
j̸=s
1 X qj . S−1 j̸=s
Proposition 1 (disagreement versus true error). ery source, |ps − qs | ≤ w, and |ds − d⋆s | ≤ 2w.
For ev-
Proof. When m(i) = y ⋆ (i), majority-disagreement and true-error indicators coincide. On each of the remaining wM identities, their difference has absolute value at most one, giving |ps − qs | ≤ w. Apply this bound to the focal source and peer mean, then use the triangle inequality. □ Thus w = 0 permits a conditional true-label interpretation on A; without a bound on w, it does not. Unanimous common-mode error is invisible, a correct minority can be the unique majority-disagreement outlier, and tied identities are excluded from the estimand.
Construction-specific consequences Every ordered-risk environment contains three constructionclean sources among four or five sources. On common support they form a correct strict majority, so A = U, w = 0, and clean sources have ps = 0. Within a fixed synthetic run, Pr{ever certifies any clean source | U, Y } ≤ δ. π
The same bound applies to ever actively excluding a clean source under the confidence gate. In the exact rotating null, each identity has a correct 3-to-1 majority and p1 = · · · = p4 = 1/4; any certificate is therefore a false action-target certification and has per-run probability at most δ. These conclusions depend on the declared synthetic construction, do not cover empirical quarantine, and do not become an across-experiment familywise guarantee when runs are repeated.
Natural Complete-Panel Stress Test Base complete-panel audit The public Bluebirds data released with CUBAM (Welinder et al. 2010) form a complete binary matrix of 108 tasks by 39 workers. We use tasks as shared identities and workers as sources; all 4,212 labels are retained. Odd panel size and binary labels give a unique strict majority on every task, hence A = U and M = 108. The focal worker remains in that majority. External ground truth is unavailable to permutation generation, warning, interval, certificate, and request predicates; it is used only for the separately labeled descriptive comparison. Before computing any worker-level or replay outcome, we fixed the dataset revision, complete-panel protocol, analysis
Bluebirds: 10,000 orderings of one fixed 108 x 39 panel A Outlier detection 0
Fraction of 4,212-label panel .25 .50 .75
B Any-null restraint 1 Any-null request fraction
Detection fraction
1.0
0
0.8 0.6 0.4 0.2 0.0 0
1053 2106 3159 Cumulative labels (39 x prefix) Frozen ALIVE
Fraction of 4,212-label panel .25 .50 .75
1
1.0 0.8 0.6 0.4 ALIVE and FPC overlap at zero
0.2 0.0
4212
0
Empirical immediate
1053 2106 3159 Cumulative labels (39 x prefix)
4212
ALIVE-FPC (sensitivity)
Descriptive only; empirical immediate is not a certificate; ALIVE-FPC is post-outcome.
Figure 1: Descriptive restraint–power trajectories across 10,000 order randomizations of one fixed 108 × 39 Bluebirds panel, not independent panels. Panel A is cumulative outlier worker–randomization detection; Panel B is the fraction of randomizations with any null-worker request. Empirical immediate is a descriptive comparator, not a statistical certificate; the Serfling/census curve is an exploratory sensitivity analysis specified after the base result, not independent confirmation. code, and seed range. The audit then consumed 10,000 fixedseed pseudorandom task permutations from PCG64 seeds 27,010,000–27,019,999. It retained δ = 0.05, τ = 1/S, evaluation from prefix eight, two fresh certificate confirmations, and the remaining-prefix guard. Label cost at prefix n is exactly 39n. The comparison rule instead maps the first empirical warning at or after prefix eight directly to a persistentrequest analogue. This is intentionally not an ALIVE variant: it tests the consequence of treating a non-latching warning as a persistent request. Neither replay includes routing, capacity activation, features, or training. The fixed panel has 16 population relative-disagreement outliers and 23 null workers under Corollary 1. ALIVE detected 0/160,000 outlier–permutation pairs and produced a null-worker request in 0/10,000 replays. The empirical analogue detected 160,000/160,000, but issued at least one nullworker request analogue in 10,000/10,000 replays; across the replays every worker received at least one analogue request. Conditional on this fixed panel, the Wilson 95% upper endpoint for 0/10,000 is 0.000384; it is a Monte Carlo descriptor rather than a replacement for the theorem’s δ bound. These 10,000 permutations probe randomization on one fixed panel; they are not independent worker panels. Majority disagreement has Spearman association 0.802 with external worker error, but majority truth accuracy is only 0.759. A full-panel binary Dawid–Skene fit (Dawid and Skene 1979) uses deterministic majority initialization, additive-one smoothing, and a 10−10 convergence criterion; it passes convergence, objective-monotonicity, deterministic-rerun, and worker-permutation gates in 65 it-
erations. Its truth accuracy is 0.889, and inferred worker quality has Spearman association −0.974 with external error. Dawid–Skene is offline, sees every label, and is neither sequential nor cost matched. The comparison both supports semantic relevance of disagreement and demonstrates why no truth interpretation follows from it.
Exploratory finite-population completion The base radius remains nonzero even at a census. After observing the base result, we specified one exploratory, nonconfirmatory diagnostic, ALIVE–FPC. For every ordinary comparable prefix before support exhaustion, it uses n−1 , ρn,|U | = 1 − |U | ) ( r (7) ρn,|U | log(2Sn(n + 1)/δ) FPC rn = min 1, . 2n All thresholds, seeds, sources, tasks, and ordinary two-fresh and horizon rules are unchanged. At support exhaustion, all comparable values have been seen and the interval is set exactly to Ls = Us = ps . A qualifying exact census may record one closure request without a duplicate census; it is not an executed capacity action. Proposition 2 (FPC coverage and exact closure). Suppose the comparable population has unknown size M ≤ |U |, and its order is induced by a uniform permutation of U . Intervals formed with Equation (7) simultaneously cover every source and ordinary comparable prefix before support exhaustion with probability at least 1 − δ. Hence any ordinary
request for a member of N has probability at most δ. At support exhaustion, an exact-census request for a member of N is deterministically impossible. Proof. For sampling without replacement, the Serfling inequality gives each tail probability at most exp[−2nϵ2 /ρn,M ] (Serfling 1974; Bardenet and Maillard 2015). For fixed n, ρn,M is nondecreasing in M ; therefore replacing unknown M ≤ |U | by |U | can only widen the interval. Substitution in the two-sided bound gives P δ/[Sn(n + 1)]. Summing over sources and prefixes uses n≥1 1/[n(n + 1)] = 1, exactly as in Theorem 1. The certificate inequalities on simultaneous coverage imply s ∈ / N ; two fresh confirmations and horizon only select a subset. Once all of U is exhausted, A, M , and every ps are observed. The strict census inequalities then hold exactly if and only if s ∈ / N. □ Before evaluating this diagnostic, implementation checks covered the conservative-M direction, closed-form radii, outward rounding, exhaustive small populations, exactcensus predicates, state transitions, horizon, and deterministic replay. Because the construction was chosen after observing the base audit, the result is exploratory rather than independent validation. ALIVE–FPC closed all 160,000 outlier–permutation pairs with zero null-worker request in 10,000 replays (Wilson 95% upper endpoint 0.000384). Ordinary two-prefix evidence accounted for 96,117 closures (60.07%); exact census accounted for 63,883 (39.93%). Median closure prefix was 105/108, and labels exposed were 4,095/4,212 (97.2% of the full panel). Thus FPC reaches finite-population completeness but generally only near census cost. Exact closure is deterministic population identification, not early prediction, capacity activation, or downstream utility.
Exploratory exact-PPR replay After the two preceding natural-panel outcomes were known, we specified a one-shot exact-PPR replay before computing its results. The analysis is therefore exploratory and non-confirmatory, although its procedure was fixed before this replay was evaluated. The panel, 10,000 PCG64 permutations, seeds 27,010,000–27,019,999, N = M = 108, S = 39, δ = 0.05, τ = 1/39, prefix-eight start, two-fresh rule, and two-group horizon are identical to the parent replay. Ground truth is not parsed by the order, interval, certificate, or closure computation. This instantiates the published without-replacement PPR confidence-sequence engine (Waudby-Smith and Ramdas 2020) as a same-target, same-FWER certificate comparator; it is not a published source-aware or downstream method. PPR and Serfling are two alternative level-0.05 procedures on the same source-by-prefix family. We do not take the union of their action events. PPR set membership is checked by exact integer arithmetic with the uniform prior on total disagreement count. An ordinary latch requires two consecutive fresh positive prefixes and is first possible at n = 9; the horizon precludes it in the final two groups. At n = 108, the PPR set is an exact singleton. A qualifying event there is recorded only as exact_census_closure: it is not an ordinary latch, capacity activation, or early detection.
Rule
Outlier worker×replay closures
ALIVE (base 0/160,000 certificate) Empirical im- 160,000/160,000 mediate ALIVE–FPC 160,000/160,000 (exploratory) PPR (ex- 160,000/160,000 ploratory)
Replays with Median first any null clo- closure presure fix / labels exposed 0/10,000
—
10,000/10,000 8/312 0/10,000
105/4,095
0/10,000
95/3,705
Table 4: Natural fixed-panel closure diagnostic on the Bluebirds crowd-label data (108 tasks, 39 workers, and the full 4,212-label panel). ALIVE–FPC (exploratory) and PPR were designed after parent outcomes and are non-confirmatory. PPR and Serfling are separate level-0.05 procedures; their actions are not unioned. Counts describe fixed-panel audit closure behavior. The median pools ordinary two-fresh requests and exact-census closures; labels exposed are complete-panel prefix accounting (39n), not query savings or downstream utility. PPR closed all 160,000 fixed outlier-worker–permutation pairs. Ordinary two-fresh evidence accounted for 128,917 closures and exact census for 31,083. There were zero ordinary null-worker closures and zero census null-worker closures in 10,000 replays; the ordinary replay-level 0/10,000 Wilson 95% upper endpoint is 0.000384. Against Serfling on exactly paired paths, PPR was no later on 160,000/160,000, strictly earlier on 128,917 (80.57%), equal on 31,083, and later on zero. The pooled median first-closure prefix was 95 versus 105, corresponding to 3,705 versus 4,095 labels exposed under complete-panel prefix accounting. These randomizations probe one fixed small panel, not independent worker populations; the comparison has no routing, training, capacity, or downstream utility. It therefore does not alter the failed downstream PPR criterion, and the official test set remains unopened.
Experimental Protocol and Decision Criteria Features, model, and environments CIFAR-100 and CIFAR-10 (Krizhevsky 2009) images are encoded once by an ImageNet-1K-pretrained (Deng et al. 2009) ResNet-18 (He et al. 2016). Each run trains on a seedspecific 10,000-example subset of a fixed 40,000-example training cache and evaluates on a fixed 10,000-example validation cache. The head is a one-hidden-layer MLP of width 256, trained with AdamW (Loshchilov and Hutter 2019), learning rate 10−3 , and weight decay 10−4 . Candidate count is 512, requested batch size 256, minimum final batch 32, evaluation interval 25, and proxy compute anchor 204,000. In the exact null, source s = i mod 4 receives (yi + 1) mod K and the other three receive yi . Each source disagrees on exactly 2,500 identities; no harmful source is designated. Null outcomes are never pooled into ordered-risk utility.
Stage
Environment
Sources
Constructed risk
Audit support
CIFAR-100 dev CIFAR-100 dev CIFAR-100 dev CIFAR-100 dev CIFAR-100 dev
e20-single e40-single e60-two-of-five e80-single e80-low-overlap
4 4 5 4 4
one persistent 0.20 one persistent 0.40 two independent persistent 0.60 one persistent 0.80 one persistent 0.80
full 10,000 full 10,000 full 10,000 full 10,000 intersection 250; ordinary support full
CIFAR-10 validation CIFAR-10 validation CIFAR-10 null
e20/e40/e80-single e60-two-of-five exact-rate-symmetricnull
4 5 4
one persistent 0.20/0.40/0.80 two independent persistent 0.60 one rotating wrong source per identity; every ps = 0.25
full 10,000 full 10,000 full 10,000
Table 5: Experimental environments. A persistent rate changes each designated source label independently at an identity-fixed rate.
Utility estimand and inference For metric g, environment e, method m, seed z, and b1 < · · · < b4 , normalized AUBC is 3
Aemz =
X gemz (bk ) + gemz (bk+1 ) 1 (bk+1 − bk ). b4 − b 1 2 k=1
Budgets are 0.05, 0.10, 0.15, 0.20. The predecessor, ALIVE– CBE, ALIVE–CBR, Full-Consensus, and switch diagnostics use seeds 40–49; the separate PPR replication uses previously unused seeds 60–69. Accuracy is primary and macro-F1 descriptive. A contrast is computed within environment–seed and then averaged equally across the fixed environments. The analysis unit is the seed aggregate (n = 10), not a budget, environment, or stored run. The reported two-sided sign-flip calculation enumerates all 210 coordinate-wise signs. It is exact only under joint sign-exchangeability of the paired seed effects (Ernst 2004); fixed seeds and deterministic cache splits do not establish that condition. We therefore call these conditional enumerated reference values. Paired-t and bootstrap intervals are descriptive. Failure to reject is never called equivalence or non-inferiority. The ALIVE–CBE Holm family contains comparisons against the coupled-action predecessor, empirical quarantine, empirical agreement, and uniform sampling; routingonly is outside that family. ALIVE–CBR has a separate three-comparison family against standalone CBR, empirical quarantine, and routing-only. The later audit-only analysis adds no hypothesis: all contrasts involving it are outcomeinformed, descriptive, and outside Holm. Full-Consensus and the switch diagnostics belong to neither family. The PPR replication likewise has a separate conjunctive criterion, and its conditional sign-flip value is not Holm-adjusted. Exploratory audit-only protocol. The diagnostic retained the ALIVE–CBR audit/controller and adaptive 12.5%/25% shadow-warning audit, but forced uniform incumbent allocation, disabled source action, and made requests and activation impossible. Its fixed grid crossed e20/e40/e60/e80, four budgets, and seeds 40–49 (160 validation runs). The seed cluster after equal-environment aggregation remained the analysis unit. It was specified before its own outcomes
but after the ALIVE–CBR results, so it can only decompose the observed chain descriptively and cannot reopen either the failed endpoint or the original Holm family. PPR replication protocol. The PPR replication crosses e20/e40/e60/e80, four budgets, seeds 60–69, and paired ALIVE–PPR/full and routing-only methods (320 utility runs). The exact rotating null crosses the same budgets and seeds for the full method only (40 runs). Ten equalenvironment seed aggregates are the primary units. Both methods share the exact PPR prefix and all pre-action behavior; only the full method may execute certified exclusion. The reused validation cache and all constructions have known N = M = 10,000 and a strict majority at every identity. An initial evaluation used the same scientific design, but its analyzer named the comparator generically while the prespecified criterion required the explicit routing-only comparator. It therefore produced no eligible primary estimate. The fresh-seed replication corrected only that interface and was fixed before its own outcomes. It is not an external replication, and the official test set remained unopened.
Decision criteria ALIVE–CBE development screen. The screen required 600 complete, finite, in-budget runs; zero selected-ineligible or certificate-reopening events; inclusion of all common truecertifying predecessor and ALIVE–CBE curves; at least 35% earlier mean first certification in e40 and e60; a positive mean against the predecessor with positive leave-one-environmentout means; and a mean against empirical quarantine above −0.0015. Missing common timing support counted as failure. ALIVE–CBE cross-dataset criterion. The 960 orderedrisk and 160 null runs had to satisfy integrity and mechanism checks, e80/e60 detection and e20 abstention, a positive Holm-rejected contrast against the predecessor with positive leave-one-environment-out effects, empirical-quarantine price bounds, a positive contrast against uniform sampling in every environment, and zero certified actions under the exact null while the empirical comparator produced at least one false isolation.
ALIVE–CBR criterion. The 640 ordered-risk and 120 null runs had analogous integrity and mechanism requirements. The contrast against standalone CBR had to be positive, Holm-rejected, positive in e40/e60/e80, and above −0.0015 in e20. Routing-only had to be positive, empirical quarantine had to satisfy the same price bounds, and the null had to show zero certified ALIVE–CBR actions but an empirical false isolation. Full-Consensus validity. Full-Consensus had no performance threshold. Reporting required 160 unique, complete, finite, in-budget runs; exact input identities; strict reconstruction of alignment, majority targets, selected indices, and ledger counts; and a complete matched ALIVE–CBE comparator. Every valid result direction was then reported. Adaptive-Switch diagnostic criteria. The exploratory criteria jointly required complete integrity, e20 parity with ALIVE–CBR, zero e20 actions, the intended one-decision delayed switch in higher-risk environments, exact-null abstention, an ALIVE–CBE improvement of at least 0.001 with conditional sign-flip value below 0.05, a positive high-risk contrast against ALIVE–CBR, and a positive all-risk contrast against CBR. Failure of any clause precluded further official evaluation. ALIVE–PPR replication criterion. The conjunctive criterion comprised: (i) exactly 320 utility and 40 null runs that were complete, finite, issue-free, and in budget; (ii) known N = 10,000, complete support, strict majorities, a uniform without-replacement prefix, and exact-integer PPR reconstruction; (iii) matched pre-action paths, zero e20/null action, and no clean-source certificate; (iv) exact harmful-set exclusion on all high-risk curves, PPR no later than the Hoeffding engine on at least 90%, strictly earlier on at least 50%, and no worse median evidence count in each high-risk environment; and (v) full minus routing-only accuracy AUBC at least +0.001, conditional two-sided sign-flip p < 0.05, a strictly positive mean in each of e40/e60/e80, a nonnegative pooled high-risk 5%-budget mean, positive leave-one-highrisk-environment-out means, and exact e20 parity. All clauses had to hold jointly. Thresholds could not be revised after observing outcomes, and the official test set remained closed on failure.
Complete Experimental Results Coupled-action predecessor and selector study The coupled-action predecessor has 1,400/1,400 valid, finite, budget-compliant runs and zero active quarantine training violations. Severe one- and two-source failures were certified/quarantined in 10/10 seed aggregates; e20 abstained in 10/10. Because e20 fixes ps = 0.20 < τ = 1/4, that abstention is the declared action-target boundary, not a failure to detect an alternative inside the target. Nevertheless, the predecessor minus empirical agreement accuracy AUBC was −0.0048277, descriptive paired-t 95% interval [−0.0056235, −0.0040318], with all ten seed effects negative. The conditional enumerated value was 0.001953, Holm-adjusted to 0.0078125, but in the unfavorable direction. The predecessor exceeded routing-only by 0.0072773.
ALIVE–CBE contrast (accuracy AUBC) minus predecessor minus empirical quarantine minus empirical agreement minus uniform minus routing-only
Mean +0.0047042 −0.0002371 +0.0001558 +0.0078183 +0.0051717
Table 6: ALIVE–CBE equal-environment seed-clustered means. The pre-specified headline criterion was therefore not met: the mechanism and integrity checks passed, whereas the required positive primary effect and every positive leave-oneenvironment-out effect failed. The selector-factor study has 1,000/1,000 valid, finite, inbudget runs. Mean accuracy AUBC was 0.499892 for CBR, 0.470603 for complete historical ALIVE, 0.468894 for CBE, 0.460334 for random, and 0.411315 for entropy. Complete ALIVE minus CBR was −0.029289, negative for all ten seeds and every environment–budget cell; complete ALIVE minus CBE was only +0.001709. This result motivated the ALIVE–CBR variant but is not evidence for its effectiveness.
ALIVE–CBE development screen The strict matrix has 600/600 complete finite runs, zero budget overruns, selected-ineligible violations, or certificate reopenings. On the 40 common true-certifying curves per environment, mean first certification moved from decision 21.9 to 12.3 in e40 (43.84% earlier) and from 12.8 to 7.3 in e60 (42.97% earlier); all 40/40 curves remained true-certifying in both. ALIVE–CBE minus the predecessor mean accuracy AUBC was +0.0023127, range [−0.0001200, +0.0065800] over ten seed aggregates. All leave-one-environmentout means were positive: +0.0031871, +0.0021100, +0.0023175, +0.0022954, +0.0016533 when excluding e20, e40, e60, e80, and e80-low-overlap, respectively. ALIVE–CBE minus empirical quarantine was −0.0007240, above the pre-specified −0.0015 tolerance. Every development condition passed, permitting the pre-specified crossdataset validation stage.
ALIVE–CBE cross-dataset validation The ordered-risk package has 960/960 and the exact null 160/160 complete finite in-budget runs, with zero selected-ineligible or reopening events. ALIVE– CBE minus the predecessor accuracy AUBC was +0.0047042, range [+0.0024958, +0.0069042], descriptive paired-t 95% interval [+0.0035212, +0.0058871], and positive for all ten seed aggregates. The conditional sign-flip value was 0.001953, Holm-adjusted to 0.0078125. All leave-one-environment-out means were positive: +0.0063961, +0.0047567, +0.0037422, +0.0039217. The empirical-quarantine contrast had descriptive paired-t interval [−0.0004144, −0.0000598], satisfying the price guard; routing-only had interval [+0.0033102, +0.0070332]. The e80, e60, and e20 joint mechanisms qualified in 10/10 seeds each. Under the
exact null, all 40 ALIVE–CBE curves had zero certificates, exclusion requests, and active quarantines, whereas empirical quarantine falsely isolated on all 40 curves. ALIVE–CBE minus empirical quarantine there was −0.0288483. The conditional headline nevertheless failed. Nine of ten conditions passed, but ALIVE–CBE minus uniform was −0.0003717 in e20, violating the requirement of a positive effect in every environment (e40 +0.0071233, e60 +0.0136117, e80 +0.0109100). Accordingly, the crossdataset headline criterion was not met; the strong primary comparison does not erase that pre-specified failure.
ALIVE–CBR: component improvement without a system-level win ALIVE–CBR has 640/640 ordered-risk and 120/120 null runs, all complete, finite, and in budget, with zero violations or reopenings. All e80/e60/e20 mechanisms qualified in 10/10 seeds. ALIVE–CBR minus CBR accuracy AUBC was +0.0019538, range [−0.0025542, +0.0063792], descriptive paired-t 95% interval [+0.0000330, +0.0038745], with seven wins and three losses. The conditional value 0.048828 became 0.097656 after its separate Holm adjustment. Environment means were e20 −0.0014233, e40 −0.0024200, e60 +0.0062083, and e80 +0.0054500. ALIVE–CBR minus routing-only was +0.0019346, interval [+0.0011720, +0.0026971], positive in all ten seeds and Holm-adjusted p = 0.005859. ALIVE–CBR minus empirical quarantine was −0.0000058, interval [−0.0002009, +0.0001892]; this is not evidence of equivalence. The null had zero ALIVE–CBR certificates, requests, activations, violations, or reopenings, while the empirical ablation falsely isolated. Eight of nine ALIVE–CBR headline conditions passed. The combined primary condition failed because the Holmadjusted value exceeded 0.05 and e40 was negative. The pre-specified system-level criterion was therefore not met; no headline claim or threshold revision was made.
Exploratory audit-only decomposition All 160/160 runs completed in budget. The behavior audit reconstructed 15,400 decisions and 1,021,396 audit slots, including 796 decisions at the 25% shadow audit rate. It found zero non-incumbent routes, requests, activations, allocation violations, cheap-evaluation events, and paired-audit mismatches. Thus the control retained the audit opportunity cost while removing both source actions. Accuracy-AUBC seed-cluster means were audit-only minus standalone CBR −0.0022742 (−0.2274 pp), routingonly minus audit-only +0.0022933 (+0.2293 pp), and ALIVE–CBR minus audit-only +0.0042279 (+0.4228 pp). Together with the original ALIVE–CBR minus routing-only effect +0.0019346 (+0.1935 pp), the chain closes to the unchanged endpoint ALIVE–CBR minus CBR +0.0019538 (+0.1954 pp). The endpoint still failed its original Holm test (adjusted p = 0.097656); no new contrast belongs to a Holm family.
Full-Consensus: valid boundary with lower utility The validity checks passed on 160/160 finite in-budget runs. Exact aligned feature reconstruction found 6,048,200 proposals, 5,896,768 unique identities, 151,432 replacement collisions (2.50375%), 25,657,120 source-label queries, and strict-majority availability for every proposal: zero abstentions and zero simulator-truth disagreements. Training used 3,007,810 selected examples over 11,830 optimizer steps. Full-Consensus mean accuracy AUBC was 0.8177667, versus 0.8237479 for ALIVE–CBE. Full-Consensus minus ALIVE–CBE was −0.0059813, with descriptive paired-t interval [−0.0081912, −0.0037714], bootstrap interval [−0.0079196, −0.0042850], zero wins in ten, and conditional descriptive signflip value 0.001953. Environment effects were −0.0011550, −0.0066817, −0.0099983, −0.0060900 for e20/e40/e60/e80. Macro-F1 differed by −0.0072104, interval [−0.0098088, −0.0046120], again zero wins. These are mandatory post-diagnostic results, not a failed promotion test. Ledger totals across all Full-Consensus runs were 513,142.4 source, 302,410.0 cheap evaluation, 256,571.2 maintenance, and 3,007,810.0 training units, totaling 4,079,933.6; fine evaluation and refresh were zero. Full consensus is thus neither a free oracle nor a superior boundary in this matrix.
Adaptive-Switch diagnostic All 160 adaptive and 40 null runs passed the behavior audit, remained in budget, and had no selected-ineligible events. In high-risk environments, ranking was CBR at the first certificate decision, entropy began exactly one decision later, all charges were reconstructed, and the switch was monotone. The exact null produced no certificates, exclusions, or entropy switches. The e20 parity check initially reported 40 selection and 40 metric mismatches. A fieldwise audit showed that selected indices, source IDs, trace lengths, ledger events, controller states, and structural diagnostics were identical in all matched cells. Differences were confined to modeldependent floating-point diagnostics: Adaptive-Switch minus ALIVE–CBR e20 accuracy had mean −2.496 × 10−6 and range [−0.0018000, +0.0013000]; macro-F1 mean was −4.773 × 10−6 , and compute and overhead metrics were exact. The most plausible explanation is GPU numerical nondeterminism because deterministic PyTorch/cuDNN/CUBLAS execution was not enforced. Thus the original parity label was overinclusive; it is not evidence that different examples or sources were selected. This clarification does not change the scientific decision. Three independent criteria failed: the pre-specified exact-parity clause; the primary conditional sign-flip requirement (mean versus ALIVE–CBE +0.0011783, but p = 0.326172); and the high-risk contrast against ALIVE– CBR (−0.0012650). The all-risk contrast against CBR was positive (+0.0009975), while the integrity, delayed-switch, e20-abstention, and null criteria passed. Adaptive-Switch therefore remains an exploratory diagnostic, and the official test set was not evaluated.
A Utility difference by risk regime
B Exact-null active hard actions
ALIVE--CBE minus standalone CBR (accuracy AUBC, pp) LOW RISK
cells with active isolation
+0.637
HIGH RISK
+0.341
+0.5
ALIVE--CBE 0/40
0
Empirical hard action
-0.135 -0.5
-1.0
40/40
-0.916 e20
0
e40
e60
20
40
e80 40 cells per method
negative at e20/e40; positive at e60/e80
Validation only. Exact-null cells are not independent replicates.
Figure 2: Matched restraint–utility evidence. Left: the ALIVE–CBE versus CBR contrast changes sign with risk. Right: exactnull active actions; cells are repeated diagnostics, not independent replicates.
Fixed-Switch diagnostic and cost disclosure
Criterion
Result
Evidence
Fixed-Switch’s 160/160 validation runs passed the no-metric behavior audit with zero issues or budget violations. The first entropy decision was exactly 8 in every run; decision counts at the four budgets were 36, 71, 106, and 141. The auditor reconstructed exactly 12,880 candidate-forward events, 9,708 quarantine-request steps, and 9,708 active-quarantine steps, with no eligibility violation. Fixed-Switch accuracy AUBC was 0.8226146 and macroF1 AUBC 0.8208347. Adaptive-Switch minus Fixed-Switch accuracy AUBC was +0.0023117, conditional descriptive sign-flip value 0.0390625, with environment effects e20 +0.0088650, e40 +0.0000567, e60 +0.0017417, and e80 −0.0014167. This has no confirmatory status or multiplicity adjustment. Fixed-Switch minus ALIVE–CBE was −0.0011333 (p = 0.185547); Fixed-Switch minus ALIVE– CBR was −0.0032679 (p = 0.001953, all ten negative); and Fixed-Switch minus CBR was −0.0013142 (p = 0.201172). The last contrast varied sharply: e20 −0.0103183, e40 −0.0031517, e60 +0.0053333, and e80 +0.0028800. Across 160 runs, mean cheap-evaluation, training, overhead, and total proxy costs were 2,060.8, 22,527.75, 2,971.68625, and 25,499.43625. Mean overhead ratio was 0.1145392 and mean compute-to-budget ratio 0.9999688. These costs and all comparisons are descriptive; they do not alter the Adaptive-Switch conclusions or justify official evaluation.
Integrity
pass
PPR validity
pass
Path matching
pass
Mechanism
pass
Utility
fail
Joint
fail
320 utility + 40 null runs complete and valid 360/360 known-population valid; zero ties 160/160 matched; e20/null zero action 120/120 correct; no later 120/120, earlier 100/120 mean = +0.224 pp, p = 0.001953; e40 = −0.056 pp e40 clause fails; official test remains unopened
ALIVE–PPR: lower evidence cost without satisfying the headline criterion The initial PPR evaluation completed 320 utility and 40 null runs, but its analysis interface did not name the routingonly comparator in the exact form required by the prespecified criterion. It therefore yielded no eligible primary estimate and is not used for scientific interpretation. The fresh-seed replication retained the method, PPR calcula-
Table 7: Pre-specified ALIVE–PPR replication criteria. The utility clause fails despite a positive pooled mean because e40 is negative; the joint criterion therefore fails.
tion, controller, trainer, cost model, comparator, environments, budgets, thresholds, and all numeric criteria; only the analyzer–comparator interface, provenance fields, and seed domain changed. The replication completed all 320 utility and 40 null runs; every trajectory was finite, in budget, and valid for the knownpopulation PPR construction. All 160 matched pairs had identical pre-action selections and ledgers. All e20 full curves and all null curves had zero certificates, exclusion requests, and active exclusions, with no clean-source violation. Every one of the 120 high-risk full curves certified exactly the constructed harmful set and activated exclusion. PPR separation was no later than Hoeffding on 120/120 and strictly earlier on 100/120 curves. Median PPR versus Hoeffding evidence counts were 96 versus 304 in e40, 62 versus 171 in e60, and 48 versus 48 in e80. The primary accuracy-AUBC difference, full minus routing-only, was +0.223541 pp: all ten seed aggregates were
Env. e20 e40 e60 e80
5%
10%
15%
20% AUBC
+0.000 +0.000 +0.000 +0.000 +0.000 +0.012 +0.008 −0.219 +0.077 −0.056 +0.736 +0.638 +0.625 +0.428 +0.615 +0.132 +0.380 +0.337 +0.442 +0.335
Table 8: ALIVE–PPR replication mean paired accuracy differences (full minus routing-only, percentage points; seeds 60–69). AUBC is the normalized trapezoidal integral; positive values favor the full method. positive, the two-sided conditional exact sign-flip value was 0.001953125, and the descriptive paired-t 95% interval was [+0.1468, +0.3003] pp. Environment means were exactly 0 in e20, −0.055500 pp in e40 (3 wins, 7 losses), +0.615000 pp in e60, and +0.334667 pp in e80. The pooled high-risk 5%-budget mean was +0.293334 pp and every leave-onehigh-risk-environment mean was positive. However, the prespecified utility criterion required a strictly positive mean in every high-risk environment; the negative e40 result violated that clause. The conjunctive headline criterion was therefore not met. Macro-F1 was descriptive rather than part of the prespecified criterion. Its aggregate difference was +0.227030 pp, with e20 0, e40 −0.043399 pp, e60 +0.605671 pp, and e80 +0.345848 pp. These secondary results do not alter the failed e40 clause. No threshold revision, official-test access, external-replication claim, or deployment claim is made; the PPR study remains a validation-adaptive, freshseed contract-repair replication.
Unavailable Resource Comparisons Alternative-resource integration requires a positive-width coordinate support common to every method–seed curve in an environment. The selector-factor study failed this condition for its global resource comparison. The ALIVE– CBE development analysis included all 600 runs but found zero common-width trainer-core-time support in e60. The ALIVE–CBR analysis found the same limitation in all four environments. Consequently, no aggregate estimate is reported for those axes; partial environment summaries or support redefinitions would change the estimand. The declared-budget results and exact ledgers remain available. Synchronized trainer-core time, where defined, covers only trainer entry through selection, training, and final evaluation with end synchronization; it excludes setup, launch, and serialization and is not complete-job wall time. The 16 fixed-trajectory repricings of fine-evaluation and maintenance costs leave trajectories fixed and are descriptive feasibility views, not evidence of behavior under an actual alternate-price rerun.
Study Sequence and Outcome Isolation The experiments were developed sequentially on a reused validation cache, so later variants are validation-adaptive rather than members of one independent confirmatory family. Table 9 reports the sequence with reader-facing names.
Internal package identifiers are included only to connect the manuscript to the anonymous code package; they are not method versions and are not used elsewhere in the exposition. ALIVE–CBE was fixed before its development outcomes. ALIVE–CBR was specified from the selector study and fixed before its own model outcomes, with a separate Holm family. Full-Consensus and the audit-only control were introduced after related results were available and are explicitly descriptive. Adaptive-Switch was motivated by opened validation evidence but fixed before its own outcomes; Fixed-Switch was specified independently of those outcomes. The initial PPR evaluation was fixed before execution. The PPR replication was specified only after its analyzer–comparator interface defect was observed; it retained the scientific design and numeric criteria while using fresh run and audit seeds. These controls provide local outcome isolation, not an externally timestamped preregistration or an independent population.
Reproducibility and Provenance The anonymous code package contains a SHA-256 manifest that binds the protocol, configuration, code, run grid, behavior-audit outputs, statistical analyses, and generated tables/figures for each package ID in Table 9. The manuscript omits the raw digest list: the machine-readable manifest is the authoritative byte-level index and can be verified directly from the extracted package. No absolute paths, host names, home directories, or machine timestamps are required. Reproduction follows the same five-stage workflow for each study: (i) verify the packaged manifest and configuration; (ii) execute the complete method–environment–budget– seed grid; (iii) reconstruct trajectories and ledger events without reading outcome metrics; (iv) compute the pre-specified statistical summaries; and (v) evaluate the corresponding decision criteria. Unknown IDs, duplicate or incomplete grid cells, non-finite metrics, type mismatches, budget violations, or digest mismatches invalidate the affected analysis. Resume operates at whole-run granularity: an incomplete run restarts from its seed and replaces the partial trace. Code-only validation covers compilation and linting, unit tests for selection, allocation, ledgers, replay, and the exact rotating-null construction. These tests verify implementation invariants; they are not conditioned on obtaining a favorable scientific result. For the PPR replication, provenance also binds the implementation diff from the initial evaluation, the fresh seed domains, and the explicit routing-only comparator used by the unchanged scientific criterion.
Limitations The theorem assumes fixed common support and potential labels, with probability taken over one ideal audit permutation. Its simultaneous guarantee applies within a run and does not extend familywise across repeated seeds, datasets, or deployments. The estimand excludes source-only records, temporal drift, call-dependent T noise, and identities without a strict majority. Since U = s Us , adding sources may reduce common support; partial overlap or changing source membership requires a revised estimand and certificate.
Study
Package ID
Coupled-action predecessor R15 and selector study ALIVE–CBE R16
ALIVE–CBR
R17
Full-Consensus
R18
Adaptive-Switch
R19
Fixed-Switch Initial PPR evaluation
D8 R20
PPR replication
R20.1
Role
Outcome status
Initial controller and incumbent-selection evidence Development screen followed by crossdataset validation
Mechanism checks passed; utility criterion failed; selector favored CBR Development screen passed; crossdataset criterion failed in e20 versus uniform Matched CBR-incumbent controller Persistent-action increment was positive; adjusted system-level criterion versus CBR was not met Post-diagnostic aggregation boundary Valid execution; lower utility than ALIVE–CBE Outcome-informed ranking diagnostic Valid execution; parity, primary significance, and high-risk improvement criteria failed Pre-fixed schedule diagnostic Descriptive only; no promotion criterion Same scientific design as the PPR replica- Analysis-interface mismatch produced tion no eligible primary estimate Fresh-seed repair of that interface Mechanism criteria passed; the e40 utility clause failed
Table 9: Study sequence, package mapping, and outcome status. All unfavorable results are retained. The official test set remained unopened. Certified means satisfying the pre-specified absolute disagreement and relative peer-separation criteria, not that a source is necessarily incorrect or harmful. The threshold τ must be chosen before label inspection, and (1/S) is an experimental setting rather than a universal corruption rate. Consensus-based certification may favor a shared incorrect majority or flag a correct minority, while monitoring and fallback mechanisms provide operational safeguards rather than downstream loss guarantees. The experiments provide controlled evidence using synthetic persistent perturbations, fixed CIFAR features, a small MLP head, four proxy budgets, and ten paired seeds per study. CIFAR-10 serves as cross-dataset validation rather than a natural multi-provider deployment. The study sequence is validation-adaptive, the official test cache remains unopened, and the proxy ledger is specific to the evaluated trajectories and costs. The Full-Consensus, Adaptive-Switch, PPR, auditonly, and Bluebirds analyses should therefore be interpreted as complementary diagnostic or exploratory evidence. The PPR replication improves the analysis interface but is not external confirmation and does not satisfy the pre-specified e40 utility condition; Bluebirds provides only a small complete panel without routing or training, where the base certificate has no detection power and the direct certificate remains unevaluated.
References Bardenet, R.; and Maillard, O.-A. 2015. Concentration Inequalities for Sampling without Replacement. Bernoulli, 21(3): 1361–1385. Dawid, A. P.; and Skene, A. M. 1979. Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm. Journal of the Royal Statistical Society: Series C (Applied Statistics), 28(1): 20–28. Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-
Fei, L. 2009. ImageNet: A Large-Scale Hierarchical Image Database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 248–255. Ernst, M. D. 2004. Permutation Methods: A Basis for Exact Inference. Statistical Science, 19(4): 676–685. He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 770–778. Hoeffding, W. 1963. Probability Inequalities for Sums of Bounded Random Variables. Journal of the American Statistical Association, 58(301): 13–30. Krizhevsky, A. 2009. Learning Multiple Layers of Features from Tiny Images. Technical report, University of Toronto. Loshchilov, I.; and Hutter, F. 2019. Decoupled Weight Decay Regularization. In International Conference on Learning Representations. Serfling, R. J. 1974. Probability Inequalities for the Sum in Sampling without Replacement. The Annals of Statistics, 2(1): 39–48. Waudby-Smith, I.; and Ramdas, A. 2020. Confidence Sequences for Sampling without Replacement. In Larochelle, H.; Ranzato, M.; Hadsell, R.; Balcan, M.-F.; and Lin, H.-T., eds., Advances in Neural Information Processing Systems, volume 33, 20204–20214. Curran Associates, Inc. Welinder, P.; Branson, S.; Belongie, S.; and Perona, P. 2010. The Multidimensional Wisdom of Crowds. In Advances in Neural Information Processing Systems, volume 23, 2424– 2432.