Conceptio › Archive › arXiv CS
arXiv CSopen access

When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making Ken Chen1 Wei Wang1 Sachith Seneviratne1 Hansani Weeratunge2 Saman Halgamuge1 1 Department of Mechanical Engineering, The University of Melbourne, Melbourne, Australia 2 Department of Mechanical Engineering, Sri Lanka Institute of Information Technology, Sri Lanka [email protected]

Abstract

solve complex tasks. One prominent paradigm extends single-model capabilities by prompting heterogeneous agents to address the same query, shifting the focus toward collaborative decision-making, where diverse outputs must be aggregated into a final consensus. This paradigm is rooted in human collective intelligence (Woolley et al., 2010; Zhou et al., 2026) and cognitive diversity, positing that a team of diversely skilled solvers can outperform a uniform group of top individuals (Hong and Page, 2004). However, diversity alone does not guarantee superior group performance. When agents disagree, the aggregation mechanism ultimately determines whether complementary insights are synthesized or shared errors are compounded. Existing collective decision-making pipelines typically rely on voting-based mechanisms (Zhao et al., 2024) or a designated LLM judge (Zhao et al., 2026). However, these approaches present distinct limitations. Voting and electoral rules-based methods rely on simple consensus and lack a reliable, objective anchor to resolve conflicts when agents disagree. Conversely, dictatorial approaches such as LLM-as-a-judge attempt to supply an external decision, but the evaluation remains grounded in the same forward-generated traces. Furthermore, they have also been observed to exhibit systematic biases, such as position bias, verbosity bias, or self-enhancement bias (Zheng et al., 2023). While multi-round debate can refine opinions (Du et al., 2024; Liang et al., 2024), it incurs high computational costs and remains vulnerable to conformity, identity-driven sycophancy, and persona instability (Baltaji et al., 2024; Choi et al., 2026; Li et al., 2026). Focusing on single-round aggregation, we argue that the critical missing ingredient is a reliable external anchor, a reference distribution that guides the system in selecting and fusing conflicting answers. The root cause of aggregation failure is directional rather than merely algorithmic. Plu-

arXiv:2609.11709v1 [cs.AI] 10 Sep 2026

When multiple LLM agents yield conflicting answers, the decision-making process dictates whether agent diversity improves performance or merely compounds shared errors. Existing collective decision-making methods, including voting, electoral rules, and LLM judges, rely on forward reasoning: they map evidence to labels in one direction. Although these methods can combine diverse forward traces, they still aggregate estimates that share this evidence-tolabel factorization and can inherit correlated errors within the forward pool. We therefore construct a reverse posterior for each instance through Bayesian backward reasoning from an explicit likelihood. The forward and reverse posteriors provide differently factorized approximations of the underlying posterior. Because estimates from different factorizations may tend to share the same error less often, we use Jensen-Shannon divergence to rank agents by cross-path consistency. This cross-path consistency signal underlies three strategies: hard selection (MinJS), soft reweighting (FwdJS), and log-linear fusion (LogLin). Evaluated on DDXPlus across five LLM backbones, our proposed strategies show consistent improvements: MinJS outperforms random selection across all backbones, FwdJS generally improves over the strongest baseline, and LogLin achieves the best performance among the evaluated methods, with its largest gains on the subset where the agents disagree. Despite its weaker standalone accuracy, the reverse posterior serves as a more useful anchor than forward-only alternatives, providing complementary information for collective decision-making. When labeled data are available, a lightweight two-stage calibration can further refine the reverse anchor and improve aggregation performance.

1

Introduction

Large Language Model (LLM)-based multi-agent systems coordinate multiple agents to collectively 1

reasoning, while each Fi represents an agent’s forward posterior. Unlike existing backward checks that merely score individual reasoning chains in math problems (Weng et al., 2023; Jiang et al., 2024), our R serves as a complete class distribution used to evaluate and fuse multiple forward agents. We leverage the Jensen–Shannon divergence, JS(Fi , R), through three progressive aggregation strategies:

rality voting, electoral rules, and LLM judges all operate exclusively within forward reasoning: they reason directly from evidence to conclusions along the exact same conditioning path as the agents they evaluate. Consequently, errors within this pool are highly collinear. If the majority hallucinates, a vote or a judge is likely to echo the same mistake, trapping the system in a collective error. To break this cycle, the system requires a reference that resides in the same class-posterior space but is differently factorized, constructed via a path distinct from standard forward elicitation. Bayes’ theorem provides two complementary factorizations of the same posterior. While forward reasoning executes one as a discriminative evidence-to-conclusion mapping, we explicitly construct the other through Bayesian backward reasoning. By defining an explicit likelihood P (e | d) over evidence e and labels d ∈ Y, together with a prior, we invert the inference once per instance, a process of Bayes-style backward reasoning that yields a shared reverse posterior R(d | x), where x denotes the observed input. Consequently, each agent’s forward posterior Fi and R serve as differently factorized, biased approximations of the same class posterior P (d | x). Our working hypothesis is that their respective errors are far less collinear than those strictly within the forward pool. Because biased estimates derived from different factorizations are less likely to converge on the same incorrect conclusion than estimates from the same factorization, cross-path agreement provides a useful consistency signal for identifying more reliable predictions. Therefore, we utilize the JensenShannon divergence, JS(Fi , R), as a ranking signal for cross-path consistency rather than a definitive certificate of correctness. This symmetric metric treats neither channel as absolute ground truth, and its bounded nature ensures comparability across instances. Crucially, measuring distance to the pool’s internal consensus, such as the majority vote or the average prediction, cannot fulfill this role: it merely rewards conformity, actively penalizing the rare agent that happens to be correct when the majority is wrong. Figure 1 summarizes our reverse-anchored framework: a shared reverse posterior serves as a reference for selecting, reweighting, and fusing the forward agents. Concretely, a single backward inference generates the shared reverse posterior R via explicit Bayesian backward reasoning over the candidate shortlist proposed by the agents’ forward

• MinJS (Hard Selection): Selects the single forward agent closest to R. While the ranking is informative, hard selection can be brittle because cross-path agreement measures compatibility with R, not an absolute guarantee of correctness. • FwdJS (Soft Reweighting): Softly reweights the forward pool based on their alignment with R. Crucially, R solely determines the aggregation weights, meaning the final prediction remains a pure mixture of the forward posteriors: R shapes the weights but contributes no probability mass of its own. • LogLin (Log-Linear Fusion): Actively blends R into the final prediction as a multiplicative factor with a small fixed weight (wR =0.2). This combination captures complementary cross-path information that the first two methods miss, as they rely solely on the forward outputs. Our main contributions are summarized as follows: 1. Cross-path reverse anchor. We introduce a reverse posterior obtained by Bayesian backward reasoning that acts as a reliable, label-free reference for multi-agent decisionmaking. We demonstrate that JS(Fi , R) effectively ranks and routes conflicting agents based on cross-path consistency. 2. Label-free aggregation suite. We propose three training-free strategies (MinJS, FwdJS, and LogLin). Extensive experiments demonstrate the progressive effectiveness of our proposed suite: MinJS consistently outperforms the selection baseline, FwdJS improves over the strongest baseline in most settings, and LogLin achieves strictly superior performance against all baselines across all scenarios. 2

3. Anchor indispensability. We establish that despite often being the weakest standalone predictor, the reverse posterior R acts as an irreplaceable anchor. Replacing it with the poolmean forward posterior (M eanF ) or a poolexternal general single-agent forward posterior (GenF ) strictly degrades both FwdJS and LogLin, confirming the unique value of crosspath diversity.

mean. RevThink (Chen et al., 2025) instead trains a model to internalize forward and backward reasoning, while inference still produces a single forward answer. These methods thus use reverse reasoning primarily as a candidate-level verification signal. Rather than using the reverse signal only to verify individual candidate-level predictions, we derive a full reverse posterior and use its cross-path consistency with the forward posteriors to guide selection, reweighting, and fusion across multiple agents.

4. Lightweight labeled calibration. When a limited labeled split is available, we apply a two-stage calibration that repairs the reverse anchor’s quality and further elevates the aggregation strategies.

2

Ensembling and unlabeled routing. Model ensembling and routing also combine information from multiple models, but they formulate the problem as prediction fusion or model selection rather than collective decision-making among explicit agents. LLM-Blender (Jiang et al., 2023) merges candidate answer texts. DeePEn (Huang et al., 2024) and PackLLM (Mavromatis et al., 2024) combine next-token predictions across models. SMOOTHIE (Guha et al., 2024) performs unlabeled routing by scoring each model’s sample-wise quality from the unlabeled outputs and sending the input to the highest-scoring model. These methods operate on forward predictions and use agreement or predictive compatibility as their quality signal. In contrast, we retain the multi-agent collectivedecision setting and introduce a reverse posterior as an external reference without relying on another forward reasoning agent or a learned judge.

Related Work

Multi-agent collective decision-making. LLM multi-agent systems commonly coordinate through debate, role specialization, or layered synthesis (Du et al., 2024; Liang et al., 2024; Wang et al., 2025). Once multiple agents produce answers to the same instance, the system must aggregate their individual decisions into a single collective outcome. Existing approaches largely do so either by aggregating agents’ ballots through majority voting or more elaborate electoral rules (Zhao et al., 2024; Ai et al., 2025), or by delegating the decision to a designated judge (Zheng et al., 2023; Liu et al., 2023; Chan et al., 2024; Zhao et al., 2026). Confidence-weighted consensus further refines the former through discussion (Chen et al., 2024), while recent analyses suggest that simple majority already captures much of the benefit attributed to multi-round debate (Choi et al., 2025). Despite these differences, both voting and judging derive their decision signal from the agents’ forward outputs. In contrast, we derive an instance-level reference from a generative reverse posterior, providing a training-free alternative to another vote or judge.

3

Method

3.1

Collective decision over agent posteriors

Consider an instance x and a finite label set Y, shown at the left of Figure 1. K heterogeneous agents A = A1 , . . . , AK each answer the same x and return a forward posterior Fi (d | x) over d ∈ Y. Our goal is to make a collective decision from the agent posteriors Fi K i=1 . We consider two forms of collective decision: selection, which adopts the posterior of one agent, and fusion, which fuses the pool into a single posterior P (d | x) and predicts arg maxd P (d | x). Both are implemented within the same label-free and training-free inference framework.

Forward–backward reasoning. A separate line of work uses backward reasoning to evaluate a proposed answer or label rather than to aggregate decisions from multiple agents. SelfVerification (Weng et al., 2023) introduces backward checks at inference time to assess candidate solutions by testing whether the candidate is consistent with the underlying conditions. FOBAR (Jiang et al., 2024) further combines Self-Consistencystyle forward answer votes (Wang et al., 2023) and backward probabilities through a geometric

3.2

Reverse posterior as a consistency anchor

The class label d ∈ Y is latent; what differs across inference paths is how the evidence is conditioned on d. Forward agents elicit Fi (d | x) through a discriminative route from the instance x to the 3

Figure 1: Reverse-anchored aggregation, left to right. Input x is decomposed into contextual evidence a and remaining evidence e under the structure a→d→e. The forward path (yellow) produces agent posteriors Fi , while the reverse path (blue) produces a shared reverse posterior R ∝ P (e | d) P (d | a). The shared anchor signal (green), Di = JS(Fi , R), guides all three heads: MinJS selects one Fi , FwdJS reweights {Fi }, and LogLin additionally incorporates R into the output. The dashed box denotes optional labeled calibration R→R′ . The dashed blue arrow indicates the direct contribution of R to LogLin.

label space. We instead construct a complementary reverse view by specifying how the observed evidence factorizes conditioned on d and then applying Bayes’ rule. We decompose the input x into two evidence components, a and e, and posit the conditionalindependence structure a → d → e. Here, a represents contextual evidence, while e provides additional evidence about the latent label d. This implies P (e | a, d) = P (e | d), and Bayes’ rule gives P (d | a, e) ∝ P (e | d) P (d | a).

shared model priors or reasoning patterns, while the reverse factorization provides a structurally different source of evidence. We therefore do not treat R as a ground-truth posterior, but as a structurally distinct consistency signal whose errors need not be fully aligned with those of the forward predictors. This motivates using R to rank or weight the forward posteriors rather than introducing another forward vote or judge. We instantiate the reverse anchor with one shared reverse inversion per instance, which uses ordinal likelihood and context-activation maps together with two LLM elicitation steps over the agents’ top-k labels, followed by exact replay on the Bayesian network. The construction is an implementation of the reverse anchor rather than a requirement of the aggregation framework itself. For the label-free experiments, the resulting reverse distribution is used without any training data.

(1)

We thus obtain a reverse likelihood P (e | d) and a label prior P (d | a), and define the resulting reverse posterior R(d | x) ∝ P (e | d) P (d | a),

(2)

normalized over Y. Forward and reverse inference therefore provide two different conditionalization perspectives on the same latent label: direct evidence-to-label prediction versus generative inversion. Neither path is expected to recover the exact posterior P (d | x). LLM elicitation and model bias make both Fi and R biased approximations rather than ground-truth posteriors. In particular, no exact Bayes identity links an elicited forward posterior Fi to our constructed reverse posterior R. The motivation for introducing R is weaker and does not require either estimate to be exact: forward agents can exhibit correlated errors due to

Because the forward and reverse estimates arise from different factorizations, their errors can be partially complementary. This motivates a productlike composition that emphasizes labels supported by both channels, without requiring the stronger assumption that either distribution is exact or that the two estimates are statistically independent.

3.3

Reverse-referenced selection and fusion

We measure the consistency between each forward posterior and the reverse anchor (cross-path consis4

tency) using the Jensen-Shannon divergence:

LogLin additionally allows the reverse anchor to reshape the final posterior (dashed arrow in Figure 1). This fixed-weight product-like form is related at the operator level to the geometric-mean composition used in FOBAR (Jiang et al., 2024), but serves a different purpose: FOBAR combines forward and backward evidence for candidate verification, whereas we use the reverse posterior as a shared anchor for multi-agent aggregation.

Di = JS(Fi , R), = 21 KL(Fi ∥Mi ) + 21 KL(R∥Mi ),

(3)

Mi = 12 (Fi + R). The reverse posterior R serves as a shared external reference, shown in green in Figure 1, and the same consistency score is reused by the three heads on the right: hard selection, soft reweighting, and forward–reverse fusion.

Summary. The reverse posterior R is not treated as another competing prediction. Instead, it provides a shared consistency reference that supports three levels of collective decision: hard selection (MinJS), soft reweighting (FwdJS), and direct forward–reverse fusion (LogLin). Importantly, cross-path consistency is a measure of compatibility rather than a certificate of correctness. Consequently, R need not be the top-1 predictor itself to be useful as an aggregation anchor.

MinJS (hard selection). MinJS selects the agent whose posterior is most consistent with the reverse anchor: i⋆ = arg min Di , i (4) ˆ d = arg max Fi⋆ (d). d

The output dˆ is obtained from the selected agent’s forward posterior Fi⋆ . Thus, R is used only to rank the agents and does not contribute probability mass to the final prediction.

3.4

The ordinal likelihood and context-activation maps used to construct R may be misspecified, making the resulting reverse posterior a biased approximation of the underlying posterior. When a labeled training split is available, we optionally calibrate the reverse anchor R while keeping MinJS, FwdJS, and LogLin fixed (dashed box in Figure 1). The low-capacity parameterization is designed to permit calibration with limited labeled data, without retraining the underlying LLM. Calibration only modifies the reverse anchor used by the aggregation heads. Stage 1 calibrates the two ordinal maps used to construct the reverse posterior: the ordinal likelihood map used to parameterize P (e | d) and the contextual-activation map fused to parameterize P (d | a). Each map converts an ordinal rank k ∈ {0, . . . , 6} into a monotone continuous value through a logit-linear curve:

FwdJS (soft reweighting). MinJS uses reverse consistency for hard selection. FwdJS converts the same signal into soft weights over the forward pool: exp(−τ Di ) wi = P , j exp(−τ Dj ) X PFwdJS (d) = wi Fi (d).

(5)

i

We fix the temperature to τ = 5.0 throughout all experiments. The collective label is arg maxd PFwdJS (d). Agents whose posteriors are more consistent with R receive larger weights, while the resulting posterior remains a convex combination of the forward posteriors. In this operator, R shapes the routing weights but does not itself contribute probability mass to the output. LogLin (forward–reverse fusion). When the forward and reverse estimates contain partially complementary errors, a product-like composition can emphasize labels supported by both channels. LogLin therefore builds on the FwdJS aggregate and directly incorporates the reverse posterior: PLogLin (d) ∝ PFwdJS (d) 1−wR R(d) wR ,

Labeled calibration of the reverse anchor

 vk = ℓ + (h − ℓ) σ a + b · k/6 ,

b > 0, (7)

where σ is the logistic function and (a, b) are learned separately for the two maps. A temperature T is additionally fitted to rescale the replayed reverse posterior. After fitting, the two maps and T are frozen, and the reverse posterior R is recomputed by replaying the Bayesian network with the calibrated parameters. Stage 2 then applies a scalar class-prior correc-

(6)

where we use wR = 0.2 by default unless stated otherwise (Appendix D). FwdJS determines how much each forward agent contributes, whereas 5

tion R(d) , R (d) ∝ m(d)γ ′

which the five forward top-1 predictions are not unanimous. Headline tables use these two slices.

(8)

where m(d) is the class marginal of R and γ is a single scalar parameter. The resulting calibrated reverse anchor R′ therefore adds only this scalar correction and does not introduce a class-specific bias bd . Fitting details and a higher-capacity variant with a class-specific bias vector b ∈ R|Y| , used as a capacity upper bound, are provided in Appendix B. 3.5

4.2

Collective-decision baselines operate entirely within the forward channel. Random Agent (Rand.) uniformly samples one of the five forward agents per case. We report the pooled accuracy over three fixed seeds. It serves as the selection control for MinJS. Voting-based aggregators follow GEDI (Zhao et al., 2024) and operate on the same five forward ballots or posteriors without reverse input. Plurality counts top-1 votes; Range voting treats each posterior as a cardinal score vector and selects the disease with the highest total score; Borda Count, Bucklin, IRV, Minimax, and Ranked Pairs apply the corresponding ordinal scoring or runoff rules. The dictatorial-based baselines are two LLM judges over the same five agents: Informed dictatorial judge (Informed Dicta.) and ACH-inspired structured judge (Zhao et al., 2026).

Computational cost

Inference requires K forward agent inferences and one shared reverse inversion per instance, followed by lightweight divergence computation and posterior aggregation. The reverse inversion is implemented with two LLM elicitation steps composed through the Bayesian network, and the same reverse posterior R is reused by all aggregation operators. No parameter updates are performed at test time; labeled calibration, when used, is fit on a training split only. Thus, the method introduces no fine-tuning cost and requires only a single shared reverse construction in addition to the K forward agent inferences.

4

Experimental Setup

4.1

Dataset

Baselines

4.3

Models and implementation details

The forward pool comprises five heterogeneous prompting strategies applied to each backbone: Tree-of-Thought (ToT), few-shot MedPrompt, vanilla chain-of-thought, common-bias, and rarebias. Each agent produces a top-5 posterior over the official 49-label set. We use unified sampling settings across all agents, with temperature = 0.6 and top-p = 0.95, and generate n=5 responses per query, one for each agent. We evaluate the resulting pools on five backbones: Qwen330B-A3B (Yang et al., 2025) (hereafter Qwen330B), Mistral-Small-3.1-24B (Mistral AI, 2025), Llama-3.1-8B (Grattafiori et al., 2024), GLM-432B (Team GLM et al., 2024), and DeepSeekV4-Flash (DeepSeek-AI et al., 2026). The primary metric is Gold Top-Pathology Accuracy at 1 (GTPA@1), where a prediction is counted as correct when the gold pathology is the arg max of the final posterior. The shared reverse posterior is the label-free inversion described in Section 3.2 and is reused by all reverse-referenced operators for each case. Our methods are MinJS selection, FwdJS reweighting with τ = 5.0, and LogLin forward– reverse fusion with fixed wR = 0.2 (Tables 1–2; see Appendix D for the wR sweep). Within each backbone, all methods are evaluated on the same fixed case set.

We evaluate on DDXPlus (Fansi Tchango et al., 2022), a synthetic diagnostic benchmark whose cases pair structured clinical evidence with a closed set of 49 disease labels. Each case records demographics and observed findings, together with contextual information that precedes the remaining evidence. For this benchmark, the contextual component corresponds to a, the remaining observed evidence to e, and the latent class d to the diagnosis. Forward agents map the full case description to a disease posterior, while the reverse procedure constructs R(d | x) ∝ P (e | d) P (d | a) once per case from the same evidence. We use a fixed 2,000-case test slice. Because every aggregator must use the same five forward posteriors and the same shared reverse anchor, we restrict evaluation to the per-backbone intersection of complete forward cases and cases with an available reverse posterior. We apply the same availability and parsing criteria to all methods within each backbone, yielding a fixed evaluation pool independent of the aggregation method. All denotes this complete perbackbone pool. Disagree is the subset of All in 6

5

Results

5.1

Main Results: Reverse-Anchored Collective Decision

line to test whether a generic single agent can replace collective aggregation, and as a replacement anchor to test whether a pool-external forward posterior can substitute for R. On All, R is the weakest standalone predictor among {R, MeanF, GenF} for every backbone, trailing M eanF by 7.8–26.4 pp, yet using R as the shared anchor yields the highest FwdJS and LogLin performance. On Disagree, R is the weakest of the three on Mistral, Llama, and DeepSeek. GLM is the exception (40.68% for R vs. 40.23% for M eanF ), while on Qwen GenF (40.30%) is slightly weaker than R (40.90%). GenF is often a stronger standalone predictor than R, yet performs worse as a replacement anchor. Using M eanF as the frozen LogLin anchor yields ∆≤0 relative to the strongest forward-only electoral baseline on every backbone, whereas R remains +1.2 to +4.7 pp above the same electoral baseline. The electoral gain is therefore attributable to the reverse channel, not the log-linear operator alone. Replacing R with either M eanF or GenF lowers both fusion heads on every backbone. The reverse construction R(d) ∝ P (e | d) P (d | a) factorizes into a likelihood term and a contextual prior. Neither one-factor variant matches the full shared R on Disagree LogLin for Qwen, Mistral, GLM, or DeepSeek. On Llama Disagree, Rprior slightly exceeds R (51.69 vs. 50.90), but remains weaker than M eanF and GenF as a standalone predictor. Appendix C extends Table 3 with the one-factor reverse anchors Rlik (likelihood only) and Rprior (contextual prior only).

A shared reverse posterior turns an unanchored forward pool into a cross-path collective decision. Tables 1 and 2 report GTPA@1 on All and Disagree. A double rule separates single-agent selection from pool aggregation. For selection methods, the best result is shown in bold; for voting, judging, and fusion methods, the best and runner-up results are shown in bold and underline, respectively. LogLin consistently outperforms the strongest electoral baseline. On Disagree, it exceeds the best forward-only electoral rule on all five backbones by 1.2–4.7 pp and also outperforms both LLM judges. On All, which includes cases where the forward agents already agree, LogLin achieves the best result for every backbone. Gains on Disagree are therefore diluted in the aggregate score. FwdJS likewise outperforms the strongest electoral rule on Disagree by 0.4–3.5 pp. The comparison on GLM illustrates that reverse-guided reweighting is not uniformly dominant: ACH reaches 44.07%, slightly above FwdJS at 42.60%, while LogLin remains best at 44.41%. The three operators also exhibit the intended progression from hard selection to soft aggregation and direct reverse fusion. MinJS already exploits reverse consistency, outperforming random agent on every backbone, but hard selection can remain below the strongest election (e.g., 71.64% vs. 72.69% for Range voting on Qwen All). FwdJS improves on this by distributing weight across agents according to their consistency with R, while LogLin further incorporates R directly into the fused posterior. Thus, the gains arise from using R as a cross-path reference for routing and fusion. 5.2

5.3

What the Reverse Anchor Adds Beyond the Forward Pool

The two analyses below examine complementary aspects of what the reverse anchor R contributes beyond the forward pool. The first measures how often a reference predictor reproduces the forward consensus’ same incorrect label, while the second evaluates whether distance to the reference provides a useful signal for ranking the five forward agents.

Anchor Utility Is Not Standalone Accuracy

Table 3 reports standalone GTPA@1 and frozenhead fusion performance when each candidate distribution is used as the shared anchor. The central pattern is that R is often weaker as a standalone predictor, yet more useful as an anchor. M eanF denotes the equal-weight average of the in-pool forward posteriors {Fi } and is used as a replacement anchor in Table 3. GenF is a generic pool-external single agent: one forward call with a neutral persona and the same top-5 output schema as the pool agents. We use it in two roles: as a standalone base-

Reverse errors collide less often on the same incorrect label. When plurality top-1 and a reference predictor are both wrong, π denotes the conditional probability that they assign the same incorrect label. Table 4 reports three matched comparisons on the same headline pool: plurality versus R, versus the pool-external GenF , and versus 7

Selection Model

Rand. MinJS

Qwen3-30B Mistral-Small-3.1-24B Llama-3.1-8B GLM-4-32B DeepSeek-V4-Flash

70.10 69.18 48.39 62.16 75.35

71.64 71.37 49.87 64.26 77.60

Voting-based

Dictatorial-based

Our method

Borda Ranked Informed Plurality Range Bucklin IRV Minimax ACH FwdJS LogLin Count Pairs Dicta. 71.54 71.37 54.11 65.32 76.84

72.69 73.20 57.90 65.12 77.80

71.34 72.54 54.16 64.31 77.60

71.94 72.19 54.42 65.42 77.29

72.14 72.29 54.01 64.82 77.19

71.54 72.03 52.75 64.66 76.99

71.69 72.29 54.47 65.57 77.14

72.84 72.39 56.39 65.93 76.58

71.84 72.13 50.38 66.78 75.87

73.85 73.51 58.15 66.18 78.61

74.25 74.58 58.46 67.04 79.07

Table 1: GTPA@1 (%) on All. ‘Rand.’ and ‘Dicta.’ denote ‘random’ and ‘dictatorial’. The double rule separates single-agent selection from pool aggregation. For selection, the best result is shown in bold; among pool-aggregation methods, the best and runner-up results are shown in bold and underlined, respectively. Selection Model

Rand. MinJS

Qwen3-30B Mistral-Small-3.1-24B Llama-3.1-8B GLM-4-32B DeepSeek-V4-Flash

40.35 40.69 35.64 33.63 41.38

45.26 46.24 37.76 38.19 48.15

Voting-based

Dictatorial-based

Our method

Borda Ranked Informed Plurality Range Bucklin IRV Minimax ACH FwdJS LogLin Count Pairs Dicta. 44.96 46.24 44.14 40.79 45.93

48.12 50.59 49.70 40.23 48.44

44.06 48.88 44.14 38.42 47.85

45.86 47.95 44.52 40.90 46.96

46.47 48.22 43.92 39.55 46.67

44.96 47.56 42.04 39.32 46.07

45.41 48.22 44.59 41.36 46.52

48.57 48.48 47.60 42.03 45.19

46.17 49.01 40.62 44.07 44.15

51.58 51.39 50.08 42.60 50.81

52.78 53.90 50.90 44.41 52.30

Table 2: GTPA@1 (%) on Disagree. Column groups and formatting match Table 1.

each in-pool agent Fi , with results pooled across the five agents. Across all five backbones, π(F, R) is the lowest of the three (0.196–0.413), below π(F, GenF ) (0.594–0.765) and π(F, Fi ) (0.680– 0.829). The same ordering holds on Disagree (Appendix E). GenF serves as an external extraforward control, while π(F, Fi ) measures an inpool echo rate because plurality is itself constructed from Fi . Thus, R shares the same mistaken label with the forward consensus less often than either a pool-external forward predictor or the agents that constitute that consensus. Appendix E further reports the 2×2 correctness grids and Matthews ϕ coefficients for the corresponding error indicators.

Qwen3-30B, Disagree

GTPA@1 (%)

45 40 35 Reverse R

30

meanF GenF Random agent

25 1

2 3 4 JS rank (1 = closest)

5

Figure 2: GTPA@1 (%) by JS rank on Disagree (Qwen3-30B). Reverse R (blue), M eanF (orange), and GenF (green). Dashed line: random agent.

JS divergence to R ranks agents by accuracy. For each case, we rank the five forward agents by JS(Fi , ·) with respect to a given anchor and evaluate each rank k by the corresponding agent’s GTPA@1. In aggregate, agents with lower JS divergence tend to be more accurate than those with higher divergence. We show Qwen as a representative case in Figure 2 using R, M eanF , and GenF as anchors, and report all backbones in Appendix A. The ranking is more consistently monotone under R than under M eanF or GenF . Under R, accuracy decreases with rank on Qwen and Mistral, is nearly tied between rank-1 and rank-2 on DeepSeek, and peaks at rank-2 and rank-3 on Llama and GLM, respectively. By comparison, the ranking under M eanF is monotone only on Llama, while the other four backbones break the rank ordering. Under GenF , only Mistral is monotone.

On Qwen, Llama, and GLM, rank-2 achieves the highest accuracy, indicating that the agent closest to GenF is not necessarily the most accurate. 5.4

Calibrating the Reverse Anchor

Section 3.4 calibrates the reverse anchor R on a labeled training split while keeping MinJS, FwdJS, and LogLin fixed. Table 5 reports test GTPA@1 gains from the calibrated reverse anchor R′ over the uncalibrated R on the same evaluation pool. On All, all three aggregation heads improve on all five backbones: LogLin by 0.05–2.83 pp, FwdJS by 0.20–1.51 pp, and MinJS by 1.06–4.64 pp. The standalone reverse top-1 accuracy also improves by 1.4–8.6 pp, while the higher-capacity class-specific 8

All

Disagree

Model

Anchor

Stand. FwdJS LogLin Stand. FwdJS LogLin

Qwen3-30B

R M eanF GenF

61.50 72.69 68.47

73.85 72.54 72.94

74.25 72.69 72.44

40.90 48.12 40.30

51.58 47.67 48.87

52.78 48.12 48.27

Mistral-S.3.1-24B

R M eanF GenF

59.96 73.26 72.24

73.51 73.00 72.90

74.58 73.10 73.10

40.82 50.73 48.75

51.39 50.07 49.80

53.90 50.33 50.20

Llama-3.1-8B

R M eanF GenF

31.50 57.90 46.64

58.15 57.45 57.60

58.46 57.60 56.28

30.63 49.70 35.96

50.08 49.02 49.25

50.90 49.25 47.52

GLM-4-32B

R M eanF GenF

57.34 65.12 61.48

66.18 65.47 65.88

67.04 65.47 64.41

40.68 40.23 34.80

42.60 41.02 41.92

44.41 41.02 39.55

R DeepSeek-V4-Flash M eanF GenF

67.76 77.80 75.22

78.61 77.90 77.65

79.07 77.80 77.75

36.59 48.44 43.85

50.81 48.74 48.00

52.30 48.44 48.44

Table 3: Anchor utility on the headline pool (GTPA@1 %). Each row uses one anchor distribution for standalone prediction and for frozen FwdJS / LogLin (wR =0.2). Bold: best of {R, MeanF, GenF} in that column, per backbone. GenF : general single agent (pool-external, neutral persona). Stand.: standalone anchor argmax.

Model

π(F, R)

π(F, GenF )

π(F, Fi )

Qwen3-30B Mistral-S.3.1-24B Llama-3.1-8B GLM-4-32B DeepSeek-V4-Flash

0.413 0.404 0.196 0.348 0.393

0.765 0.725 0.594 0.658 0.691

0.829 0.775 0.680 0.750 0.781

(MinJS), reverse-guided reweighting (FwdJS), and lightweight forward–reverse fusion (LogLin). Across five LLM backbones on DDXPlus, reverse-anchored aggregation improves over baseline methods, with the largest gains on Disagree. Notably, R is often weaker as a standalone predictor, yet replacing it with either M eanF or GenF degrades FwdJS and LogLin, highlighting the distinction between predictive accuracy and anchor utility. When both the forward consensus and a reference are wrong, R is also less likely to share the same incorrect label with the forward consensus than GenF or an in-pool agent. Two-stage calibration further improves the frozen aggregation rules by refining the reverse anchor. Together, these results support reverse consistency as a useful complementary signal for multi-agent collective decision-making.

Table 4: Label-collision rates on All (headline pool). π=P (predF =predX | both wrong) for plurality versus X ∈ {R, GenF, Fi }. π(F, Fi ) is pooled over the five pool agents. Bold: lowest π per row. Disagree π, 2×2 grids, and ϕ are in Appendix E.

variant is reported separately in Appendix B. The gains are concentrated on Disagree: LogLin improves by 0.59–4.51 pp on four backbones, with Mistral the sole exception (−0.26 pp). Adding a class-specific bias b ∈ R|Y| further improves fusion performance, but increases reverse top-1 accuracy by up to 26 pp on Disagree. This pattern is consistent with the additional capacity primarily capturing class-prior effects rather than providing a stronger routing signal.

6

Limitations The reverse anchor is defined over a finite label set. It inverts an explicit likelihood P (e | d) over a closed candidate set, using agents’ shortlisted candidates, and therefore targets evidence-to-label aggregation rather than open-ended generation without a discrete candidate space. Constructing R also incurs one additional generative inversion per case on top of the K forward agents. Future work could characterize how the reliability of R interacts with aggregation performance, including the point at which a sufficiently weak reverse posterior begins to hurt aggregation. For the labeled extension, studying calibration under

Conclusion

LLM-based multi-agent systems aggregate heterogeneous answers, yet voting and LLM judges remain confined to the same evidence-to-conclusion path and can inherit correlated errors from the agent pool. We introduce a shared reverse posterior R by Bayesian backward reasoning as a structurally distinct reference and use JS(Fi , R) to guide collective decision-making. This signal drives three training-free aggregators: hard selection 9

All (∆ pp vs. uncalibrated)

Disagree (∆ pp vs. uncalibrated)

Backbone

∆rev ∆minJS ∆FwdJS ∆log-lin.

∆rev ∆minJS ∆FwdJS ∆log-lin.

Qwen3-30B Mistral-S.3.1-24B Llama-3.1-8B GLM-4-32B DeepSeek-V4-Flash

+4.97 +3.06 +8.63 +2.27 +1.42

+7.22 +3.96 +5.93 +2.82 +3.56

+1.86 +1.53 +4.64 +1.06 +1.22

+0.80 +0.20 +1.51 +0.20 +0.25

+1.61 +0.05 +2.83 +0.35 +0.10

+5.41 +3.96 +6.91 +2.49 +3.41

+2.41 +0.53 +2.25 +0.45 +0.74

+4.51 −0.26 +3.30 +2.03 +0.59

Table 5: Test GTPA@1 gain (pp) of calibrated R over uncalibrated R. MinJS, FwdJS, and LogLin heads are frozen. The same case pools as in Tables 1–2 are used.

progressively smaller labeled sets would clarify its sample efficiency and establish how much supervision is sufficient for reliable anchor refinement. Our evaluation focuses on a single-round, samebackbone setting. The forward pool consists of different prompting strategies applied to one LLM backbone at a time, leaving mixed-backbone pools and multi-round debate as natural extensions. The label-collision statistic π captures how often two incorrect predictors assign the same wrong label, while future work could further investigate the mechanisms underlying such error overlap. Finally, all experiments are conducted on DDXPlus, a synthetic closed-set diagnostic benchmark with 49 diseases. Evaluating the framework across additional evidence-to-label domains would therefore be an important next step, and the proposed aggregators should not be interpreted as clinical decisionsupport systems.

Papers), pages 7066–7085, Bangkok, Thailand. Association for Computational Linguistics. Justin Chen, Zifeng Wang, Hamid Palangi, Rujun Han, Sayna Ebrahimi, Long Le, Vincent Perot, Swaroop Mishra, Mohit Bansal, Chen-Yu Lee, and Tomas Pfister. 2025. Reverse thinking makes LLMs stronger reasoners. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8611–8630, Albuquerque, New Mexico. Association for Computational Linguistics. Hyeong Kyu Choi, Jerry Zhu, and Sharon Li. 2026. When identity skews debate: Anonymization for biasreduced multi-agent reasoning. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14284–14311, San Diego, California, United States. Association for Computational Linguistics. Hyeong Kyu Choi, Xiaojin Zhu, and Sharon Li. 2025. Debate or vote: Which yields better decisions in multi-agent large language models? In Advances in Neural Information Processing Systems. DeepSeek-AI, Anyi Xu, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chenchen Ling, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chengyu Hou, Chenhao Xu, Chenze Shao, Chong Ruan, Conner Sun, and 300 others. 2026. Deepseek-v4: Towards highly efficient million-token context intelligence. arXiv preprint arXiv:2606.19348.

References Rui Ai, Yuqi Pan, David Simchi-Levi, Milind Tambe, and Haifeng Xu. 2025. Beyond majority voting: LLM aggregation by leveraging higher-order information. arXiv preprint arXiv:2510.01499. Razan Baltaji, Babak Hemmatian, and Lav Varshney. 2024. Conformity, confabulation, and impersonation: Persona inconstancy in multi-agent LLM collaboration. In Proceedings of the 2nd Workshop on Cross-Cultural Considerations in NLP, pages 17–31, Bangkok, Thailand. Association for Computational Linguistics.

Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch. 2024. Improving factuality and reasoning in language models through multiagent debate. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 11733–11763. PMLR.

Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2024. ChatEval: Towards better LLM-based evaluators through multi-agent debate. In The Twelfth International Conference on Learning Representations.

Arsene Fansi Tchango, Rishab Goel, Zhi Wen, Julien Martel, and Joumana Ghosn. 2022. Ddxplus: A new dataset for automatic medical diagnosis. Advances in neural information processing systems, 35:31306– 31318.

Justin Chen, Swarnadeep Saha, and Mohit Bansal. 2024. ReConcile: Round-table conference improves reasoning via consensus among diverse LLMs. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long

Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad AlDahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh

10

Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, and 542 others. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783.

Mistral AI. 2025. Mistral small 3.1. https://mistral. ai/news/mistral-small-3-1/. Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Dan Zhang, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, Hao Yu, Hongning Wang, Jiadai Sun, Jiajie Zhang, Jiale Cheng, Jiayi Gui, Jie Tang, Jing Zhang, and 39 others. 2024. Chatglm: A family of large language models from glm-130b to glm-4 all tools. arXiv preprint arXiv:2406.12793.

Neel Guha, Mayee F. Chen, Trevor Chow, Ishan S. Khare, and Christopher Ré. 2024. Smoothie: Label free language model routing. In Advances in Neural Information Processing Systems. Lu Hong and Scott E. Page. 2004. Groups of diverse problem solvers can outperform groups of highability problem solvers. Proceedings of the National Academy of Sciences, 101(46):16385–16389.

Junlin Wang, Jue Wang, Ben Athiwaratkun, Ce Zhang, and James Zou. 2025. Mixture-of-agents enhances large language model capabilities. In The Thirteenth International Conference on Learning Representations.

Yichong Huang, Xiaocheng Feng, Baohang Li, Yang Xiang, Hui Wang, Ting Liu, and Bing Qin. 2024. Ensemble learning for heterogeneous large language models with deep parallel collaboration. In Advances in Neural Information Processing Systems.

Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations.

Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin. 2023. LLM-blender: Ensembling large language models with pairwise ranking and generative fusion. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14165–14178, Toronto, Canada. Association for Computational Linguistics.

Yixuan Weng, Minjun Zhu, Fei Xia, Bin Li, Shizhu He, Shengping Liu, Bin Sun, Kang Liu, and Jun Zhao. 2023. Large language models are better reasoners with self-verification. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 2550–2575, Singapore. Association for Computational Linguistics.

Weisen Jiang, Han Shi, Longhui Yu, Zhengying Liu, Yu Zhang, Zhenguo Li, and James Kwok. 2024. Forward-backward reasoning in large language models for mathematical verification. In Findings of the Association for Computational Linguistics: ACL 2024, pages 6647–6661, Bangkok, Thailand. Association for Computational Linguistics.

Anita Williams Woolley, Christopher F. Chabris, Alex Pentland, Nada Hashmi, and Thomas W. Malone. 2010. Evidence for a collective intelligence factor in the performance of human groups. Science, 330(6004):686–688.

Jiayi Li, Xiao Liu, and Yansong Feng. 2026. From single to societal: Analyzing persona-induced bias in multi-agent interactions. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 31609–31617.

An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, and 41 others. 2025. Qwen3 technical report. arXiv preprint arXiv:2505.09388.

Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Shuming Shi, and Zhaopeng Tu. 2024. Encouraging divergent thinking in large language models through multi-agent debate. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 17889–17904, Miami, Florida, USA. Association for Computational Linguistics.

Xiutian Zhao, Ke Wang, and Wei Peng. 2024. An electoral approach to diversify LLM-based multiagent collective decision-making. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 2712–2727, Miami, Florida, USA. Association for Computational Linguistics.

Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-eval: NLG evaluation using GPT-4 with better human alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2511–2522, Singapore. Association for Computational Linguistics.

Xuyang Zhao, Shiwan Zhao, Hualong Yu, Liting Zhang, and Qicheng Li. 2026. AgentCDM: Enhancing multi-agent collaborative decision-making via ACHinspired structured reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 34985–34993. AAAI Press.

Costas Mavromatis, Petros Karypis, and George Karypis. 2024. Pack of LLMs: Model fusion at testtime via perplexity optimization. In First Conference on Language Modeling.

Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot

11

of the reverse posterior R obtained by replaying the Bayesian network on the training cases. After fitting, the two maps and the temperature T are frozen. Stage 2 fits the scalar γ in (8) by k-fold minimization of the same NLL on the training fit split. Evaluation then applies the frozen MinJS, FwdJS, and LogLin heads on the same evaluation pool used for the label-free results.

arena. Advances in neural information processing systems, 36:46595–46623. Zhilun Zhou, Zihan Liu, Jiahe Liu, Yihan Wang, Qingyu Shao, Fengli Xu, Depeng Jin, and Yong Li. 2026. Identifying collective intelligence factor in LLM agent groups for generalizable multi-agent system design. In Findings of the Association for Computational Linguistics: ACL 2026, pages 12827–12842, San Diego, California, United States. Association for Computational Linguistics.

B.2 Capacity upper bound (L5 CURVE _L1 GB)

A

JS-based agent ranking across backbones

Table 7 reports the same frozen-head evaluation as Section 5.4, but gives each label its own bias term bd in addition to the scalar correction γ:

Table 6 reports GTPA@1 after ranking the five forward agents by JS(Fi , R) on Disagree. Figure 2 in the main text shows the Qwen Disagree slice and includes M eanF and GenF as alternative anchors. Figure 3 summarizes the remaining backbone–subset curves for both All and Disagree. Under R, the Disagree rank curve is strictly monotone on Qwen and Mistral. Llama and GLM attain their highest accuracy after rank-1, while rank-1 and rank-2 on DeepSeek differ by only one case. The ranking is less consistently monotone under M eanF and GenF . Model

Rank-1 Rank-5 Rand.

Qwen3-30B Mistral-S.3.1-24B Llama-3.1-8B GLM-4-32B DeepSeek-V4-Flash

45.26 46.24 37.76 38.19 48.15

27.52 29.19 26.73 24.63 25.48

R′ (d) ∝ R(d) m(d)−γ ebd .

(9)

This gives the calibration more flexibility and further improves both the standalone accuracy of R and the downstream fusion results. However, the large gain in reverse top-1 accuracy suggests that much of this improvement may come from adjusting label-specific prior preferences rather than from making R a better routing signal. We therefore use this higher-capacity version only as an upperbound reference, rather than as the primary calibrated method.

39.88 40.42 35.36 34.21 41.19

Backbone

All (∆ pp vs. uncalibrated)

Disagree (∆ pp vs. uncalibrated)

∆rev ∆minJS ∆FwdJS ∆log-lin.

∆rev ∆minJS ∆FwdJS ∆log-lin.

Qwen3-30B +12.05 Mistral-S.3.1-24B +7.90 Llama-3.1-8B +16.15 GLM-4-32B +10.70 DeepSeek-V4-Flash +12.52

+3.71 +2.29 +7.32 +3.33 +4.36

+2.06 +0.97 +3.33 +1.41 +1.32

+5.77 +19.10 +10.98 +2.75 +11.76 +5.55 +6.81 +14.49 +10.96 +4.24 +12.66 +7.57 +3.75 +26.52 +12.59

+6.17 +2.51 +4.95 +3.16 +3.85

+13.23 +5.81 +8.86 +8.70 +11.11

Table 6: GTPA@1 (%) by reverse JS rank on Disagree. Rank-k: the forward agent with the k-th lowest JS(Fi , R) on each case. Rand.: mean GTPA@1 across the five ranks.

Table 7: Test GTPA@1 gain (pp) from class-specific calibration of R over uncalibrated R. Same evaluation pool as Table 5.

B

C

Labeled calibration: fitting and capacity upper bound

Table 3 in the main text compares the full reverse posterior R with the forward-pool anchors M eanF and GenF . Table 8 uses the same evaluation pools and frozen FwdJS/LogLin heads, but replaces R with one-factor reverse constructions: Rlik uses only the likelihood term P (e | d), whereas Rprior uses only the contextual prior P (d | a). On Disagree with LogLin, neither one-factor construction matches the full R on Qwen, Mistral, GLM, or DeepSeek. Llama is the exception: Rprior reaches 51.69%, slightly above R at 50.90%, but remains weaker than M eanF and GenF as a standalone predictor.

Section 3.4 defines the two-stage calibration of R. This appendix details the fitting protocol for the calibrated reverse anchor and the higher-capacity class-specific bias variant. B.1

Likelihood-only and prior-only reverse anchors

Fitting protocol

The reverse inversion uses ordinal evidencelikelihood and context-activation bins. Each axis in (7) has seven ordinal ranks k ∈ 0, . . . , 6, with axis-specific clamps ℓ, h. Stage 1 fits four map parameters (as , bs , aa , ba ) together with a temperature T for rescaling the replayed reverse posterior. The objective is to minimize the label NLL 12

All Qwen3-30B, All

55

70.0 67.5 Reverse R meanF GenF Random agent

65.0 62.5 1

70.0 67.5 Reverse R

65.0

meanF GenF Random agent

62.5

2 3 4 JS rank (1 = closest)

5

1

GLM-4-32B, All

67.5

GTPA@1 (%)

GTPA@1 (%)

GTPA@1 (%)

72.5

2 3 4 JS rank (1 = closest)

50 45 Reverse R meanF GenF Random agent

40

1

5

2 3 4 JS rank (1 = closest)

5

DeepSeek-V4-Flash, All 80.0 GTPA@1 (%)

65.0 GTPA@1 (%)

Llama-3.1-8B, All

Mistral-S.3.1-24B, All 72.5

62.5 60.0 Reverse R

57.5

meanF GenF Random agent

55.0 1

77.5 75.0 72.5

Reverse R meanF GenF Random agent

70.0 67.5

2 3 4 JS rank (1 = closest)

5

1

2 3 4 JS rank (1 = closest)

5

Disagree (Qwen in Figure 2) Mistral-S.3.1-24B, Disagree

Llama-3.1-8B, Disagree

GLM-4-32B, Disagree

45

40 35

Reverse R meanF GenF Random agent

30 1

35 30 Reverse R

25

meanF GenF Random agent

20

2 3 4 JS rank (1 = closest)

GTPA@1 (%)

40 GTPA@1 (%)

GTPA@1 (%)

45

40

5

1

35 30 Reverse R

25

meanF GenF Random agent

20

2 3 4 JS rank (1 = closest)

5

1

2 3 4 JS rank (1 = closest)

5

DeepSeek-V4-Flash, Disagree

GTPA@1 (%)

50 45 40 35 Reverse R

30

meanF GenF Random agent

25 1

2 3 4 JS rank (1 = closest)

5

Figure 3: GTPA@1 (%) by JS rank on the evaluation pool. Reverse R, M eanF , and GenF are used as anchors. Dashed line: random agent.

All

cause FwdJS has already incorporated R through JS-based reweighting, we use wR =0.2 as a light direct contribution from the reverse anchor rather than selecting the weight separately for each backbone on the test set. On Disagree subset, larger wR can yield higher accuracy on some backbones. For example, GLM improves from 44.41% at wR =0.2 to 47.57% at wR =0.5.

Disagree

Model

Anchor Stand. FwdJS LogLin Stand. FwdJS LogLin

Qwen3-30B

Rlik Rprior

51.53 43.69

72.90 73.20

72.90 30.23 73.40 31.58

48.87 49.77

48.87 50.38

Mistral-S.3.1-24B

Rlik Rprior

44.74 43.21

73.06 73.21

74.23 34.44 73.83 26.23

50.20 50.60

53.11 52.19

Llama-3.1-8B

Rlik Rprior

19.67 30.59

57.89 58.29

56.88 16.98 59.20 29.53

49.74 50.34

48.53 51.69

GLM-4-32B

Rlik Rprior

47.32 39.89

64.86 65.57

65.12 33.75 66.03 25.37

39.64 41.22

40.32 42.13

DeepSeek-V4-Flash

Rlik Rprior

59.49 48.93

78.02 77.92

77.87 32.89 78.58 23.41

49.19 48.89

48.74 50.81 All

Table 8: One-factor reverse anchors (GTPA@1 (%)) on the same evaluation pools as Table 3. Rlik : likelihoodonly, using P (e | d); Rprior : prior-only, using P (d | a).

D

Model

0.2

0.3

Disagree 0.4

0.5

0.2

0.3

0.4

0.5

Qwen3-30B 74.25 74.25 74.25 73.85 52.78 53.08 53.38 53.08 Mistral-Small-3.1-24B 74.58 74.73 73.71 73.10 53.90 54.29 52.97 51.78 Llama-3.1-8B 58.46 57.40 55.83 53.36 50.90 49.92 49.40 47.97 GLM-4-32B 67.04 66.99 66.89 66.89 44.41 45.31 45.88 47.57 DeepSeek-V4-Flash 79.07 79.02 78.97 78.71 52.30 52.00 52.15 52.00

Log-linear wR sweep

Table 9: Log-linear GTPA@1 (%) as a function of reverse weight wR on the evaluation pools. Bold denotes the best wR within each block. The main tables use the fixed default wR =0.2.

Table 9 reports LogLin GTPA@1 for wR ∈ 0.2, 0.3, 0.4, 0.5 on the same evaluation pools as Tables 1, 2. The wR =0.2 column matches the LogLin results reported in the main tables. Be13

E

Label collision, co-occurrence, and error-indicator ϕ

All Model Qwen3-30B Mistral-S.3.1-24B Llama-3.1-8B GLM-4-32B DeepSeek-V4-Flash

Table 10 reports the matched-unit π statistic from Table 4 on Disagree subset. Across all five backbones, π(F, R) remains the lowest of the three comparisons, consistent with the All results in the main text. Tables 11 and 12 provide complementary views that are not used for the main same-incorrect-label claim: a 2×2 correctness grid and the Matthews ϕ coefficient computed from binary error indicators. The grid uses the same plurality GTPA@1 indicator as Tables 1–2, so Both✓+F ✓R× matches the plurality column of those tables. Here, ϕ measures the association between whether plurality is wrong and whether the named reference is wrong. Unlike π, it does not measure whether the two predictors assign the same incorrect label. π(F, R)

π(F, GenF )

π(F, Fi )

Qwen3-30B Mistral-S.3.1-24B Llama-3.1-8B GLM-4-32B DeepSeek-V4-Flash

0.333 0.332 0.162 0.296 0.352

0.639 0.621 0.497 0.588 0.609

0.717 0.677 0.600 0.663 0.716

Table 10: Label-collision rates on Disagree. Definitions match Table 4. Bold denotes the lowest π in each row.

All

Disagree

Both✓ F✓R× F×R✓ Both× Both✓ F✓R× F×R✓ Both×

Qwen3-30B Mistral-S.3.1-24B Llama-3.1-8B GLM-4-32B DeepSeek-V4-Flash

50.8 50.8 21.6 45.3 62.0

20.8 20.6 32.6 20.0 14.9

10.5 9.6 11.1 13.3 6.8

17.9 19.1 34.8 21.4 16.4

22.6 25.4 18.0 21.4 23.4

22.4 20.9 26.1 19.4 22.5

20.3 17.7 13.8 21.4 15.6

0.31 0.35 0.16 0.30 0.47

0.73 0.79 0.61 0.74 0.71

0.81 0.81 0.71 0.82 0.78

0.13 0.24 0.17 0.16 0.23

0.41 0.61 0.45 0.53 0.46

0.53 0.60 0.56 0.60 0.52

Table 12: Matthews ϕ on binary error indicators. Each ϕ measures the association between whether plurality is wrong and whether the named reference is wrong, rather than whether they assign the same incorrect label. Unlike Table 11, the error indicators here use stringlevel top-1 equality with the gold label, not plurality GTPA@1. ϕ(F, Fi ) is pooled across the five in-pool agents. Bold denotes the lowest ϕ within each All or Disagree block.

Model

Model

Disagree

ϕ(F, R) ϕ(F, GenF ) ϕ(F, Fi ) ϕ(F, R) ϕ(F, GenF ) ϕ(F, Fi )

34.7 36.1 42.0 37.9 38.5

Table 11: Forward/reverse correctness co-occurrence (% of the slice). Columns denote Both✓, plurality correct and R wrong, plurality wrong and R correct (recovery), and Both×. Plurality correctness is the same GTPA@1 indicator as in Tables 1–2.

14

Record · ID 673548 · SHA-256 9233f9d393724d82
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.