Conceptio › Archive › arXiv CS
arXiv CSopen access

A Closed-Form Estimator and Diagnostic Battery for Anchor-Judge Error Correlation, Under a Single-Common-Factor Model

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

A Closed-Form Estimator and Diagnostic Battery for Anchor–Judge Error Correlation, Under a Single-CommonFactor Model Veerendra Kumar Sunkavalli

[email protected]

Independent Researcher

arXiv:2609.08826v1 [stat.ME] 8 Sep 2026

Abstract When an external reference set (an anchor) is used to decompose an LLM-judge panel’s error into a quality signal and a shared common-mode error, standard practice assumes the anchor is uncontaminated: its error uncorrelated with the judges’ shared error. We study when that assumption can be dropped and replaced by an estimate. Under a single-common-factor model, ≥ 2 judges and ≥ 2 anchors point-identify the quality variance, the common-mode variance, and each anchor’s contamination correlation ρk in closed form, with an exact peranchor-pair failure boundary; a designated clean-anchor estimator, by contrast, reports a contaminated companion anchor as fully clean once its trusted anchor is itself contaminated. Because the single-common-factor assumption is itself untestable, the estimator ships gated behind a calibrated diagnostic battery (judge-covariance dispersion; over-identification; a family-block test from judge metadata, with a family-blocked estimator that removes familylevel shared-residual bias exactly), bootstrap confidence intervals with measured coverage, and a weak-identification screen. A proposition maps which violations bias ρk , in which direction, and which evade detection. For ordinal scores we show an identification hierarchy: with all variables ordinal, ρk is not identified at any number of anchors; with ordinal judges and ≥ 3 continuous anchors it is, and we give an estimator for that case. On real data the validation is asymmetric, and we say so plainly: the diagnostics are validated in the rejecting direction (both real panels we test, human raters and a six-provider LLM panel, are correctly rejected by the model-adequacy pre-test, Test A; specificity on a real adequate panel is untested, none being available), while the estimator is validated in simulation and stress-tested semi-synthetically under oracle calibration (a robustness observation, not a clean-model validation; Section 7); no real panel has yet passed that adequacy pre-test, so the estimator has not been validly applied to a real panel; the pre-test exists precisely to say so. All results replay offline from shipped, checksummed artifacts; the pre-registration and its amendments are included. Relative to MTMM method-factor models, identifiability is not new; the closed form, the exact boundary, the violation taxonomy, and the calibrated battery for the LLM-judge setting are the contribution.

1

Introduction

LLM-as-a-judge panels score model outputs at scale, and a recurring worry is that judges make correlated errors from shared pretraining, style preferences, or prompt framing, so averaging does not cancel the error; measured directly, a nine-judge panel can carry only about two independent votes’ worth of information (Kohli, 2026). A common remedy anchors the judges to an external reference set and estimates the shared error against it (Zhao et al., 2026). Existing anchored decompositions treat the anchor as clean: its error is assumed uncorrelated with the judges’ shared error. If the anchor’s builders share the judges’ biases, the decomposition is silently wrong, with no check. This paper asks when that clean-anchor assumption can be replaced by a measurement of the anchor’s contamination ρ. One status statement up front: the diagnostics we ship are validated on real panels; the 1

estimator is not yet, because every real panel we tested fails the model pre-test. Its current status is a correct tool awaiting a qualifying panel, and Section 8 states what qualifying requires. Identifiability of a free-loading method factor from multiple indicators is classical (Section 2); we do not claim it. What we contribute, under an explicitly stated single-common-factor model, is a closed-form estimator specialised to the judge/anchor structure; an exact characterisation of when it fails; a proposition delimiting which model violations bias the estimate; and, because the single-common-factor assumption is itself untestable, a calibrated diagnostic battery that detects the violations that matter. Contributions. 1. Closed-form estimator (Theorem 1). Under the single-common-factor model, (σt2 , σc2 , {ρk }) from a ≥ 2-judge panel and ≥ 2 anchors, with no clean-anchor and no anchor-difference assumption. Identifiability itself is inherited from MTMM; the closed form and the identifying role of the judge panel are what we add. 2. Exact non-identification boundary (Corollary 1). A measure-zero condition βk = σc2 , with a characterised positive-measure weak-identification neighbourhood, the practically relevant failure. 3. Which violations matter (Proposition 1). A second common factor loading uniformly across judges and anchors is observationally equivalent to inflating σt2 and leaves ρk unbiased; only asymmetric loadings bias ρk (Proposition 1); the harmful shared-residual case is treated in contribution 4. 4. A calibrated three-test diagnostic battery. A judge-covariance dispersion test (≥ 3 judges; non-uniform judge loadings), a ≥ 3-anchor over-identification test (non-uniform anchor loadings), and a family-block test (Test C: shared family-level judge residuals, from judge metadata alone, at any anchor count), each with a null-calibrated 5% false-positive threshold; the first two with reported power curves. A family-blocked variant of the estimator removes the shared-residual bias exactly when the panel spans ≥ 2 families. 5. Inference: item-bootstrap confidence intervals with measured coverage at the nominal level, and a weak-identification screen (a studentized denominator, in the spirit of weak-instrument first-stage diagnostics) that flags boundary proximity from the data alone. 6. An ordinal identification hierarchy (Section 7): with all variables ordinal, ρk is not identified at any number of anchors; with ordinal judges and ≥ 3 continuous-scored anchors it is, and we give a working estimator for that case. 7. Positioning against standard estimation: on identical data the closed form and full-information ML (fitted with a standard SEM package) are statistically indistinguishable; the closed form’s value is transparency and the exact boundary diagnosis, not efficiency. 8. An honest failure map: finite-sample behaviour, weak identification near the boundary, and a finite-sample upward bias of the mixed-scale estimator near ρ = 0 at small N .

2

Related work and positioning

The identifiability is not new; we state the difference. Our model (Jj = t+c+ej , Ak = t+λk c+uk ) is a two-factor model with a shared common-mode/method factor and correlated errors across methods; point-identifying factor variances and cross-loadings from ≥ 2 indicators per factor is standard in multitrait– multimethod (MTMM) common-method-factor models with correlated errors (for the model family, its identifiability, and the critique of its variants, see Campbell & Fiske, 1959; Lance et al., 2002; Bollen, 1989; Anderson & Rubin, 1956), and the one-clean-anchor special case closely parallels the reference-method CT-C(M −1) design (Eid, 2000; Eid et al., 2003). Relative to that literature our additions are: (a) a closedform estimator for the unit-loading judge parameterisation; (b) the exact non-identification boundary; (c) Proposition 1 on which violations bias ρk ; and (d) the calibrated diagnostic battery for the judge/anchor setting. Recent LLM-judge aggregation work models quality plus shared confounders and recovers quality 2

without ground truth; it uses any human anchor only as a trusted labelling device and has no parameter for anchor–judge error correlation (Zhao et al., 2026), and its omitted-confounder bias result overlaps our misspecification analysis, which we cite (Zhao et al., 2026) rather than claim. The diagnostic-testing literature on estimating error rates without a gold standard (Hui & Walter, 1980; Dawid & Skene, 1979) and on relaxing conditional independence between tests (Dendukuri & Joseph, 2001), whose random-effects dependence models play the role our shared-residual analysis plays here, estimated there by Bayesian machinery with informative priors where we give the closed-form bias expression and a metadata-based test, established that such correlations can be weakly identified and model-dependent (Albert & Dodd, 2004); our boundary and weak-identification results are the continuous-score, judge-panel analogue, and we credit that literature for the phenomenon. A contaminated anchor is an invalid instrument; our two-anchor system is an over-identified moment system (Conley et al., 2012).

3

Model and identification

Assumption 1 (Single common factor). All variables are mean-zero. There is one latent quality t (Var = σt2 ) and one common-mode factor c (Var = σc2 ), t ⊥ c. Judges and anchors are Jj = t + c + ej (j = 1, . . . , p) and Ak = t + βσk2 c + uk (k = 1, . . . , m), with ej , uk mutually independent and independent of (t, c). This is the c load-bearing and untestable assumption; Sections 5–8 study its violation. Assumption 2 (Conditional judge-residual independence). The judge residuals ej are mutually uncorrelated: all shared judge variation is carried by the single common-mode factor c. This names, as its own assumption, the judge-residual component already contained in Assumption 1’s mutual-independence clause; we state it separately because it is the empirically fragile part and the one whose violation this paper studies: LLM judges built on a shared base model, prompt template, or preference-tuning lineage plausibly share residual correlation beyond c, violating it. A homogeneous such residual (loading uniformly across judges) is misattributed by the estimator. The exact bias follows from Theorem 1 with K shifted by the residual variance: with at least one clean anchor the effect is pure attenuation, |ρ̂k | < |ρk | (a false “clean”); with both anchors contaminated the direction is configuration-dependent: attenuation in the moderate-contamination regime (simulation: a uniform judge residual of strength 0.6 drives ρ̂2 down, 0.701 → 0.565, while inflating σ̂c2 , 0.80 → 1.13), but inflation of ρ̂k itself, even past |ρ| = 1, near the identification boundary. Because it is uniform across judges, the dispersion test (Test A) does not detect it; it is the judge-side analogue of the asymmetric second factor of Proposition 1 and is visible in principle to the over-identification test (Test B) when the anchors are unaffected, but only weakly at realistic scale: measured power at residual strength 0.6 is 0.06 at N =500 (its own false-positive rate), 0.16 at N =2000, and 0.99 only at N =104 (Section 5), so in practice the family-blocked machinery of Section 5 is the operative remedy. We foreground this as the empirically dominant failure mode in LLM evaluation. Section 5 gives a partial remedy: when the residual is family-level and the panel spans ≥ 2 families, a family-blocked estimator removes the bias and a third test detects the residual from judge metadata alone; the panel-wide case remains open. Here βk := Cov(Ak , c) and the contamination is ρk := βk /(σak σc ), σa2k the total error variance of anchor k; ρk = 0 is the clean anchor. The unit loadings on the judges fix the scale of t. Observable second ¯ Ak ) = σ 2 + βk ; Pkℓ = Cov(Ak , Aℓ ) = σ 2 + βk βℓ /σ 2 ; moments: K = Cov(Ji , Jj )i̸=j = σt2 + σc2 ; Mk = Cov(J, t t c 2 2 Var(Ak ) = σt + σak . Theorem 1 (Closed-form identification under Assumptions 1 and 2). Under Assumptions 1 and 2 (the latter contained in the former’s independence clause, and doing the identifying work in K = σt2 + σc2 ), with p ≥ 2 judges and m ≥ 2 anchors, whenever (K + P12 ) − (M1 + M2 ) ̸= 0, σt2 =

KP12 − M1 M2 , (K + P12 ) − (M1 + M2 )

(1)

and σc2 = K − σt2 , βk = Mk − σt2 , σa2k = Var(Ak ) − σt2 , ρk = βk /(σak σc ). No clean-anchor and no anchor-difference assumption is required; with p < 2 the quantity K is undefined and the parameters are not identified. Corollary 1 (Exact boundary and weak-identification neighbourhood). The denominator of (1) equals (σc2 − β1 )(σc2 − β2 )/σc2 ; identification through this anchor pair fails exactly at βk = σc2 (measure zero); with 3

Table 1: Approaching the boundary (σc2 = 0.8, σa1 = 0.9). ρ1

β1

ρ̂1 (sd)

0.90 0.95 0.98 0.99

0.724 0.765 0.789 0.797

0.900 ± 0.004 0.949 ± 0.003 0.979 ± 0.004 0.955 ± 0.061 (variance blow-up)

m > 2 anchors the closed form averages over pairs, and identification of the remaining parameters fails only when every anchor sits at the boundary, since any non-degenerate pair recovers σt2 and hence each ρk . In a positive-measure neighbourhood the estimator is weakly identified and its variance grows (Table 1). Proposition 1 (Which violations bias ρk ). Add a second common factor d ⊥ (t, c) with variance σd2 , loading g on every judge and h on every anchor. Then, for m = 2, there exists a single-common-factor model reproducing the observable moments (K, Mk , P12 ) exactly (for m ≥ 3 an asymmetric second factor is generically not reproducible by any single-factor model, which is exactly why the over-identification test can detect it); when g = h the equivalent parameters satisfy βk′ = βk , σc2′ = σc2 , and σt2′ = σt2 + g 2 σd2 . Hence a uniformly loaded second factor leaves ρk unbiased (it is absorbed into the quality variance), and ρk is biased only to the extent the second factor loads asymmetrically across judges versus anchors. The g = h case is, in effect, definitional: a factor loading identically on every judge and anchor is indistinguishable from the unit-loading quality signal t, so “harmless” here means “reparameterises t”, not “detected and corrected”. A genuinely shared bias of exactly this symmetric shape would be silently credited as quality, and σ̂t2 is inflated by g 2 σd2 with no warning: users of the quality variance itself (e.g. for signal-to-noise assessments) inherit that bias even though ρk does not. Proof (SymPy-verified) in Appendix A; the bias-vs-asymmetry relationship is confirmed in simulation (Sec/ (0, K) are out-of-range (Heywood-type) and reported as such, not tion 5). Estimates with σc2 ≤ 0 or σt2 ∈ clipped. The closed form is not range-restricted: |ρ̂k | > 1 is possible and is itself evidence of misspecification or weak identification (the raw real panel of Section 7 produces 1.026); any use of the estimator should be gated behind the adequacy pre-test (Test A) and the battery.

4

Simulation study

Deterministic, seeds {11, . . . , 88} (one run per seed); each entry is mean±standard deviation across the eight seeds, emitted to p4_core_results.json, p4_extended_results.json, p4_v2_results.json, p4_opchar_results.json. Quantities are simulated unless marked [analytic]; these are estimatorcorrectness and finite-sample characterisations, not evidence about real judges. Recovery under contamination (N =40,000). With neither anchor clean and anchor reliabilities unknown, recovery is within 0.02, including statistically identical anchors: true (ρ1 , ρ2 ) = (0.3, 0.7) → (0.299 ± 0.017, 0.701 ± 0.007); (0.1, 0.9) → (0.096 ± 0.020, 0.900 ± 0.003); (0.5, 0.5) → (0.500 ± 0.012, 0.502 ± 0.009); (0.2, 0.8) → (0.198 ± 0.019, 0.801 ± 0.005). This confirms the closed form inverts the moments; it is not evidence the model holds for real judges. Finite sample and out-of-range. Recovery of ρ2 =0.7: 0.718 ± 0.067 (N =100), 0.680 ± 0.055 (500), 0.690 ± 0.024 (2000), 0.701 ± 0.007 (40000); no out-of-range estimates arose here. At small N the estimator is high-variance and should be used with interval estimates. Weak identification near the boundary (Table 1). As β1 approaches σc2 the denominator of Theorem 1 shrinks: point estimates remain approximately unbiased through ρ1 = 0.98, and at ρ1 = 0.99 the standard deviation grows fifteen-fold (0.955 ± 0.061), the practical signature of the measure-zero boundary. The weak-identification screen of Section 6 is designed to flag exactly this regime. 4

Table 2: Bias in ρ̂2 (true 0.7) vs. loading asymmetry |g − h| (judge g=0.6). h |g − h| ρ̂2

0.6 0.00 0.701

0.5 0.10 0.761

0.4 0.20 0.780

0.3 0.30 0.769

0.0 0.60 0.565

Table 3: Battery operating characteristics. FP = false-positive rate under the correct model; power at second-factor strength 0.4 / 0.6.

5

N

FP (A, B)

Test A power (0.4/0.6)

Test B power (0.4/0.6)

500 2000 10000

0.05 / 0.05 0.05 / 0.05 0.05 / 0.05

0.23 / 0.91 0.93 / 1.00 1.00 / 1.00

0.50 / 0.99 0.99 / 1.00 1.00 / 1.00

Misspecification diagnostic battery

By Proposition 1, only asymmetric second factors bias ρk . Table 2 confirms it: holding judge loading g = 0.6 and varying anchor loading h, the bias in ρ̂2 is 0.001 at h = g; away from h = g it is non-monotone in |g − h| and changes sign at large asymmetry (Table 2: ρ̂2 rises 0.761 → 0.780, dips to 0.769, then crosses to 0.565 against true 0.7), so asymmetry determines that the estimate is biased, not the direction or size. We therefore target asymmetry with two complementary, null-calibrated tests. Test A: judge-covariance dispersion (needs p ≥ 3). Under Assumption 1 all off-diagonal judge covariances are equal; a second factor with non-uniform judge loadings makes them unequal. The statistic is the coefficient of variation of the off-diagonal judge covariances. It requires ≥ 3 judges (with p=2 there is a single off-diagonal), a caveat we flag since identification itself needs only p ≥ 2. Test B: ≥ 3-anchor over-identification. With m ≥ 3, σt2 is estimated from each anchor pair; under Assumption 1 all agree. Non-uniform anchor loadings break the agreement; the statistic is the spread of σ̂t2 across pairs, and it fires precisely where Test A is blind. Operating characteristics (Table 3). Thresholds are set at the 95th percentile of each statistic’s null distribution (correctly specified model, 1000 seeds), giving a 5% false-positive rate by construction at every N tested; this is a correct-model, Gaussian property (heavy tails inflate it to 0.083, Section 6, and specificity on real adequate panels is untested). Power (probability of exceeding threshold) rises with both signal strength and N , from a floor that is genuinely uninformative: at spread 0.2 Test A’s power is 0.04 at N =500 and 0.11 at N =2000, at or below its own false-positive rate, so mild violations at realistic N are invisible to it. Test B’s power is measured against its advertised target: a second factor shared by ≥ 2 anchors with non-uniform loadings (h = (s, s/2, 0) at strength s); a factor loading on a single anchor is not a violation at all (it is absorbed into that anchor’s idiosyncratic variance, and measured power equals the false-positive rate there). Two blind spots, stated plainly. First: by Proposition 1 a uniformly loaded second factor (g = h) is invisible to both tests and to any second-moment statistic, but it does not bias ρk ; there is nothing to correct when it occurs. Second, and harmful: a shared judge residual (Assumption 2) biases ρ̂k (attenuation toward a false “clean” in the typical regime; inflation near the boundary), Test A cannot see it, and Test B responds only weakly below N ≈ 104 (judge-side power 0.06/0.16/0.99 at N =500/2000/104 , strength 0.6, versus 0.99 at N =2000 for its anchor-side target; Table 3). The family-blocked estimator below removes this bias exactly when the residual is family-level and the panel spans ≥ 2 families; what survives is the panel-wide residual (a shared prompt template affecting every judge), which no within-panel statistic can separate from the common mode. Practitioners with single-family panels should treat small ρ̂k as unverified, not as clearance; and a panel generated under a shared prompting protocol or template induces exactly 5

this undetectable panel-wide case, so on such panels the estimator should not be used at all, whatever the diagnostics say. Family-blocked estimator, and Test C. Judge lineage is observable metadata. Model the familylevel residual explicitly, Jj = t + c + fb(j) + ej with fb independent across families: then cross-family judge covariances equal σt2 +σc2 exactly (the residual cancels), so computing K from cross-family pairs only restores the closed form unchanged [analytic; SymPy-verified], while the within-minus-cross gap estimates the familyresidual variance directly, Kwithin − Kcross = σf2 , giving a third diagnostic (Test C) that needs ≥ 2 families with at least one family containing ≥ 2 judges (so p ≥ 3); no extra anchors, no over-identification machinery. Singleton families’ residuals enter only through the cross terms. The cross-family covariance contains no family terms at all, so the blocked estimator’s exactness does not depend on family sizes; with familyspecific variances σf2b the within-minus-cross gap estimates their pair-weighted mean. In simulation (six judges in three families, N =4000): the naive estimator drifts toward “clean” as the family residual grows (ρ̂2 0.697 → 0.650 at residual s.d. 0.7) while the family-blocked estimator stays at 0.696–0.697 throughout; over 100 replicates per cell, Test C flags 4/100 null datasets (5% nominal) and 100/100 at every tested strength, and recovers σf2 to the second decimal (0.088/0.248/0.488 against the design values 0.32 /0.52 /0.72 ). Score types. Under heteroscedastic noise recovery is essentially unaffected (ρ̂2 = 0.703 ± 0.008). Ordinal (Likert) scores are a deeper matter than a bias: they change what is identifiable at all. Section 7 treats this in full.

6

Inference, and a comparison with maximum likelihood

Confidence intervals with measured coverage. We use a percentile bootstrap over items (300 resamples). Measured coverage of the nominal 95% interval for ρ2 , over 150 Monte-Carlo replicates: 0.953 at N =2000 (mean width 0.118) and 0.953 at N =500 (width 0.244; both are 143/150, the granularity of the replicate count). Near the boundary the intervals stay conservative (coverage 0.96 at ρ1 =0.90; 0.992 at ρ1 =0.97, where only 118 of 150 replicates return an estimate at all; the other 32 produce no admissible value and are excluded, so the 0.992 is coverage conditional on estimability) but widen to 0.993 and 1.548: they correctly report that the data carry little information there. A weak-identification screen. Practitioners cannot know ex ante whether their system sits near the boundary of Corollary 1. We therefore add to the battery a studentized-denominator screen, in the spirit of d d weak-instrument first-stage diagnostics: with T = |den|/SD boot (den), flag weak identification when T < 4. In simulation the flag fires on 0% of datasets at mid parameters (both N =500 and N =2000), on 46% at ρ1 =0.90, and on 100% at ρ1 =0.97. An estimate that arrives flagged should be reported as an interval only. Against full-information ML. These models are ordinarily fit by ML in SEM software; the comparison is owed. On identical simulated data, the closed form and full-information ML (Igolkina & Meshcheryakov, 2020) are statistically indistinguishable: at mid parameters both give ρ̂2 = 0.704 ± 0.016, and near the boundary 0.517 ± 0.265 (closed form) versus 0.520 ± 0.237 (ML). The closed form is faster (roughly 0.2ms versus 4ms per fit on our hardware; per-fit timings are emitted with the results) but both are trivial at this scale. The comparison extends beyond the just-identified design: with m=3 anchors (over-identified) the two remain nearly identical (closed 0.699 vs ML 0.702 on ρ2 ), and under a misspecified model (shared judge residual of strength 0.5) both are biased almost identically (closed 0.585 vs ML 0.597 against true 0.7): ML confers no robustness to the violations that matter here. Against the closest deployed alternative, the comparison is not a tie: a CT-C(M −1)-style estimator that designates one anchor as clean recovers ρ2 correctly only while that trust is justified, and degrades linearly as the designated anchor’s true contamination grows (ρ̂2 = 0.502, 0.411, 0.270, −0.002 at ρ1 = 0, 0.2, 0.4, 0.6; true 0.5), reporting a heavily contaminated companion anchor as fully clean at the end of that range, while our estimator stays at 0.501 throughout: the no-clean-anchor property is the delta, made quantitative. (CARE-class aggregation (Zhao et al., 2026) recovers quality under shared confounding but has no anchor-contamination parameter: it treats any reference set as a trusted labelling device, so there is no CARE estimate of ρk to compare against; what a reader loses in exchange for our estimand is CARE’s freedom from anchors altogether. Bayesian latent-class 6

dependence models target the binary/ordinal regime and carry prior sensitivity that our closed form avoids.) We conclude the closed form sacrifices nothing statistically in this design family; its value is transparency: the exact boundary diagnosis of Corollary 1 and the screen above fall out of the formula, not out of an optimizer trace. Non-Gaussian robustness. The estimator is moment-based and needs no Gaussianity for consistency; the calibrations might. Re-running the machinery with skewed (standardized lognormal) and heavy-tailed (t4 ) latents and residuals: CI coverage 0.975 and 0.908 (nominal 0.95), the weak-identification screen stays quiet (1–3% false flags), and the Gaussian-calibrated tests show only mild false-positive inflation (Test A 0.083/0.058, Test C 0.042/0.083 against nominal 0.05). For real panels we recommend recalibrating the null on matched marginals, as done for the real-panel test in Section 7. What your configuration buys you. At the realistic evaluation scale of N ≈ 500: the estimator is usable with honest intervals (95% CI coverage 0.953, mean width 0.244 at mid parameters), Test A has power 0.23/0.91 at loading-spread 0.4/0.6, Test B has power 0.50 at loading-spread 0.4 (and 0.99 at 0.6) against shared non-uniform anchor factors, and the ordinal mixed-scale estimator should be treated as ordering-only (Section 7). As a decision aid:

7

Configuration

ρ̂k + CIs

Diagnostics

Unguarded failure modes

2 judges, 2 anchors ≥ 3 judges, 2 anchors ≥ 3 judges (≥ 2 families), 2 anchors ≥ 3 judges (≥ 2 fam.), ≥ 3 anchors, N ≳ 104

yes yes yes

weak-ID screen only + Test A + Test C, blocked est.

all misspecification anchor-side; shared residual anchor-side; panel-wide residual

yes

full battery

panel-wide residual; uniform factor (harmless for ρk )

Ordinal scores: an identification hierarchy

LLM judges commonly emit Likert scores. Ordinal observation is not a small perturbation of the continuous theory; it changes what is identifiable, because thresholds absorb every latent location and scale, and ρk = βk /(σak σc ) needs the scale of c relative to the anchor errors. Negative result. With all variables ordinal, the identifiable information is the latent (polychoric) correlation structure. A rank analysis of the moment map (numerically: Jacobian plus null-space test at interior parameter points, the standard local-identification criterion) shows ρk is not identified at any m. The rank computation retains the model’s unit-loading constraint throughout (the null space arises from the ordinal observation map, which forgets scale, not from freeing loadings). The argument is analytic, not enumerative: the equivalence transformation acts on each judge–anchor and anchor–anchor pair separately (a per-pair identity in the latent scales), so it maps solutions to solutions for every m simultaneously; adding ordinal anchors adds equations and unknown scales at exactly the rate that preserves the two-dimensional null family. We verified the rank computation numerically at m ∈ {2, 3, 4, 6} and exhibit explicit equivalent parameterisations with ρ1 = 0.30 and 0.80 at m = 2. One structured exception: under exchangeable anchors (equal reliabilities imposed), identification is restored at m ≥ 3 by the same rank test; we do not build on it because exchangeability is itself untestable in this setting. Of the two null directions, one is the common scale, along which ρk is invariant; the second moves ρk . A constructive version of the same fact: distinct parameter vectors reproducing the observed correlations exactly can carry ρ1 = 0.30 or ρ1 = 0.80. More ordinal anchors do not help, because each new anchor brings its own unknown scale. A known-clean anchor does not restore identification of the other anchor’s ρ either. Positive result, and an estimator. With ordinal judges but m ≥ 3 continuous-scored anchors (a realistic configuration: Likert LLM judges, finely-scored human reference sets), ρ is locally identified, with one over-identifying restriction. The estimator: polychoric correlations among judges, polyserial judge–anchor correlations scaled by the observed anchor standard deviations, and the raw anchor covariance block; the anchor tetrads (in the sense of confirmatory tetrad analysis, Bollen & Ting, 1993) give wk (σt2 ) = |βk |/σc in 7

closed form (signs recovered from Mk∗ − σt2 ), and a one-dimensional search over σt2 minimizing the full overidentified misfit completes the solve. With six 5-level judges and three continuous anchors at N =4000 (true ρ = (0.3, 0.7, 0.5)), it recovers (0.282 ± 0.129, 0.670 ± 0.078, 0.484 ± 0.103) with σ̂t2 = 0.996 (true 1.0); results are essentially unchanged at 3-level and 7-level discretization. Treating the ordinal codes as continuous and running the covariance estimator, by contrast, is erratic across discretizations. The information cost of ordinal judges is real and visible in the standard deviations: roughly an order of magnitude more variance than the continuous case at the same N . Validation on real rater textures. Gaussian simulations cannot certify behaviour under real Likert data: real raters have skewed marginals, ties, and idiosyncratic thresholds. We therefore validate on the HANNA story-evaluation benchmark (Chhun et al., 2022): 431 stories jointly rated (coherence, 1–5) by the same three human raters, whose marginals are extreme (one rater places 62% of mass on category 5; another 41% on category 1). We keep the latent scaffolding synthetic, so injected contamination is exact by construction, and make the noise real: each synthetic judge draws its idiosyncratic errors by bootstrap from one real rater’s item-centred residuals and discretizes with that rater’s own empirical thresholds. Injected ρ = (0.0, 0.4, 0.7); recovered, at the real panel size N =431: (0.372 ± 0.200, 0.661 ± 0.189, 0.712 ± 0.105); at N =2000: (0.242 ± 0.175, 0.562 ± 0.157, 0.714 ± 0.054), monotone in the means (per-replicate strict ordering holds in 8/24 runs at N =431 and 14/24 at N =2000; the means, not the individual replicates, carry the ordering claim at these sizes). Two honest readings. High contamination is recovered well under fully real textures. Near ρ = 0 the mixed-scale estimator carries a finite-sample upward bias at small N with p=3 judges; an ablation shows this is a property of the estimator, not of the real textures (skewed and balanced marginals give the same numbers). Detecting a clean anchor from three Likert judges needs either N ≳ 2000 or more judges. A parametric-bootstrap bias correction (refit on data simulated from the fitted model, subtract the measured bias) cuts the near-zero bias by two thirds at N =431 (+0.248 → +0.092; measured in the correction harness with synthetic thresholds, hence the different baseline from the real-texture +0.372 above) at the cost of variance (0.128 → 0.210 at ρ = 0, more at high ρ); we recommend it when the question is specifically whether an anchor is clean, and the raw estimator otherwise. Test A on the real panel. Applying the dispersion test directly to the real HANNA panel (three raters, N =431, threshold from a matched null: single common factor, p=3, the raters’ own marginals) gives a statistic of 0.589 against a 95th-percentile null threshold of 0.247 (p < 0.0005): the test fires. Real raters violate the equal-loading structure, and the battery detects it from the data alone. This is the intended use: the diagnostics are a pre-test that tells a practitioner when this paper’s model, and therefore its estimator, should not be trusted on their panel; on the one real panel we examined, they correctly said no. Real LLM judges with injected contamination. This study is a robustness observation under oracle calibration, not a clean-model validation; the reader should carry that label through the paragraph. We ran a synthetic-anchor variant of the validation protocol of Section 8 (the variant uses a key-based quality construct in place of held-out human consensus; the full protocol is stated in Section 8): six instruction-tuned judges from six providers (Amazon Nova Lite, Llama-3-70B, Mixtral-8x7B, GPT-OSS-120B, Qwen3-Next80B, DeepSeek-V3.2, temperature 0, K=2) scored 499 complete-case items from a fresh 500-item keyed pool (GSM8K and SciQ; 270 correct, 230 wrong by construction): 6,000 calls, 5,996 parsed verdicts in an append-only, checksummed cache (the four failures span three items; one item lost both repetitions of one judge cell and was dropped). Anchors were constructed with injected ρ = (0.0, 0.4, 0.7) against a key-based quality construct, with the common-deviation signal measured on one half of the panel and estimation run on the disjoint other half. The run supports the paper in two ways: it validates the pre-test on real judges, and it provides a robustness observation for the calibrated estimator. First, the raw panel fails the pre-test: Test A statistic 0.175 against a matched-null 95th-percentile threshold of 0.064 for this configuration (p=6, N =499; the judges’ quality loadings, the per-judge regression coefficients on the known quality construct, which Assumption 1 fixes at unity, span 0.324–0.870). Estimates computed in defiance of the pre-test are badly inflated, (0.887, 0.960, 1.026) against the injected (0.0, 0.4, 0.7) (the naive moment estimator is not range-restricted), so the diagnostics did their job on real judges. Second, after oracle per-judge loading calibration (each judge is divided by its loading on the known quality construct, an operation possible only in a validation harness, since it conditions on the very quantity the method exists to estimate; the pre-test’s 8

point), the estimator recovers (+0.024, +0.268, +0.558) against the injected (0.0, 0.4, 0.7): point-estimate ordering exact (S1 was registered on point estimates; the intervals overlap, so the ordering is directional evidence, not a significant separation), the clean anchor’s interval covers zero (S3), as does the 0.4-anchor’s ([−0.56, 0.71]), so at this N a moderately contaminated anchor is not statistically distinguishable from a clean one, and the high-contamination anchor lands inside the amended cross-panel attenuation band [0.4, 0.85] (S2′ ; the original magnitude criterion failed and was replaced, with the amendment logged before the final analysis; both versions and the first analysis’s numbers are in the shipped pre-registration). Three honest qualifications. The calibration fixes quality loadings only: the calibrated panel still fails Test A (0.343 against the 0.064 threshold: dividing by small quality loadings amplifies the relative heterogeneity of the common-mode loadings, so the statistic rises even as quality loadings equalise), so this is recovery under detectably imperfect conditions, a robustness observation, not a clean-model validation. The injection is defined against panel A’s common deviation while the estimator measures against panel B’s; the two correlate at 0.534, predicting attenuated targets (0, 0.214, 0.374); the recovered values sit at or between those and the nominal injections, within the intervals. The predicted high-anchor target (0.374) lies below the S2′ band, so the S2′ pass depends on the observed attenuation being milder than the cross-panel prediction, and the interval extends below the band edge: we treat S2′ as directionally satisfied, not robustly. And the anchors are synthetic; only the judges are real. The pre-registration, its amendments, and the cache ship with the artifact.

8

What the study does and does not establish

Established: the closed form inverts the model’s moments (estimator correctness); the exact identifiability boundary and its weak-identification neighbourhood; which violations bias ρk (Proposition 1); that identification is supplied by the ≥ 2-judge panel; and the calibrated detection profile of the battery. Established also: at least one real LLM panel detectably violates the unit-loading structure, and the pre-test catches it (Section 7). Not established: that real LLM-judge and human-anchor errors follow Assumption 1 (passing the battery is necessary, not sufficient). When Test A fires on a real panel, the unit-loading closed form should not be used; with ≥ 3 judges, free-loading CFA identification is classical (Bollen, 1989) and an ML fit is the natural fallback, at the price of the transparency this paper trades on. Evaluated on the six-judge panel of Section 7, however, the free-loading fit is weakly identified (non-positive-definite information matrix) and does not recover the injection (a single real data point; a simulation comparison of free-loading CFA under this paper’s violation taxonomy is future work and the natural next step for the qualifying-panel question): when the pre-test fires, no within-panel method evaluated here is validated, and the honest options are to fix the panel or to treat any ρ̂k as unverified. An empirical validation would require a setting with independently known contamination. A concrete protocol: run ≥ 3 LLM judges plus held-out human raters on a keyed item set; construct anchors Ak = H + gk ĉ + εk with H the held-out human consensus and ĉ the measured judge common deviation, giving known injected ρk ; check recovery and battery behaviour. Section 7 reports a synthetic-anchor variant of this protocol on six real judges (key-based construct in place of held-out human consensus, with the deltas stated there); the full human-consensus version remains to be run, and we provide the protocol for it. Ordinal scores restrict identification (Section 7): with all variables ordinal, ρ is not identifiable at all, and the mixed-scale estimator that handles ordinal judges needs ≥ 3 continuous anchors and carries a finite-sample upward bias near ρ = 0 at small N . Small N yields high variance throughout; the weak-identification screen of Section 6 tells the practitioner when intervals are the only honest report. We therefore present a measurement tool with a stated domain of validity, explicitly characterised blind spots (one harmless, one remediable by family blocking, one, the panel-wide residual, open), and calibrated diagnostics. It is not a validated description of real judges. Proposition 1 and Assumption 2 express the same fact: the estimator cannot separate any signal that shares the quality factor’s unit-loading pattern (symmetric across judges and anchors, or homogeneous across judges) from quality itself; such a signal is silently folded into σt2 or σc2 . A genuine shared bias of that exact shape would therefore be miscredited; the battery detects only its asymmetric departures. 9

9

Conclusion

Under a single-common-factor model, anchor contamination can be estimated in closed form from a judge panel and two anchors, with an exact identifiability boundary, a proposition delimiting which model violations bias the estimate, and a calibrated three-test diagnostic battery for the violations that do. Whether the assumption holds for a given judge/anchor system is an empirical question we leave open; our contribution is the estimator, its boundary, the harmless-symmetric-factor result, and the diagnostics with which a practitioner can probe the assumption.

Reproducibility and data statement All code is deterministic and seeded; one command regenerates every number within numeric tolerance from a pinned environment, replaying two frozen inputs: the public HANNA human-evaluation benchmark (story ratings; public license) and our own six-judge LLM verdict cache (6,000 calls, 5,996 parsed verdicts, SHA256-checksummed; collected once on personal cloud infrastructure for the injected-contamination validation and shipped with the artifact; reproduction replays the cache and requires no model access). Everything else is simulation. The pre-registration and its amendments are included. An anonymized code archive accompanies submission.

References Paul S. Albert and Lori E. Dodd. A cautionary note on the robustness of latent class models for estimating diagnostic error without a gold standard. Biometrics, 60(2):427–435, 2004. doi:10.1111/j.0006-341X. 2004.00187.x. Theodore W. Anderson and Herman Rubin. Statistical inference in factor analysis. In Proceedings of the Third Berkeley Symposium on Mathematical Statistics and Probability, volume 5, pp. 111–150. University of California Press, 1956. Kenneth A. Bollen. Structural Equations with Latent Variables. Wiley, New York, 1989. Kenneth A. Bollen and Kwok-fai Ting. Confirmatory tetrad analysis. Sociological Methodology, 23:147–175, 1993. doi:10.2307/271009. Donald T. Campbell and Donald W. Fiske. Convergent and discriminant validation by the multitraitmultimethod matrix. Psychological Bulletin, 56(2):81–105, 1959. Cyril Chhun, Pierre Colombo, Chloé Clavel, and Fabian M. Suchanek. Of human criteria and automatic metrics: A benchmark of the evaluation of story generation. In Proceedings of the 29th International Conference on Computational Linguistics (COLING), 2022. Timothy G. Conley, Christian B. Hansen, and Peter E. Rossi. Plausibly exogenous. Review of Economics and Statistics, 94(1):260–272, 2012. doi:10.1162/REST_a_00139. A. P. Dawid and A. M. Skene. Maximum likelihood estimation of observer error-rates using the EM algorithm. Journal of the Royal Statistical Society: Series C (Applied Statistics), 28(1):20–28, 1979. doi:10.2307/ 2346806. Nandini Dendukuri and Lawrence Joseph. Bayesian approaches to modeling the conditional dependence between multiple diagnostic tests. Biometrics, 57(1):158–167, 2001. doi:10.1111/j.0006-341X.2001. 00158.x. Michael Eid. A multitrait-multimethod model with minimal assumptions. Psychometrika, 65(2):241–261, 2000. doi:10.1007/BF02294377. Michael Eid, Tanja Lischetzke, Fridtjof W. Nussbeck, and Lisa I. Trierweiler. Separating trait effects from trait-specific method effects in multitrait-multimethod models: A multiple-indicator CT-C(M-1) model. Psychological Methods, 8(1):38–60, 2003. doi:10.1037/1082-989X.8.1.38. 10

S. L. Hui and S. D. Walter. Estimating the error rates of diagnostic tests. Biometrics, 36(1):167–171, 1980. doi:10.2307/2530508. Anna A. Igolkina and Georgy Meshcheryakov. semopy: A python package for structural equation modeling. Structural Equation Modeling: A Multidisciplinary Journal, 27(6):952–963, 2020. doi:10.1080/10705511. 2019.1704289. Guneet Kohli. Nine judges, two effective votes: Correlated errors undermine LLM evaluation panels. arXiv preprint arXiv:2605.29800, 2026. Charles E. Lance, Carrie L. Noble, and Steven E. Scullen. A critique of the correlated trait–correlated method and correlated uniqueness models for multitrait-multimethod data. Psychological Methods, 7(2): 228–244, 2002. doi:10.1037/1082-989X.7.2.228. Jitian Zhao, Changho Shin, Tzu-Heng Huang, Satya Sai Srinath Namburi GNVV, and Frederic Sala. CARE: Confounder-aware aggregation for reliable LLM evaluation. arXiv preprint arXiv:2603.00039, 2026.

A

Proofs

Theorem 1. Expanding (P12 − x)(K − x) = (M1 − x)(M2 − x), the x2 terms cancel, giving x[(M1 + M2 ) − (P12 + K)] = M1 M2 − P12 K, i.e. (1). Substituting the structural moments, the denominator is (σc2 − β1 )(σc2 − β2 )/σc2 [analytic; SymPy-verified]. Proposition 1. With a uniform second factor (g on judges, h on anchors), the observable moments become K = σt2 + σc2 + g 2 σd2 , Mk = σt2 + βk + ghσd2 , P12 = σt2 + β1 β2 /σc2 + h2 σd2 . Solving for an equivalent singlecommon-factor parameterisation and setting g = h yields βk′ = βk , σc2′ = σc2 , σt2′ = σt2 + g 2 σd2 [analytic; SymPy-verified], so ρk is unchanged.

11

Record · ID 668048 · SHA-256 da0f9c6dce1ed538
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.