Deposon: An Auditable, Conservation-Guaranteed, Game-Theoretically Tested Scattering Layer over LLM Reasoning Paths Qihao Yuan School of Chemistry and Life Resources, Renmin University of China github.com/zeroandcat/Deposon
arXiv:2609.09001v1 [cs.AI] 8 Sep 2026
∗
2026-09-08 arXiv: cs.AI Repository: https://github.com/zeroandcat/Deposon License: CC BY 4.0
Date: 2026-09-08
Abstract Multi-step LLM reasoning lacks a machine-recheckable ledger: discarded reasoning paths leave no auditable record. We propose the Deposon scattering layer, which binds each node of an LLM-generated concept-decomposition graph to a two-parameter Deposon state; paths undergo three-channel scattering—transmission, reflection, irreversible dissipation—obeying T+R+A=1 for arbitrary parameters, with a maximum per-path energy-audit deviation of 2.2×10−16 (machine epsilon). We report all three evidence tiers honestly. On synthetic trap benchmarks the path-filtering gain is closed (pre-registered): unified reaches 100% versus a decoy-capture baseline at 7%/10%. On real benchmarks the layer is indistinguishable from a trivial six-keyword rule filter (GSM8K 0.87 ≥ 0.85, McNemar p=0.5; StrategyQA 0.899 = 0.899); no difference is detected here, so we sharpen the claim to "the differential value lies solely in machine verifiability." Fusion yields a second negative result: convex combinations with a semantic prior never improve (physics 0.484→0.452), and the apparent λ=2 gain is an anti-field artifact; any fusion gain must be nonlinear. Modeling the reverse dynamics as a potential game on the graph, we evidence an auditable scalar’s monotonicity and neargradientness and quantify the empirical coordination ratio (ECR). The three formalized dynamical-equivalence propositions (P1a/P1b/T-P1c) are falsified under the pre-registered kill protocol, and the potential-game claim is downgraded to approximate (cyclic-graph median residual 0.669): only consistency-level evidence survives at the dynamical level. Code: github.com/zeroandcat/Deposon. Keywords: auditability; conservation identity; physics-constrained layer; potential game; pre-registration; negative results
1
Introduction
Chain-of-thought [1] and its descendants—search-based decomposition [2] and self-consistency voting [3]—organize LLM reasoning into explicit intermediate steps and expand them into a space of candidate paths; yet chain-of-thought text need not be faithful to the internal computation [4], and faithfulness requires dedicated evaluation [5]. Three structural difficulties afflict the reliability of multi-step reasoning: superficially related decoys in the problem statement ∗ Project version: Deposon v2. The Deposon (author-provided) is a single quasiparticle whose v1 (blocking) and v2 (tunneling) are two limits distinguished only by whether energy dissipates into the infinite-dimensional orthogonal aether, with T+R+A=1 conservation per path.
1
systematically attract search branches; errors in intermediate steps have no intrinsic absorption mechanism and compound along the chain; and the basis for path selection is uninterpretable, so that after the fact one cannot answer "why should this path have been discarded." Whatever the organizational scheme, one structural question remains open: when a path is eliminated, the system cannot produce a recheckable account. Pruning, voting, and heuristic scoring all make elimination decisions, yet the elimination itself leaves no conserved record—how much budget the discarded branch consumed, why it deserved discarding, and whether a third party could recompute the decision step by step have no answer within the framework. Algorithmic auditing research supplies accountability frameworks at the organizational level [6]; verifiable computation at the cryptographic layer [7] and the tightening of differential-privacy auditing [8] demonstrate that verifiability itself can be a first-class design goal. But at the level of LLM reasoning paths, an empty layer sits between post-hoc provenance and cryptographic proof— run-time, per-instance invariants—and we are not aware of prior work occupying this layer. This paper occupies that empty layer with a physical construction. The vocabulary can be borrowed directly from scattering theory, and we fix it once here: the three-state verdict (transmit / block / tunnel) corresponds to the three scattering channels; dissipation corresponds to a commitment device—once energy condenses into the infinite-dimensional aether it cannot flow back, and the budget of an eliminated path cannot be resurrected; conservation corresponds to the accounting identity—every unit of energy has exactly one destination. The design goals of this project’s precursor design phase (preceding the frozen pipeline build) supplied the embryo of this construction: a two-parameter Deposon state (path coupling g_couple, aether coupling g_aether), a scattering form inspired by Feshbach resonance, and the axiomatic stipulation "dissipation = irreversible condensation." The route from that definition to the conservation guarantees and auditable representation of §2, and then to the game-theoretic formulation of the reverse dynamics in §3, is two successive tightenings of the same entity; the details are left to those sections. Writing physical properties into algorithms is not new, but most precedents inject soft constraints: residuals are suppressed by gradients yet can still be violated at inference time. Our construction sits at the opposite pole—conservation holds by construction, independent of training and of tuning, and an auditor needs only double-precision arithmetic to recheck it. This determines the shape of our evidence: the strongest claims (conservation, auditability) rest on no benchmark accuracy whatsoever, while the weakest claim (the benefit of the dissipation channel) is explicitly demoted by our own experiments to "motivation only." One declaration belongs up front: the output of this research line includes, alongside the mechanism itself, an honest boundary drawn by a set of negative results; the two carry equal weight. Pre-registered controlled experiments show that the scattering layer’s apparent accuracy advantage is partly an artifact of benchmark construction (once decoy edges are flattened to equal weight, the attribution of the advantage changes), and that on two real benchmarks the layer is indistinguishable from a trivial rule filter. Such "effect size goes to zero" corrections are nothing to be ashamed of—underclaiming is as harmful as overclaiming [9], and disciplined reporting norms [10, 11] together with systematic records of underspecification effects [12] are exactly the genre this paper follows. We therefore locate the value proposition in auditable representation and conservation guarantees, and label every claim with one of the three strength tiers. Our contributions: • C1 Auditable representation and the accounting identity (closed (preregistered)): the constructive definition of three-channel scattering makes T+R+A=1 hold for arbitrary parameters; the per-path energy-audit residual is 2.2×10−16 ; elimination decisions are attributable node by node (§2). • C2 Three-tier placement of evidence strength (closed (pre-registered) / consistency / motivation, side by side): the path-filtering gain on synthetic benchmarks, 2
the tie with a rule filter on real benchmarks (E9.5), and the zero-benefit statement for the dissipation channel are reported in parallel, without blending tiers (§2.3–§2.4). • C3 An exclusionary conclusion on fusion dilution (closed (pre-registered) on the measured λ settings): the convex-combination hybrid never exceeds either single arm on any measured setting; the apparent λ=2 gain is an anti-field artifact; any fusion gain can only come from a nonlinear mechanism (§2.5). • C4 Game-theoretic evidence (closed (pre-registered)): the reverse dynamics admits a scalar (the potential) that can be audited against a benchmark; the empirical coordination ratio (ECR) quantifies "how far from optimum"; the temperature frontier demarcates the audit boundary (§3). Roadmap: §2 gives the representation construction, the conservation identity, per-path auditing, and all three evidence-strength tiers of the static line; §3 gives the game-theoretic evidence chain; §4 collects the honest boundaries; §5 concludes; Appendix A gives the number-traceability table.
2
Auditable Representation and Conservation Guarantees
This section presents in a self-contained manner the construction and conservation properties of the scattering layer and reports evidence strength as mechanically evaluated against preregistered verdicts. §2.1 constructs the representation; §2.2 states the accounting identity and the per-path audit; §2.3 reports the path-filtering gain together with its attribution boundary; §2.4 reports the tie against a rule filter and the resulting relocation of the value proposition; §2.5 reports fusion dilution.
2.1
Representation: from concept graphs to Deposon states
Given a natural-language question, an LLM backend first decomposes it into a directed concept graph G=(V,E): nodes include numbers, operations, traps (superficially related but semantically irrelevant decoy candidates), answer candidates, and general concepts; BFS generates the set of candidate paths from the start node to an answer node. The scattering layer binds each node v to a Deposon state characterized by three parameters: a path-coupling strength g_couple≥0, an aether-coupling strength g_aether≥0, and a resonance energy E0 . Parameters are not assigned ad hoc after the fact but bound to node-type semantics: trap nodes take (5.0, 0.0) (strong scattering centers), operation nodes take (0.3, 0.2), all other nodes take (0.05, 0.05), and g_couple is further multiplied by a small-degree correction. The detuning δ is defined as the difference between the path energy and the node’s resonance energy; a Lorentzian factor 1/(1+δ 2 ) modulates the effective coupling g_eff=g_couple/(1+δ 2 ). The scattering formulas are inspired by the Feshbach-resonance form, but we state explicitly: the T/R/A weights below are constructive definitions, and their normalization is a design choice, not a consequence of any S-matrix. In the main experiments every node’s resonance energy identically equals its own energy, so δ≡0 and the resonance channel is dormant—all measured effects are driven by the type-wise contrast of (g_couple, g_aether). We spell this out so that readers do not overestimate the current role of the resonance mechanism. The two working modes of the initial definition thereby unify as limiting states of a single entity: at g_aether=0 energy splits only between transmission and reflection (the blocking state, in which erroneous paths decay by local reflection); at g_aether≫0 erroneous energy condenses into the aether while correct paths transmit almost losslessly (the tunneling state); at general parameters all three channels are open and behavior interpolates continuously in the ratio η=g_aether/g_couple. The content of this unification goes no further than "the limiting 3
behavior of a two-parameter continuous family at its parameter boundaries," and its consistency check is that energy allocation at three parameter points must agree qualitatively with the two limits (verified in the text of §2.2).
2.2
The accounting identity and per-path auditing
Define the normalization constant Λ=1+g_eff+g_aether. The three-channel energy-allocation weights are T=1/Λ, R=g_eff/Λ, A=g_aether/Λ, corresponding to transmission, reflection, and irreversible dissipation into the aether. By construction T+R+A=1 immediately: with the dissipation channel explicitly included, the total energy of system plus environment is strictly conserved. Transmitted energy recurses along a path as E(i) =E(i−1) T_i, while cumulative reflection and cumulative dissipation are sums of per-node shares; for any path and any parameters, E(n) +E_refl+E_diss=E(0) (a one-step induction). The point of this identity is not mathematical difficulty but auditability: every unit of energy either reaches the endpoint, is reflected, or condenses into the aether—exactly one of the three—and a third party can independently recompute this for each path. The audit is measured, not promised. Across all variants, all 200 synthetic problems, and all candidate paths tested scattering by scattering, the maximum deviation of |T + R + A − 1| is 2.220446049250313×10−16 —exactly the scale of double-precision machine epsilon and far below the 10−6 implementation tolerance (closed (pre-registered); physics_audit field of results/deposon_v19_benchmark_fixes.json, passed=true). On the trap benchmark, the per-path average dissipation of the three limiting states is 0 / 3.63 / 0.358 respectively (energy units; deposon_benchmark_v1_3_traps.json → variant_results.v1_blocking,v2_ tunneling,unified.avg_ether_dissipated), qualitatively matching theoretical prediction: the blocking state dissipates nothing (closed system), the tunneling state dissipates heavily (open system, erroneous energy condensing in bulk), and the mixed state dissipates moderately. The aether channel is one-directional at the interface level: dissipate() atomically accumulates energy into a monotone counter, and the system exposes no recover(); this corresponds to the physical argument that the premise of Poincaré recurrence fails in an infinite-dimensional orthogonal environment. Honest disclosure: in a finite-dimensional software implementation, "irreversibility" is a combination of an engineering lock and an asymptotic property; the infinitedimensional approximation cannot be verified numerically and remains an irreducible idealization. The conservation audit also played a guarantor’s role in one error correction. In the frozen pipeline build the high_couple variant was a pure alias of v1_blocking (a configuration bug), so its then-reported GSM8K figure of 0.86 was wrong; after a true fix (g_couple×5 over the whole field, g_aether=0) an offline rerun gave 0.82, with 4 problems flipped and McNemar versus v1_blocking p=0.125, not significant (E9.3, deposon_v19_benchmark_fixes.json); on StrategyQA the post-fix prediction vector was bitwise identical to v1_blocking (p=1.0)—the physical perturbation genuinely occurred but does not alter greedy ranking on shallow 3-step graphs. Both reruns passed the conservation audit at 2.2×10−16 : the same ledger constrains the physics layer and our own correction process.
2.3
Path-filtering gain and its attribution boundary
On two controlled synthetic benchmarks of 100 problems each (seed=42, real LLM-backend decomposition, zero fallback in the final run), the full pipeline (unified variant) reaches 100% on both the simple set and the trap set, while the same-graph field-free greedy baseline reaches only 7%/10% (variant_results of deposon_benchmark_v1_3_simple.json / _traps.json). Two boundary statements must travel with these numbers. First, the baseline is a decoy-capture baseline: decoy edges are deliberately weighted 0.9 (above the correct operation edges’ 0.6) at graph-construction time, and all 93 failed problems select the decoy path; the effect sizes 4
+0.93/+0.90 measure same-graph path-filtering increment on an adversarially weighted graph, not a general capability improvement. Second, a three-way ablation shows that the increment comes from the combination of labels and dynamics: with type labels randomly permuted, trapset accuracy falls to 17.2%±6.4% (± is the sample standard deviation over 5 seeds, ddof=1; t95 half-width 7.9pp), and with uniform parameters on all nodes, trap-set accuracy degenerates to 10% (numerically coinciding with the field-free baseline; deposon_benchmark_v1_3_ labelshuffle.json → uniform_params.accuracy). The most accurate reading of the scattering layer is therefore a transducer: it converts semantic type labels into auditable energy decisions, and its increment is conditional on label quality. The picture on real benchmarks is more restrained. On the GSM8K subset (n=100, seed=42): CoT baseline 97.0%, unified 85.0%, v1_blocking 86.0%, and exact McNemar for unified versus CoT gives p=4.9×10−4 —CoT is significantly better, and we report the negative conclusion as it stands (deposon_benchmark_v1_4_gsm8k.json). The 95% confidence interval for the 85% versus 97% difference is [−20.5pp, −4.1pp] (unpaired Newcombe hybrid interval, same caliber as §2.4; deposon_v22_e95ci.json → unified_vs_cot), and the binomial CIs are computed by the Newcombe-Wilson method. On StrategyQA (n=99): unified 89.9% versus CoT 92.9%, p=0.549, no significant difference (deposon_benchmark_v1_4_strategyqa.json); the three arms v1_blocking, high_couple, and unified all score 89.9%—the constraint layer produced no differential action whatsoever on this task, a "constraint-layer inertia" that we record honestly. Read together, the two real benchmarks reveal a task-dependent constraint–fidelity trade-off: on GSM8K’s clean long chains the cost of information loss dominates, while on StrategyQA’s short chains of implicit reasoning the constraint layer ties CoT; the layer’s cost varies with chain length and decomposition fidelity. The equal-weight decoy control (E9.4, pre-registered) further severs misattribution: after flattening every edge weight of the same cached concept graphs to 0.7, the unified advantage does not vanish (GSM8K 0.85 versus no_deposon 0.04, p=1.7×10−23 ; StrategyQA 0.899 versus 0.202, p=7.5×10−15 ; E9.4 field of deposon_v19_benchmark_fixes.json), but its source is located in BFS’s shortest-path-first ordering and the type=’trap’ labels the physical layer gets for free—not in the scattering mechanism itself. The apparent effect size of unified versus no_deposon under the main protocol is therefore an unattributable number under structural bias, and this paper does not cite it as anti-capture value.
2.4
The E9.5 tie: differential value lies only in machine verifiability
The sharpest control is a trivial baseline. We construct a purely deterministic rule filter: after the same greedy path generation, discard any path passing through a node whose label hits the six-keyword list trap, dead, end, impossible, guess, wrong—reading only label strings, never type metadata, and involving no scattering mechanism at all. The result (E9.5, mechanically judged after pre-registration): on GSM8K the rule filter scores 0.87 ≥ unified 0.85 (McNemar b=0, c=2, p=0.5); on StrategyQA 0.899 = 0.899 (b=0, c=0, p=1.0). On two real benchmarks, the three-channel scattering pipeline is indistinguishable from a six-keyword filter on the accuracy dimension. The power-honest statement: no difference is detected at this sample size—the GSM8K difference is −2pp (95% CI [−11.8pp, +7.8pp], Newcombe-Wilson) and the StrategyQA difference is 0pp (95% CI [−8.8pp, +8.8pp])—both are unpaired Newcombe hybrid intervals (conservative bound; paired discordant pairs are only 0/2 and 0/0, so paired intervals degenerate; computation artifact deposon_v22_e95ci.json); with only 2/0 discordant pairs, the detectabledifference threshold is about ±10pp, so smaller increments cannot be excluded. The claim is therefore sharpened to "the differential value lies solely in machine verifiability," supported by the E9.4 attribution mechanism rather than by the strong form of this test. The mechanism behind the tie is no mystery: the scattering parameters are driven by nodetype labels, and the rule filter reads the string form of those same labels—same information source, same discriminative power. The scattering layer does not squeeze more discriminative information out of the labels; what it does is transcribe the same information into a representa5
tion that carries a conservation ledger. This negative result is not a footnote of failure but the basis for relocating the value proposition; the paper executes the pre-registered kill rules and claims nothing beyond the data [9]. What the rule filter lacks and the scattering layer alone possesses is: a per-path T+R+A=1 conservation ledger at arbitrary parameters (machine-precision residual), a recheckable energy record left by every block, and elimination decisions attributable node by node to concrete (T,R,A) shares. In other words, the differential value of the scattering layer lies only in machine verifiability, not in filtering performance. A parallel zero-benefit statement: the irreversible dissipation channel has never outperformed its disabled counterpart on any task measured so far (v1_blocking and unified tie on the synthetic benchmarks; GSM8K 86.0% ≥ 85.0%; all three arms tie at 89.9% on StrategyQA); its benefit claim currently has only theoretical motivation— erroneous energy cannot be resurrected within a single run—and validation requires iterative re-search scenarios, which we list as future work. The conditional value of LLM semantic signals on graph tasks aligns with existing chains of evidence [13, 14, 15]; that concept-map evaluation is protocol-sensitive is a thirty-year-old lesson [16, 17], of which E9.4/E9.5 are contemporary instances.
2.5
Fusion dilution: convex combination improves on no setting; gains can only come from nonlinear fusion
Beyond the scattering layer, the same concept-graph completion task admits a semantic-prior arm (a labels-only LLM prior with zero leakage: the prompt contains only node labels). A natural hypothesis holds that the two are complementary—the field handles structure, the prior handles semantics—so a convex combination hybrid=λ·field+(1−λ)·prior should take the best of both. A pre-registered scan refutes the weakest operational implication of this hypothesis. In the four-λ single-graph scan with λ∈0.25,0.5,1,2 all four settings are identical (dimensional saturation already at λ=0.25): named Hits@3 is constant at 0.294 and any_ lambda_pass=false (success_evaluation of deposon_v16_llm_prior.json). In the familyL four-graph all-candidates protocol at λ=0.5 (deposon_v20_crossval.json, hybrid_lambda_ convex=0.5): physics 0.484→0.452, historical 0.783→0.739, and the other two graphs flat—the hybrid never exceeds the prior-only arm on any measured setting; the field only dilutes a genuine semantic prior. A further ablation exposes how an apparent gain can be faked. The λ=2 setting (field coefficient −1, prior blank rows effectively reverse-sorted by field score) once produced an apparent "fusion gain" of named Hits@3=0.471 under a normalized variant (the [email protected] arm of deposon_v17_fusion_fix.json); the E9.6 null ablation, which fixes the endpoints and shuffles only the confidence values, yields named scores exactly equal to the real prior (0.1176=0.1176, with the random-edge null hypothesis scoring 0 on all 5 runs; E9_6c_lambda2_null_ablation of deposon_v19_quickwins.json)—that hit is not attributable to semantic confidence, only to endpoint position and the anti-field artifact, and the per-edge annotations are on record. Together the two lines of evidence form an exclusionary conclusion (closed (pre-registered) on the measured λ settings; we do not extrapolate it to a theorem over the whole λ space): the convexcombination channel is closed, and if field–prior gains exist at all, they can only come from nonlinear fusion mechanisms. The GNN literature already answers systematically how value domains divide between structural and semantic signals—the effective domain of structural bias is determined by graph properties, and label signals take over on heterophilous graphs [18, 19]. Our dilution conclusion points the same way and pushes the demarcation to the sharper setting of "extreme signal pairs that are mutually blind at the mechanism level."
6
3
The Game-Theoretic Formulation: Potential, the Empirical Coordination Ratio (ECR), and the Audit Boundary
The conservation ledger of §2 is verified only statically: it proves that the energy of each scattering event goes where it should, but it does not answer the dynamical question—does the reverse evolution of the field, taken as a whole process, possess a scalar against which "where to step next, how far one has come, and how far one remains from optimum" can all be audited on the same ledger? This section models the reverse dynamics as a potential game on the graph and answers under a single pre-registered protocol on 22 controlled concept graphs. We fix the reporting convention first: the positive narrative of this section adopts the consistency register— the evidence is consistent with a potential-game reading, i.e., consistency evidence rather than formal proof; wherever a pre-registered verdict closed (GT-5b, GT-6) or a formalization kill-test closed (GT_FORMAL, T-P1c), we label it separately as "closed (pre-registered)" and never blend tiers. The section gives its operational definitions in a self-contained manner and relies on no other document. Task and field are defined as follows. The task is concept-graph completion: leave-one-out prediction over each gold edge, candidates are all graph nodes, and the metric is named Hits@3. For a leave-one-out task (s,t), the scattering layer defines a field-guided energy on the graph’s adjacency weight matrix W (row-stochastic): E(s,t)=−log(Σ_p Π_e t_e)+λ_smooth·Σ W2 _ij, where the first term is the negative log of the aggregate transmissivity over all s→t paths and the second is a smoothness regularizer. The reverse process performs simplex naturalgradient descent on masked positions under an annealing schedule (β linear, 50 steps, lr=0.1), formalized as the update rule w_t+1∝(1−lr)·w_t◦exp(lr·w_t◦∇Φ), projected back onto the simplex at each step; upon termination the distribution over the masked row ranks the candidate targets, yielding field scores. The deterministic mean-field reverse (field_mean) takes the mean of a Dirichlet starting distribution and involves no sampling noise anywhere; the control arm introduces Dirichlet sampling at the start, whose concentration parameter α plays the role of an effective temperature. The corpus is designed along two families, "structural negation × genuine semantics": family S (16 graphs, synthetic placeholder labels decoupled from structure) negates the strong claim "field = universal skeleton detector"; family L (6 graphs, genuine domain concept labels, LLM-generated, 30–45 node DAGs) carries the semantic-arm tests.
3.1
The model: a potential game on a finite graph
The modeling quadruple is fixed in one pass: each leave-one-out prediction edge is a player; strategies are candidate target nodes; utilities are field scores; and the negative physical energy Φ=−E is the potential-function candidate. Classical existence conditions for potential games come from Monderer & Shapley [20] and Rosenthal’s congestion games [21]; Sandholm’s population best-response dynamics [22] prove that deterministic best-response dynamics ascend the potential gradient in the population limit—the nearest theoretical anchor for reading "mean-field deterministic reverse as noiseless best-response dynamics"; Candogan et al.’s flow decomposition of games [23] provides the operative analogy for "field = graph-flow component" (we decompose the edge-utility vector and do not integrate along trajectories). The decomposition "each task is an independent player" is a modeling choice, not a unique one; Φ=−E has analytic grounding—the reverse annealing gradient is exactly −∇E. One point must be stated plainly: the leave-one-out tasks are mutually independent (each has its own (s,t) and its own energy function), so the existence of an exact potential at the game level is trivial—Φ=Σu_i already proves it, and no test is needed. GT-6’s contribution is therefore not nontrivial evidence for game-potential existence but the near-gradient property of the per-task energy function as a scalar ledger: in the edge-utility vector space (Euclidean inner product), the edge-utility vector F is projected onto the gradient subspace, p=pinv(B)F, with residual ratio r=∥F−Bp∥/∥F∥;
7
r≈0 means the scalar ledger explains the edge utilities almost completely (§3.2). The model lends three concepts new semantics. First, dissipation is an endogenous commitment device: at g_a>0 energy condenses into the aether and cannot flow back, the budget of an eliminated branch cannot be resurrected, and no player can undo an elimination by rewriting the books after the fact—commitment is not an external rule but a property of the dynamics itself. Second, auditability upgrades from stepwise compliance checking to a scalar ledger: if a potential exists, the entire trajectory can be audited as the stepwise ascent of a scalar curve Φ(t) (equivalently the stepwise descent of E(t)); a third party need not recompute every energy component at every step, only recheck whether one scalar sequence is monotone—the dynamical-layer counterpart of the conservation ledger. Third, "how far from the coordination optimum" becomes computable: we define the empirical coordination ratio (ECR) to quantify the gap between self-interested dynamics and the field benchmark. ECR is the distribution-level operational counterpart of the classical worst-case PoA (the Koutsoupias & Papadimitriou sense [24]; the affine-congestion PoA=4/3 bound of Roughgarden & Tardos [25]), not the same metric; to our knowledge, such a distribution-level operational counterpart has rarely been reported systematically (distribution-level reporting has precedent in [26], and classical equilibrium-efficiency analysis in [27]). Scope qualifications are in §3.2.
3.2
Decisive evidence: monotonicity, near-gradientness, and quantification of the auditable scalar (closed (pre-registered) tier)
Two pre-registered verdicts closed and one closed under restrictions; together they answer "does the scalar exist, and to what quantitative precision." GT-5b potential-trajectory monotonicity (closed (pre-registered), narrowed claim): the Φ trajectory of mean-field reverse is monotone non-decreasing on 22/22 graphs, a monotonicity rate of 100%, above the pre-registered line of 80%; the kill line—any single triggering graph kills—did not trigger, and the verdict supports_narrowed_monotonicity closed (deposon_v20_gt5b.json → per_graph_summary.*.meanfield_monotone_rate=1.0). The claim is the narrowed version: the original GT-5 endpoint condition failed (the noise arm’s terminal Φ overtook on 3/4 graphs, judged inconclusive; see §3.5), the monotonicity claim closed independently under a new pre-registration, and the endpoint reversal remains on record, unrewritten. The scope qualification must travel along: 22/22 monotonicity is a frozen fact under the "single-edge leave-out mask × corpus graphs" protocol and is not extrapolated into betterresponse evidence—under general mask structures that dynamics has already been killed (§3.3). GT-6 non-potential residual (closed (pre-registered), near-gradientness register): with the residual ratio r as defined in §3.1 (projection of the edge-utility vector F onto the gradient subspace; cited here, not restated), the median non-potential residual is 1.594×10−29 , far below the pre-registered line of 0.10; the verdict potential_game_explanation_complete closed (deposon_v20_gt6.json → verdict.median_residual_ratio). Relocation note (see §3.1): since game-level exact-potential existence is trivial under independent tasks (Φ=Σu_i suffices), what GT-6 tests is not that, but the near-gradient property of the per-task energy function as a scalar ledger. Two honest disclosures. First, three graphs with cyclic structures exceed the line—S4=0.148, L_algorithm_process=0.136, S5=0.121—on which the potential explanation is only approximate. Second, the numerical-precision qualification: the median residual sits at the scale of double-precision floating-point underflow and reflects the structural fact that "the edge-utility vector lies almost entirely within the gradient subspace"; it should not be read as a physical quantity by its absolute magnitude, and the 0.12–0.15 residuals of the three exceptional graphs are the informative non-potential components. This decomposition is the operative analog of Candogan’s flow decomposition [23]. GT-4 empirical coordination ratio ECR (restricted-consistent, scope qualification attached): the operational definition is ECR=field_mean/max(self-interested arm), with self-interested arm set random, degree (the pre-registered llm_prior arm is unavailable on fam8
Figure 1: Distribution-level empirical coordination ratio (ECR) across all graphs (bars include the red-boxed ECR<1 graphs and the separately counted ∞ cases; median 1.3333 and the pass line 1.2 are both read from the frozen verdict fields) ily S; the denominator can only be smaller and the ECR only larger—disclosed as-is). The main report covers all 17 finite-valued graphs: median ECR=1.333 > the pre-registered line 1.2; the family-S subset of 13 finite-valued graphs (median 1.5) is listed as a supplementary register; the 3 graphs with ECR=∞ are, per the standing handling rule, counted separately and excluded from the median (self-interested arm named=0 while field>0). GT-4 covers 20/22 graphs: L_geography_world and L_project_management were not evaluated in the frozen run (arm data absent)—neither finite nor ∞; disclosed as-is, the median caliber is unaffected. Two qualifications. First, ECR is the distribution-level operational counterpart of the classical worst-case-NE/social-optimum ratio (PoA), not the same metric: the affine/separable cost-structure premises are not closed, we report at the distribution level only, and we draw no numerical juxtaposition against 4/3 or any other classical worst-case bound—with different metrics, numerical proximity is not evidence. Second, all ECR<1 cases are disclosed in parallel: two in family L (L_historical_causality 0.5, L_physics_concepts 0.75) and, elsewhere in family S, S2_n45=0.5—the field is negatively coordinating in the semantic domain, which is self-consistent with the division-of-labor boundary of §3.4: the field creates coordination value only in the structural domain, the applicable domain of the potential-game reading coincides with that boundary, and the theory does not contradict itself. Frozen-convention note: the field names in the verdict JSON (GT4_price_of_anarchy, field_coordination_value_supported) are frozen pre-run conventions and are not renamed; the text uses ECR throughout. (Note on figure order: Figure 5 is first cited here, ahead of Figures 1–4 in §3.4—a forward cross-reference; figure numbers are kept identical to the frozen figure-file names and are not renumbered.) Taken together: each step of the evolution is auditable as potential ascent (GT-5b + GT6, closed (pre-registered) under their respective protocols), and the distance of self-interested dynamics from the coordination optimum is computable (GT-4, restricted-consistent). The empirical case for the monotonicity and near-gradientness of the auditable scalar now stands.
9
3.3
Formalization kill-tests: all three tiers of dynamical equivalence falsified (closed (pre-registered), systematically sampled graph families × exhaustive-state protocol)
The model of §3.1 contains an interpretive correspondence: deterministic mean-field reverse = noiseless best-response dynamics. The strongest formalized version of this correspondence is systematically refuted by the GT_FORMAL kill-test (deposon_v21_gtformal.json, seed=210021; the verdict function was committed as a pure function before any run): 61 systematically sampled small graphs with n≤8 (five families: chain/star/tree/random DAG/cyclic) × the 338 single-node all-candidate mask tasks on them × 20 mean-field steps = a full enumeration of 6760 kill states—i.e., the exhaustion operates at the task/state level, while the graph families are systematically sampled. The strength tier is "closed (pre-registered) (systematically sampled graph families × exhaustive-state protocol)"—what is closed (pre-registered) is the negative fact itself, that each strong formulation fails under this protocol; the kill is the answer, and the strength is not rounded upward. The kill logic does not depend on enumerative completeness: one counterexample kills a universal claim, and 10/338 violations are more than sufficient. • P1a strong dynamical equivalence: falsified. max∥T−BR∥∞=0.8569; an lr scan shows the directional deviation 1−cos≈0.91 does not vanish with lr—an O(1) deviation rather than an O(lr2 ) discretization error, not a step-size problem. • T-P1b better-response version: killed. min directional cosine −1.0, min −2 ∆Φ=−1.2424×10 (kill line −1e−9), with 10/338 tasks in violation, all on support graphs with non-empty cycle space. The mechanism in one sentence: the fixed point of the composite operator T=Π◦C◦M is not a row-constrained maximizer of Φ; the trajectory keeps advancing past the peak (overshoot), and ∆Φ<0 beyond the maximum. • P1c residual entropy-regularization characterization: killed (T-P1c, deposon_ v22_p1c.json, same sampled-family/exhaustive-state protocol, verdict function pre-registered before any run). The residual claim was "the first-order update direction = the mirror-ascent direction of the entropy-regularized potential Φ_τ =Φ+τ H" (at τ >0 the entropy term might repair overshoot). Two kill lines, both struck: the strong form requires min directional cosine ≥0.999 under a globally uniform τ —an 81-point grid over τ ∈[0,4] fails everywhere, and at the best global τ the min cos=−1.0; the weak form allows τ to be chosen per state (τ *) and requires min cos≥0.99—min cos=−1.0, i.e., overshoot states exist whose direction is strictly opposite to every entropy-regularized mirror direction; the median per-state τ * is 0 (85.98% of states have τ *=0), so the entropy term has no reparative power. The verdict function is locked by 11 tests, and a same-seed rerun is byte-identical. With all three tiers—P1a strong form, P1b better-response, P1c first-order entropyregularized characterization—killed, the game-theoretic main line retains no residual claim of dynamical equivalence. The protocol relationship needs clarifying: GT-5b’s 22/22 monotonicity (single-edge leave-out mask × corpus graphs) is unaffected and remains a closed frozen fact; the present test proves, under the strengthened protocol "all-candidate mask × systematically sampled small graphs," that under general mask structures the dynamics is not better-response. Both protocols stand on record; neither cancels the other. Pre-registration time anchors. The verdict pure functions and SPECs were frozen before any run, with SHA-256 anchors (first 12 hex digits): GT_FORMALIZATION_v1.md = aeefb8ef6972; run_v21_gtformal.py = 9bbe43f41fa8; run_v22_p1c.py = 6e9673205dc0; SPEC_GT2B = 68a5b08ef007; SPEC_GT8C = 6b09de9911c0. Anyone can verify the frozen versions against these anchors and mechanically rerun every verdict.
10
P2 potential completeness: downgraded to approximate potential game. Acyclic (undirected-forest) support graphs have residual r≤5.9×10−16 (numerical zero); support graphs containing cycle space have median r=0.669, and the fraction with r>0.30 is 0.924 > the pre-registered 1/3 downgrade line, triggering the downgrade verdict downgraded_to_approximate_potential_game—same direction as the three cyclic-graph exceptions of GT-6 and of higher magnitude (all-candidate-mask support graphs are denser). P3 dissipation channel: dead in both directions, but yielding the first mechanistic premise evidence. "Dissipation is merely a reparameterization of payoffs" is killed (max|r(0.1) − r(0)|=0.1585 > 1e−9); "the equilibrium set is invariant" is likewise killed (maximum difference of best-response fixed points across three g_a settings: 0.8442)—dissipation is no pure reparameterization; it genuinely moves equilibrium positions. An incidental finding changed the evidential landscape of the dissipation channel: in 58/338 tasks with g_a=0 and cycles in the support, ρ(G)=1 makes (I−G) singular and the closed-form dynamics diverges; the spectral condition for convergence, ρ(G)<1, is guaranteed precisely by dissipation g_a>0—this is a built-in property of the g_a>0 construction and is archived as mechanistic premise evidence (previously only the negative "no rent" register existed; see §4).
3.4
Demarcation and boundary regularities: where the auditable advantage holds (directional evidence, not upgraded)
Once the auditable scalar exists, the next question is "on which graphs does it hold." This subsection states the demarcation regularities with all their qualifications, positioned as observational regularities (directional evidence), not as an established discriminator contribution. Contraction of the validity domain (pre-registered verdict): the headline benchmark H-A1 (field_mean > random) is killed—a 22-graph sign test gives 16+/4−/2 ties, p=0.0118, passing Holm, but the pre-registered kill line is a disjunctive rule (non-significance or ≥3 reversed graphs), and with 4 reversed graphs (L_historical_causality, L_physics_concepts, L_project_management, S2_n45) the kill line triggers and the headline claim dies. The surviving statement is H-A2 (field_mean > degree), robust across protocols: 22 graphs, 19+/1−/2 ties, p=4.0×10−5 , Holm-passed; on the 20-graph subset, Wilcoxon (p=0.0031, |r|=0.83, with √ r=Z/ N) and a paired t-test (p<0.0001, d=2.05, Cohen’s dz) confirm a large effect; since Hits@3 is a bounded discrete quantity, Wilcoxon is the primary criterion and the paired t-test is reported only as a reference. Multiplicity statement: the Holm correction is applied only within the H-A1/H-A2 family; each GT experiment was independently pre-registered and independently judged, with no pooled correction across families—stated as-is. The field’s validity domain thereby contracts to "robust advantage over trivial structural baselines + local advantage on high-hub graphs." Cross-vendor robustness (GT-3b, restricted-consistent): five evaluators from three model families reproduce the prior advantage on Kimi-generated graphs—doubao 4/4 and deepseek 6/6 pass the criterion, the three model families combined show 0 domains where prior ≤ field, and Kendall’s W=1.0 across all ok domains (per-domain rankings in complete agreement). The hypothesis "the prior advantage is a same-vendor same-source contamination artifact" is substantially weakened; the residual limitation: all three families are Chinese-optimized large models, and shared Chinese corpora cannot be ruled out. Demarcation regularities: an exploratory regression (n=20) shows hub_concentration (max in-degree / edge count) is positively associated and the strongest correlate of field efficacy (positive direction, p=2.8×10−4 ), while real_semantics significantly suppresses field performance (negative direction, p=0.012), R2 =0.628; the coefficient point estimates are archived in Appendix A and the text keeps only direction and significance; two qualifications—n=20 is a small sample and the regression is labeled exploratory, and the features are themselves corpus-design variables, creating quasi-circularity, so the coefficients indicate directional association only. Out-of-corpus replication: on the hub axis, both paired new-graph pairs (2/2) 11
Figure 2: Overview of the division-of-labor boundary (22-graph field–prior win/loss map; fieldarm bars plus family-L prior-arm bars plus random/degree reference scatter; the per-graph winner is marked with ⋆) moved in the same direction (direction only; power suffices only to resolve very large effects); on the real_semantics axis, 3/4 domains in total satisfy "prior stronger"—GT-8b’s two new domains pass the 0.6/0.2 pre-registered thresholds 2/2 (chinese_dynasties prior 0.7805 versus field 0.0732; chemical_elements 0.6429 versus 0.1429; an initial inconclusive was turned positive by a pre-registered amendment adding data, with the chain on record), while GT-8c’s cross-backend retest is judged mixed—biological_taxonomy prior 1.000 versus field 0.075 (diff +0.925) crosses the line, programming_concepts prior 0.500 versus field 0.3333 (diff +0.1667) misses the margin threshold, and in GT-8c the graph-generation arm and the prior arm share the same backend and the same model, so same-source contamination risk is on record. The regularities read: high hub_concentration → structural signal is usable overall (not exclusive to the field); real_semantics=1 → use the semantic prior. Cross-domain heterogeneity is on record; the strength tier remains directional evidence, not extrapolated to tasks beyond family L or to larger graphs.
3.5
Consistency evidence and the inconclusive archive
Beyond the main-line verdicts, four weaker pieces of evidence are archived honestly—neither upgraded nor deleted. GT-1 (consistency, weak discriminative power): Dirichlet-noise reverse is uniformly worse than the deterministic limit under the hit-rate register—dirichlet mean 0.10 versus meanfield 0.40, gap=0.30 ≥ the pre-registered line 0.2, strictly worse on 20/20 runs. A statement of discriminative power: any noise-induced degradation satisfies this pattern, so this evidence serves only as background consistency and does not independently support the potential-game reading. GT-2 (restricted-consistent, no_separation): an attacker who knows the keyword list generates rule-evading semantic trap labels 100% of the time (evasion_rate=1.000, 4/4 graphs); after injecting 10 trap nodes per graph, rule_filter drops −7.5pp on average (−10pp on three graphs, 0pp on the fourth), while field_mean drops −0.0pp (zero on all four graphs). The pre-registered mechanical verdict is no_separation: the rule collapse does not reach the 20pp threshold and the attack strength is not decisive, so this is not upgraded to a "rule defense fails" claim; the directional signal that the field does not read labels and is mechanistically immune to semantic traps is on record. The literature chain on the failure of rule-based defenses [28, 29, 30, 31, 32] and the methodology of adaptive attacks [33] point in the same direction,
12
Figure 3: The H-A1 kill line and the distribution of reversed graphs (22-graph sign-test scatter, n_pos=16/n_neg=4/n_tie=2, the four reversed graphs marked with red boxes) but we do not draw strength from them here. GT-5 and GT-2B (inconclusive archive): the GT-5 endpoint condition failed—the noise arm’s terminal Φ exceeded mean-field on 3/4 graphs (S6 gap=−0.31), judged inconclusive and left unrewritten; we read this as "noise explores, mean-field exploits," consistent with the classical division of labor between log-linear learning [34] and deterministic best response, whereby the scattering layer’s "temperature" acquires game-theoretic semantics rather than serving as a tuning knob; the narrowed monotonicity claim has since closed independently as GT-5b (§3.2). GT-2B’s multi-trap strength escalation (T∈1,2,3) is judged inconclusive: rule_filter scores 0.150/0.275/0.200, non-monotone, and the fixed four-option design makes the number of in-graph candidates vary inversely with T, so the field-immunity criterion is contaminated by option composition (field accuracy rises mechanically at 0.375/0.525/1.000)—the design lesson is archived as "immunity criteria must be robust to option degrees of freedom." GT-7 temperature frontier (mixed, the audit boundary): α∈0.3,. . . ,20 × 4 graphs × 5 seeds. The GT-5 reversal reproduces and is systematic (the same 3/4 graphs show terminalΦ overtaking in high-temperature settings; Φ gains concentrate at the high-temperature end, corr=−0.87), so temperature genuinely controls the global exploration benefit of the potential; but a "win-win frontier" does not hold—on S6 at high temperature Φ rises while hit rate falls 0.4→0.08. Potential and hit rate are two objectives, and the audit promise covers only the former: raising temperature improves potential exploration only, not hit rate—the audit bound13
Figure 4: Scatter of the division-of-labor regularities (hub_concentration × real_semantics, colored by field_named; 20-graph sample; regression coefficients read from frozen fields) ary is thereby drawn. Boundary-case disclosure: graphs with hit rate = 0 make the condition degenerately true; the rule was not rewritten after the fact.
4
Honest Boundaries
The three tiers of evidence strength are never blended in this paper. The limitations are declared centrally here, each traceable to a verdict recorded above. First, all three tiers of formalized equivalence are killed—the clearest limitation. The strong formulations of "mean-field reverse = noiseless best-response dynamics" are explicitly refuted under the sampled-family/exhaustive-state kill-test (61 graphs / 338 tasks / 6760 states): P1a shows an O(1) deviation of 0.8569; P1b has min cosine −1.0 (the overshoot mechanism is on record); T-P1c is doubly killed both on the 81-point grid τ ∈[0,4] and under per-state self-chosen τ * (min cos=−1.0, median τ *=0). This means none of the game-theoretic readings in this paper may rest on "the dynamics implements best response"; outside the original GT-5b/GT-6 protocols no residual claim is retained, and the positive narrative of the game-theoretic line lives only in the consistency register. The kill direction is not softened: what is refuted is not some parameter configuration but the correspondence itself. Second, the consistency register. Apart from the closed pre-registered verdicts (GT-5b, GT-6) and the kill conclusions, the positive evidence is consistency evidence, not formal proof— trajectories agreeing with the potential-game reading does not amount to a proved potentialfunction theorem. The register is not relaxed, but neither is it upgraded to a theorem; readers should calibrate their trust in the conjunctive statements of §3.2 accordingly. Third, family-L graphs are generated by a single vendor’s LLM. The cross-vendor evaluator question has been addressed by GT-3b (0 defeats, W=1.0), but the single-vendor graph-generator question is not closed: the graphs’ structure and label distributions carry the generating model’s preferences. All three tested families are Chinese-optimized large models, 14
Figure 5: Per-graph shape of the GT-7 temperature frontier (dual axes for hit rate and terminal Φ; 4 graphs × 6 temperature settings × 5 seeds; verdict=mixed) same-source contamination through shared Chinese corpora cannot be ruled out, and testing non-Chinese model families and human-annotated graphs remains an open limitation. Fourth, the question-bank track has small samples at n=40 per cell. ±1 problem = ±2.5pp, the same order as chance noise; every number on that track (including prior 92.5%, field 52.5%, rule 27.5%) is to be read as a small-sample wide interval and is not cited as an effect size. Fifth, the demarcation regularities are exploratory evidence. The n=20 regression exhibits feature–design circularity (the predictors themselves participated in corpus design); the hub axis has 2 paired pairs giving direction only, with power sufficient only for very large effects; the real_semantics axis holds in 3/4 domains but cross-domain heterogeneity is on record (GT-8c mixed, programming_concepts below the line). The regularities are not offered as a discriminator contribution and are not extrapolated beyond family-L tasks or to larger graphs. One more: the benefit of the dissipation channel remains motivation only. GT_FORMAL’s incidental finding supplies mechanistic premise evidence for dissipation—a built-in property of the g_a>0 construction (it guarantees the spectral condition ρ<1, avoiding divergence in 58/338 tasks)—but task-level rent remains negative on every task measured— v1_blocking and unified tie on the synthetic benchmarks, GSM8K 86.0%≥85.0%, and all three arms tie at 89.9% on StrategyQA. Mechanistic premise (guaranteeing that convergence exists, a built-in property of the construction) and task-level rent (delivering accuracy gains) are two different things: the former is closed (pre-registered); the latter is zero on every task measured in this paper, and its validation requires iterative re-search scenarios, listed as future work.
15
5
Conclusion
This paper addresses one empty layer—run-time, per-instance invariants—along two equally weighted lines of results. Auditable representation and conservation guarantees (§2): the constructive definition of three-channel scattering makes T+R+A=1 hold for arbitrary parameters, with a per-path audit residual of 2.2×10−16 and node-by-node attributable elimination decisions; at the same time, pre-registered controls killed the attributability of the apparent accuracy advantage (E9.4/E9.5), relocating the value to machine verifiability. The game-theoretic formulation (§3): the empirical case for the auditable scalar closed (GT-5b 22/22 monotonicity, GT-6 median residual 1.594×10−29 ), the distribution-level empirical coordination ratio ECR=1.333 quantifies "how far from the coordination optimum," and the temperature frontier demarcates the audit boundary; at the same time, the clearest kill conclusions sealed off all three tiers of the formalized "dynamics = best response" claim, and the P2 downgrade together with the two-directional P3 kills delineate the applicable domain of the potential explanation, incidentally yielding mechanistic premise evidence that dissipation guarantees convergence (a built-in property of the g_a>0 construction). The genre discipline common to both lines: kill conclusions carry the same weight as closed (pre-registered) conclusions, the three strength tiers are never blended, and every negative result is archived without cosmetic repair. The scattering layer shows no detected accuracy difference at this sample size (E9.5: GSM8K −2pp, 95% CI [−11.8pp, +7.8pp]; StrategyQA 0pp, 95% CI [−8.8pp, +8.8pp], both unpaired Newcombe hybrid intervals (conservative bound; paired discordant pairs are only 0/2 and 0/0, so paired intervals degenerate; deposon_v22_e95ci.json); increments below about ±10pp cannot be excluded, and the differential-value claim rests on the E9.4 attribution mechanism); convex-combination fusion only dilutes on the measured λ settings (E9.6); "dynamics = best response" is killed by the sampled-family/exhaustive-state test—these three negatives are conclusions of this paper just as much as the three positives of the conservation ledger, potential monotonicity, and ECR quantification, not footnotes. The value proposition of the scattering layer thus compresses to one sentence: it is not more accurate, but it keeps the books, and the books can be recomputed step by step by any third party with double-precision arithmetic. Future work, exhaustively listed: trainable KGE baselines (TransE/ComplEx/RotatE, scheduled at the 20-graph scale); re-estimating the exploratory regression after corpus expansion; upgrading the hub axis to ≥6–8 paired graph pairs for power; the question-bank track control SPEC_GT2C and the monotone-endpoint retest SPEC_GT5C are pre-registered and await execution. All three tiers of dynamical equivalence have been killed and are not listed as future work.
Acknowledgments The author thanks colleagues for discussions and feedback on early drafts.
6
Data availability and evidence-strength conventions
Every experimental number is traceable to a specific field of a frozen JSON under results/, as enumerated in Appendix A. Every statistical verdict is the mechanical evaluation of a preregistered decision rule (a pure function frozen before the run that reads the JSON and returns a boolean). Evidence strength is labeled in three tiers throughout: closed (pre-registered) (the verdict rule was frozen as a pure function before any run and mechanically evaluated to closure), consistency evidence (direction agrees but the pre-registered strong criterion was not met), and motivation only (theoretical intuition, no empirical support).
16
7
Artifact availability
All code, frozen JSONs, verdict pure functions, and the test suite ship inside the repository at https://github.com/zeroandcat/Deposon; frozen versions can be checked against the SHA256 anchors of §3.3, and every verdict can be mechanically rerun via the bundled scripts. No proprietary dependencies are required to reproduce any number in this paper.
8
AI use disclosure
The Deposon framework definition and the project’s research direction were provided by the corresponding author. The core algorithm, the experimental pipeline, the data-analysis scripts, and the first draft of this manuscript were produced with the assistance of KIMI-K3 (Moonshot AI) under the author’s direct instruction. The author reviewed and edited every section, verified all numerical claims against the frozen JSONs in the repository, and takes full responsibility for the entire content of this paper under arXiv’s authorship and AI-use policy. No content was generated by the AI without human review, and no proprietary model weights are required to reproduce any result in this paper.
A
Number Traceability Table
Every key number maps to a specific field of a frozen JSON under results/ (verdicts are read mechanically by scripts; hand-copying is forbidden). Number
Source JSON → field path
Conservation-audit residual 2.220446049250313×10−16 , passed=true, tolerance 1e-6 E9.3 post-fix GSM8K 0.82, 4 problems flipped, McNemar p= 0.125; StrategyQA p=1.0 E9.4 equal-weight control GSM8K 0.85 vs 0.04 (p= 1.7e-23), StrategyQA 0.899 vs 0.202 (p=7.5e-15) E9.5 rule filter GSM8K 0.87 vs 0.85 (b=0,c=2,p=0.5), StrategyQA 0.899=0.899 (p=1.0) E9.5 difference 95% CIs: GSM8K [−11.8,+7.8]pp /StrategyQA [−8.8,+8.8]pp (unpaired Newcombe, conservative caliber); unified vs CoT (GSM8K) −12pp, CI [−20.5,−4.1]pp Synthetic benchmarks unified 100%/100% vs decoy-capture baseline 7%/10% (seed=42) Label-permutation ablation 17.2%±6.4%; uniform-parameter degeneration 10% (trap set)
deposon_v19_benchmark_fixes.json → audit.t_plus_r_plus_a_max_deviation, audit.passed, physics_audit.tolerance deposon_v19_benchmark_fixes.json experiments[’E9.3_high_couple_fix’]
physics_ physics_ →
same → experiments[’E9.4_equal_weight_decoy_ control’].benchmarks
same → baseline’].benchmarks
experiments[’E9.5_rule_
deposon_v22_e95ci.json gsm8k/strategyqa/unified_vs_cot (method/ci)
→
deposon_benchmark_v1_3_simple.json / deposon_ benchmark_v1_3_traps.json → variant_results deposon_benchmark_v1_3_labelshuffle.json label_shuffle / uniform_params
17
→
Number
Source JSON → field path
GSM8K: CoT 97.0%, unified 85.0%, v1_blocking 86.0%, p= 4.9e-4 StrategyQA: unified 89.9% vs CoT 92.9% (p=0.549), three arms tied at 89.9% Fusion dilution: physics 0.484→ 0.452, historical 0.783→0.739 (λ=0.5 convex combination) Four-λ single-graph scan: all four settings identical, named Hits@3=0.294, any_lambda_pass=false λ=2 normalized-variant apparent 0.471; E9.6 null ablation 0.1176= 0.1176, random edges 0 on all 5 runs H-A1 killed: 16+/4−/2, p= 0.0118, kill line triggered, 4 reversed graphs H-A2 survives: 19+/1−/2, p= 4.0e-5 Wilcoxon p=0.0031/|r| =0.83; paired t p<0.0001/d=2.05
deposon_benchmark_v1_4_gsm8k.json
Regression β(hub)=2.12 (p= 2.8e-4), β(real_sem)=−0.16 (p= 0.012), R2 =0.628, n=20 GT-1 gap=0.30 (0.10 vs 0.40), strictly worse on 20/20 runs GT-4 ECR median 1.333 (17 graphs) /1.5 (family S, 13 graphs); family L 0.5/0.75; ∞×3 counted separately, excluded from the median; S2_n45=0.5 (field names GT4_price_of_anarchy and field_coordination_value_ supported are frozen conventions, not renamed; coverage 20/ 22 graphs: L_geography_world and L_project_management not evaluated in the frozen run—arm data absent, neither finite nor ∞) GT-5b 22/22 monotone (monotonicity rate 100%, pre-registered line 80%) GT-5 endpoint reversal S6 gap= −0.31 (3/4 graphs inconclusive)
deposon_benchmark_v1_4_strategyqa.json
deposon_v20_crossval.json → per-graph fields of hybrid_lambda_convex=0.5 deposon_v16_llm_prior.json evaluation
→
success_
deposon_v17_fusion_fix.json → [email protected]; deposon_v19_quickwins.json → E9_6c_lambda2_ null_ablation deposon_v20_corpus_eval.json → verdicts.H_A1_ field_mean_gt_random.sign_test, verdicts.kill_ lines.H_A_dead same → verdicts.H_A2_field_mean_gt_ degree.sign_test.p_exact=4.005e-05 v20_statcheck_fm_vs_rand.json (p_value=0.003052), v20_statcheck_fm_vs_deg.json (p=2.13e-08, d=2.0478) v20_regression_field_v2.json → coefficients.*, r_squared=0.628226, n_observations=20 deposon_v20_gt.json → GT1_potential_game_ convergence.verdict same → GT4_price_of_anarchy.verdict.poa_per_ graph_finite, n_poa_inf=3
deposon_v20_gt5b.json → per_graph_ summary.*.meanfield_monotone_rate=1.0 deposon_v20_gt5.json → per_graph_detail.S6
18
Number
Source JSON → field path
GT-6 median residual 1.594e29; exceptions S4=0.148 / L_algorithm_process=0.136 / S5=0.121 GT-7 mixed: corr −0.87; S6 hit rate 0.4→0.08
deposon_v20_gt6.json → verdict.median_ residual_ratio, per_graph_summary.*.residual_ ratio_mean
GT-2 no_separation: rule −7.5pp (0.1/0.0/0.1/0.1), field −0.0pp, evasion=1.0 GT-2B inconclusive: rule 0.150/ 0.275/0.200; field 0.375/0.525/ 1.000 GT-3b: doubao 4/4, deepseek 6/ 6, 0 defeats, Kendall W=1.0 GT-8 2/2 same direction (pair A +0.7917>+0.1333; pair B +0.0526>−0.0833) GT-8b supports_H_GT8B: chinese_dynasties 0.7805 vs 0.0732 (diff +0.7073); chemical_elements 0.6429 vs 0.1429 (diff +0.5000) GT-8c mixed: biological_taxonomy 1.000 vs 0.075 (diff +0.925) crosses; programming_concepts 0.500 vs 0.3333 (diff +0.1667) misses; vendor= volces_ark_bytedance GT_FORMAL: P1a max∥T−BR∥∞=0.8569, 1−cos≈0.91 does not vanish with lr; P1b min cos=−1.0, min ∆Φ=−1.2424e−2; P2 acyclic r≤5.9e−16, cyclic median 0.669, fraction 0.924; P3 max r difference 0.1585, fixed-point difference 0.8442, g_a=0 divergence in 58 tasks; 61 graphs/338 tasks/6760 states T-P1c killed: 81-point grid over τ ∈[0,4] fails everywhere, min cos=−1.0 (both strong form and per-state τ *), median τ *=0 (frac_at_tau0=0.8598) Three limiting-state per-path average dissipation 0 /3.63 /0.358 (energy units)
deposon_v20_gt7.json → per_graph (corr −0.87 is the document-level summary value (GT_RECONSTRUCTION §7), not a direct field read) deposon_v20_crossval.json → gt2_verdict, gt2_ attacker_meta.*.evasion_rate=1.0 deposon_v20_gt2b.json → verdict, per_T.*.per_ domain.*.accuracy deposon_v20_gt3.json → verdict.H_GT3_ supported=true, top-level kendall_W=1.0 deposon_v20_gt8.json → verdict.verdict="supports_H_GT8", per_pair.* deposon_v20_gt8b.json → domain.*.named_summary, field_named
gt8b_verdict, per_ prior_named_minus_
deposon_v20_gt8c.json → gt8c_verdict, domain.*.named_summary, backend.*
per_
deposon_v21_gtformal.json (seed=210021) → verdict.T_P1b, verdict.P1a_deviation, verdict.T_P2, verdict.T_P3, residuals.dag, residuals.cyclic, n_graphs/n_tasks/n_states
deposon_v22_p1c.json (seed=210021) → verdict, descriptive.tau_star_per_state, tau_grid
deposon_benchmark_v1_3_traps.json → variant_ results.v1_blocking,v2_tunneling,unified.avg_ ether_dissipated
19
References [1] Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., and Zhou, D. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS, 2022 arXiv:2201.11903. [2] Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. NeurIPS, 2023 arXiv:2305.10601. [3] Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D. Self-Consistency Improves Chain of Thought Reasoning in Language Models. ICLR, 2023 arXiv:2203.11171. [4] Turpin, M., Michael, J., Perez, E., and Bowman, S. R. Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. NeurIPS, 2023 arXiv:2305.04388. [5] Lanham, T., Chen, M., Gopalan, A., Jia, R., Forde, C., Ewart, T., and Hadfield-Menell, D. Measuring Faithfulness in Chain-of-Thought Reasoning. TMLR, 2023 arXiv:2307.13702. [6] Raji, I. D., Smart, A., White, R. N., and Mitchell, M. Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic Auditing. FAccT, 2020 arXiv:2001.00973. [7] Jia, H., Yaghini, M., Reber, A., Hovsepian, S., Cherubin, G., and Traynor, P. Proof-ofLearning: Definitions and Practice. IEEE S&P, 2021 arXiv:2103.05635. [8] Nasr, M., Hayes, J., Steinke, T., Balle, B., Tramèr, F., Jagielski, M., Carlini, N., and Terzis, A. Tight Auditing of Differentially Private Machine Learning. USENIX Security, 2023, pp. 1631–1648. [9] Bowman, S. R. The Dangers of Underclaiming: Reasons for Caution When Reporting How NLP Systems Fail. ACL, 2022 arXiv:2110.08300. [10] Lipton, Z. C., and Steinhardt, J. Troubling Trends in Machine Learning Scholarship. Communications of the ACM, 2019. [11] Dodge, J., Gururangan, S., Card, D., Schwartz, R., and Smith, N. A. Show Your Work: Improved Reporting of Intermediary Results in Empirical NLP. EMNLP, 2019 arXiv:1904.02624. [12] D’Amour, A., Heller, K., Moldovan, D., Adlam, E., Alipanahi, B., Beutel, A., and others. Underspecification Presents Challenges for Credibility in Modern Machine Learning. Journal of Machine Learning Research, 2022 arXiv:2011.03395. [13] Yao, J., and others. KG-LLM: A Framework for Knowledge Graph Construction from Large Language Models. ICASSP, 2025. [14] Wei, Y., Huang, Q., Zhang, Y., and Kwok, J. KICGPT: Large Language Model with Knowledge in Context for Knowledge Graph Completion. Findings of EMNLP, 2023, pp. 8667– 8683 arXiv:2402.02389. [15] Wadhwa, S., Amir, S., and Wallace, B. C. Investigating the Use of Large Language Models for Knowledge Graph Construction from Text. ACL Workshop, 2023.
20
[16] Ruiz-Primo, M. A., and Shavelson, R. J. Problems and Issues in the Use of Concept Maps in Science Assessment. Journal of Research in Science Teaching, 1996. [17] Novak, J. D., and Cañas, A. J. The Theory Underlying Concept Maps and How to Construct and Use Them. Technical Report IHMC CmapTools, 2008. [18] McPherson, M., Smith-Lovin, L., and Cook, J. M. Birds of a Feather: Homophily in Social Networks. Annual Review of Sociology, 2001, vol. 27, pp. 415–444. [19] Zheng, X., Liu, Y., Pan, S., Zhang, M., Jin, D., and Yu, P. S. Graph Neural Networks for Graphs with Heterophily: A Survey. arXiv preprint, 2022 arXiv:2202.07082. [20] Monderer, D., and Shapley, L. S. Potential Games. Games and Economic Behavior, 1996. [21] Rosenthal, R. W. A Class of Games Possessing Pure-Strategy Nash Equilibria. International Journal of Game Theory, 1973. [22] Sandholm, W. H. Population Games and Evolutionary Dynamics. MIT Press, 2010. [23] Candogan, O., Bimpikis, K., and Ozdaglar, A. Flows and Decompositions of Games: Potential, Harmonic, and Decomposition Games. Mathematics of Operations Research, 2011. [24] Koutsoupias, E., and Papadimitriou, C. Worst-Case Equilibria. STACS, 1999. [25] Roughgarden, T., and Tardos, É. How Bad Is Selfish Routing?. Journal of the ACM, 2002. [26] Benita, F. On the Price of Anarchy of the Nash Equilibrium for Atomic Splittable Routing. Operations Research Letters, 2020. [27] Christodoulou, G., Koutsoupias, E., and Spirakis, P. G. On the Performance of Approximate Equilibria in Congestion Games. Algorithmica, 2014. [28] Gröndahl, T., Pajola, L., Juuti, M., Conti, M., and Asokan, N. All You Need Is "Love": Evading Hate Speech Detection. AISec Workshop, 2018 arXiv:1808.09115. [29] Hosseini, H., Kannan, S., Zhang, B., and Poovendran, R. Deceiving Google’s Perspective API Built for Detecting Toxic Comments. arXiv preprint, 2017 arXiv:1702.08138. [30] Kahu, S. Y., and Ahuja, S. A Survey on Evasion Attacks Against Text Classifiers. ACM Computing Surveys, 2025. [31] HateBench Team. HateBench: A Benchmark for Hate Speech Detection. USENIX Security, 2025. [32] Jain, N., and others. Adversarial Attacks on NLP Models: A Survey. ACM Computing Surveys, 2023 arXiv:2304.08583. [33] Tramèr, F., Kurakin, A., Papernot, N., Goodfellow, I., Boneh, D., and McDaniel, P. Ensemble Adversarial Training: Attacks and Defenses. NeurIPS, 2020 arXiv:1705.07204. [34] Blume, L. E. The Statistical Mechanics of Strategic Interaction. Games and Economic Behavior, 1993.
21