Conceptio › Archive › arXiv CS
arXiv CSopen access

A Memorization Floor for LLM Refinement of Decompiled Code

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
software-architecturesoftware-engineeringtesting
software engineering, software architecture, testing

A Memorization Floor for LLM Refinement of Decompiled Code Muhammad Asjad1

arXiv:2609.17236v1 [cs.SE] 15 Sep 2026

1

School of Electrical Engineering and Computer Science (SEECS), National University of Sciences and Technology (NUST), Islamabad, Pakistan.

Contributing authors: [email protected]; Abstract We introduce a memorization floor : a within-item control separating what LLM refinement of decompiler output recovers from its input from what it recovers from its prior. Refine a function, then refine it again from an input whose identifiers have been destroyed, and measure what survives. Because the comparison is within-item, corpus difficulty cannot contribute; it costs twenty API calls. Applied to functions written after our analysis plan was committed, so no released model could have memorized them, it reports two things. Recovery is real: refined output sits +0.072 to +0.137 above an arm-matched permutation null built from its own output vocabulary. But it does not depend on the input we ablate: destroying the input’s dataflow changes the naming gain by +0.001 (95% CI [−0.026, +0.026]), and removing type prefixes or permuting names changes it by no more. A second refiner from another vendor, registered in advance and given byte-identical inputs, reproduces this — twelve contrasts, two models, twelve nulls. Readability stays at ceiling throughout, so a reader is given no signal. The null is bounded, not absolute: contributions under 0.056 are invisible, and the ablation leaves operations intact, so naming from those alone remains a competing reading. No registered hypothesis was confirmed, and we report the five instrument failures behind that in full, including a reassembly harness biased against the treated arm and an equivalence checker we registered without checking it worked on our inputs. Keywords: Decompilation, Large language models, Pre-registration, Memorization, Measurement validity, Reverse engineering

1

1 Introduction Decompiled C is hard to read. A stripped binary run through Ghidra yields functions named FUN 00101169, parameters named param 1, locals named iVar1, and types named undefined4. A large language model asked to refine that output produces something that looks like source code: descriptive names, explanatory comments, plausible types. The transformation is dramatic, and close to saturated — a blind judge rates raw Ghidra output 1.46 out of 5 for usefulness and the refined output 4.89. This is what makes refinement dangerous. Refined output carries no marker distinguishing a name recovered by analyzing dataflow from a name recalled because the function’s shape is familiar. Both arrive as confident, well-commented C. A reverse engineer has no way to tell which they are reading, and the readability gain actively encourages trust. Readable-but-wrong is the most dangerous output a reverse engineer can receive, and readability metrics by construction cannot detect it. The concern is concrete. Decompilation benchmarks derived from HumanEval are documented as contaminated in code-LLM training data (Riddell et al. 2024), and the functions used to evaluate refinement are overwhelmingly canonical: Fibonacci, binary search, matrix multiply. In an earlier pilot of this pipeline we watched the refinement pass recover the identifier n for the parameter of a recursive Fibonacci function. That is either competent reverse engineering or recall of a memorized idiom cued by a recognizable shape, and no aggregate metric can separate the two — because contamination and analysis predict the same aggregate. A model that has memorized a reference solution can emit correct, readable C from a maximally mangled -O3 input without performing any repair at all, and since higher optimization degrades the input more, memorization contributes proportionally more exactly where the measured effect is largest. Contamination does not merely inflate such an effect; it can manufacture it.

The control this needs, and what it shows. We introduce a within-item ablation we call a memorization floor. Refine a function normally; then refine it again from an input whose identifiers have been destroyed, and measure how much of the naming gain survives. Because the comparison is withinitem, corpus difficulty cannot contribute to it by construction. It costs twenty API calls. Applied to functions written after our analysis plan was committed — so that no released model could have memorized them — the floor reports two things at once, and both matter. 1. Recovery is real. Refinement recovers identifier signal well above an arm-matched chance baseline built from its own output vocabulary: +0.072 on novel functions and +0.110 on textbook ones for the first refiner, +0.105 and +0.137 for a second refiner from a different vendor. Whatever refinement is doing, it is not producing noise that happens to score. 2. But it does not depend on the input we ablate. Destroying the input’s dataflow changes the naming gain by +0.001 (95% CI [−0.026, +0.026]); removing Ghidra’s type prefixes or permuting every identifier changes it by no more. At the full corpus of each tier, none of the three input components has a detectable 2

contribution. A second refiner from a different vendor, given byte-identical inputs under a registration committed before any call, reproduces that on every contrast: twelve contrasts across two models and two tiers, twelve nulls. Readability stays near ceiling under every condition, so a reader is given no signal that anything has changed. That combination — real recovery, no measurable dependence on the ablated input, and undiminished apparent quality — is a property this evaluation paradigm has to control for, and the floor is how we propose controlling for it.

Contributions. 1. The memorization floor, a within-item ablation that separates input-driven recovery from prior-driven recovery, immune to corpus-difficulty confounds by construction and costing twenty API calls (§3.2). We give its manipulation check, without which an ablation is a claim about information that has not been measured. 2. Refinement recovers real signal above chance on uncontaminated functions, replicated across two vendors on byte-identical inputs (§4.1, §4.3). The margin is modest and we report it as such. 3. A bounded null on input dependence: twelve contrasts, two models, two tiers, twelve nulls, with the detection bound stated rather than implied (§4.2). 4. Arm-matched chance baselines for generative output. A null built from reference-vs-reference pairings is the wrong null for any generated arm; ours sat 0.074 too high and made real signal read as chance. The correction costs zero API calls (§3.3). 5. Execution-specific precision. A registered re-execution from byte-identical inputs reproduced every verdict while inflating the null’s detection bound by 44% — a distinction nobody in this literature reports, because nobody re-runs (§5). What this paper does not claim. No registered hypothesis was confirmed. The readability-gradient hypothesis was refuted; three correctness hypotheses became untestable because an instrument we registered did not discriminate on our inputs; and the two primary hypotheses are undecided at the sample sizes the corpus permits. §5 reports the five instrument failures behind those outcomes in full, including two that were ours rather than the instruments’. We put them in one place, late, because they are the part most likely to transfer and the part least likely to be read if scattered.

2 Related Work Decompiler fidelity. Decompiled C is systematically degraded relative to source, and not merely cosmetically. Dramko et al. (2024) give the canonical taxonomy: 15 top-level categories of fidelity defect identified by open coding across four decompilers. They also observe that establishing correspondence between source and decompiled code cannot be done with line numbers or naive heuristics, because decompilers split expressions, inline 3

statements, and drop code — an observation that motivates the alignment-based instrument we adopt. Cao et al. (2024) report that Hex-Rays output recompiles at 30–50% while other decompilers manage under 6%, with many failures semantically rooted, and argue on that basis that recompile-and-run methods alone are unsuitable for assessing semantic accuracy. D-Helix (Zou et al. 2024) shows Ghidra is sometimes outright wrong, flagging 4,515 incorrectly decompiled functions among roughly 93K and surfacing 17 previously unknown decompiler bugs. Together these establish both the premise — Ghidra output leaves real repair work — and a constraint on evaluation: Ghidra’s own output cannot serve as ground truth.

Optimization as a fidelity variable. It is important to be specific about which axis degrades. Structural fidelity clearly does: SAILR (Basque et al. 2024) traces spurious goto emission to roughly nine specific compiler transformations, most introduced at -O2, and finds structural damage at every level, with 17% of spurious gotos present even at -O0; Zhou (2025) corroborates the mechanism cross-language in Rust. Semantic fidelity does not measurably degrade: Cao et al. (2024) evaluate four decompilers across -O0–-O3 and -Os and find they ”all demonstrate certain robustness across compiler optimization, not showing a significant drop in effectiveness with the optimization level increasing.” Consistent with this, Agent4Decompile’s raw-Ghidra re-executability baseline rises from 18.9% at -O0 to 27.5% at -O3 (Zhang et al. 2026). Every degradation claim in this paper is therefore scoped to the structural axis. SAILR also cautions against equating goto-elimination with quality, since original source contains gotos — a warning that applies directly to refinement instructed simply to ”improve readability.” LLM refinement. The pipeline shape we study — sound lifting by a traditional decompiler, then LLM repair — has an established lineage. DecGPT (Wong et al. 2023) introduced the compiler-in-the-loop pattern over IDA Pro output, raising recompilation success from 45% to 75%. DeGPT (Hu et al. 2024) targets Ghidra with a three-role architecture, reporting a 24.4% reduction in measured cognitive burden; its MSSC component symbolically checks that each refinement preserves value behavior and rejects edits that do not, machinery whose existence is itself evidence that unconstrained refinement introduces semantic drift, a failure mode also targeted by feedback-driven multi-turn refinement (Liu et al. 2026). LLM4Decompile-Ref (Tan et al. 2024) is the bestresourced instance, and its framing — Ghidra mangles syntax but preserves underlying logic — is the load-bearing assumption of the whole paradigm. D-LiFT (Zou et al. 2025) contributes the metric structure we adopt: readability credit is awarded only to output that first passes an accuracy gate. Recent work sharpens the tension. CoDe-R (Zhang and Li 2026) identifies degraded control flow at -O3 as the mechanism that limits refinement, while reporting that its own method still improves over baseline there. Agent4Decompile (Zhang et al. 2026) reports a refinement-baseline table over identical Ghidra input in which singlepass refinement’s re-execution advantage of roughly +24 points at -O0 becomes a deficit of −5 at -O3; its own iterative multi-agent method narrows without reversing; 4

CodeInverter (Liu et al. 2025) reports the same high-at--O0, low-at--O3 gradient. Our enhancement arm is single-pass, so the reversing row is the matching comparator. Meanwhile the readability tables in Decompile-Bench (Tan et al. 2025) trend the opposite way: Ghidra’s R2I score falls from -O0 to -O3 while LLM systems’ scores hold or rise. That table must be cited with a caveat we verified against the paper: no system it evaluates refines Ghidra or IDA output — the LLM systems are end-toend neural decompilers consuming assembly, so it is cross-architecture evidence, not same-shape evidence.1

Symbol recovery and measurement. Identifier recovery is the sub-problem our task generalizes, from DIRE (Lacomis et al. 2019) and VarBERT (Banerjee et al. 2021), which reports up to 84.15% sourceidentical variable-name prediction, to ReSym (Xie et al. 2024). Because the readability of decompiled code is to a large degree the readability of its identifiers, we report naming and structural gains separately. For readability, R2I (Eom et al. 2024) is the only metric purpose-built for decompiled code; it is relative and cannot score standalone output, which is why we use a rubric-based judge and report its limitations rather than claiming a validated absolute scale. For semantic equivalence we adopt codealign (Dramko et al. 2025), which computes an equivalence alignment over SSA form and control-dependence information and yields a graded measure rather than a binary verdict — pass/fail execution cannot detect drift on untested paths. §5.3 reports that this instrument did not work on our inputs.

Contamination. HumanEval-derived decompilation benchmarks are documented as contaminated (Riddell et al. 2024). To our knowledge no prior refinement evaluation separates memorization from repair by design. That gap is what this paper’s floor design addresses.

Human factors. Readability gains must matter to analysts or the metric measures something without stakes. Yakdan et al. (2016) provide the strongest evidence, with user-study participants solving 3× more tasks under readability-oriented transformations; observational studies of reverse-engineering practice (Votipka et al. 2020) ground rubric design. Our judge is an LLM, not a human, and we do not claim otherwise.

1 We checked this directly against the Decompile-Bench paper rather than inheriting it from a secondary summary, because it is an unusual and load-bearing correction to how that table is normally read. LLM4Decompile-End and LLM4Decompile-DCBench consume assembly; Ghidra and IDA appear as standalone traditional-decompiler baselines, never as an input stage feeding an LLM. The full audit, including the two claims we could not verify and dropped, is in docs/CITATION VERIFICATION.md in the artifact. The Cao et al. (2024) robustness quotation in this section was verified the same way, against the paper’s own contributions list.

5

3 Method Two roles for language models, kept separate. Language models appear here in two unrelated capacities. As the object of study they are the refiner and the judge: what they produce is the data, and every detail of how they were invoked is specified below. As an authoring aid they were used during manuscript preparation and in developing analysis scripts, including adversarial review passes over near-final drafts; where such a pass changed a result or its interpretation, the change is recorded with its date and reasoning in the project log and reported here in the same terms as every other correction. The author directed the work, verified every reported figure against the committed artifacts mapped in NUMBERS.md, and is accountable for the content and the conclusions.

3.1 Pipeline and corpora Each corpus item is one self-contained C program: a target function plus a main() that is its self-test, printing PASS and exiting 0 when behavior is correct. Items are compiled with gcc 13.3.0 on x86-64 at -O0, -O1, -O2, -O3 and -Os, stripped, and decompiled with Ghidra 12.1 headless. Each decompiled target function is then refined by a single stateless API call to Claude Sonnet 5 under a fixed prompt template, which asks the model to replace generic Ghidra identifiers with descriptive names, restore types where recoverable, and add explanatory comments. Both arms — raw Ghidra and refined — are evaluated identically. Sonnet 5 is the refiner throughout except in §4.3, which repeats the floor design against a second model from a different vendor. Naming similarity compares recovered identifiers against ground truth, aligned by declaration order, scored by exact-match rate and by cosine similarity of all-MiniLM-L6-v2 embeddings; §3.3 establishes what a point of it is worth. Readability is scored by a blind, randomized-order LLM judge (Claude Haiku 4.5, temperature 0) on a four-part 1–5 rubric. Functional correctness is measured by reassembling each arm’s functions into one translation unit, recompiling and running the self-test.

Two tiers, and why the second exists. T1 is 25 hand-written textbook functions — the kind of corpus this literature evaluates on, and the kind a released model has plausibly seen. T3 is 20 functions written after the analysis plan was committed and kept unpublished until this paper, so that no model trained before their release could have memorized them. Their pre-run existence is verifiable against a hash manifest committed before the run. T3 is the tier that carries every claim about uncontaminated recovery; T1 is the comparison. Tiers are never pooled: no table, figure or abstract number in this paper averages the two. Because the tiers must be comparable on difficulty for the cross-tier comparison to mean anything, T3 was built to a difficulty profile matched against T1 on control-flow depth, loop nesting and operation mix (Table 1). 6

Property

T1 (n=25)

T3 (n=20)

diff

overlap

p

statement count cyclomatic complexity max loop-nesting depth number of parameters number of local variables distinct identifiers source LOC

6.2 ± 4.1 4.6 ± 3.2 0.7 ± 0.5 2.5 ± 1.8 1.1 ± 1.0 6.6 ± 3.4 12.4 ± 6.8

5.0 ± 1.4 3.5 ± 1.2 0.7 ± 0.5 1.9 ± 0.6 1.5 ± 0.8 6.7 ± 1.7 9.1 ± 2.9

−1.2 −1.1 −0.0 −0.7 +0.4 +0.1 −3.4

1.00 0.90 1.00 1.00 1.00 1.00 1.00

0.94 0.46 0.90 0.60 0.09 0.65 0.15

Table 1 Difficulty-matching outcome (mean ± SD). “Overlap” is the fraction of T3 items inside T1’s observed range; p is Mann–Whitney, two-sided. No property separates the tiers.

3.2 The memorization floor The floor is the disambiguating design, and it is within-item, so cross-tier difficulty differences cannot contribute to it by construction. For each item we re-run refinement on an ablated input in which Ghidra’s synthetic identifiers have been replaced. Two transformations answer to that description and they are not the same experiment; we state both precisely, because the distinction is the difference between a result and a null.

Alpha-renaming. Build a single bijection over the set of synthetic identifiers appearing in the function and apply it at every occurrence: draw a permutation π of the identifier set once, and rewrite every occurrence of v as π (v ). Only the identity permutation is rejected, so individual names may map to themselves. Every def-use edge survives — the result is the same program under a consistent renaming, with Ghidra’s typeencoding prefixes destroyed. Scrambling. For each occurrence of each synthetic identifier, draw a replacement uniformly at random from the identifier set, independently per occurrence. The same original name now maps to different targets at different sites, which is what severs def-use chains. Three properties of both transforms, stated because Figure 1 exhibits all of them. The identifier set is everything matched by the extraction pattern — variables, parameters, and function and data symbols — so a function symbol can replace a variable. Substitution is textual and applies at every occurrence including declarations and the parameter list. And draws cross prefix classes by design; the paragraph on type preservation below records the measurement that forced that choice.

The manipulation check, without which this is not a measurement. An ablation is a claim about what information was removed, and a claim about information is a measurement. Define a co-reference edge as an unordered pair of identifier occurrences naming the same value in the intact input; every def-use edge is a coreference edge, so a transform preserving few co-reference edges cannot preserve many 7

Intact Ghidra output char cVar1; undefined4 local_10; undefined4 local_c; local_10 = 0; for (local_c = 0; *(char *)(param_1 + local_c) != ’\0’; local_c = local_c + 1) { cVar1 = *(char *)(param_1 + local_c);

Alpha-renamed (what Run D executed): still a coherent loop char local_c; undefined4 local_10; undefined4 param_1; local_10 = 0; for (param_1 = 0; *(char *)(cVar1 + param_1) != ’\0’; param_1 = param_1 + 1) { local_c = *(char *)(cVar1 + param_1);

Scrambled (corrected transform): references severed char param_1; undefined4 local_10; undefined4 FUN_00101169; local_10 = 0; for (local_10 = 0; *(char *)(local_10 + local_10) != ’\0’; cVar1 = FUN_00101169 + 1) { cVar1 = *(char *)(cVar1 + param_1);

Fig. 1 The two transformations on real Ghidra output (04 count vowels, -O0). Under alpharenaming the loop counter is consistently param 1 and the loop still reads as a loop: an analyst — or a model — can follow it. Under scrambling the induction variable, its guard and its increment name three different values, and there is no chain left to follow. Both draws exhibit the properties stated in §3.2: the function symbol FUN 00101169 replaces a variable (function symbols are in the draw set), param 1 reappears as a local declaration (substitution covers declarations), and local 10 is a fixed point of the alpha-renaming permutation (only the identity permutation is rejected).

Transform Alpha-renaming (as executed) Scrambling, type-preserving Scrambling, cross-class (registered)

def-use survival

chance

1.000 0.784 0.284

— 0.797 0.324

Table 2 Manipulation check over the 10 Run D item-levels. The executed transform preserved every edge; the corrected one sits at chance.

def-use edges. The check is the fraction of those edges whose two occurrences still carry the same name after the transform. Table 2 gives the result: alpha-renaming preserves every edge, as a bijection must, and the registered cross-class scramble sits at chance. We recommend this check as mandatory rather than optional. It cost forty lines over an existing parser, and §5 reports what happened in its absence. 8

Type preservation, decided by measurement. Ghidra’s prefixes encode type and storage: iVar (signed int), uVar (unsigned), pcVar (char pointer), local (stack), param (argument). Drawing replacements only within a prefix class would preserve that information and ablate dataflow alone, which is the cleaner experiment. On this corpus it does not work: these functions carry 2–8 distinct synthetic identifiers over 2–4 classes, so the modal class is a singleton and admits no substitution. Type-preserving scrambling leaves 0.784 of def-use edges intact and, on 6 of 10 item-levels, leaves all of them. We therefore draw across classes, and state the consequence rather than hide it: the ablation removes type and storage information along with dataflow. What the ablation does not remove, and how that bounds the claim. Scrambling rewrites identifiers. It leaves the operation vocabulary and the control structure completely intact — dereferences, comparison guards, arithmetic, call structure and literal constants all survive verbatim, and so does the shape of every loop and branch. Figure 1’s scrambled listing is still visibly a loop walking a pointer to a NUL byte. The floor is therefore a floor against “the input carries no identifier coreference and no type or storage prefixes”, not against “the input carries no analyzable structure”. This leaves a competing non-memorization reading we cannot exclude: the model may be naming from surviving operations alone — inferring count from an incremented accumulator — in which case identifier and type information were never what it used, and removing them costs nothing without memorization entering. Distinguishing that needs a fourth condition perturbing operations while preserving names, which we did not run and register as future work rather than gesture at. What the present design establishes stands either way: the two information sources most likely to carry a memorized idiom’s surface can both be destroyed with no measurable effect. Three conditions, not two. Because alpha-renaming destroys type prefixes while preserving dataflow, keeping it as a third condition turns a two-way comparison into a decomposition: Condition

dataflow

type prefixes

Intact Alpha-renamed Scrambled

preserved preserved destroyed

preserved destroyed destroyed

Two contrasts follow and answer different questions. Intact minus alpharenamed is what Ghidra’s type prefixes buy with reference structure held fixed. Alpha-renamed minus scrambled is the clean estimate of the dataflow contribution: both conditions carry equally arbitrary names and differ only in whether the references cohere. Formally, per tier: ∆intact − ∆alpha = the type-prefix contribution 9

∆alpha − ∆scramble = the dataflow contribution ∆scramble = prior-matching alone. A scrambled-arm gain that stays high means the model’s default naming prior happens to match ground truth. A contribution near zero means the removed information bought nothing.

3.3 What a cosine point is worth A number like +0.060 is uninterpretable without a scale, and short identifier embeddings are not well spread. We fix reference points before reporting any result. Chance — which must be arm-matched. The operative null for each arm is a permutation null over that arm’s own output : its recovered names for item i scored against the ground truth of items j ̸= i, 1,000 resamples, seed 0. This matters more than it sounds. Our first chance line paired each ground-truth identifier with one drawn from a different item’s ground truth — two arbitrary real C identifiers score 0.306, nowhere near 0 — but no observed arm pairs ground truth with ground truth, and model-generated names do not embed like source-idiomatic ones. The arm-matched lines sit substantially lower (Table 3), and the difference is not cosmetic: it is what separates “at chance” from “clearly above chance” for the refined arm. §5 reports how we found this. A useful ceiling. Each ground-truth identifier is paired with a hand-written acceptable synonym (count/counter, sum/total), written before inspecting any model output. These score 0.428 on T1. This is the top of the useful range, not of the scale: string identity would score 1.0, but a reviewer would accept any of these synonyms. The ceiling shares the vocabulary asymmetry the chance line had — it pairs short source-idiomatic ground truth with short source-idiomatic synonyms, and the refined arm’s compound names may be unable to reach it even when semantically correct — so the band widths below are, if anything, too wide, which makes every “fraction of the band” claim in this paper conservative. Every delta in this paper should be read against a useful band of roughly 0.17–0.20, not against 1.0.

3.4 Pre-registration The analysis plan was committed before any API call; git timestamps are the registration and the full dated log of every deviation, correction and superseded number ships with the artifact. Three commitments constrain what follows. Fix rules were written to the log before the code implementing them. Superseded numbers are never overwritten: where an extension changed a conclusion, both are reported side by side. And instrument parameters are frozen at their documented defaults, so that no setting can be chosen after seeing which value flatters a hypothesis. Two honest qualifications. This is a pre-registration, not a Registered Report: no reviewer saw the plan before the data existed, so it constrains us but was not externally validated. And the commit history was rewritten once, after all analysis was complete, to normalize authorship; content, commit count and every author timestamp 10

Reference point Raw Ghidra output vs. ground truth arm-matched chance, Ghidra’s names Refined output vs. ground truth arm-matched chance, refined names Chance, ground truth vs. unrelated ground truth Useful ceiling (accepted synonyms)

T1 (textbook)

T3 (novel)

0.239 0.229 0.367 0.257 0.306 0.428

0.222 0.212 0.309 0.237 0.311 0.435

Table 3 Reference scale for all-MiniLM-L6-v2 identifier similarity, computed per tier because the scale is a property of the corpus, not of the metric. The useful band for the refined arm — from its own chance line to the accepted-synonym ceiling — is 0.171 wide on T1 and 0.198 wide on T3, so every delta in this paper should be read against roughly 0.17–0.20 and not against 1.0. Each tier is summarised over its own optimization levels (T1: -O0–-Os; T3: -O0, -O3). Each arm carries its own chance line — that arm’s actual output vocabulary scored against ground truth from the wrong item (1,000 resamples, seed 0; registered before computation, producer eval/permutation null.py) — because model-generated and ground-truth identifiers embed differently, so a single ground-truthvs-ground-truth line (0.306/0.311, retained for comparison) is the correct null for no observed arm. Every arm is above its own chance line; the text below takes up how this corrects an earlier reading. The refined arm’s null on T3 has 95% interval [0.223, 0.253]. The two refined rows are the first refiner’s: “its own vocabulary” means the second refiner needs its own line too, and Table 4 gives it. The raw-Ghidra rows are shared by both, since that arm does not depend on which refiner ran. The synonym ceiling covers only 47% of ground-truth names (39 had no hand-written synonym) and is correspondingly less reliable than T1’s, which covers all 87. T3’s refined figure is from the registered re-execution; the raw-Ghidra figure is deterministic from the committed decompilations.

are unchanged, an old-to-new hash mapping ships with the artifact, and the Zenodo deposit provides third-party timestamping the rewrite cannot touch.

3.5 A second refiner A floor measured on one model bounds that model. Whether the result is a property of the paradigm or a quirk of one refiner is not a question one refiner’s data can answer, so we ran the entire floor design again against gemini-3.6-flash: three conditions × two levels × 45 items = 270 calls at temperature 0.

Its registration status, which is weaker than the rest of the paper’s. The four research questions were fixed before any data existed. This extension was not: it was registered after the first refiner’s results were known, in the project log, before any code for it was written and before any call was made. It is a registered replication, and we label it that way throughout rather than let it inherit the original plan’s standing. What the registration fixes is what it always fixes — the design, the analysis, and a commitment to report either outcome — against a result already known. 11

Tier

Arm

observed

arm-matched null

margin

T1 T1 T1

raw Ghidra (shared) refined, Sonnet 5 refined, Gemini

0.239 0.367 0.392

0.229 0.257 0.255

+0.010 +0.110 +0.137

T3 T3 T3

raw Ghidra (shared) refined, Sonnet 5 refined, Gemini

0.222 0.309 0.328

0.212 0.237 0.223

+0.010 +0.072 +0.105

Table 4 Identifier recovery against each arm’s own permutation null (1,000 resamples, seed 0). The raw-Ghidra rows are shared by construction — that arm does not depend on which refiner ran, and both received identical inputs — so their exact agreement is a consistency check on the cross-refiner apparatus rather than a finding.

Byte-identical inputs, which is what makes this matched pairs. Both refiners received exactly the same decompiler output. This is checkable rather than asserted: both runs’ raw request snapshots are committed, and all 40 T3 prompts sent to the second refiner are byte-for-byte identical to those sent to the first. The Ghidra arm is therefore shared between the two analyzes — literally the same numbers — so a cross-refiner difference cannot be a difference in inputs, and the comparison is paired at the item level rather than being two independent samples. Each model at its own floor, not a common setting. Sonnet 5 rejects sampling parameters outright, so its variance control is thinking disabled; Gemini accepts temperature and is set to 0. For inference-time compute, Sonnet is disabled and Gemini is at its lowest accepted level, which is not the same thing. Every call records its thinking-token count and all 270 report zero, so the claim is the measured one. We compare two refiners each at its own lowest available setting, and say so rather than describe the runs as identically configured.

4 Results Runs A (T1, five levels), C (T3), D/D3/D4/D5 (the floor, extended to both full corpora), a registered T3 re-execution and the 270-call second-refiner replication completed with zero call failures, at a total measured API spend of $6.06 against a $20 budget — the corpus, not the budget, is what caps every sample size here. Every figure maps to a committed producer via NUMBERS.md in the artifact.

4.1 Refinement recovers real signal, above chance The floor is a null, and a null is only informative if the thing being ablated was measurable to begin with. So we start with what refinement does recover. Each arm is scored against its own arm-matched permutation null (§3.3), recomputed per refiner from that refiner’s own output vocabulary — reusing the first refiner’s would repeat the mis-specification §5 reports. 12

Mean identifier cosine similarity

Refined, observed

Its own chance null

Raw Ghidra (shared)

0.4

0.2

0 T1 Sonnet

T1 Gemini

T3 Sonnet

T3 Gemini

Fig. 2 Refinement recovers identifier signal above chance in both models and on both tiers. Each refined bar is scored against its own arm-matched permutation null, recomputed from that refiner’s output vocabulary (§3.3); a null borrowed from another arm is the wrong comparison and would move these margins by up to 0.074. The raw-Ghidra bars are shared between refiners by construction — that arm does not depend on which model ran — so their agreement is a consistency check rather than a finding. The margins are +0.110 and +0.137 on textbook functions and +0.072 and +0.105 on functions written after the analysis plan was committed, against a useful band roughly 0.17–0.20 wide (Table 3). Producer: eval/permutation null.py.

Figure 2 and Table 4 give it. Both refiners clear their own chance lines on both tiers, and by margins that are modest rather than dramatic: the refined arm sits +0.072 to +0.137 above chance against a useful band roughly 0.17–0.20 wide, so refinement buys something in the region of half that band. The raw-Ghidra rows are shared by construction — that arm does not depend on which refiner ran, and both received identical inputs — so their exact agreement is a consistency check on the cross-refiner apparatus rather than a finding. Two things follow. On functions written after the analysis plan was committed, which no released model could have memorized, refinement still recovers identifier signal well above chance. And the tier ordering points the same way for both models: novel functions sit below textbook ones (+0.072 against +0.110; +0.105 against +0.137), the direction RQ4 predicted, observed twice on the same items. That is a different measure from the registered tier comparison in §4.4, which tests the naming gap and remains undecided; the agreement is directional, not numerical. Figure 3 shows the same picture per optimization level: the refined arm sits clearly above its own chance line everywhere, the raw arm barely above its own, and the gap between them does not widen with optimization.

4.2 The floor: what that recovery does not depend on Having established that there is signal to ablate, we ablate it. At the full corpus of each tier, none of the three input components has a detectable contribution on either tier — six contrasts, six nulls. Destroying dataflow and type prefixes together leaves the naming gain statistically unchanged, and so does each component alone. 13

Mean identifier cosine similarity

0.4

0.3

Refined Raw Ghidra Refined arm’s chance null Ghidra arm’s chance null

0.2

O0

O1

O2

O3

Os

Compiler optimization level

Fig. 3 Identifier recovery by optimization level on T1 (n = 25 per level), measured as mean cosine similarity of recovered identifiers to ground truth. Dotted lines are the arm-matched permutation nulls of §3.3 (0.2569 refined, 0.2285 raw), each built from that arm’s own output vocabulary scored against mismatched references. The superseded ground-truth-versus-ground-truth chance line of 0.311 would sit above the raw arm entirely, which is the error described in §3.3. The gap between arms is flat across levels

Contrast

registered n=5

pooled n=12

full corpus

+0.060 [+0.007, +0.132] −0.032 [−0.074, +0.012] +0.028 [−0.037, +0.085] 0.101 0.071

+0.054 [+0.007, +0.103]∗ −0.051 [−0.153, +0.069] +0.003 [−0.082, +0.095] 0.133 0.167

−0.005 [−0.052, +0.039] −0.003 [−0.064, +0.064] −0.008 [−0.053, +0.043] 0.071 0.094

T3 (novel), full n=20; re-execution throughout type prefixes +0.049 [+0.014, +0.084] dataflow −0.009 [−0.048, +0.019] both +0.040 [−0.000, +0.088] realized MDE, “both” 0.071 realized MDE, “dataflow” 0.055

+0.031 [+0.005, +0.056]† −0.022 [−0.052, +0.007] +0.008 [−0.031, +0.044] 0.056 0.044

+0.019 [−0.004, +0.042] −0.020 [−0.049, +0.008] −0.001 [−0.034, +0.030] 0.047 0.041

T1 (textbook), full n=25 type prefixes dataflow both realized MDE, “both” realized MDE, “dataflow”

Table 5 The floor at the full corpus, beside the registered and pooled analyzes rather than in place of them. Selection was exhaustive — every item not already in the n = 12 subset — and registered before the run. Manipulation check on the added item-levels: alpha-rename def-use survival 1.000 exactly on all 42, scramble 0.191 (T1) and 0.187 (T3) against chance rates of 0.182 and 0.153. All six contrasts are null at the full corpus. ∗ Wilcoxon p = 0.021 at n = 12, p = 0.65 at n = 25. † Wilcoxon p = 0.034, the re-execution’s nominally significant T3 prefix contrast; p = 0.123 at n = 20. The T3 block is the re-execution at all three n, so its n = 5 and n = 12 columns are the C1R figures of docs/RERUN COMPARISON.md rather than the recorded ones recorded at run time — the recorded T3 outputs were destroyed and cannot be extended, so only the re-execution supports a like-for-like comparison across n.

The floor null is bounded rather than merely unmeasured. On T3, intact minus scrambled is −0.001 [−0.034, +0.030] with a realised detection bound of 0.047, so the full corpus excludes input contributions above roughly 24% of the refined arm’s band. We nonetheless claim the wider 0.056 bound of the n = 12 re-execution, for a reason §5 establishes: the only two executions we have of this measurement differed by 44% 14

Ground truth

Intact

Alpha-renamed

Scrambled

T3 01 zigzag reset accumulate, -O0 (re-execution) values (par.) param 1 arr arr count (par.) param 2 count count total running sum running total result index index i i — term term term T1 04 count vowels, -O0 (recorded) text (par.) str ptr str ptr count vowelCount vowel count i index index c currentChar current char

str vowel count index current char

Fig. 4 What the model produced for the same function under all three input conditions, quoted verbatim from enhanced.c in declaration order. Items were selected by rule, not by inspection: the T3 item is the first in sorted filename order with all three arms present in the re-execution trees, and the T1 item is Figure 1’s exemplar, so the reader sees the same function whose scrambled input appears there. The floor result is visible directly. On T3 the ablated arms name the parameters (arr, count) that the intact arm leaves as param 1 and param 2, and the scrambled arm recovers count and a plausible result/i pair from an input with every co-reference severed. On T1 all three conditions produce essentially the same four names. Neither tier shows the degradation under ablation that a dataflow-driven account predicts. The T3 item reads better than its tier: two exact matches and three defensible synonyms in the scrambled column is well above the tier’s 3–5% exact-match rates. The selection rule was fixed before inspection and we did not swap the item it picked; the tier means, not this item, are the evidence. Producer: eval/output examples.py.

in precision, the full-corpus figure rests on a single execution, and we would rather quote a bound that survived a replication than a tighter one that has not faced one. Figure 4 shows the result at the level of individual identifiers rather than means, on two rule-selected items.

Two effects that did not survive the full corpus. Both of this paper’s nominally significant contrasts appear at intermediate sample sizes and disappear at the corpus maximum: T1’s type-prefix effect is +0.054 with p = 0.021 at n = 12 and −0.005 with p = 0.65 at n = 25, reversing sign; T3’s prefix contrast is +0.031, p = 0.034 on re-execution at n = 12 and +0.019, p = 0.123 at n = 20. Our registration for the extension committed in advance to reporting a changed conclusion as prominently as a confirmed one, so we state it plainly: both are gone, not merely unconfirmed. Neither was a registered primary contrast, and neither survives a Bonferroni correction across the six. Power, stated rather than implied. The registered design’s decision rule for the floor required Wilcoxon p < 0.05 from a signed-rank test on n = 5 pairs, whose minimum attainable p is 2/25 = 0.0625. The criterion was unreachable at the registered sample size regardless of the effect — a defect in our pre-registration, not a property of the data, and the reason the extensions to n = 12 and then to the full corpus were registered and run. Table 6 gives the power analysis. Any replication needs n ≥ 6 per tier for p < 0.05 to be 15

n per tier

projected MDE (T1)

projected MDE (T3)

min. attainable p

5 (registered) 6 10 12

0.101 0.092 0.071 0.065

0.046 0.042 0.032 0.030

0.0625 0.0312 0.0020 0.0005

25 / 20 (full corpus, realized)

0.071

0.047

< 10−6

Table 6 Power for the floor comparison, projected from the registered n = 5 variance. MDE is the paired minimum detectable effect at 80% power, α = 0.05 two-sided, from the per-item SDs observed at n = 5 (0.080 for T1, 0.037 for T3). Minimum attainable p for the signed-rank test is 21−n . The realized MDEs at n = 12 are worse than projected (0.133 for T1, 0.039 for T3), because the items added by the extension rule were more variable than the registered five — a reminder that power projections from a small pilot are optimiztic by construction. The final row is realized rather than projected: at the full corpus the “both” MDE is 0.071 on T1 and 0.047 on T3 (Table 5) — better than n = 12 on both tiers, but T1 is still short of the 0.065 the n = 5 variance predicted for n = 12 alone. The corpus caps n at 25 and 20, so these are close to the best this design can do and the tier comparison’s power ceiling stands where §6 puts it.

attainable at all. The corpus caps n at 25 and 20, so the realised bounds in Table 5 are close to the best this design can support; a materially tighter bound needs more items, not more calls.

4.3 The floor replicates across vendors The entire design was executed a second time against gemini-3.6-flash under the registration and caveats of §3.5. Table 7 reports both refiners side by side. Twelve contrasts across two refiners and two tiers, twelve nulls. No component of the input we can remove has a detectable contribution for either model. The result is not a property of one vendor’s model. The replication is the weaker of the two measurements, and reading it the other way round would be the natural error. Every one of the second refiner’s intervals is wider than the first’s and its realised bounds are larger on both tiers (0.098 against 0.071 on T1, 0.053 against 0.047 on T3). Two models agreeing does not tighten a bound; it adds a second, looser bound that happens to agree. The bound this paper claims is therefore unchanged — the first refiner’s 0.056 — and the second refiner corroborates the verdict, not the precision.

What we decline to conclude. The second refiner’s margins are larger and its T1 recovery higher (0.392 against 0.367). We do not read this as one model being better at reverse engineering. Two refiners at one prompt each, with no per-model prompt tuning and configurations that are each model’s own floor rather than a matched setting, do not support a ranking. What the second refiner is evidence for is the thing it was registered to decide, and on that it is informative precisely because the answer does not depend on which model is stronger. 16

Contrast (what is removed)

Refiner 1 (Sonnet 5)

Refiner 2 (Gemini)

T1 (textbook), n = 25 type prefixes dataflow both realized MDE, “both”

−0.005 [−0.052, +0.039] −0.003 [−0.064, +0.064] −0.008 [−0.053, +0.043] 0.071

−0.026 [−0.085, +0.034] +0.015 [−0.042, +0.079] −0.011 [−0.081, +0.057] 0.098

T3 (novel), n = 20 type prefixes dataflow both realized MDE, “both”

+0.019 [−0.004, +0.042] −0.020 [−0.049, +0.008] −0.001 [−0.034, +0.030] 0.047

+0.009 [−0.033, +0.062] −0.002 [−0.056, +0.051] +0.008 [−0.030, +0.042] 0.053

Table 7 The floor under two refiners, at the full corpus of each tier. Percentile bootstrap 95% CIs, 10,000 resamples, seed 0. Refiner 1’s column is Table 5’s, unchanged. The two columns reach the same statistics by different entry points into the same producer — the registered full-corpus path for refiner 1, an explicit-arm-files path for refiner 2, whose trees the former cannot name — and both delegate to the same bootstrap, Wilcoxon and MDE code. That the entry points agree is checked rather than assumed: driven through the explicit-arm path, refiner 1’s recorded arms reproduce its registered T3 column here on all nine statistics, to every stored digit. Both refiners received byte-identical inputs (§3.5), so the Ghidra arm is shared and the comparison is paired at the item level. All twelve contrasts span zero. Wilcoxon p for refiner 2: T1 0.36 / 0.84 / 0.72, T3 0.65 / 0.99 / 0.43.

Tier

n

mean per-item gap

sd

median

T1 T3

25 20

+0.1442 +0.0768

0.1495 0.0514

+0.1595 +0.0646

Table 8 H3-tier, common levels -O0 and -O3 only. Tiers never pooled.

4.4 Tier comparison, readability and correctness Table 8 reports the per-item gaps. The registered tier ordering holds — T3’s per-item naming gap is 53% of T1’s, a difference of +0.0674 — but neither deciding test clears 0.05: Kruskal–Wallis H = 3.093, p = 0.0786; Mann–Whitney U = 327.0, p = 0.0806. A percentile bootstrap CI on the difference is [+0.0066, +0.1290] and excludes zero. We report that tension and do not resolve it in the bootstrap’s favour: the plan names the two rank tests as deciding, and switching because the third clears the bar would be choosing the test after seeing the result. The verdict is undecided. Power caps what a null here could mean anyway — the minimum detectable effect is 0.100, about 69% of T1’s observed gap, and T1’s n = 25 floors it at 58% no matter how many T3 items are written.

Readability: the gradient hypothesis is refuted. H1-readability predicted the refined-versus-raw readability gap would widen with optimization, on the reasoning that worse input leaves more room to improve. It does not. Figure 5 is two near-horizontal lines separated by about 3.4 points at every level. 17

Judge overall usefulness (1–5)

5

Refined Raw Ghidra

4

3

2

1 O0

O1

O2

O3

Os

Compiler optimization level

Fig. 5 Blind LLM-judge overall usefulness by optimization level on T1 (n = 25 per level, ungated scores). The gap is large and near-saturated but does not widen with optimization, which is why H1readability is refuted rather than confirmed. §5.2 reports that a hypothesis-blind human rater agrees on the ordering while placing the gap 2.8× smaller, so the vertical distance here should be read as contested in magnitude and not as a validated scale

recompiles

test passes

Level

Ghidra

LLM

Ghidra

LLM

-O0 -O1 -O2 -O3 -Os

25 22 20 19 14

24 22 20 18 14

17 16 17 17 10

17 16 16 16 10

Total

100

98

77

75

Table 9 Functional correctness, n = 25 per level. 4 discordant pairs of 125.

Notably T3’s readability gain is undiminished (1.23 → 4.92) while its naming gain is roughly half T1’s — consistent with refinement’s readability effect being independent of familiarity while its identifier-recovery effect is not. This is the mechanism that makes the floor matter: the output looks equally good either way.

Correctness: a null we cannot interpret. With the reassembly harness corrected (§5), test-pass is 77/125 for the raw arm against 75/125 for the refined arm, with 4 discordant pairs. Table 9 gives the breakdown by level and Figure 6 the same data by arm. Four discordant pairs in 125 is not demonstrated equivalence; it is a near-degenerate axis with almost no power to detect a real effect of moderate size. We report it descriptively and draw no conclusion from it. 18

100

Raw Ghidra Refined

Test-pass rate (%)

80 60 40 20 0 O0

O1

O2

O3

Os

Compiler optimization level

Fig. 6 Functional correctness by optimization level on T1, as the percentage of the 25 items per level whose reassembled output recompiles and passes its own self-test. Unconditional rates: an item that fails to recompile counts as a failure rather than being excluded. The arms differ by a single item at -O2 and at -O3, and not at all at the other three levels, totalling 77 versus 75 of 125 with four discordant pairs. These are the corrected figures; §5.1 reports the harness bugs that had made this comparison read as a 20-point deficit against refinement

RQ

Hypothesis

Verdict

Basis

RQ4

H3-tier (primary)

Undecided

RQ3

H3-floor (primary)

Undecided

RQ1 RQ2 RQ2 RQ2

H1-readability H1-correctness H1-divergence H2

Refuted Untestable Untestable Untestable

Ordering holds (+0.144 vs +0.077) but Kruskal–Wallis p = 0.079, Mann– Whitney p = 0.081; CI spans the MDE Registered criterion unattainable at n = 5 Gap is flat across levels, not widening codealign I = 0.092 < 0.20 Requires H1-correctness Same instrument

Table 10 All registered hypotheses, in the plan’s order, with the research question each operationalizes. Run B (T2) was deferred, so the monotone three-tier ordering was never evaluated. RQ5 appears in no row: it was not registered, and no hypothesis was written for it.

5 Instrument Validity This study set out to measure refinement and spent most of its evidence measuring its own instruments. Five measurements we registered turned out to be measuring something other than what we intended. We report them together, here, rather than scattered through the results, because they are the part most likely to transfer to other pipelines and the part least likely to be read if distributed. Two of the five are failures of instruments; two are failures of ours; one is a property of the pipeline nobody in this literature measures. Table 10 states the consequence in the plan’s own terms. No registered hypothesis was confirmed. 19

Bug

Penalized the model for

Pairs

Reassembler could not link a renamed function extract prototype flattened the leading block comment, leaving it unterminated and swallowing following definitions Fence stripping left stray ‘‘‘ lines

obeying prompt instruction 1 (rename identifiers) obeying prompt instruction 3 (add comments)

25

emitting markdown

4

13

Table 11 Harness fragility is directional: it runs systematically against the behaviors refinement is asked to produce. Each bug damaged only the refined arm, and each did so because the model complied with a prompt instruction. None touched the control arm.

5.1 Harness bias is directional, and it favours the raw arm Our first correctness measurement showed the refined arm losing to raw Ghidra output by roughly twenty points. It was an artifact of our own reassembly harness, and the failure modes were not random: every one of them penalised the model for following the registered prompt. Table 11 lists them. The prompt asks for descriptive names, so the reassembler could not link a renamed function; it asks for explanatory comments, so a prototype extractor that flattened block comments truncated declarations; the model emitted markdown fences, which stripping missed. Each bug fires only when refinement does its job. Two further bugs surfaced later, one arm-neutral and therefore invisible to every between-arm check we had registered. The transferable point is the shape rather than the count. A reassembly harness sits downstream of exactly the behaviors refinement is asked to produce, so its failures are systematically anti-correlated with treatment success. A between-arm difference is the natural check and it cannot catch this; what caught ours was a manipulation check on absolute plausibility — asking whether a 20-point deficit was credible at all, rather than whether it was larger than zero. We recommend reassembly harnesses be treated as a primary suspect for correctness deficits in refinement studies, and that absolute plausibility be checked before any between-arm comparison is believed.

5.2 An LLM judge tracks a human’s ranking but not the size of the gap Table 12 gives the comparison. Against one human rater blind to our hypotheses, the judge ranks arms as the human does (ρ = 0.77; both prefer refinement in 18–20 of 20 pairs) but reports a gap 2.8× larger, and the disagreement sits almost entirely in how the two score raw decompiler output rather than refined output. Neither instrument is validated, so this is a measured divergence rather than a demonstrated judge error — but at least one of the two is wrong about that baseline, and rank correlation alone would have hidden it either way. If the human is closer to right, every comparison drawn against a decompiler baseline in this literature is inflated. LLM judges should be calibrated on effect size, not only on ranking. 20

mean rating refined

gap

Spearman

mean |diff|

Human rater (blind, n = 20 pairs) Naming clarity 2.05 4.60 Structural clarity 3.65 4.25 Comment usefulness 1.45 4.90 Overall usefulness 3.00 4.25

+2.55 +0.60 +3.45 +1.25

0.84 0.37 0.91 0.77

0.65 0.88 0.30 1.20

LLM judge, same 20 pairs Naming clarity 1.10 Structural clarity 3.15 Comment usefulness 1.00 Overall usefulness 1.30

+3.55 +1.35 +3.85 +3.45

Rubric dimension

raw

4.65 4.50 4.85 4.75

Table 12 Human calibration of the LLM judge. Spearman and mean absolute difference are human-vs-judge, over 40 ratings per dimension. Test–retest reliability from the four repeated pairs: Spearman 0.80, mean |diff| 0.94, exact agreement 25%.

Metric Applicability (both arms scorable) Informative-pair rate I Ceiling rate C

Value

Registered threshold

109/125 = 87.2% 0.092 0.028

≥ 80%: passes ≥ 0.20: fails < 0.60

Table 13 codealign gate. Applicability passes; discrimination does not.

The calibration is one rater on 20 pairs and unregistered; per-item disagreement is close to that rater’s own test–retest noise, so we rest nothing on individual items and report only the systematic gap difference, consistent across 17/20 pairs and all four rubric dimensions.

5.3 We registered a correctness instrument without checking it worked on our inputs Table 13 gives the registered gate. codealign (Dramko et al. 2025) is a graded instruction-level equivalence checker, and we registered it as our primary correctness measure. Its informative-pair rate on our data is 0.092 against a threshold we registered at 0.20 before seeing any data, and mean coverage falls from 0.31 at -O0 to 0.007 at -O3. Three of our six hypotheses became untestable as a pre-registered consequence. Every failure we hit lies inside limitations codealign’s own paper documents, so this is a finding about our procedure rather than about the tool. Its paper excludes from its own evaluation “functions which contain features codealign does not currently support, such as #ifndef macros and goto statements”, and states that “instructions must have at most one control dependency”. Ghidra emits goto 21

and labels by design when structured control flow cannot be recovered. The authors exclude such functions; we did not. The collapse itself is a control-dependence cascade rather than a gradual loss of alignable structure: loop rotation at -O1 and above introduces a guarding if in the candidate with no counterpart in the reference, it fails to align, and every instruction control-dependent on it fails with it. One item shows it in isolation — a five-line while loop computing a GCD aligns 6/6 reference instructions at -O0 and 0/6 at -O3, both decompilations correct and differing only in loop shape.2 Because the mechanism implicates a registered parameter, we checked whether the verdict depends on it: re-running all 250 pairs with control dependence disabled raises the informative-pair rate from 0.092 to 0.193 — still below the 0.20 registered before any data existed, by one pair of 109. The verdict is unchanged. We report this as an exploratory sensitivity analysis and not as a substitute for the registered figure. The generalization is cheap to state and would have been cheap to act on. Applicability is a property of an instrument on an input class, and a registration that does not establish it is registering a hope. One -O3 item, run before registration, would have exposed all of this in an afternoon — at which point the design could have excluded goto-bearing functions as codealign’s own evaluation does, chosen a different instrument, or registered the pass/fail axis as primary instead.

5.4 Two errors of ours, and how each was caught A chance baseline built from the wrong vocabulary. Our first null paired ground-truth identifiers with ground-truth identifiers from other items. No observed arm does that, and model-generated names embed with lower mean similarity to everything, so the correct arm-matched null sits 0.074 lower. A released draft read genuinely above-chance recovery as chance on that basis. The correction costs zero API calls and changes a headline in either direction, which is why we argue it is the only defensible chance line for generative output. An ablation that did not ablate. Run D as first implemented was a consistent bijection, which preserves every def-use edge; measured def-use survival was 1.000. A full draft of this paper reported a dataflow result on that basis. The manipulation check of §3.2, which the design had omitted, is what caught it. The correction cuts both ways: it rescued the novel-tier claim, now supported by an ablation that demonstrably removes what it says it removes, and it cost the textbook-tier claim, which did not survive the same treatment. Every ablation should ship a manipulation check quantifying what it removed, reported next to the effect it licenses. The failure mode neither pre-registration nor re-execution can catch. Both of these were caught by outside reads of near-final drafts, not by anything the registered study contained. That is not coincidence but a structural property. 2 Mechanism established in correspondence with codealign’s authors, whom we thank. A standalone reproduction is available with the artifact.

22

Pre-registration and re-execution are both faithful mechanisms: registration fixes the question before the data can bias the answer, and re-execution confirms the pipeline computes what it says. A faithful mechanism cannot detect that the question itself is wrong, because it reproduces the wrong comparison exactly, now with authority. Every number in this paper was correctly computed, and two of them were correct answers to the wrong question. Against that failure mode we know of one defence — a reader who has not spent months inside the design’s assumptions — and we would now budget for such a read as deliberately as we budget for a re-execution.

5.5 Reported precision belongs to one execution After the T3-side per-item artifacts were destroyed by a file-handling accident after analysis was complete, we re-executed the entire T3 side — 128 calls — from inputs verified byte-identical, under a registration committed before any call, with recorded and re-run numbers reported side by side regardless of agreement. The outcome cuts both ways. Every registered hypothesis verdict reproduces; the central null reproduces in sign and nullity at all three sample sizes; and the T1 control recomputes to four decimals, confirming an identical measurement pipeline. But the precision does not reproduce: the pooled detection bound inflates from 0.039 to 0.056, a 44% move, and a secondary contrast crosses nominal significance in one execution and not the other. At the registered n = 5 the same contrast moves by 0.049 between executions — comparable to the 0.056 bound this paper claims — shrinking with n and washing out at the full corpus. An LLM-in-the-loop pipeline re-executed from identical inputs thus reproduced every pre-registered verdict while moving estimates enough to break a claimed bound. That is simultaneously a validation of verdict-level pre-registration and a caution against quoting tight bounds on nulls, or unregistered significance, from a single execution. Ours cost $1.19. Nobody in this literature re-runs; this is what one re-run showed.

6 Threats to Validity Scope. Single compiler (gcc 13.3.0), single decompiler (Ghidra 12.1), single architecture (x8664), one prompt per model, textbook-scale single functions with self-tests. Results may not transfer to real-world binaries, other decompilers, or larger functions. -Os is reported descriptively and excluded from every trend test. The floor is the one finding measured on two refinement models; everything else rests on the first refiner alone. The corpus is small, and the corpus is the binding constraint. The floor’s detection bounds are 0.047 (T3) and 0.071 (T1) at the corpus maximum, roughly a quarter of the refined arm’s useful band. The corpus caps n at 20 and 25, so a materially tighter bound needs more items rather than more calls — the full-corpus extension cost $1.04. Budget was never the binding constraint and we would rather say so than let a reader supply a worse explanation. 23

The ablation conflates its target with idiom recognition. Scrambling destroys dataflow, prefix types and the surface cues that would let a model recognize a memorized idiom, all at once, so a contrast is not attributable to any one of them alone. This bounds the T1 comparison. The T3 null does not depend on it, since there nothing is detectable under any condition. The competing “naming from operations alone” reading (§3.2) remains open and needs a fourth condition we did not run. Two models is a small sample of the population that matters. The replication distinguishes “one model’s quirk” from “not one model’s quirk” and no more. Both models ran at one prompt each and at each model’s own lowest inferencetime setting rather than a common one, so neither prompt sensitivity nor reasoningbudget sensitivity is bounded by anything here. The extension was also registered after the first refiner’s results were known, which is why we call it a registered replication rather than folding it into the pre-registration. We wrote the T3 items ourselves, knowing the hypothesis. §3.1 addresses the difficulty confound; it does not address demand characteristics. An author who expects to show that refinement cannot name unfamiliar functions may, without intending to, write functions whose names are hard to guess from behavior. Nothing in our design excludes this. The cheap empirical bound — have readers blind to the source propose names from behavior alone, and compare the rate across tiers — is not run here; the stronger fix is to have probe items written by someone who does not know the hypothesis, and we recommend it for any replication. The judge is an LLM and shares a developer with the model under evaluation. Blinding and order randomization address presentation bias, not that. §5.2 reports a measured divergence against one unregistered volunteer rater without reverseengineering experience; an experienced panel is the proper version and was not run. Readability magnitudes should be read as contested rather than as bounded in a known direction. Run B was deferred. No T2 corpus was built, so the registered monotone three-tier ordering was never evaluated. ∆(T 1) > ∆(T 3) must not be read as if a monotone ordering had been tested.

7 Conclusion We set out to measure how the readability/correctness trade-off in LLM refinement of decompiler output scales with compiler optimization. We did not succeed: the graded semantic instrument we registered does not discriminate on our inputs, which made three of six hypotheses untestable, and the readability hypothesis was refuted because the gap is flat and near-saturated rather than widening. 24

What the study does establish is narrower and, we think, more useful. On functions written after the analysis plan was committed, refinement recovers identifier signal well above an arm-matched chance line — and that recovery does not measurably depend on the input we can ablate. Destroying dataflow changes the naming gain by +0.001 (95% CI [−0.026, +0.026]); removing type prefixes or permuting names changes it by no more; and a second refiner from a different vendor, on byte-identical inputs, reproduces the null on all six of its contrasts. Readability stays near ceiling throughout, so a reader is given no way to tell. Since refined output is what an analyst actually reads, and since it looks equally confident either way, this is a property the evaluation paradigm has to control for. We are deliberate about what this does not show. It is a bounded null, not a demonstration that nothing is happening: contributions smaller than 0.056 are invisible to us, and the ablation leaves the operation vocabulary intact, so naming from operations alone remains a live competing reading. The replication does not tighten anything — its intervals are wider throughout, so the bound we claim remains the first refiner’s — and one prompt per model leaves prompt sensitivity untested.

Five recommendations. A memorization floor should be a standard control in refinement evaluation: it is within-item and therefore immune to corpus-difficulty confounds, it costs twenty API calls, no aggregate metric substitutes for it, and where a floor result is the headline it should be run against a second model from a different vendor on byte-identical inputs, which cost us $0.74. Every ablation should ship a manipulation check quantifying what it removed, reported beside the effect it licenses; ours cost forty lines and caught a defect that had survived an entire study and a full draft. Chance baselines for generative output must be arm-matched, built from the system’s own output vocabulary rather than from reference-versus-reference pairings; ours differed by 0.074 and the error survived every earlier review. Graded instruments should carry an applicability report on the specific inputs used, registered in advance — a null result is otherwise indistinguishable from an instrument that never had signal on those inputs, and one item run before registration would have told us. And LLM-in-the-loop evaluations should re-execute at least once under a registration committed before the re-run: ours cost $1.19, reproduced every verdict, and showed that the precision we had been quoting belonged to one execution rather than to the pipeline. Artifact. The corpus (both tiers), pipeline, pre-registered analysis plan, the complete dated research log including every correction and deviation, raw API responses for the T1 runs, the registered T3 re-execution and all 270 calls of the second-refiner replication, and a per-number provenance map are archived at https://doi.org/10.5281/zenodo. 21968878. The original T3-side raw responses and per-item result files were destroyed by a file-handling accident after analysis; the log records the loss and exactly what survives, and the re-execution’s side-by-side comparison against every recorded T3 figure ships with it. The T3 probe items are published here, and their pre-run existence 25

is verifiable against a hash manifest committed before the run. Publishing them is one-way: they cannot serve as an uncontaminated probe for any later-trained model, and a replication wanting a clean probe needs new items and a new commitment.

Declarations Funding. No funding was received for conducting this study. The author received no grant, studentship, internship, institutional support, or vendor credits of any kind. The study’s entire computational cost was $6.06 in API charges, paid by the author personally. The affiliation above identifies where the author is enrolled as a student; the institution had no role in the design, execution, analysis, or reporting of this work, which the author conducted independently. Competing interests. The author declares no competing interests, financial or non-financial. The study evaluates commercial models produced by Anthropic and by Google; the author has no employment, consulting, financial or other relationship with either company, received no support, discounts or credits from either, and paid list price for all API usage reported here. The second refiner (§3.5) was registered to run on Google’s free tier and moved to the paid tier by a filed amendment when the free tier’s daily quota proved incompatible with the registered design; that run was billed at list price like every other. Ethics approval and consent to participate. The judge-calibration analysis (§5.2) involved one adult human participant rating twenty pairs of source code. No institutional ethics review was sought for this study. The author states the protocol so that a reader may judge that decision: the study was non-interventional; it collected no personal, sensitive or demographic data; it involved a single adult volunteer personally known to the author, who gave informed consent to participate and to publication and was free to withdraw at any point; the released dataset contains rating values keyed to pair identifiers and nothing that could identify him; and the only information withheld was the study’s hypotheses, which the blinding required. Consent for publication. The participant consented to publication of his aggregated and per-pair ratings in anonymous form. Data and code availability. The corpus (both tiers), the full pipeline, the pre-registered analysis plan, the complete dated research log, per-run configuration and raw API responses for every run that survives, and a per-number provenance map are available in the replication package archived at https://doi.org/10.5281/zenodo.21968878 (all versions; v1.1.0 current at 26

the time of writing), released under CC BY 4.0 for data and prose and the MIT licence for code. Three limits on that package are stated so a reader knows what is not in it. The manuscript source is excluded and its rights reserved, the paper being under submission when the deposit was made; the provenance map is included, since this paper cites it. The original T3-side raw responses and per-item result files were destroyed by a file-handling accident after analysis and before write-up (§7), so T3 figures rest on the registered re-execution rather than on the original run. And publishing the T3 probe corpus is irreversible: those items cannot serve as an uncontaminated probe for any model trained after their release.

Author contributions. The single author designed the study, wrote the pre-registration, built the pipeline and evaluation code, executed all runs, performed the analysis, and wrote the manuscript. Use of generative AI in the preparation of this manuscript. Generative AI was used in two distinct roles, which we separate to avoid conflating them. First, as the object of study : the refinement and judging models are the subject of the experiments and are documented in §3. Second, as an authoring aid : large language models were used during manuscript preparation and analysis-script development, including review passes over near-final drafts. The author directed all of this work, verified every number against the committed artifacts listed in NUMBERS.md, and takes full responsibility for the content, the analysis and the conclusions. No AI system is or could be an author of this paper.

References Banerjee P, Pal KK, Wang F, et al (2021) Variable name recovery in decompiled binary code. arXiv preprint arXiv:210312801 Basque ZL, Bajaj AP, Gibbs W, et al (2024) Ahoy SAILR! there is no need to DREAM of C. In: USENIX Security Symposium Cao Y, Liang R, Yang K, et al (2024) Evaluating the effectiveness of decompilers. In: ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA) Dramko L, Lacomis J, Hu EJ, et al (2024) A taxonomy of C decompiler fidelity issues. In: USENIX Security Symposium Dramko L, Le Goues C, Schwartz EJ (2025) Fast, fine-grained equivalence checking for neural decompilers. ACM Transactions on Software Engineering and Methodology (TOSEM) Eom H, Kim D, Lim S, et al (2024) R2I: A relative readability metric for decompiled code. In: ACM International Conference on the Foundations of Software Engineering (FSE) 27

Hu P, Liang R, Chen K (2024) DeGPT: Optimizing decompiler output with LLM. In: Network and Distributed System Security Symposium (NDSS) Lacomis J, Yin P, Schwartz EJ, et al (2019) DIRE: A neural approach to decompiled identifier naming. In: IEEE/ACM International Conference on Automated Software Engineering (ASE) Liu P, Sun J, Sun R, et al (2025) The CodeInverter suite: Control-flow and data-mapping augmented binary decompilation with LLMs. arXiv preprint arXiv:250307215 Liu P, Sun J, Xing M, et al (2026) Binary decompilation LLM with feedback-driven multi-turn refinement. arXiv preprint arXiv:260616162 Riddell M, Ni A, Cohan A (2024) Quantifying contamination in evaluating code generation capabilities of language models. Association for Computational Linguistics Tan H, Luo Q, Li J, et al (2024) LLM4Decompile: Decompiling binary code with large language models. In: Conference on Empirical Methods in Natural Language Processing (EMNLP) Tan H, Tian X, Qi H, et al (2025) Decompile-Bench: Million-scale binary-source function pairs for real-world binary decompilation. arXiv preprint arXiv:250512668 Votipka D, Rabin S, Micinski K, et al (2020) An observational investigation of reverse engineers’ processes. In: USENIX Security Symposium Wong WK, Wang H, Li Z, et al (2023) Refining decompiled C code with large language models. arXiv preprint arXiv:231006530 Xie D, Zhang Z, Jiang N, et al (2024) ReSym: Harnessing LLMs to recover variable and data structure symbols from stripped binaries. In: ACM SIGSAC Conference on Computer and Communications Security (CCS) Yakdan K, Dechand S, Gerhards-Padilla E, et al (2016) Helping Johnny to analyze malware: A usability-optimized decompiler and malware analysis user study. In: IEEE Symposium on Security and Privacy Zhang Q, Li Z (2026) CoDe-R: Refining decompiler output with LLMs via rationale guidance and adaptive inference. arXiv preprint arXiv:260412913 Zhang Y, Wang X, Zhang Y, et al (2026) Constraint-guided multi-agent decompilation for executable binary recovery. arXiv preprint arXiv:260423940 Zhou Z (2025) Decompiling Rust: An empirical study of compiler optimizations and reverse engineering challenges. arXiv preprint arXiv:250718792

28

Zou M, Ding A, Xu D, et al (2024) D-Helix: A generic decompiler testing framework using symbolic differentiation. In: USENIX Security Symposium Zou M, Cai H, Wu H, et al (2025) D-LiFT: Improving LLM-based decompiler backend via code quality-driven fine-tuning. arXiv preprint arXiv:250610125

29

Record · ID 919462 · SHA-256 89b739ce928e4ad3
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.