Conceptio › Archive › arXiv CS
arXiv CSopen access

Efficiency Hallucination: Formalizing and Measuring Behavioral Calibration in LLM-Based Code Optimization

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
software-architecturesoftware-engineeringtesting
software engineering, software architecture, testing

arXiv:2609.14839v1 [cs.SE] 13 Sep 2026

Efficiency Hallucination: Formalizing and Measuring Behavioral Calibration in LLM-Based Code Optimization Sarah Wilson*

Gail Kaiser

Patrick Musau

Columbia University [email protected]

Columbia University [email protected]

Vanderbilt University [email protected]

Abstract—The integration of Large Language Models (LLMs) into automated code optimization introduces a critical reliability risk we term the Efficiency Hallucination: an LLM’s tendency to issue non-functional mutations with unsubstantiated performance claims on already-optimized code. This is driven by the Evaluation Trap, wherein binary benchmarks incentivize unnecessary modifications over safely abstaining. We present a validation framework using classification penalty methods, evaluated across 180 optimization runs on nine models (GPT, Claude, Gemini) using EffiBench. Under standard prompts, models exhibit a 100% over-edit rate on optimal code. Our guardrail raises correct abstention from 0% to to 44.4%, preserving a 100% edit rate on sub-optimal code with zero false abstentions. Calibration is uneven: GPT-5.4 Mini approaches near-perfect abstention, and simple code is recognized more reliably than complex code. Our framework offers a training-free mechanism to mitigate LLM overconfidence before deployment in production. Index Terms—large language models, hallucination, code optimization, code generation.

I. I NTRODUCTION The automation of code optimization represents one of the highest-stakes applications of Large Language Models (LLMs) in software engineering. Classical compilers—from production toolchains like GCC [1] and LLVM/Clang [2] to research systems such as CETUS [3] and PLUTO [4]—operate through deterministic, rule-based static transformations. Modern LLM-based frameworks—AlphaEvolve [5], LLaMoCo [6], PerfRL [7]—go substantially further, leveraging deep semantic understanding to identify algorithmic improvements such as loop restructuring, data structure substitution, and closed-form mathematical replacements that static tools cannot discover. In some cases, these systems have surpassed prior human records in algorithmic problem solving [5]. Yet this semantic fluency introduces a subtle and underexamined reliability failure. When presented with code that is already optimal, LLMs do not admit uncertainty: as our pilot study confirms (Section IV), they invent optimizations. We term this the Efficiency Hallucination—a model’s generation of plausible code edits accompanied by confident, unverifiable performance claims. Unlike functional hallucinations that cause test failures, Efficiency Hallucinations are largely invisible to per-commit test suites that gate everyday *Current affiliation: Google.

development—the code still compiles, passes tests, and may even contain comments confidently asserting a speedup, yet no actual performance improvement occurs. Even dedicated performance-regression suites, which typically run only before a release rather than per-commit, may not catch them until much later. In production pipelines, such edits cost review time and erode trust in automated tooling. The root cause is what we call the Evaluation Trap. Optimization pipelines using benchmarks like PIE [8] and EffiBench [9] evaluate success mainly via functional correctness and execution runtime, leaving unpenalized any hallucinated speedup claims a model generates in its textual rationale. This binary incentive creates an asymmetric risk structure: for a model encountering already-optimal code, admitting optimality yields zero optimization credit, while generating a plausible bluff is a zero-risk gamble that may capitalize on runtime variance. We formalize this mathematically through the Is-It-Valid (IIV) classification framework of Kalai et al. [10], which demonstrates that generative error rates are mathematically guaranteed to be at least double the underlying misclassification rate. Without an explicit penalty for overoptimization, bluffing becomes the dominant inference strategy for models encountering performance ceilings. This paper makes four contributions. First, we formally define a taxonomy of Efficiency Hallucinations, distinguishing over-edits (hallucinated optimizations on already-optimal code) from false abstentions (failure to improve genuinely improvable code), and from the complementary true outcomes of correct edits and correct abstentions. Second, we provide the first application of the IIV framework to code optimization, explaining why current benchmark designs are mathematically biased toward elevated bluffing rates. Third, we introduce the Optimal Baseline evaluation methodology, which measures an agent’s Behavioral Calibration—its wisdom to abstain— rather than raw optimization speedup. Fourth, we conduct a controlled 180-trial pilot study across nine models from three LLM families, validating these predictions empirically and revealing per-model, per-family, and per-problem variation with direct implications for the reliability engineering of AIassisted software pipelines.

II. R ELATED W ORK A. LLM-Based Code Optimization The integration of LLMs into automated performance engineering spans three main directions: architectural specialization, agentic search strategies, and evaluation and benchmarking. Architectural Specialization. The LLaMoCo framework [6] addresses a core failure mode of general-purpose LLMs in optimization: their tendency toward API hallucinations, syntax errors, and hyper-parameter misinterpretations. LLaMoCo fine-tunes in two stages, first teaching the model to recognize what type of optimization problem it faces, then how to solve it—reducing the kind of confident misjudgment that, in our own results, surfaces as treating visually complex code as inherently improvable. The PerfRL framework [7] fine-tunes a small model (CodeT5) with reinforcement learning from unit-test and compiler feedback; despite far fewer parameters, it matches or exceeds the larger CodeGen-2B baseline on speedup and runtime-reduction scores. These works demonstrate that specialization reduces the hallucination patterns above, motivating our investigation of whether it also improves behavioral calibration. Agentic Search Strategies. AlphaEvolve [5] orchestrates autonomous evolutionary search over entire codebases, pairing high-throughput generation with strong-reasoner verification to discover improvements to established algorithms (e.g., a new bound for matrix multiplication). The CSE framework [11] adds hierarchical evolution memory for inter-task generalization, and Artemis [12] automates agent configuration from execution logs. While these frameworks improve optimization capability, their iterative pressure creates the exact conditions under which Efficiency Hallucinations are most likely: models are incentivized to propose improvements in every round, even when the code has reached its ceiling. Evaluation and Benchmarking. Recent work reveals a persistent gap between functional correctness and efficiency. Peng et al. [13] find AI patches 18% less likely than human patches to include performance validation, with some claiming extreme speedups (7400×) without benchmarks; ENAMEL [14] quantifies the gap (GPT-4: 0.831 functional pass rate, 0.454 efficiency score). SWE-Perf [15] scales evaluation to 140 real GitHub performance PRs, yet—like ENAMEL—tests only genuinely improvable code; by construction it cannot surface the Evaluation Trap on already-optimal code, the failure mode we study here. Nogueira et al. [16] confirm the same pattern in code generation—pass/fail metrics miss structural quality failures—making this cross-domain and motivating the Efficiency Hallucination. B. LLM Hallucination in Code CECoder [17] reduces hallucinations in repository-level code generation through fine-grained code element retrieval; unlike retrieval-based grounding, our work targets the distinct failure mode where models understand the code but fabricate optimization improvements regardless. CodeMirage [18]

identifies three code-specific hallucination subtypes: API misuse, logic errors, and efficiency misrepresentation (claiming superior complexity while producing sub-optimal code)—the last being most directly related to our work. Hallucinations in repository-level code are driven by cross-file dependency neglect, stale documentation retrieval, and over-generalization of local patterns [19]. EffiBench [9] specifically benchmarks efficiency claims, finding that even when LLMs produce correct solutions, they tend to use sub-optimal algorithms where human experts choose optimal ones. The IIV framework of Kalai et al. [10] provides the mathematical foundation for this work. They demonstrate that text generation amplifies misclassification errors: if a model’s misclassification rate on training facts is ϵ, its generative error rate is at least 2ϵ—each misclassification compounds at generation because the model emits the incorrect pattern as if it were correct. At a 10% misclassification rate, merging 1,000 LLM optimization PRs guarantees at least 200 unverified performance claims that may pass CI yet introduce latent complexity. This amplification is especially severe for patterns appearing only once in training data—singletons— because models cannot distinguish these from statistical noise. Optimization micro-patterns (e.g., a specific bitwise trick valid only for a narrow input range) are prototypical singletons, making Efficiency Hallucinations a mathematically predictable failure mode rather than a correctable bug. III. T HE E FFICIENCY H ALLUCINATION F RAMEWORK A. The Evaluation Trap and IIV Formalization Following the IIV analysis [10], we model the optimization decision as a binary classification problem over two example types: • Valid (+): Edits that maintain functional correctness while achieving measurable runtime improvements. • Error (-): Edits that maintain functional correctness but provide no runtime improvement, degrade performance, or introduce functional regressions. To the best of our knowledge, current benchmarks provide no negative signal for the Error class during training or evaluation, removing the only mechanism that could suppress this mathematically guaranteed bluffing rate. B. Behavioral Calibration: Taxonomy We define four disjoint outcome categories for any model response to an optimization request. Let C denote code at its performance ceiling (optimal) and S denote genuinely suboptimal code. • Edit (True Positive): Model modifies S. Correct outcome. • Abstention (True Negative): Model signals C is already optimal (e.g., outputs ALREADY_OPTIMAL). Correct outcome. • Over-Edit (False Positive): Model modifies C. This is an Efficiency Hallucination. • False Abstention (False Negative): Model declines to improve S. Pathological refusal.

A perfectly calibrated agent achieves edit rate = 1.0 on S and abstention rate = 1.0 on C simultaneously. We define the Bluff Rate as the over-edit frequency on optimal code, and the False Abstention Rate as the abstention frequency on sub-optimal code. Perfect calibration requires both to be zero. Under the Evaluation Trap, the Bluff Rate is guaranteed to be positive and approaches 1.0 as the performance ceiling becomes harder to verify statically. C. The Optimal Baseline Strategy Standard benchmarks measure capability by testing agents exclusively on improvable code. Our methodology adds a complementary evaluation axis: testing agents on top-percentile human solutions from EffiBench [9] that have effectively reached their theoretical performance ceiling. This Optimal Baseline design isolates the Bluff Rate from optimization capability, enabling independent measurement of each. We operationalize calibration via an IIV-derived Penalty Prompt: “Only suggest an edit if you are >90% confident it improves execution speed; otherwise, output ALREADY_OPTIMAL.” The threshold exploits a static-verification asymmetry: on sub-optimal code, a model can often confirm confidence exceeds 90% from complexity analysis alone (replacing O(n2 ) with O(n log n) requires no runtime clock); on alreadyoptimal code, no such static proof exists, so an honestly calibrated model must fall below threshold and abstain. This explicit 90% confidence threshold, grounded in the singleton-avoidance mechanism of the IIV framework, introduces a non-zero inference-time cost for the Error class—the missing signal that current benchmarks omit. By comparing model behavior under the control condition (no penalty, simply “Optimize this code for execution speed.”) versus the IIV penalty condition, we can directly measure whether models are capable of calibrated abstention and whether that capability degrades their optimization performance on improvable code. IV. P ILOT S TUDY A. Experimental Setup We constructed a 180-trial dataset spanning two snippet types and two prompt conditions. Problem Pairs. Five LeetCode problems were selected from the EffiBench benchmark suite [9], each yielding a paired optimal/sub-optimal snippet. The optimal variant is the EffiBench top-percentile solution; the sub-optimal variant is a functionally correct but algorithmically degraded version generated by Gemini-3.5-Flash and human-verified. The five pairs are: (1) Combination Sum II (optimal: sorted DFS with duplicate skipping; sub-optimal: brute-force DFS without pruning), (2) Remove Duplicates from Sorted Array II (optimal: O(n) two-pointer; sub-optimal: O(n2 ) list.count() loop), (3) Is Same Tree (optimal: recursive structural comparison with short-circuit; sub-optimal: string serialization), (4) Finding 3-Digit Even Numbers (optimal: O(1000) iteration with Counter; sub-optimal: O(n3 ) nested loops), and (5)

TABLE I P ILOT S TUDY R ESULTS BY P ROMPT C ONDITION AND S NIPPET T YPE (N = 180) Condition

Snippet

Edit

Abs.

O-Edit

F.Ab.

Control Control IIV Penalty IIV Penalty

Optimal Sub-Optimal Optimal Sub-Optimal

1.000 1.000 — 1.000

0.000 0.000 0.444 0.000

— — 0.556 0.000

0.000 0.000 0.000 0.000

All values are proportions. Over-Edit and Edit are mutually exclusive by condition: any modification of optimal code under the penalty condition is classified as Over-Edit (not Edit). Over-Edit is undefined (—) under the control condition since no penalty is active. F. Ab. = False Abstention.

Min Operations to Reduce an Integer to 0 (optimal: bitmanipulation; sub-optimal: naive step-by-step simulation). Using paired problems controls for problem-specific familiarity effects. All code and data are available at https://github.com/ sarah-wilsxn/efficiency-hallucinations. Models and Conditions. Nine models from three LLM families were evaluated via direct API—not through codingagent wrappers (e.g., Claude Code, Codex CLI) whose refinement loops may partially suppress bluffing; agent evaluation is designated future work. Two conditions were used: control (standard optimization request) and iiv penalty (the Penalty Prompt from Section III). Model families: Claude (claude-opus-4.7-fast, claude-opus-4.8, claude-sonnet-4.5; n = 60), Gemini (gemini-3-flash-preview, gemini-3.1-pro-preview, gemini-3.5-flash; n = 60), and GPT (gpt-5-mini, gpt-5.4-mini, gpt-5.4; n = 60). Each model received 20 trials: 5 optimal × 2 conditions + 5 sub-optimal × 2 conditions, yielding 9 × 20 = 180 total. B. Result I: The Evaluation Trap Confirmed Table I presents results disaggregated by prompt condition and snippet type. Under the control condition, the edit rate on optimal code is 100% across all four model families, with zero abstentions. Every model, in every trial, modified already-optimal code without hesitation, producing confident suggestions for changes that, by construction, should not improve execution speed. The edit rate on sub-optimal code is also 100% under control, confirming that baseline optimization capability is uniformly high. This 100% Bluff Rate confirms the Evaluation Trap: absent an explicit penalty, no tested model recognizes when code is already optimal, validating the IIV prediction that bluffing is the dominant strategy at the performance ceiling. C. Result II: Behavioral Calibration under IIV Penalty Introducing the Penalty Prompt produces a clear directional shift. On optimal code, the abstention rate rises from 0% to 44.4% while the over-edit (Bluff) rate falls from 100% to 55.6%. Equally critical: the edit rate on sub-optimal code remains at 100% under penalty and the False Abstention Rate is 0.0% across all models. The IIV penalty does not

TABLE II P ER -M ODEL C ALIBRATED A BSTENTION ON O PTIMAL C ODE (IIV P ENALTY, n = 5 PER MODEL , N = 180) Family

Model

Abstain

O-Edit

Avg.

Claude

Opus 4.7-Fast Opus 4.8 Sonnet 4.5

3/5 (60%) 2/5 (40%) 3/5 (60%)

2/5 3/5 2/5

53%

Gemini

3 Flash Preview 3.1 Pro Preview 3.5 Flash

3/5 (60%) 1/5 (20%) 0/5 (0%)

2/5 4/5 5/5

27%

GPT

GPT-5 Mini GPT-5.4 Mini GPT-5.4

2/5 (40%) 5/5 (100%) 1/5 (20%)

3/5 0/5 4/5

53%

All models achieved 5/5 correct edits on sub-optimal snippets under penalty (0% False Abstention, omitted for brevity). O-Edit = over-edit. Avg. = family average. Bold denotes best result per column.

cause models to refuse all edits: it specifically reduces overediting on already-optimal code while leaving the 100% edit rate on sub-optimal code intact. Calibration and capability are behaviorally independent. The persistence of a 55.6% Bluff Rate follows the IIV framework: the penalty asks a model to report confidence above 90%, but that self-assessment is the unreliable judgment IIV describes—the threshold changes the decision rule, not underlying calibration [10]. External runtime verification is the definitive solution, motivating our execution-based pipeline (Section VII). D. Result III: Cross-Family and Within-Family Variation Table II reports per-model abstention rates on optimal code under the IIV penalty condition. Aggregate family-level averages mask substantial within-family heterogeneity. Under the penalty condition, all three families achieve nonzero calibration (Figure 1). Claude and GPT reach equal family-level abstention at 53% (8 of 15 trials each), with Gemini at 27% (4 of 15). Within-family variation is substantial. GPT-5.4-mini achieves perfect calibration (5/5, 100%), the only model in the study to do so, while GPT-5-mini and GPT5.4 achieve 40% and 20% respectively. Gemini spans the full range: 3 Flash Preview (60%), 3.1 Pro Preview (20%), and 3.5 Flash (0%). Claude’s three models cluster tightly (60%, 40%, 60%), suggesting consistent behavior across model generations. Capability-Calibration Inversion. A striking pattern emerges in the GPT and Gemini families: larger, more expensive variants exhibit lower calibration than their lighter, cheaper counterparts. GPT-5.4-Mini (100%) outperforms GPT5.4 (20%) by 80 percentage points; Gemini-3 Flash Preview (60%) similarly outperforms Gemini-3.1 Pro Preview (20%). For industry practitioners, this inverts the typical cost-quality assumption: lighter, lower-cost models achieve superior calibration on simple algorithmic problems, making them highly effective for localized or block-level optimization pipelines.

# EffiBench top-percentile k = 0 for x in nums: if k < 2 or x != nums[k - 2]: nums[k] = x; k += 1 return k # Gemini-3.5 Flash: "An elegant way to optimize this is to... allows us to bypass the first two elements without slicing... keeps the loop body extremely minimal and fast. Here is the optimized code:" if len(nums) <= 2: return len(nums) k = 2; it = iter(nums); next(it); next(it) for x in it: if x != nums[k - 2]: nums[k] = x; k += 1 return k Listing 1. Efficiency Hallucination: Gemini-3.5 Flash on already-optimal code (id 7, IIV penalty condition, Potential Over-Edit).

This pattern suggests capability-focused training reinforces compulsive-edit behavior. Claude shows a weaker version— Sonnet 4.5 (60%) vs. Opus 4.8 (40%)—though its tighter cluster suggests uniform constraints attenuate the effect. Example: Phantom Bottleneck. Listing 1 illustrates the behavioral signature of an Efficiency Hallucination. Presented with already-optimal O(N) code under the IIV penalty, Gemini-3.5 Flash diagnoses O(N) list slicing as the bottleneck and proposes an iterator rewrite that is “extremely minimal and fast”—stated confidently, not as a hypothesis. The original never slices the list; the bottleneck does not exist. Whether the rewrite is faster was never measured. A qualitative analysis of Claude responses reveals a chainof-thought calibration pattern: models such as Opus 4.7Fast and Sonnet 4.5 explicitly enumerate and verify candidate micro-optimizations before concluding “Already Optimal”— a deliberative, auditable reasoning chain rather than a bare abstention signal. A residual compulsive-edit pattern persists in the remaining over-edits across all families, suggesting RLHF-driven actionability biases are only partially suppressed by inference-time constraints. E. Result IV: Per-Problem Calibration Difficulty Table III disaggregates optimal-code abstention rates by benchmark problem across all nine models under the IIV penalty (n = 9 per problem). Variation is substantial, ranging from 89% to 11% abstention. Remove Duplicates from Sorted Array II—a two-pointer sweep with O(n) complexity—is easiest to verify: 8/9 models abstained correctly. Is Same Tree (56%) proved similarly tractable due to its short recursive structure. Combination Sum II (22%) was harder: the backtracking structure invites spurious pruning suggestions models cannot rule out without execution. Min Operations (44%) benefited from CoT reasoning: models verified bit-level invariants before concluding optimality. Conversely, Finding 3-Digit Even Numbers generated the highest Bluff Rate (89%) despite its O(1000) Counter

TABLE III P ER -P ROBLEM C ALIBRATED A BSTENTION ON O PTIMAL C ODE (IIV P ENALTY, A LL M ODELS , n = 9 PER P ROBLEM ) Optimal Benchmark Problem

Abstentions

Rate

Remove Duplicates from Sorted Array II Is Same Tree Min Operations to Reduce an Integer to 0 Combination Sum II Finding 3-Digit Even Numbers

8/9 5/9 4/9 2/9 1/9

89% 56% 44% 22% 11%

Overall

20/45

44.4%

Problems sorted by descending abstention rate. Find Even Numbers generated the highest Bluff Rate (89%), including among models that calibrated successfully on simpler problems. Per-Model Abstention Rate on Optimal Code (IIV Penalty Condition) 100%

GPT-5.4 Mini

20%

GPT-5.4

40%

GPT-5 Mini Gemini 3.5 Flash 0%

20%

Gemini 3.1 Pro Preview

60%

Gemini 3 Flash Preview

60%

Claude Sonnet 4.5

40%

Claude Opus 4.8

60%

Claude Opus 4.7-Fast 0

20

40

60

Claude Gemini GPT Overall Avg. (44.4%)

80

Calibrated Abstention Rate (%)

100

Fig. 1. Per-model calibrated abstention rates under the IIV penalty condition on optimal code (n = 5 per model). The dashed vertical line marks the overall average (44.4%). All nine models achieve 100% edit rate on suboptimal code under the same penalty condition (not shown), confirming zero capability degradation.

solution: nested comprehensions and Counter objects create a complex surface models misinterpret as improvable. These results confirm surface complexity as an independent risk factor for Efficiency Hallucinations. V. D ISCUSSION A. Calibration as an Orthogonal Reliability Dimension Our results establish Behavioral Calibration as a modelspecific reliability property orthogonal to optimization capability. The universally zero False Abstention Rate—across all models and sub-optimal problems—demonstrates the penalty constraint does not inhibit correct optimizations. Critically, this independence means Efficiency Hallucinations are invisible to standard suites: code passes all functional tests, compiles, and may claim a speedup—yet introduces no actual improvement. Behavioral Calibration is therefore a reliability dimension not captured by current benchmarks. B. The Calibration Pareto Frontier The IIV penalty at 90% confidence represents one operating point on a calibration frontier. Claude and GPT achieve equal

family-average abstention (53% each), with Gemini at 27%, all with 0% false abstentions. As penalty thresholds decrease, false abstention rates are expected to rise (pathological refusal); as thresholds increase, behavior may return toward the 100% over-edit control state. The safe operating region— where abstention is non-trivial and false abstentions remain zero—defines the calibration frontier. The 44.4% aggregate abstention achieved here demonstrates that this frontier is navigable at the 90% threshold for the majority of models, while the residual 55.6% over-edit rate motivates threshold tuning and fine-tuning studies to push the Pareto point further. C. Practical Deployment Guidance We suggest three actionable guidelines for integrating LLM optimization into CI/CD pipelines. First, apply the IIV penalty prompt as a zero-overhead guardrail: it dropped the over-edit rate by 55.6% with zero false abstentions in our study. Second, given the Capability-Calibration Inversion, prefer lighter models (e.g., GPT-5.4-Mini, Gemini Flash) over premium counterparts for isolated, function-level tasks—achieving superior calibration at lower cost. Third, flag syntactically complex code (Counter objects, nested comprehensions, backtracking) for human review: surface complexity reliably predicts elevated Bluff Rates across all model families. VI. L IMITATIONS Internal Validity. The five sub-optimal problems are drawn from well-known LeetCode problems with widely-documented optimal solutions. This may artificially inflate the edit rate under penalty if models have memorized the target optimal forms from pre-training data, making the correct edit trivial rather than genuinely reasoned. Sub-optimal snippets were generated by Gemini, which may introduce a self-recognition confound for the Gemini family. Both effects would work in the direction of higher edit rates (lower false abstention), making our zero false-abstention result a conservative bound. External Validity. This pilot establishes directional validity but has limited statistical power per model (n = 5). As a foundational case study formalizing Efficiency Hallucinations, we prioritize isolating behavioral calibration on algorithmic problems rather than benchmarking repo-scale software noise. Nevertheless, our findings demonstrate that the underlying Evaluation Trap is universal across all tested SOTA families. Whether these calibration dynamics transfer to complex systems with cross-file dependencies remains an open question [19]. Construct Validity. The Optimal Baseline assumes EffiBench top-percentile solutions represent performance ceilings. Some solutions admit further source-level microoptimizations (cache-line alignment, bitwise tricks), and compiler/runtime optimizations (e.g., vectorization, JIT) may widen the source-level/execution-level gap. If any optimal snippet was truly improvable, the 55.6% Bluff Rate is an upper bound; execution-based verification may resolve this.

VII. C ONCLUSION AND F UTURE W ORK We have formalized the Efficiency Hallucination as a failure of Behavioral Calibration rooted in the binary incentive structure of current optimization benchmarks. The IIV-theoretic analysis demonstrates that hallucination at performance ceilings is not a correctable model bug but a mathematically predictable output of current training frameworks encountering already-optimal inputs. Our pilot study confirms the Evaluation Trap (100% Bluff Rate under control across all families), demonstrates that IIV penalty prompts yield meaningful calibration gains (44.4% abstention on optimal code) without capability degradation (100% edit on sub-optimal code, 0% false abstentions), uncovers within-family variation including a capability-calibration inversion finding, and identifies surface code complexity as an independent risk factor for Efficiency Hallucinations. Macro-Scale Automated Pipeline. Immediate next steps are scaling to all 1,000 EffiBench problems with executionbased verification, as well as extending to production software such as via SWE-Perf’s real-world GitHub performance PRs [15] and open-source repositories (e.g., Abseil, RocksDB) to validate whether calibration findings transfer beyond algorithmic-contest problems. Mapping the Calibration Frontier. Future work must systematically vary the IIV penalty threshold and map the resulting Pareto frontier between Bluff Rate and False Abstention Rate. The optimal threshold should reflect real-world code criticality—high-frequency, latency-sensitive paths may warrant more conservative thresholds than rarely-executed cases— and context-specific calibration will provide practitioners with evidence-based deployment guidelines. Positioning our contribution relative to concurrent benchmarks, FrontierCode [20] demonstrates that binary verifiers misclassify code quality broadly; our IIV penalty can be read as the optimization-specific instance of this same problem, and a natural next step is integrating execution-based performance verification as an explicit blocking criterion in rubric-based benchmarks of this kind. Attention-Based Self-Calibration. Integrating RAUQ-style uncertainty signals [21] as live feedback into agentic optimization loops offers a more principled long-term solution than static prompt engineering: when attention heads signal hallucination onset, the agent can automatically escalate confidence thresholds or trigger human-in-the-loop review without requiring offline calibration studies. Until evaluation frameworks explicitly reward the wisdom to abstain, the Evaluation Trap will continue to induce Efficiency Hallucinations—reforming these benchmarks will enable LLMs to become reliable optimization partners for software engineering teams. ACKNOWLEDGMENTS The first author thanks Yusuf Simonson, Ameya Shringi, Fred Lewis, Aditya Patil, Chris Kennelly, and Google’s AI & Infrastructure team for the internship that inspired this

research. Kaiser’s research is supported in part by NSF CNS2247370 and NSF CCF-2313055. R EFERENCES [1] Free Software Foundation, “Gnu compiler collection (gcc),” 2026. [2] C. Lattner and V. Adve, “Llvm: a compilation framework for lifelong program analysis & transformation,” in International Symposium on Code Generation and Optimization, 2004. CGO 2004., 2004, pp. 75–86. [3] H. Bae, D. Mustafa, J.-W. Lee, R. von Hanxleden, S. P. Midkiff, and R. Eigenmann, “The cetus source-to-source compiler infrastructure: Overview and evaluation,” International Journal of Parallel Programming, vol. 41, no. 6, pp. 753–767, 2013. [4] U. Bondhugula, A. Hartono, J. Ramanujam, and P. Sadayappan, “A practical automatic polyhedral parallelizer and locality optimizer,” SIGPLAN Not., vol. 43, no. 6, p. 101–113, Jun. 2008. [5] A. Novikov, N. Vũ, M. Eisenberger, E. Dupont, P.-S. Huang, A. Z. Wagner, S. Shirobokov, B. Kozlovskii, F. J. R. Ruiz, A. Mehrabian, M. P. Kumar, A. See, S. Chaudhuri, G. Holland, A. Davies, S. Nowozin, P. Kohli, and M. Balog, “Alphaevolve: A coding agent for scientific and algorithmic discovery,” 2025. [6] Z. Ma, Y.-J. Gong, H. Guo, J. Chen, Y. Ma, Z. Cao, and J. Zhang, “Llamoco: Instruction tuning of large language models for optimization code generation,” IEEE Transactions on Evolutionary Computation, pp. 1–1, 2026. [7] S. Duan, N. Kanakaris, X. Xiao, H. Ping, C. Zhou, N. K. Ahmed, G. Ma, M. Capota, T. L. Willke, S. Nazarian, and P. Bogdan, “PerfRL: A Small Language Model Framework for Efficient Code Optimization,” arXiv preprint arXiv:2312.05657, 2025. [8] A. G. Shypula, A. Madaan, Y. Zeng, U. Alon, J. R. Gardner, Y. Yang, M. Hashemi, G. Neubig, P. Ranganathan, O. Bastani, and A. Yazdanbakhsh, “Learning performance-improving code edits,” in The Twelfth International Conference on Learning Representations, 2024. [9] D. Huang, Y. Qing, W. Shang, H. Cui, and J. M. Zhang, “Effibench: Benchmarking the efficiency of automatically generated code,” 2025. [10] A. T. Kalai, O. Nachum, S. S. Vempala, and E. Zhang, “Why language models hallucinate,” 2025. [11] T. Hu, R. Chen, S. Zhang, J. Yin, M. X. Feng, J. Liu, S. Zhang, W. Jiang, Y. Fang, S. Hu, H. Wang, and Y. Xu, “Controlled self-evolution for algorithmic code optimization,” 2026. [12] P. Brookes, V. Voskanyan, R. Giavrimis, M. Truscott, M. Ilieva, C. Pavlou, A. Staicu, M. Adham, W. Evers-Hood, J. Gong, K. Zhang, M. Fedoseev, V. Sharma, R. Bauer, Z. Wang, H. Nair, W. Jie, T. Xu, A. Constantin, L. Kanthan, and M. Basios, “Evolving excellence: Automated optimization of llm-based agents,” 2025. [13] H. Peng, A. Zhong, R. A. C. Méndez, K. G. Kalu, and J. C. Davis, “How do agents perform code optimization? an empirical study,” 2025. [14] R. Qiu, W. Zeng, J. Ezick, C. Lott, and H. Tong, “How efficient is llm-generated code? a rigorous & high-standard benchmark,” in International Conference on Learning Representations, vol. 2025, 2025, pp. 2233–2261. [15] X. He, Q. Liu, M. Du, L. Yan, Z. Fan, Y. Huang, Z. Yuan, and Z. Ma, “Swe-perf: Can language models optimize code performance on realworld repositories?” 2025. [16] R. P. Nogueira, M. Vieira, and J. R. Campos, “Beyond functional correctness: An empirical evaluation of large language models for textto-code generation,” in 2025 IEEE 36th International Symposium on Software Reliability Engineering (ISSRE), 2025, pp. 264–275. [17] Y. He, J. Wu, X. Ling, T. Luo, M. Yang, and C. Zhao, “Cecoder: Finegrained code element retrieval for repository-level code generation,” in 2025 IEEE 36th International Symposium on Software Reliability Engineering Workshops (ISSREW), 2025, pp. 89–94. [18] V. Agarwal, Y. Pei, S. Alamir, and X. Liu, “Codemirage: Hallucinations in code generated by large language models,” 2025. [19] Z. Zhang, Y. Wang, C. Wang, J. Chen, and Z. Zheng, “Llm hallucinations in practical code generation: Phenomena, mechanism, and mitigation,” 2025. [20] E. Lu, B. Pan, D. Birlikci, S. Lee, R. Wang, R. Choudhury, F. Ma, T. Qin, C. Baronio, and S. Alberti, “Introducing frontiercode,” Cognition AI, 2026. [21] A. Vazhentsev, L. Rvanova, G. Kuzmin, E. Fadeeva, I. Lazichny, A. Panchenko, M. Panov, T. Baldwin, M. Sachan, P. Nakov, and A. Shelmanov, “Efficient hallucination detection for LLMs using uncertaintyaware attention heads,” 2026.

Record · ID 919495 · SHA-256 e593e4cb25256c73
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.