ConceptioArchivearXiv CS
arXiv CSopen access

Beyond Pass Rate: A Multilingual, Execution-Grounded Evaluation of Open Code LLMs

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
softwarearchitecturesoftwareengineeringtesting
software engineering, software architecture, testing

Beyond Pass Rate: A Multilingual, Execution-Grounded Evaluation of Open Code LLMs Sayed Erfan Arefin

arXiv:2606.08840v1 [cs.AI] 7 Jun 2026

Department of Computer Science University of Dayton Dayton, OH 45469, USA Email: [email protected]

Abstract—Code generation models are typically compared using compact execution benchmarks and aggregate pass rates, but such summaries obscure how performance varies across programming languages, problem families, and failure modes. We present a largescale, execution-grounded evaluation of 9 openly accessible LLMs specialized for coding on 2,707 free LeetCode problems across 12 programming languages. Our corpus contains 325,343 problem– model–language jobs, each linked to prompt metadata, extracted code, LeetCode execution outcomes, and static-analysis signals. The results show that current open models remain far from the human acceptance reference: the best model, Yi-Coder-9B-Chat, reaches 23.64% mean correctness, compared with a 57.2% human acceptance baseline. Rankings are also slice-dependent: Qwen2.5Coder-14B-Instruct is strongest on hard problems and distinctproblem coverage, while Gemma-2-27B-IT achieves the highest alllanguage lint pass rate. Failure analysis shows that compile errors account for 63.25% of non-accepted best submissions, indicating that many failures occur before semantic correctness can be tested. Static quality further diverges from functional correctness. Together, these findings show that multilingual, artifact-preserving evaluation reveals tradeoffs hidden by single-language or singlemetric leaderboards. Index Terms—code generation, large language models, execution-grounded evaluation, multilingual benchmarks, LeetCode

I. I NTRODUCTION Large language models (LLMs) for code are increasingly used as programming assistants, but standard evaluations often rely on compact function-level benchmarks, single-language settings, or aggregate pass@1-style scores. These evaluations are useful, yet they obscure properties that matter in practice: whether a model is reliable across languages, how performance changes with task difficulty, whether failures are syntactic or semantic, and whether accepted code is also maintainable. We study these questions using a large execution-grounded archive of LeetCode-derived tasks. The corpus contains 2,707 free LeetCode problems, 325,343 cleaned problem–model– language jobs across 12 programming languages and 9 open or openly accessible code LLMs. Each job preserves the prompt, target language, model identity, raw response, extracted code, official LeetCode result, and static-analysis signals, enabling analysis beyond binary accepted/not-accepted outcomes. Our contributions are three fold: • We construct and analyze a large scale execution archive covering 2,707 problems, 12 programming languages and 9 models.

We evaluate open code LLMs across mean correctness, accepted-problem coverage, difficulty, language, topic, pairwise head-to-head outcomes, and failure composition. • We jointly analyze functional correctness and static quality, showing that execution success and maintainabilityoriented signals can diverge substantially. Our results show that model comparison is multidimensional. Yi-Coder-9B-Chat achieves the highest mean correctness and strongest average language rank, while Qwen2.5-Coder-14BInstruct performs best on hard problems and accepted-problem coverage. All models remain far below the human acceptance reference, compile errors dominate non-accepted outcomes, and static quality does not follow the same ordering as functional correctness. These findings motivate code-model evaluation that preserves execution, language coverage, failure modes, and quality signals instead of collapsing performance into a single score. •

II. R ELATED R ESEARCH Execution-based evaluation remains a central method for studying code generation because it checks whether generated programs actually run and satisfy task-level tests. Chen et al. [1] evaluated large language models trained on code using executable programming tasks, while Austin et al. [2] studied program synthesis with large language models and Hendrycks et al. [3] measured coding-challenge competence using APPSstyle programming problems. Li et al. [4], [5] extended this direction to competition-level code generation, and Lu et al. [6] introduced CodeXGLUE as a multi-task benchmark suite for code understanding and generation. CodeBLEU prposed by Ren et al. [7] as an automatic code-evaluation metric that incorporates lexical, syntactic, and semantic signals, but execution-based judging remains important because it directly measures functional behavior. Subsequent work broadened code-model evaluation across languages, task types, and freshness settings. Cassano et al. [8], Zheng et al. [9], Athiwaratkun et al. [10], and Jain et al. [11] studied multilingual or contamination-aware evaluation settings, showing that code-model performance should not be treated as a single-language or single-benchmark property. Lai et al. [12], Du et al. [13], Gu et al. [14], and Zhuo et al. [15] further expanded evaluation toward data-science code, class-level generation, code reasoning, pragmatic code generation, and

complex function-call tasks. Puri et al. [16] introduced a largescale code dataset covering diverse programming tasks, while Roziere et al. [17] and Lachaux et al. [17] studied unsupervised translation between programming languages. Together, these studies motivate evaluating code models across languages, problem categories, and execution outcomes rather than relying on one narrow benchmark. Repository-level and context-rich benchmarks form another related thread. Liu et al. [18] studied repository-level code completion, Ding et al. [19] focused on cross-file and multilingual code-completion evaluation, Jimenez et al. [20] evaluated whether language models can resolve real GitHub issues, and Zan et al. [21], [22] extended issue-resolution evaluation toward Java and multilingual settings. These studies are complementary to competitive-programming evaluations: repository benchmarks test tool use, retrieval, and patch integration, whereas LeetCode-style tasks stress algorithmic reasoning, language-specific syntax, and executable correctness under standardized problem specifications. Open model development evolved alongside these benchmarks. Feng et al. [23], Guo et al. [24] and Wang et al. [25] developed representation-learning or encoder-decoder models for program understanding and generation.Subsequent larger generative systems further advanced open code generation. Nijkamp et al. [26], Fried et al. [27], Allal et al. [28], Luo et al. [29], Roziere et al. [30], and Guo et al. [31] introduced open or openly described code-generation models with increasingly strong multilingual and instruction-following capabilities. Work on evaluation rigor also showed that benchmark conclusions can be fragile. Liu et al. [32] showed that weak test suites can materially overestimate correctness. Another line of work focuses on how code models fail rather than only how often they pass. Tian et al. [33], Arefin et al. [34], Heya et al. [35], Sobania et al. [36], and Wang et al. [37] examined issues such as wrong answers, runtime failures, debugging behavior, repair loops, hallucinations, API/tool use, repeated-query stability, and deprecated API usage. Chen et al. [38], Shinn et al. [39], and Yao et al. [40] studied test generation, feedback, and reasoning-action loops as mechanisms for improving model outputs. These methods motivate evaluating not only final pass/fail outcomes but also the failure states that would inform repair, debugging, or agentic follow-up systems. Code quality is a smaller but growing literature. Tian et al. [33] used Pylint-style analysis when studying ChatGPTgenerated Python solutions, while Simoes and Venson [41] compared large language models as code-quality judges against static-analysis tools. However, most existing code-generation benchmarks emphasize functional correctness, and most staticquality studies focus on narrower language slices or smaller model sets. This leaves a gap between execution-oriented benchmarking and maintainability-oriented static analysis. This study is unique in combining these lines at scale. Rather than evaluating a small function-level benchmark, a single language, a single proprietary model, we analyze a large archive of multiple open coding-LLM submissions across many

LeetCode problems, and multiple programming languages. The study jointly preserves functional execution outcomes, difficulty/topic/language slices, accepted-problem coverage, non-accepted failure composition, partial-correctness signals, and static-quality/linter evidence. This makes the work distinct from prior benchmarks, where it does not only ask which model solves more tasks, but also where each model succeeds, which languages and topics change the ranking, how failures manifest, and whether executable success aligns with maintainable code quality. III. M ETHODOLOGY A. Problem Collection and Corpus Construction We constructed the benchmark from publicly available, free LeetCode problems. The collection process first enumerated all non-premium problems available through LeetCode and stored the problem statements, metadata, slugs, difficulty labels, topic annotations, and acceptance-rate fields as local text and structured records. We used only free problems so that the resulting benchmark would be reproducible without relying on paid problem access. Each problem statement was preserved before model prompting so that all models received the same task specification for a given problem–language pair. The benchmark targets 12 programming languages: C, C++, C#, Go, Java, JavaScript, PHP, Python, Ruby, Scala, Swift, and TypeScript. For each problem, we generated language-specific prompts that instructed the model to produce a LeetCodecompatible solution in the target language. These prompts were stored in a MySQL database as individual execution jobs. Each job MYSQL entry stored one problem–language– model attempt and contains the prompt text, problem identifier, language identifier, model identifier, job status, and execution metadata. The database therefore acted as the central source of data for the experiments, supporting reproducible scheduling, retries in case of any operational errors, and later analysis. B. Distributed Prompt Serving and Model Inference To execute the large number of problem–language–model jobs, we implemented a server-backed job queue. The server exposed pending prompts to client workers and tracked each job through a finite set of operational states, including pending, completed, and operational-error states. Client workers repeatedly requested available jobs, submitted the prompt to the selected local model through LM Studio, stored the raw model response, and returned completion metadata to the server. Table I summarizes the nine evaluated open-weight or openly accessible instruction-tuned models. We used the default setting of each models in our experiments. The machine that was used in these experiments is a Lambda Vector One machine with a AMD Ryzen 7 CPU, 128 GB Memoery, 4 TB of SSD and Nvidia RTX 4090 GPU with 24 GB of memoery. Operational failures were handled separately from model failures. Network interruptions, local inference failures, API timeouts, database errors, or browser automation errors were marked as operational errors rather than treated as incorrect

TABLE I: Evaluated models and high-level model characteristics. The descriptions are used only to identify the model family and intended emphasis; model performance is measured empirically in our benchmark rather than inferred from these descriptions.

Model

Producer

Notable characteristic

Codestral-22B-v0.1 DeepSeek-Coder-V2-Lite-Instruct Gemma-2-27B-IT Granite-3.1-8B-Instruct Meta-Llama-3.1-8B-Instruct Phi-4 Qwen2.5-Coder-14B-Instruct Stable-Code-Instruct-3B Yi-Coder-9B-Chat

Mistral AI DeepSeek-AI Google IBM Meta Microsoft Alibaba/Qwen Stability AI 01.AI

Code-specialized 22B-parameter model Lite variant from the DeepSeek-Coder-V2 family. 27B-parameter member of the Gemma 2 family. 8B-parameter model from IBM’s Granite family. 8B-parameter Llama 3.1 model. 14B-parameter small language model. 14B-parameter instruction model. Small 3B-parameter model for code-completion. 9B-parameter code model.

TABLE II: Corpus summary after cleanup. Statistic Problems Models Programming languages Total programming jobs Accepted best submissions

Count 2,707 9 12 325,343 38,761

model outputs. After each execution round, jobs in operationalerror states were reset and re-queued in later rounds. This retry policy was designed to prevent infrastructure instability from being conflated with model quality. Completed jobs retained the original prompt, raw model response, extracted code artifact, model identifier, language identifier, and timestamps needed for later auditing.

difficulty, topic, and acceptance-rate coverage; functional metrics for correctness, accepted coverage, language, difficulty, topic, and language rank; failure metrics for status and wronganswer severity; comparative metrics for head-to-head outcomes and shared-task significance tests; and static-quality metrics for complexity, function length, accepted-code linting, and alllanguage ast-grep pass rates. LeetCode acceptance rate is used only as a problem-level human-reference signal. Runtime and memory fields are present but too inconsistently preserved for a defensible efficiency comparison. IV. E XPERIMENTAL R ESULTS A. Problem Corpus Characteristics

The benchmark is large and heterogeneous, but it is not uniform. Medium problems dominate the corpus, followed by C. Code Extraction and LeetCode Submission easy and hard problems. Figure 1b summarizes the difficulty Model outputs often contained explanations, markdown distribution of the cleaned archive. The distribution contains formatting, or surrounding commentary in addition to code. 715 easy problems, 1,368 medium problems, and 624 hard We therefore applied a post-processing step to extract the code problems. This imbalance matters analytically because medium block from each response, specifically the portion enclosed problems dominate aggregate scores, while the smaller hard in triple quotes, and saved it as the candidate submission for subset tests whether model performance holds under higher the corresponding problem and language. The raw response algorithmic difficulty. was retained for traceability. Candidate solutions were then Language coverage is broad but not perfectly uniform. Some evaluated using LeetCode’s official execution environment. A problems do not have support for all languages. JavaScript Python automation script used Selenium with a Chromium and TypeScript each cover 2,629 unique problems; C++ and command-line browser session to load each extracted solution Java each cover 2,598; C# covers 2,597; Go covers 2,596; C and submit it to LeetCode. For each submission, the script covers 2,592; and PHP, Ruby, Scala, and Swift each cover recorded the returned submission ID and queried LeetCode’s 2,591. Python has lower retained coverage, with 1,728 unique GraphQL API to retrieve the official judge result, including problems. Human acceptance rates provide a second view of labels such as Accepted, Wrong Answer, Compile Error, benchmark difficulty. Figure 1a shows the distribution of human Runtime Error, Time Limit Exceeded, and Memory Limit acceptance rates over the problem set. The fitted distribution has Exceeded. This ensured that correctness was measured by a mean of 57.51% and a standard deviation of 16.47 percentage LeetCode’s own judge rather than a local approximation of its points, indicating substantial variation in how difficult problems test suite. Table II summarizes the final analysis set. are for human solvers. Most problems cluster in the mid-range, Because jobs could be retried after operational resets, we with many falling between roughly 40% and 75% acceptance, avoid double-counting by selecting one best submission per while fewer problems appear in the very low-acceptance or very job: highest correctness ratio, then accepted status, then latest high-acceptance tails. This distribution provides context for timestamp. Correctness is 1.0 for accepted submissions, the interpreting model correctness because model performance can passed/total testcase ratio when available, and 0.0 otherwise. be compared against a human-centered difficulty signal rather We report five metric families: corpus-side metrics for language, than only against aggregate pass rates. The topic distribution

(a) Human acceptance-rate distribution across benchmark problems, summarized by a fitted normal curve with mean 57.51% and standard deviation 16.47 percentage points.

(b) Problem difficulty distribution in the cleaned benchmark. Medium problems form the largest slice, followed by easy and hard problems.

Fig. 1: Benchmark problem characteristics: human acceptance-rate distribution and problem-topic distribution. and more about whether the model can find at least one accepted solution for many different problems. Difficulty-conditioned performance shows that model rankings are slice-dependent. Figure 4 shows Yi-Coder leading on easy and medium problems, while Qwen is strongest on hard problems. The drop from easy to hard is steep for every model, confirming that the benchmark is not saturated. This suggests that model comparisons should not be reduced to a Fig. 2: Problem-topic distribution. Array, string, hash-table, single global score when the downstream workload emphasizes dynamic-programming, and math problems dominate the topic harder tasks. distribution. The language-conditioned results reveal another source of heterogeneity. Figure 5 reports correctness by target language reflects the composition of the available LeetCode problem and model, showing that model quality is not language-invariant. set used in this study. Figure 2 shows that arrays are the Some models are comparatively strong in conventional Leetmost common topic, followed by strings, hash tables, dynamic Code languages such as C++, Java, and Python, while others programming, and math. Consequently, aggregate results are are more competitive in languages such as C#, Swift, Ruby, or shaped more by high-frequency topics than by lower-frequency TypeScript. The heatmap also makes the relative winners easier graph, heap, or advanced data-structure topics. Arrays remain to compare across language–model pairs. Yi-Coder is strongest the dominant topic. for several conventional compiled or scripting targets, Qwen is especially competitive on C#, JavaScript, Ruby, Swift, and B. Model Performance TypeScript, and Gemma is strongest on Go and Scala. Thus, Figure 3a compares global model correctness with the human model ranking depends substantially on the target programming acceptance baseline. Yi-Coder-9B-Chat has the strongest mean language, reinforcing the central premise of this study that correctness, with Qwen2.5-Coder-14B-Instruct and Gemma-2- multilingual coding benchmarks can yield conclusions different 27B-IT forming the next tier. The absolute gap to the human from single-language evaluations. Topic-conditioned performance shows similar non-uniformity. acceptance reference remains large for every model, showing that open coding models still solve only a minority of the Table III reports correctness across the 20 most common LeetCode problems under this single-pass setting. The figure LeetCode topic tags. Yi-Coder is consistently strong across many high-frequency topics, while Qwen remains highly is therefore best read as a high-level leaderboard. Figure 3b reports accepted-problem coverage, which mea- competitive and leads selected algorithmic categories. The sures breadth rather than average correctness across all jobs. table also shows that topic-level rankings do not preserve a This view changes the ordering, where Qwen covers the largest single global ordering: for example, two-pointers, stack, binaryshare of accepted problems, indicating that a model can solve tree, and graph-search topics expose different relative strengths a broader set of distinct problems even when its average job- across models. level correctness is slightly lower than the top mean performer. Figure 6 shows that human acceptance rate provides a Coverage is useful because a downstream user may care less useful signal of model success. The strongest associations about repeated success on the same problem across languages are observed for Yi-Coder 9B and Qwen2.5 14B, both with

(a) Overall functional performance. Mean correctness is shown for each model and compared with the human acceptance baseline.

(b) Accepted-problem coverage by model. Qwen covers the broadest set of problems despite not having the highest mean correctness.

Fig. 3: Comparison of overall functional performance and accepted-problem coverage across models.

Fig. 4: Model correctness stratified by problem difficulty. Performance declines sharply from easy to hard problems, and the strongest model differs by slice. Pearson correlations of 0.36, followed by Gemma 2 27B at 0.31. This indicates that these models are more likely to succeed on problems that are also easier for human LeetCode users. However, all correlations remain moderate, ranging from 0.13 to 0.36, which shows that human acceptance rate does not fully capture model difficulty. Therefore, acceptance-rate metadata is informative for interpreting benchmark behavior. Pairwise comparisons give a direct view of relative model dominance on shared tasks. Figure 7 shows that Yi-Coder and Qwen dominate the head-to-head matrix, while weaker systems rarely beat the leading models. This reinforces the global ordering while still preserving the asymmetric nature of pairwise wins: a model can lose globally yet retain pockets of strength against another model on particular problem–language combinations. C. Failure Modes In this part of the study, we analyze non-accepted submissions to understand where generated programs break down.

Model-specific breakdowns further show that failure profiles differ across systems. Figure 8a shows that compile errors are the leading non-accepted outcome for every model, while stronger models tend to reduce this share and expose relatively more runtime or wrong-answer cases. This shift is important because it separates failures that prevent execution from failures that occur after the program runs. From an operational perspective, improving coding models should therefore mean not only increasing accepted outputs, but also moving residual failures toward more diagnosable runtime or semantic errors. The wrong-answer and partial-correctness results also reveal a limitation of the stored execution traces. Wrong-answer cases are concentrated in the zero-percent passed-testcase bin across models, with little evidence of intermediate testcase success. Figure 8b similarly shows that partial correctness is concentrated near the extremes rather than distributed across near-miss regions. Thus, the archive is useful for characterizing broad failure modes, but less effective for distinguishing almostcorrect wrong answers from fully incorrect ones.

Fig. 5: Heatmap of model performance by programming language. The heatmap captures language-conditioned correctness, highlights cross-language variation, and makes model-specific strengths easier to compare than global averages. TABLE III: Correctness by topic for the 20 most common topics. Values are percentages. Topic array string hash-table dp math sorting greedy binary-search depth-first-search matrix bit-manipulation breadth-first-search tree two-pointers prefix-sum heap-priority-queue simulation counting stack binary-tree

Yi-Coder

Qwen2.5

Gemma

DeepSeek

Codestral

Phi-4

Llama3.1

Granite3.1

21.9 23.7 23.0 18.9 21.2 21.6 18.3 21.1 26.6 22.2 18.2 25.1 24.9 31.6 15.1 14.7 23.3 24.0 28.9 30.1

21.4 20.3 22.8 15.3 19.7 22.4 18.6 20.3 23.5 21.9 18.8 23.7 22.0 26.4 17.2 15.3 23.1 23.1 22.3 25.7

17.7 18.5 16.1 14.2 16.3 16.9 12.5 18.6 15.5 18.7 14.1 15.6 12.9 27.5 12.4 9.8 19.4 17.2 18.5 16.1

11.9 12.0 12.7 10.2 10.6 11.7 8.6 12.7 13.2 13.1 11.0 14.0 11.2 17.0 8.8 9.3 10.2 11.3 14.0 13.2

10.3 11.9 11.8 8.5 10.6 9.3 6.3 12.3 16.3 12.2 10.7 15.0 15.2 18.5 7.7 6.2 10.0 8.4 15.2 18.7

9.6 9.2 10.1 7.2 8.5 8.6 6.1 9.9 11.6 11.2 8.7 11.2 10.4 13.8 8.0 5.5 10.3 7.9 12.1 13.2

7.3 8.2 6.5 5.1 7.5 6.4 5.2 7.4 7.9 7.1 6.5 7.4 8.4 13.6 5.9 2.6 8.3 7.1 9.9 11.3

2.1 2.4 2.2 1.1 2.0 2.1 0.7 2.6 1.9 2.2 2.2 1.4 2.0 4.1 1.6 1.1 2.4 2.9 1.5 2.6

D. Code Quality In this part we analyze the code quality generated by the LLMs. Figure 9 reports the structural quality score, where Stable-Code-Instruct-3B obtains the strongest score despite the weakest functional performance in the benchmark. This shows that structural regularity does not imply executable correctness, since simple, low-complexity code may still fail the task. Figure 10a shows that Gemma, Yi-Coder, and Codestral

StableCode Human 0.4 0.6 0.4 0.1 0.5 0.3 0.1 0.4 0.0 0.1 0.7 0.0 0.1 1.3 0.8 0.1 0.5 0.4 0.3 0.2

56.0 57.2 56.9 49.4 54.9 56.8 53.7 47.8 60.8 61.2 57.8 59.3 63.4 57.4 53.7 54.0 65.9 60.2 58.2 66.9

achieve the strongest all-language lint pass rates, while Qwen passes lint rules less often despite strong execution performance. This confirms that functional correctness and static cleanliness measure different properties. Figure 11 further shows that static quality is language-dependent, since a model can produce cleaner code in one language than another. Accepted-code linter coverage is limited to JavaScript, TypeScript, Java, C++, Go, C, and Python, all with 100% coverage, while Ruby, C#,

Fig. 6: Pearson correlation between model correctness and human acceptance rate. Stronger models correlate more with human acceptance, but the correlations remain moderate.

Fig. 7: Pairwise head-to-head win rates between models. Leading systems win more shared comparisons, while weaker models have low off-diagonal win rates.

TABLE IV: Model summary. Correctness, hard-slice correctness, coverage, and lint pass are percentages. Quality is reported on a 0–1 scale. Model Yi-Coder 9B Qwen2.5 14B Gemma 2 27B DeepSeek Lite Codestral 22B Phi-4 Llama 3.1 8B Granite 3.1 8B Stable-Code 3B

Correctness

Hard

Coverage

Avg. Lang. Rank

Quality

All-Lang. Lint Pass

23.6 22.2 18.9 13.1 12.3 10.3 8.1 2.6 0.7

10.1 11.1 5.5 6.1 4.6 3.7 1.5 0.3 0.0

66.0 77.1 52.0 58.1 48.5 43.8 34.3 14.6 3.6

2.33 2.67 3.50 4.75 4.25 5.33 5.42 7.92 8.83

0.519 0.490 0.532 0.455 0.516 0.461 0.469 0.476 0.588

85.7 49.3 86.6 59.8 85.3 45.0 53.6 79.9 83.2

(a) Non-accepted failure composition by model. Compile errors dominate the (b) Partial-correctness distribution by model. The archive non-accepted subset, with smaller shares of wrong-answer, runtime, memory- mainly separates zero-output wrong answers from accepted limit, time-limit, no-output, and output-limit outcomes. outputs, providing limited information about near-correct executions.

Fig. 8: Failure-mode composition and partial-correctness distribution across models.

Fig. 9: Overall structural code-quality score by model.

issues are concentrated in namespace and wildcard-import style patterns, with using-namespace-std accounting for roughly 48% of convention findings, wildcard import about 21%, force unwrap about 17%, var about 6%, global about 4%, and global-variable/var-declaration patterns making up most of the remaining mass. Error findings are dominated by AST-grep failures at about 50%, followed by throw-error patterns at about 15%, fatal-error and throw-exception patterns at roughly 7% each, panic at about 6%, and raise/throw-exception variants in the lower single digits. These patterns are more actionable than a single lint score because they indicate what kind of cleanup a developer or repair system would need to perform.

Warnings are concentrated in output and logging patterns. Print-related findings form the largest warning subtype at about Scala, PHP, and Swift are uncovered due to usable linters. 25% of warning issues, followed by system-output at about Thus, accepted-code linter conclusions apply only to the seven 12%, println at about 10%, console-log variants at roughly covered languages. 9–10% each, console-write-line and debug-print near 8% each, Figures 10b and 12 analyze this covered subset. These results var-dump around 4%, and smaller print or debug-print variants estimate review burden among code that already solves the task, below 2%. Refactor findings are rare in absolute terms, with showing that a functionally strong model may still require more 402 total cases, and are dominated by boolean-comparison cleanup. The language-specific view complements execution suggestions. The largest boolean-comparison subtype accounts scores by showing where accepted solutions remain stylistically for roughly 37% of refactor findings, followed by two additional or structurally weaker. boolean-comparison subtypes near 22% each, while the remainAll-language lint results show that the issue burden is ing variants appear only in small single-digit shares. These dominated by warnings and conventions rather than hard errors findings show that much of the all-language lint burden may not or refactor findings. The aggregate category counts contain prevent execution, but it remains relevant for maintainability, 141,376 warning findings, 45,537 convention findings, 2,280 error findings, and 402 refactor findings. Thus, the all-language output discipline, and production readiness. static-analysis burden is driven primarily by warning-style The gap between execution behavior and static analysis is and convention-style issues rather than refactor suggestions. one of the main findings. A model can generate code that Model-level category rates further show that models differ is short, regular, and structurally tidy while still failing to not only in lint-pass outcomes, but also in the types of static- compile or solve the task. Conversely, a model that solves analysis findings they accumulate. Warning rates are highest more problems may accumulate more lint findings because it for Qwen2.5-Coder-14B-Instruct, Phi-4, and Meta-Llama-3.1- produces more ambitious code or relies on conventions that 8B-Instruct, convention rates are more moderate, and refactor work operationally but violate the shared rule set. Evaluation rates remain near zero across all models. should therefore preserve both execution and static-quality Table V summarizes the dominant convention and error evidence. The best model depends on the downstream goal, subtypes previously reported as separate figures. Convention whether that goal is maximizing solved tasks, minimizing

(a) All-language lint pass rate across best-job code artifacts. Lint cleanliness and functional success produce different model rankings.

(b) Accepted-code linter score and lint pass rate by model. This view estimates review burden among submissions that already achieved accepted status.

Fig. 10: Linter-based evaluation across all best-job code artifacts and accepted-code submissions. TABLE V: Dominant all-language linter issue shares. Shares are within each linter category and replace the separate convention/error share plots.

Category

Issue subtype

Convention Convention Convention Convention Convention Convention

using-namespace-std wildcard import force unwrap var global global-variable / var-declaration

≈ 48% ≈ 21% ≈ 17% ≈ 6% ≈ 4% small remainder

Error Error Error Error Error Error

AST-grep failure throw-error fatal-error throw-exception panic raise / throw-exception variants

≈ 50% ≈ 15% ≈ 7% ≈ 7% ≈ 6% lower single digits

Fig. 11: Language-conditioned structural code quality by model, showing substantial variation in static quality across programming languages. review burden, or balancing both.

Share

Fig. 12: Accepted-code linter score by language and model. Linter quality varies across both model family and programming language. V. C ONCLUSION We presented a large-scale, multilingual, execution-grounded evaluation of openly accessible LLMs on LeetCode-derived programming tasks. By preserving prompts, raw responses, extracted code, official execution outcomes, and static-analysis signals, the benchmark supports analysis beyond aggregate pass rates. The results show that model comparison is multi-

dimensional. Yi-Coder-9B-Chat achieves the strongest mean correctness and average language rank, Qwen2.5-Coder-14BInstruct performs best on hard problems and accepted-problem coverage, and Gemma-2-27B-IT is strongest under the alllanguage lint pass-rate metric. Thus, the preferred model depends on whether the objective is average correctness, coverage, hard-task success, or reduced review burden. The evaluation also exposes persistent reliability gaps. All models remain below the human acceptance reference, with compile errors dominating non-accepted best submissions. Overall, this study shows that evaluating code-generating LLMs requires more than identifying a single best model. Model performance depends on programming language, problem difficulty, topic composition, execution behavior, and code quality signals. By preserving the full generation-to-execution pipeline, this benchmark provides a reproducible basis for understanding not only when models succeed, but also how and why they fail. These findings highlight the need for executiongrounded, multilingual, and artifact-preserving evaluations as coding models continue to be used in increasingly diverse programming contexts. R EFERENCES [1] M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba, “Evaluating large language models trained on code,” 2021. [Online]. Available: https://arxiv.org/abs/2107.03374 [2] J. Austin, A. Odena, M. I. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. J. Cai, M. Terry, Q. V. Le, and C. Sutton, “Program synthesis with large language models,” CoRR, vol. abs/2108.07732, 2021. [Online]. Available: https://arxiv.org/abs/2108.07732 [3] D. Hendrycks, S. Basart, S. Kadavath, M. Mazeika, A. Arora, E. Guo, C. Burns, S. Puranik, H. He, D. Song, and J. Steinhardt, “Measuring coding challenge competence with apps,” NeurIPS, 2021. [Online]. Available: https://arxiv.org/abs/2105.09938 [4] Y. Li, D. Choi, J. Chung, N. Kushman, J. Schrittwieser, R. Leblond, T. Eccles, J. Keeling, F. Gimeno, A. Dal Lago, T. Hubert, P. Choy, C. de Masson d’Autume, I. Babuschkin, X. Chen, P.-S. Huang, J. Welbl, S. Gowal, A. Cherepanov, J. Molloy, D. J. Mankowitz, E. Sutherland Robson, P. Kohli, N. de Freitas, K. Kavukcuoglu, and O. Vinyals, “Competition-level code generation with alphacode,” Science, vol. 378, no. 6624, p. 1092–1097, Dec. 2022. [Online]. Available: http://dx.doi.org/10.1126/science.abq1158 [5] Y. Li, D. Choi, J. Chung, N. Kushman, J. Schrittwieser, R. Leblond, T. Eccles, J. Keeling, F. Gimeno, A. Dal Lago, T. Hubert, P. Choy, C. de Masson d’Autume, I. Babuschkin, X. Chen, P.-S. Huang, J. Welbl, S. Gowal, A. Cherepanov, J. Molloy, D. Mankowitz, E. Robson, P. Kohli, N. de Freitas, K. Kavukcuoglu, and O. Vinyals, “Competition-level code generation with AlphaCode,” Science, vol. 378, no. 6624, pp. 1092–1097, 2022. [6] S. Lu, D. Guo, S. Ren, J. Huang, A. Svyatkovskiy, A. Blanco, C. Clement, D. Drain, D. Jiang, D. Tang, G. Li, L. Zhou, L. Shou, L. Zhou, M. Tufano, M. Gong, M. Zhou, N. Duan, N. Sundaresan, S. K. Deng, S. Fu, and S. Liu, “Codexglue: A machine learning benchmark dataset for code understanding and generation,” 2021. [Online]. Available: https://arxiv.org/abs/2102.04664

[7] S. Ren, D. Guo, S. Lu, L. Zhou, S. Liu, D. Tang, N. Sundaresan, M. Zhou, A. Blanco, and S. Ma, “Codebleu: a method for automatic evaluation of code synthesis,” 2020. [Online]. Available: https://arxiv.org/abs/2009.10297 [8] F. Cassano, J. Gouwar, D. Nguyen, S. Nguyen, L. Phipps-Costin, D. Pinckney, M.-H. Yee, Y. Zi, C. J. Anderson, M. Q. Feldman, A. Guha, M. Greenberg, and A. Jangda, “Multipl-e: A scalable and polyglot approach to benchmarking neural code generation,” IEEE Trans. Softw. Eng., vol. 49, no. 7, p. 3675–3691, Jul. 2023. [9] Q. Zheng, X. Xia, X. Zou, Y. Dong, S. Wang, Y. Xue, L. Shen, Z. Wang, A. Wang, Y. Li, T. Su, Z. Yang, and J. Tang, “Codegeex: A pre-trained model for code generation with multilingual benchmarking on humanevalx,” in Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, ser. KDD ’23. New York, NY, USA: Association for Computing Machinery, 2023, p. 5673–5684. [10] B. Athiwaratkun, S. K. Gouda, Z. Wang, X. Li, Y. Tian, M. Tan, W. U. Ahmad, S. Wang, Q. Sun, M. Shang, S. K. Gonugondla, H. Ding, V. Kumar, N. Fulton, A. Farahani, S. Jain, R. Giaquinto, H. Qian, M. K. Ramanathan, R. Nallapati, B. Ray, P. Bhatia, S. Sengupta, D. Roth, and B. Xiang, “Multi-lingual evaluation of code generation models,” 2023. [Online]. Available: https://arxiv.org/abs/2210.14868 [11] N. Jain, Han, A. Gu, W.-D. Li, F. Yan, T. Zhang, S. Wang, A. SolarLezama, K. Sen, and I. Stoica, “Livecodebench: Holistic and contamination free evaluation of large language models for code,” in International Conference on Learning Representations, Y. Yue, A. Garg, N. Peng, F. Sha, and R. Yu, Eds., vol. 2025, 2025, pp. 58 791–58 831. [12] Y. Lai, C. Li, Y. Wang, T. Zhang, R. Zhong, L. Zettlemoyer, W.-t. Yih, D. Fried, S. Wang, and T. Yu, “Ds-1000: a natural and reliable benchmark for data science code generation,” in Proceedings of the 40th International Conference on Machine Learning, ser. ICML’23, 2023. [13] X. Du, M. Liu, K. Wang, H. Wang, J. Liu, Y. Chen, J. Feng, C. Sha, X. Peng, and Y. Lou, “Classeval: A manually-crafted benchmark for evaluating llms on class-level code generation,” 2023. [Online]. Available: https://arxiv.org/abs/2308.01861 [14] A. Gu, B. Rozière, H. Leather, A. Solar-Lezama, G. Synnaeve, and S. I. Wang, “Cruxeval: a benchmark for code reasoning, understanding and execution,” in Proceedings of the 41st International Conference on Machine Learning, ser. ICML’24, 2024. [15] T. Y. Zhuo, V. M. Chien, J. Chim, H. Hu, W. Yu, R. Widyasari, I. N. B. Yusuf, H. Zhan, J. He, I. Paul, S. Brunner, C. GONG, J. Hoang, A. R. Zebaze, X. Hong, W.-D. Li, J. Kaddour, M. Xu, Z. Zhang, P. Yadav, N. Jain, A. Gu, Z. Cheng, J. Liu, Q. Liu, Z. Wang, D. Lo, B. Hui, N. Muennighoff, D. Fried, X. Du, H. de Vries, and L. V. Werra, “Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions,” in The Thirteenth International Conference on Learning Representations, 2025. [16] R. Puri, D. S. Kung, G. Janssen, W. Zhang, G. Domeniconi, V. Zolotov, J. Dolby, J. Chen, M. Choudhury, L. Decker, V. Thost, L. Buratti, S. Pujar, S. Ramji, U. Finkler, S. Malaika, and F. Reiss, “Codenet: A large-scale ai for code dataset for learning a diversity of coding tasks,” 2021. [Online]. Available: https://arxiv.org/abs/2105.12655 [17] B. Roziere, M.-A. Lachaux, L. Chanussot, and G. Lample, “Unsupervised translation of programming languages,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33, 2020, pp. 20 601–20 611. [18] T. Liu, C. Xu, and J. McAuley, “Repobench: Benchmarking repositorylevel code auto-completion systems,” ArXiv, vol. abs/2306.03091, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:259075246 [19] Y. Ding, Z. Wang, W. U. Ahmad, H. Ding, M. Tan, N. Jain, M. K. Ramanathan, R. Nallapati, P. Bhatia, D. Roth, and B. Xiang, “Crosscodeeval: a diverse and multilingual benchmark for cross-file code completion,” in Proceedings of the 37th International Conference on Neural Information Processing Systems, ser. NIPS ’23. Red Hook, NY, USA: Curran Associates Inc., 2023. [20] C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan, “Swe-bench: Can language models resolve real-world github issues?” ArXiv, vol. abs/2310.06770, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:263829697 [21] D. Zan, Z. Huang, A. Yu, S. Lin, Y. Shi, W. Liu, D. Chen, Z. Qi, H. Yu, L. Yu, and et al., “SWE-bench-java: A GitHub Issue Resolving Benchmark for Java,” arXiv e-prints, p. arXiv:2408.14354, Aug. 2024. [22] D. Zan, Z. Huang, W. Liu, H. Chen, S. Xin, L. Zhang, Q. Liu, L. Aoyan, L. Chen, X. Zhong, S. Liu, Y. Xiao, L. Chen, Y. Zhang, J. Su, T. Liu, R. LONG, M. Ding, and l. xiang, “Multi-swe-bench: A multilingual

benchmark for issue resolving,” in Advances in Neural Information Processing Systems, D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen, Eds., vol. 38, 2025. [23] Z. Feng, D. Guo, D. Tang, N. Duan, X. Feng, M. Gong, L. Shou, B. Qin, T. Liu, D. Jiang, and M. Zhou, “Codebert: A pre-trained model for programming and natural languages.” in EMNLP (Findings), ser. Findings of ACL, T. Cohn, Y. He, and Y. Liu, Eds., vol. EMNLP 2020, 2020, pp. 1536–1547. [24] D. Guo, S. Ren, S. Lu, Z. Feng, D. Tang, S. Liu, L. Zhou, N. Duan, A. Svyatkovskiy, S. Fu, M. Tufano, S. K. Deng, C. Clement, D. Drain, N. Sundaresan, J. Yin, D. Jiang, and M. Zhou, “Graphcodebert: Pre-training code representations with data flow,” 2021. [Online]. Available: https://arxiv.org/abs/2009.08366 [25] Y. Wang, W. Wang, S. Joty, and S. C. Hoi, “CodeT5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M.-F. Moens, X. Huang, L. Specia, and S. W.-t. Yih, Eds. Online and Punta Cana, Dominican Republic: Association for Computational Linguistics, Nov. 2021, pp. 8696–8708. [26] E. Nijkamp, B. Pang, H. Hayashi, L. Tu, H. Wang, Y. Zhou, S. Savarese, and C. Xiong, “Codegen: An open large language model for code with multi-turn program synthesis,” in International Conference on Learning Representations, 2022. [27] D. Fried, A. Aghajanyan, J. Lin, S. I. Wang, E. Wallace, F. Shi, R. Zhong, W. tau Yih, L. Zettlemoyer, and M. Lewis, “Incoder: A generative model for code infilling and synthesis,” ArXiv, vol. abs/2204.05999, 2022. [28] L. Ben Allal, R. Li, D. Kocetkov, C. Mou, C. Akiki, C. Munoz Ferrandis, N. Muennighoff, M. Mishra, A. Gu, M. Dey, and et al., “SantaCoder: don’t reach for the stars!” arXiv e-prints, p. arXiv:2301.03988, Jan. 2023. [29] Z. Luo, C. Xu, P. Zhao, Q. Sun, X. Geng, W. Hu, C. Tao, J. Ma, Q. Lin, and D. Jiang, “Wizardcoder: Empowering code large language models with evol-instruct,” in The Twelfth International Conference on Learning Representations, 2024. [30] B. Rozière, J. Gehring, F. Gloeckle, S. Sootla, I. Gat et al., “Code llama: Open foundation models for code,” arXiv preprint arXiv:2308.12950, 2023. [Online]. Available: https://arxiv.org/abs/2308.12950 [31] D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu, Y. K. Li, F. Luo, Y. Xiong, and W. Liang, “Deepseek-coder: When the large language model meets programming - the rise of code intelligence,” ArXiv, vol. abs/2401.14196, 2024. [32] J. Liu, C. S. Xia, Y. Wang, and L. Zhang, “Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation,” in Proceedings of the 37th International Conference on Neural Information Processing Systems, ser. NIPS ’23. Red Hook, NY, USA: Curran Associates Inc., 2023. [33] H. Tian, W. Lu, T. O. Li, X. Tang, S.-C. Cheung, J. Klein, and T. F. Bissyandé, “Is chatgpt the ultimate programming assistant – how far is it?” 2023. [Online]. Available: https://arxiv.org/abs/2304.11938 [34] S. E. Arefin, T. A. Heya, H. Al-Qudah, Y. Ineza, and A. Serwadda, “Unmasking the giant: A comprehensive evaluation of ChatGPT’s proficiency in coding algorithms and data structures,” in Proceedings of the 16th International Conference on Agents and Artificial Intelligence, vol. 1, 2024, pp. 412–419. [Online]. Available: https: //arxiv.org/abs/2307.05360 [35] T. A. Heya, Y. Ineza, S. E. Arefin, G. Uzor, and A. Serwadda, “Stable or shaky? the semantics of chatgpt’s behavior under repeated queries,” in 2024 IEEE 18th International Conference on Semantic Computing (ICSC), 2024, pp. 110–116. [36] D. Sobania, M. Briesch, C. Hanna, and J. Petke, “An analysis of the automatic bug fixing performance of chatgpt,” 2023 IEEE/ACM International Workshop on Automated Program Repair (APR), pp. 23–30, 2023. [37] C. Wang, K. Huang, J. Zhang, Y. Feng, L. Zhang, Y. Liu, and X. Peng, “Llms meet library evolution: Evaluating deprecated api usage in llm-based code completion,” in Proceedings of the IEEE/ACM 47th International Conference on Software Engineering, ser. ICSE ’25, 2025, p. 885–897. [Online]. Available: https: //doi.org/10.1109/ICSE55347.2025.00245 [38] B. Chen, F. Zhang, A. Nguyen, D. Zan, Z. Lin, J.-G. Lou, and W. Chen, “Evaluating large language models trained on code with humaneval and generated tests,” arXiv preprint arXiv:2207.10397, 2022. [Online]. Available: https://arxiv.org/abs/2207.10397 [39] N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: language agents with verbal reinforcement learning,” in

Proceedings of the 37th International Conference on Neural Information Processing Systems, ser. NIPS ’23. Red Hook, NY, USA: Curran Associates Inc., 2023. [40] S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, “React: Synergizing reasoning and acting in language models,” 2023, publisher Copyright: © 2023 11th International Conference on Learning Representations, ICLR 2023. All rights reserved.; 11th International Conference on Learning Representations, ICLR 2023 ; Conference date: 01-05-2023 Through 05-05-2023. [41] I. R. d. S. Simões and E. Venson, “Evaluating source code quality with large language models: a comparative study,” in Proceedings of the XXIII Brazilian Symposium on Software Quality, ser. SBQS ’24. New York, NY, USA: Association for Computing Machinery, 2024, p. 103–113. [Online]. Available: https://doi.org/10.1145/3701625.3701650

Related documents

Record · ID 267715 · SHA-256 e09cf04951ffc393
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.