Conceptio › Archive › arXiv CS
arXiv CSopen access

Specification Before Generation: A Pre-Registered, Five-Model Paired Evaluation of a Specification Frame for LLM-Generated Code in Money, Time, Idempotency, and Access Tasks

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

Specification Before Generation: A Pre-Registered, Five-Model Paired Evaluation of a Specification Frame for LLM-Generated Code in Money, Time, Idempotency, and Access Tasks Sandeep Dhuri

arXiv:2609.23270v1 [cs.SE] 20 Sep 2026

Independent Researcher, Edison, NJ, USA ORCID 0009-0008-1829-2431

security-specific prompting, so they measure what models do by default. Academic measurements of the same default point the same way: Pearce et al. found roughly 40 percent of Copilot completions vulnerable across 89 scenarios [2], and Tihanyi et al. found at least 62 percent of 331,000 C programs generated by nine models vulnerable under formal verification [3]. The engineering response has been largely informal. Teams write project instruction files (AGENTS.md, CLAUDE.md, cursor rules) on the belief that telling the model how the team works will make its output safer. The evidence for that belief is thin. The largest controlled study of instruction files we are aware of, a 438-task evaluation of agent instruction files by an ETH Zurich group, found that neither human-written nor model-generated instruction files improved task success, that they could even reduce it, and that they added more than 20 percent to cost [4]. Developer surveys add a second signal: in the 2025 Stack Overflow survey, 46 percent of respondents distrusted the accuracy of AI tools against 33 percent who trusted it, and 66 percent named solutions that are almost right but not quite as their leading frustration [5]. Almost right is the expensive kind of wrong in a payments system. This paper argues that the instruction-file null and the “almost right” complaint have a common cause: the prompt is missing a specification. A style guide tells the model how to write. A specification tells the model what must be true of the result. In regulated backend work the second category is small, stable, and well known to any senior engineer: money is not a float, I. I NTRODUCTION timestamps carry a zone, a retried request must not doubleBy Veracode’s account, AI now authors about half of charge, secrets do not live in source, passwords are stored only committed code in the organizations it measures [1]. The as slow salted hashes, and so on. The author’s book Delta: reliability of that code has not kept pace with its volume. Closing the Specification Gap [6] collected these rules into a Veracode’s 2026 GenAI Code Security Report, covering more fixed prompt preamble it calls the Specification Frame, on the than one hundred models across four years of releases, found argument that the distance between what a developer knows that generated code passed its security tests 56 percent of the and what a developer types into the prompt is where generated time and failed 44 percent of the time, with the failure rate code fails. The book presented the argument. It did not measure essentially unmoved since the first measurements [1]. Pass it. This paper does. rates on some individual models are higher, but the aggregate has not improved with model scale. Veracode’s tests use no A. The software engineering problem

Abstract—Code generated by large language models passes security checks at a rate that has barely moved in four years. In regulated backends, the defect classes that matter most are money arithmetic, time handling, retry safety, and access control. Teams answer with instruction files, yet the largest controlled study of instruction files to date found no benefit. This paper tests a narrower idea: generated code improves when the prompt carries a specification, a fixed preamble stating what must be true of the result. We pre-registered hypotheses, refuters, analysis code, and a one-shot generation rule, then ran 50 realistic backend tasks from finance, healthcare, and insurance practice through five frontier models from five vendor lineages, each task twice: bare, and preceded by a 267-word filled specification frame. Nine deterministic AST-based checkers scored the outputs. The Bandit security scanner, which knows nothing of the frame, scored them independently. The frame reduced defects in all five models (mean reduction 0.16 to 0.70 findings per task, every Holm-adjusted sign test significant, every bootstrap confidence interval excluding zero). Where the arms differed, the frame arm won 95 of 100 times. It never made any model worse in any domain. Bandit found 53 medium-or-high issues in the bare arm and 11 in the frame arm, in the same direction for every model. The effect was largest where a model’s unprompted defaults were weakest: the frame supplies the discipline a model lacks. All 500 outputs, prompts, checkers, scoring code, and the pre-registration are published with a DOI, so any team can re-derive the result without trusting the author. Index Terms—LLM code generation, prompt engineering, specification, software security, pre-registration, empirical software engineering, regulated industries

Preprint, September 2026. Prepared for submission to a peer-reviewed venue. Dataset: 10.5281/zenodo.22850887 (v1.1).

We state the software engineering problem explicitly. The artifact under study is a prompt preamble, a software engi-

neering artifact that is checked into repositories and evolved alongside the code it generates. The task is code generation for maintenance-heavy backend systems in domains where defects have regulatory and financial consequences. The question is whether a fixed specification artifact, prepended to task prompts without any per-task tuning, measurably reduces the named defect classes in generated code across current models, and whether the effect is consistent enough that a practitioner could adopt the artifact on the evidence. This is a question about a software engineering practice (specifying before generating) and a software engineering artifact (the frame), not about model internals. B. Contributions 1) A pre-registered paired evaluation of a specification preamble across five frontier models from five vendor lineages, with hypotheses, refuters, and analysis code frozen before generation and a one-shot rule that forbade regenerating any successful output. We are not aware of a prior pre-registered evaluation of a prompt-level specification artifact for code generation. 2) A consistent, corroborated effect. The frame reduced defects in all five models with Holm-adjusted significance, never made any model worse in any domain, and was independently corroborated by a security scanner that has no knowledge of the frame. 3) A gradient finding with practical consequences. The effect size tracked the weakness of each model’s unprompted defaults, from 0.16 findings per task in the most disciplined model to 0.70 in the least. The frame acts as a floor. 4) A fully re-scorable artifact. All 500 outputs, 100 prompts, 9 checkers, scoring code, seeds, per-file checksums, and the pre-registration are published under a DOI, and the scoring code re-derives every number in this paper from the raw bytes. The remainder of the paper covers background (Section II), the instrument (Section III), the study design (Section IV), results (Section V), discussion (Section VI), threats to validity (Section VII), deviations from the pre-registration (Section VIII), and conclusions, followed by the required Data Availability statement. II. BACKGROUND AND R ELATED W ORK Security of generated code. Measurements of LLM code security have been consistent in direction. Pearce et al.’s early study of Copilot [2] produced 1,689 programs across 89 CWEoriented scenarios and found about 40 percent vulnerable. Tihanyi et al.’s FormAI-v2 corpus [3] found at least 62 percent of 331,000 generated C programs vulnerable under formal verification with ESBMC in an unbounded, k-induction setting, with only minor differences between models. Blain and Noiseux [13] measured 55.8 percent across 3,500 artifacts from seven models using SMT-proven witnesses, and their ablation is directly relevant here: a generic security-explicit system prompt reduced the vulnerability rate by only four points.

RealSec-bench [18], 105 security instances drawn from real Java repositories, adds a warning: generic security-guideline prompting can lower compilation success without reliably preventing vulnerabilities, so any prompt-side intervention has to show that functionality survives. BaxBench [20] makes the same point from the benchmark side by scoring functional correctness and security of generated backend applications together. Veracode’s 2026 industry report [1] is the largest longitudinal measurement: a 56 percent pass rate that has not improved across four testing snapshots and more than 100 models, with Java the weakest language at roughly 30 percent. Our study does not replicate these measurements. It asks whether a specific kind of intervention, a specification rather than a generic security instruction, moves a specific subset of the failure classes they document. Instruction files and prompt artifacts. The closest prior work is the 438-task evaluation of agent instruction files [4], which found no improvement from either human-written or generated instruction files and a cost penalty above 20 percent. The authors conclude that instruction files should carry minimal requirements rather than volume. Our reading of that result is that it tests the wrong artifact class: instruction files as written in the wild are mostly process and style guidance. The artifact tested here is a specification, and the study is designed to isolate that difference. Related work on context engineering shows that coding agents favor retrieval recall over precision and use only a fraction of what they retrieve [7]. A multivocal review [8] names the resulting pattern a productivity-reliability paradox, with telemetry across more than 10,000 developers showing 98 percent more pull requests alongside 91 percent longer review times, and attributes it in part to insufficient specification discipline. Stoica et al. argue that specifications are the missing link that would make LLM system development an engineering discipline [11], and Patil et al. show that formally correct embedded code can be generated from specifications alone in automotive case studies [12]. Closest in spirit, Rosa et al. registered a Stage 1 report for a human-in-the-loop workflow in which developers refine a specification and tests inside an IDE before generating code [9]. That work treats specification as an interactive human activity. The present study measures specification as a fixed repository artifact, applied before any human iteration, which is the form in which most teams could adopt it. Spec-driven development evidence. Spec-driven development entered wide practice in 2025 through repository constitution files and tooling such as GitHub’s Spec Kit, and its evidence base is thin. Marri’s constitutional spec-driven development case study [14] is the closest in intent to this work: fifteen CWE-mapped principles in a versioned constitution reduced detected CWE violations from 11 to 3 in a banking application, but with one developer, one model, one project, and no statistical control, as its threats section states. Garg’s drift-review study [15] adds two cautions that shape how our result should be read: on easy tasks, much of the apparent gain from specifying first was reproduced by a reason-first control with no specification at all, and delivering a specification as

a governing artifact in a staged generation step outperformed and secrets, PII, PHI, and full card numbers are never logged. the same text inline. Related measurements of prompt form Task tells the model to implement exactly what the task asks on defect induction [16] and of prompting and fine-tuning and nothing more. Constraints adds seven rules, the first three strategies for secure generation [17] find that phrasing changes of which read verbatim: defect rates, though none isolates a fixed specification preamble Money arithmetic in decimal.Decimal or integer cents. Round half up only at the final step. Never compare money across vendors. Our study is the controlled, multi-vendor, prewith floating-point equality. registered measurement these works call for, on a narrower Any handler that can be retried (webhooks, payment calls) question: whether a fixed specification, delivered inline, moves is idempotent: it requires and checks a caller-supplied named defect classes. idempotency key before applying effects. Pre-registration. Registered reports are established in emPasswords are stored only as salted, slow hashes (for expirical software engineering, including at the major software ample hashlib.scrypt or pbkdf2_hmac with high iterations). engineering conferences. We adopted the practice for a compuNever plaintext, never fast unsalted digests. tational experiment because it removes the author’s ability to The remaining four constraints require redirect targets to be choose the analysis after seeing the data, which for a singlevalidated against an allow-list, file paths built from user input author, self-funded study is the most important threat to address. to be resolved and confirmed inside the intended base directory, Regulated-domain defect classes. The defect classes chosen inputs to be validated at the boundary with fail-closed errors, for this study are not novel. Floating-point money arithmetic and ambiguous requirements to be read in the strictest way maps to CWE-682 (incorrect calculation), floating-point equalconsistent with the task without inventing scope. ity on money to CWE-697 (incorrect comparison), timezoneThe frame states constraints, not procedures. It does not naive timestamps in business logic to no single CWE (the tell the model how to structure files, what style to use, how closest parent is CWE-682, incorrect calculation), hard-coded to name variables, or how to run tests, which is where most credentials to CWE-798, plaintext or fast unsalted password instruction files spend their words. Two design choices matter digests to CWE-916 and CWE-328, SQL built by string for interpreting the result. The frame is short, so the cost formatting to CWE-89, path traversal to CWE-22, and open concern raised for instruction files [4] does not arise at the redirects to CWE-601. Missing idempotency protection on same scale. And it is task-agnostic by construction, so a team retryable side-effecting handlers has no single CWE and is can adopt it as a repository-level artifact without per-task effort. defined as a domain rule in the study’s methods document. It is The trade-off is that it cannot encode task-specific requirements, a familiar incident class in payment and claims systems. They and the study makes no claim that it substitutes for them. were chosen because they are the classes a senior engineer in these domains would name unprompted. The frame writes IV. S TUDY D ESIGN down what the engineer already knows. A. Pre-registration and the one-shot rule III. T HE I NSTRUMENT: T HE S PECIFICATION F RAME The pre-registration was registered on 18 August 2026, The Specification Frame [6] is a fixed-structure prompt before any experimental run, and frozen at first commit, with preamble whose public template is published verbatim and is amendments permitted only as dated addenda. It fixed three independent of any task. Addendum D of the pre-registration hypotheses, three refuters, the protocol (single-turn generation, fixed how it enters the study: the template was instantiated temperature 0 requested, identical prompts across arms apart once, before generation, into a filled frame of 267 words from the frame, task order fixed by a hash of the task identifier, (FRAME_FILLED.md) for a shared regulated-Python context, scorers frozen and unit-tested before generation, all raw outputs with no placeholders and no links. Its SHA-256 is recorded in published with a hashed manifest), and the analysis. Four dated the prompt manifest of the dataset. The same bytes preceded addenda followed, all before generation: A (28 August) made every treatment-arm prompt for every task and every model. per-domain results descriptive only and fixed the reporting order, refuter verdicts first. B (1 September) extended the design to a Nothing in the frame refers to any task. The filled frame has four labeled sections, Role, Context, model roster with byte-identical prompts, made the per-model Task, and Constraints, and closes with the line “The task:” analysis primary, and declared pooled significance pseudoafter which the task text follows. Role places the model as replication. C (3 September) opened the roster to any provider a senior backend engineer in a regulated financial-services with keys, minimum three. D (5 September) raised n from and healthcare codebase whose code is reviewed against 40 to 50, fixed the frame instantiation, added the task-validity PCI DSS and HIPAA obligations before merge. Context clause (fixture proofs, 50 of 50 at freeze), the one-shot rule, fixes the stack (Python 3.12, standard library unless the task the seeded 10 percent adjudication clause, the nondeterminism says otherwise) and five domain rules: monetary amounts statement, and Holm-Bonferroni correction, and recorded the use decimal.Decimal or integer cents and never binary three pre-generation prompt edits. floating point, all datetimes are timezone-aware UTC and naive The one-shot rule states that after generation begins no datetimes are forbidden, SQL goes only through parameterized output is regenerated for any reason. Provider errors and empty placeholders, credentials and API keys come only from responses write no file and are retried under a documented environment variables and are never literals in source or logged, resume protocol until every task in every arm has exactly one

output. Any change after generation is a re-scoring, recorded in a dated verifier log with a full re-score, never a re-generation. The rule exists because a paired prompt study is easy to bias by regenerating outputs that disappoint, and hard to bias if every output is the first and only one.

TABLE I TASK INVENTORY. Domain

Identifiers

Tasks Registered checkers

Money

MO01 to MO12 ID01 to ID12 DA01 to DA10, DT11 to DT12 AU01 to AU14

12

Idempotency

B. Hypotheses and refuters Date and

12 12

money_float (12), exact_compare (11) idempotency_key (12), hardcoded_secret (1) naive_datetime (12), exact_compare (1)

Three hypotheses were registered. time H1 (primary). The frame arm yields fewer deterministiccheck findings per task than the bare arm. Under Addendum Authentication 14 hardcoded_secret (11), B this is tested within each model on the paired difference ∆ and access sql_param (12), password_hash (11), = A − B, where A is the bare arm and B is the frame arm. open_redirect (11), H2. The frame arm yields fewer Bandit findings of medium path_traversal (11), or higher severity per task than the bare arm. naive_datetime (1) Total 50 105 registrations H3. The effect concentrates in the money and datetime domains, where the frame’s exactness rules bind most directly. Under Addendum A this is evaluated descriptively only. Three refuters were registered as statements of what would detail. Each task carries a provenance row linking it to the count against the claim, with the commitment to publish them constraint it exercises. Every prompt opens with an imperative (“Write a function prominently and first. In the original wording: no significant that ...”, “Write an HTTP webhook handler that ...”) and paired difference on H1 means the measurable-safety claim closes with the same two sentences: “Use Python. Return is unsupported at this scale and is published as such. The complete, runnable code only.” Prompt scaffolding is therefore frame arm being worse in any domain is reported in the first a controlled variable rather than a source of variance. Tasks section of the results, not buried. H1 holding while H2 does not do not name the frame’s constraints. A money task asks for a narrows the public claim to exactness discipline, and the source refund computation and never mentions Decimal. text’s stronger phrasing goes to a public errata page. Under Each task registers the checkers its weak fixture must Addendum B the refuters were operationalized per model and trigger (Table I, Section IV-F). Registrations follow the domain: answered in writing by the author before any other reading of money tasks register the two money checkers (eleven of twelve the results, as three questions: register both), datetime tasks the datetime checker, idempotency • R1. Was the frame arm significantly worse than the bare tasks the idempotency checker, and the first ten authentication arm for any model or in any domain? tasks the full five-checker security panel, with the four added • R2. Did any pre-registered hypothesis fail at α = 0.05, authentication tasks registering subsets. Three tasks carry one after multiple-comparison correction, in a majority of cross-domain registration each: DT12 (float equality), ID11 models? (hard-coded secret), and AU14 (an audit-log search by date • R3. Where H2 could be scored, was its direction consistent range, naive datetime). There are 106 registrations in total, 11 with H1? to 13 per checker. A word-by-word editorial pass over all 50 prompts before C. Tasks generation, recorded in Addendum D, made three wording Fifty backend coding tasks were written in four domains changes: MO12 was replaced because its scenario near(Table I): money handling (identifiers MO01 to MO12), idempo- duplicated MO04 (Jaccard similarity 0.67), DT12 was anchored tency and retry safety (ID01 to ID12), date and time handling to a named time zone, and MO11’s percentage style was unified. (DA01 to DA10 plus DT11 and DT12), and authentication After the pass, the maximum pairwise similarity across tasks and access (AU01 to AU14). The two DT tasks are the two was below 0.30. No checker registration changed and no prompt datetime tasks added in Addendum D and carry a different changed after generation began. identifier prefix from the original ten. The split is an artifact of the expansion. The results tables report the two groups D. Arms separately because the scoring code groups by identifier prefix. Each task produced two prompts. Arm A (bare) is the task Each task is a realistic unit of backend work in these domains: text alone. Arm B (frame) is the 267-word filled frame, a blank a partial refund with proration, a monthly bill from tiered line, and the same task text, sent as a single user message. No usage rates, a webhook handler that records a payment event, system message was used in either arm. The pre-registration’s a login check against stored users, a password-reset token flow, “identical system message” was realized as the absence of one a meeting scheduled across IANA time zones, a nightly job in both arms, and its “both arms in the same session per task” at 02:30 in America/New_York that must survive daylight- as two consecutive independent calls within the same run, with saving transitions. The author wrote the tasks from practice in no conversational state shared between arms. The two prompts regulated backend systems. They contain no employer-specific for each task differ only by the frame. Both arms were sent to

the same model with identical request parameters. A prompt manifest records the SHA-256 of all 100 prompts and of the frame, so that a reader can verify byte identity of the task text across arms without trusting the author.

trigger and eleven known-good fixtures must stay clean. Second, fifty task fixture proofs: each task’s weak fixture triggers its registered checkers and its strong fixture passes all nine, machine-run through the same test file, 50 of 50 green at freeze and again on 19 September against the deposited checkers. E. Models Third, the pre-registered manual adjudication of a seeded 10 Five models from five vendor lineages were run (Table II). percent sample of real findings after the run (Section IV-H). The roster rule, fixed in Addendum C, was one flagship-tier The verifier log also records two checker defects found during model per provider, with the exact API identifier recorded development (a float-evidence miss in exact_compare and at run time and the served model identifier returned by the a missed case in sql_param, fixed with a light AST taint provider recorded per output. One provider’s flagship changed analysis) as permanent regression tests. These are reported on run day: the OpenAI model that leads Veracode’s Summer because a study of code quality should show its own. 2026 leaderboard at a 68 percent pass rate [1] was superseded, The pre-registration named and answered the obvious threat and we ran its successor. We state this plainly. One open-weight in advance: the frame states constraint classes that the checkers model on the planned roster was decommissioned by its host also measure. This is not teaching to the test. It is the before generation and was replaced by a hosted open-weight intervention itself. The claim under test is precisely that model from a different lineage. stating these constraints changes the code. Checkers score Temperature 0 was requested from every provider, as the code behavior, not keyword echo, identically and blindly on pre-registration states. Two vendor endpoints do not accept the both arms. A model that echoes the constraints and still writes parameter for the model families used, and it was omitted for float money fails the checker. them, so those two models ran at the provider default. This affects both arms of each model identically and is discussed in G. Secondary measurement: an instrument-blind scanner Sections VII and VIII. Generation was single-turn, one stateless Bandit 1.9.4, a widely used Python security scanner, was call per prompt, with no tools, no retrieval, and no multi-turn run on all 500 outputs. Bandit was chosen because it was interaction. Requests that returned a retryable status (429 or developed with no knowledge of the frame, the book, or this 5xx) were retried, up to seven attempts in total, with waits of study, and because its rule set overlaps only partly with the at least twenty seconds. Requests that returned an error or an nine checkers. It is therefore an independent measurement empty body wrote no file. Each model produced 100 outputs rather than a second reading of the same one. Only medium(50 tasks × 2 arms), for 500 outputs in total. Generation ran on or-high severity findings were counted. H2 is scored on this 5 and 6 September 2026 from a single operator machine. Total count. Per the checker validity document, Bandit is a secondary API cost was under twenty US dollars. Task order within each descriptive measurement and never enters the primary statistic. model run was fixed by the SHA-256 of the task identifier. F. Primary measurement: nine deterministic checkers

H. Manual adjudication

Each output was scored by nine deterministic checkers operating on the Python abstract syntax tree of the returned code. Where a model wrapped its code in Markdown fences, the scorer parsed the code out and recorded a stripped flag per file, per Addendum B. Raw outputs are published untouched. Each checker targets one named failure class (Table III). The design stance is precision-first: a checker flags only patterns defensible line by line to a hostile reviewer, and recall is deliberately sacrificed. A finding is a demonstrated instance of the failure class on the flagged line, not a style opinion. Each output is scored only by the checkers registered for its task (Table I), and a checker contributes at most one finding per output, so a file with three floating-point money operations counts once. The per-task score is therefore the number of registered checkers that fired, between zero and five, and ∆ for a task is the bare-arm count minus the frame-arm count. This scoring is conservative in both arms: defects outside a task’s registered classes are not counted, and repeated instances of one class are not counted twice. Checker validity was established in three layers, all shipped in the dataset from version 1.1. First, a unit gauntlet per checker, run before generation: eleven known-bad fixtures must

The pre-registration committed to a manual audit of a seeded 10 percent sample of all checker findings, to answer the objection that AST heuristics inflate counts. After scoring, 18 findings were drawn with seed 20260905 and adjudicated by the author against a strict standard: a finding is a true positive only if the flagged construct is the named defect in context, not merely a pattern match. Results are in Section V-E. I. Statistics The primary analysis is per model. For each model we report the mean paired difference ∆, a 95 percent confidence interval from a seeded bootstrap of the paired differences (seed 20260905), an exact two-sided sign test on the nonzero pairs, and a Wilcoxon signed-rank test. Because five per-model tests are run, sign-test p-values are Holm-Bonferroni adjusted and the adjusted values are reported as primary. The pooled analysis across 250 pairs is reported as descriptive only, per Addendum B. Per-domain results are descriptive only, per Addendum A. All analysis code was frozen before generation and is included in the dataset.

TABLE II M ODEL ROSTER AND GENERATION SETTINGS AS SENT. Lineage

Model (API identifier)

Endpoint

Output cap (tokens)

Temperature as sent

Anthropic

claude-sonnet-5

vendor Messages API

16000

OpenAI

gpt-5.6-sol

16000

Google

gemini-3.1-propreview deepseek-v4-pro

vendor Chat Completions API, max_completion_tokens vendor generateContent, v1beta

not sent (parameter deprecated for this family) not sent (parameter unsupported for this family) 0

vendor OpenAI-compatible endpoint

6000 final, see Section VIII 6000

DeepSeek Alibaba (Qwen)

qwen/qwen3.8-27b

Groq OpenAI-compatible endpoint

TABLE III C HECKER CLASSES , FROM THE DATASET ’ S CHECKER VALIDITY DOCUMENT. Checker

Failure class

Industry mapping

money_float

binary floating point on monetary values floating-point equality on money timezone-naive timestamps in business logic

CWE-682

exact_compare naive_datetime sql_param

string-built SQL reaching execute open_redirect redirect target taken from raw user input path_traversal user input joined into filesystem paths unresolved hardcoded_secret credential literals in source password_hash plaintext or fast unsalted digests for passwords idempotency_key retryable side-effecting handler with no deduplication

CWE-697 no single CWE (closest: CWE-682) CWE-89 CWE-601 CWE-22 CWE-798 CWE-916, CWE-328 domain rule, no single CWE

V. R ESULTS We report the refuter verdicts first, as the pre-registration requires, then the primary result, the domain breakdown, the independent scanner, the adjudication, and one finding that was not hypothesized. A. Refuter verdicts R1: was the frame arm significantly worse anywhere? No. In no model and in no domain did the frame arm record more findings than the bare arm. Of the 250 task pairs, 150 were ties and 100 differed. Of the 100 that differed, the frame arm had fewer findings in 95 and more in 5. R2: did any hypothesis fail at α = 0.05 in a majority of models? No. H1 held in all five models with Holm-adjusted p below 0.05. R3: was H2 consistent with H1? Yes. Bandit found fewer medium-or-high issues in the frame arm for every model, 53 versus 11 in aggregate. H3 (descriptive). Of the 125 findings removed across all models, money accounted for 58 (46 percent), idempotency for 30 (24 percent), the DA datetime tasks for 26 (21 percent), and

8192

0 0

claude-sonnet-5 deepseek-v4-pro gemini-3.1-pro-preview qwen3.8-27b gpt-5.6-sol 0.0 0.2 0.4 0.6 0.8 1.0 Mean paired difference (findings per task) Fig. 1. Mean paired difference ∆ (bare minus frame findings per task) per model, with seeded bootstrap 95 percent confidence intervals. All intervals exclude zero.

authentication for 11 (9 percent). Money and datetime together account for 67 percent of the reduction, so the hypothesis holds in the sense registered. It was incomplete: idempotency, which H3 did not name, contributed as much as datetime did. B. Primary result: per-model paired differences Table IV gives the per-model result. Every model shows a positive mean ∆ with a bootstrap confidence interval that excludes zero, and every Holm-adjusted sign test is significant. Fig. 1 plots the per-model ∆ with confidence intervals. Across all 250 pairs, the frame arm carried 23 findings against 148 in the bare arm, an 84 percent reduction in the named defect classes. The pooled figure is descriptive only, per Addendum B. C. Domain breakdown (descriptive) Table V shows findings by domain and model. Four patterns are visible. First, money is where the bare arm fails most and where the frame does the most work. Two models produced 25 and 23 money findings unprompted and dropped to 4 and 3 with the frame. Two others dropped to zero. The money rules in the frame occupy four sentences. Second, idempotency is the second-largest effect. Without the frame, the domain recorded 32 findings across its 60 barearm outputs. With it, two. This defect class has no CWE and

TABLE IV P RIMARY RESULT PER MODEL . ∆ = BARE - ARM FINDINGS MINUS FRAME - ARM FINDINGS PER TASK , N = 50 PAIRS PER MODEL . “I MPROVED ” COUNTS PAIRS WHERE THE FRAME ARM HAD FEWER FINDINGS , OVER PAIRS THAT DIFFERED . Model

Findings A (bare)

Findings B (frame)

Mean ∆

95% CI (bootstrap)

Improved / differing

Sign test p

Holm p

Wilcoxon p

claude-sonnet-5 deepseek-v4-pro gemini-3.1-propreview qwen3.8-27b gpt-5.6-sol All (descriptive)

40 22 44

6 2 9

0.68 0.40 0.70

[0.38, 1.02] [0.24, 0.58] [0.38, 1.08]

21 / 22 17 / 17 24 / 26

1.1e-5 1.5e-5 1.0e-5

4e-5 4e-5 4e-5

6.3e-5 8.0e-5 3.1e-5

30 12 148

2 4 23

0.56 0.16 0.50

[0.36, 0.76] [0.04, 0.28]

24 / 25 9 / 10 95 / 100

2.0e-6 0.021

1e-5 0.021

1.3e-5 0.021

TABLE V F INDINGS BY DOMAIN AND MODEL , BARE ARM (A) VERSUS FRAME ARM (B). D ESCRIPTIVE ONLY. Domain (pairs)

claude A/B

deepseek A/B

gemini A/B

qwen A/B

gpt-5.6 A/B

Total A/B

Money (12) Idempotency (12) Date and time, DA (10) Date and time, DT (2) Authentication and access (14)

25 / 4 5/1 6/0 0/0 4/1

8/0 8/0 5/1 0/0 1/1

23 / 3 10 / 1 6/1 0/0 5/4

8/0 7/0 8/0 0/0 7/2

2/1 2/0 6/3 0/0 2/0

66 / 8 32 / 2 31 / 5 0/0 19 / 8

lies outside the rule sets of security scanners such as Bandit, Bandit found 11 issues in the bare arm and none in the frame which is one reason tooling has not caught it. arm. Bandit’s rules were written by people who have never seen Third, the two DT tasks produced zero findings in every model and both arms. Both prompts foreground time zones in the frame. A reader who suspects that the nine checkers were the task text itself: a billing cutoff for a company operating designed to detect exactly what the frame tells the model to across US time zones, and a nightly job at 02:30 in Amer- avoid (which is true, and is the design) has in Bandit a second ica/New_York that must survive daylight-saving transitions. instrument with no such coupling, pointing the same way. When the task itself carries the constraint, the frame has nothing E. Manual adjudication to add, which is what a specification account predicts. We offer The seeded sample of 18 checker findings (8 money_float, this as an observation rather than a test, since three of the ten 4 naive_datetime, 3 exact_compare, 1 hardcoded_secret, 1 DA tasks also mention time zones and we did not analyze open_redirect, 1 idempotency_key) was adjudicated under findings per task. The DT null is a clean result. the strict standard described in Section IV-H. All 18 were The authentication and access domain shows the smallest true positives. The sample is small and was sized by the preand least consistent reduction. Two effects appear to combine. registered 10 percent clause, so we do not claim a precision Frontier models appear to be trained toward secure defaults for figure beyond “no false positive found in a 10 percent seeded the most publicized classes in this domain (password hashing, audit.” The adjudication records, including the flagged code SQL parameterization), so the bare arm is relatively clean. for each finding, are in the dataset. Veracode’s CWE breakdown points the same way: models pass SQL injection tasks 83 percent of the time, against 15 percent F. An unhypothesized finding: the effect tracks baseline discifor cross-site scripting [1]. And the domain’s remaining defects pline (open redirects, path traversal) are the kind that a generic The five models order by effect size in the reverse of their constraint may not fully reach. bare-arm cleanliness. The smallest effect (0.16) belongs to the model with the cleanest bare arm (12 findings). The two largest D. Independent scanner (H2) effects (0.68 and 0.70) belong to the models with the weakest Bandit reported 53 medium-or-high severity issues across the bare arms (40 and 44 findings). The relationship across all five 250 bare-arm outputs and 11 across the 250 frame-arm outputs. is monotone in rank. Fig. 2 plots bare-arm findings against ∆. Per model, the number of task pairs in which the bare arm had The ordering agrees with the one external measurement of the more Bandit issues than the frame arm, against the reverse, same lineages: on Veracode’s Summer 2026 leaderboard [1], was 10 to 0 (claude-sonnet-5), 8 to 2 (deepseek-v4-pro), 8 to Gemini 3.1 Pro sits at a 52 percent security pass rate, near the 2 (gemini-3.1-pro-preview), 7 to 1 (qwen3.8-27b), and 9 to bottom, while the OpenAI flagship leads at 68 percent. 0 (gpt-5.6-sol). The direction is consistent with H1 in every We did not hypothesize this and we report it as observational. model. For gpt-5.6-sol, the model with the smallest H1 effect, But it has a direct practical reading. The frame does not

Mean paired difference

0.8 0.6

gemini-3.1-pro-preview qwen3.8-27b

0.4

claude-sonnet-5

deepseek-v4-pro

magnitudes across studies, but the pattern is consistent: generic instruction does little, stated constraints do a lot. Our proposed synthesis, which we intend to test directly in follow-on work by delivering the frame through an AGENTS.md file in an agentic loop, is that instruction files do not help unless they are specifications. C. Implications for practice

Three implications follow for the industrial audience. First, the cheapest intervention available to a regulatedindustry team is to write down what its senior engineers already know as constraints and put that text in front of every generation. 0.0 The frame is 267 words. Its cost per call is a rounding error 0 10 20 30 40 50 against the cost of a single production incident in a payments Bare-arm findings, 50 tasks or claims system. Second, the defect classes with the largest effect, money Fig. 2. Bare-arm findings (50 tasks) against mean paired difference ∆, one arithmetic and idempotency, are outside the rule sets of security point per model. The effect is monotone in the weakness of the unprompted scanners. Teams that measure generated-code quality by scanner baseline. findings alone will not see these defects and will not see the frame’s largest benefit. Domain-specific checkers of the kind add discipline on top of a disciplined model. It supplies used here are cheap to write and should sit beside the scanner. the discipline a model lacks, and it does so proportionally. Third, the gradient finding argues for keeping the artifact Vendors rotate models beneath fixed product names, so most in the repository, where the team controls it, instead of in the teams do not control which model their tooling routes to. A prompt. Model routing is increasingly outside the developer’s repository-level specification artifact gives such a team a floor control. A specification that travels with the code is the only on engineering discipline that does not depend on the model part of the generation context the team owns. behind the interface. D. Implications for research VI. D ISCUSSION The study is small in one dimension (50 tasks) and unusual A. What the result does and does not say in another (five vendor lineages, pre-registered, one-shot). The result: a fixed, short, task-agnostic specification preamble Prompt-artifact studies that regenerate until the result looks reduced the named defect classes in code from five current right, or that choose the analysis after the data, cannot answer frontier models, consistently, and an independent instrument whether an artifact works. Pre-registration is inexpensive for a agreed. The frame never hurt. The effect is largest where computational study and should be the default for this line of work. models are weakest. The DT null is also a research signal. The two tasks that Three things the result does not show. It does not show that code generated with the frame is correct: the checkers carry their time-zone constraint in the task text produced no measure nine named classes, and a module can pass all nine findings in any model or arm. A specification account predicts and still be wrong. It does not show that the frame substitutes that the frame’s effect is a function of how much constraint for task-specific requirements, review, or tests. And it does not the task text already carries. A follow-on study could vary that show that the effect will hold for models released after the run, directly, and should also test the delivery question Garg raises though the gradient finding suggests that as models improve, [15]: the same frame supplied inline, as here, against the same the frame’s marginal effect shrinks toward the floor set by the frame supplied as a governing artifact in a staged generation step. most disciplined model and does not vanish.

0.2

gpt-5.6-sol

B. Specification versus instruction

VII. T HREATS TO VALIDITY

The instruction-file study [4] and this study are not in conflict. Construct validity. The nine checkers measure named defect They tested different artifact classes. A typical instruction file classes, not correctness or security in general. They were tells the model how the team works. The frame tells the model designed by the same author who designed the frame, so what must be true of the output. The 438-task result is that a checker could in principle detect the frame’s phrasing rather the first kind of artifact does not help and costs tokens. Blain than the defect. Three mitigations are in place: checkers operate and Noiseux’s ablation [13] points the same way for generic on the AST and never on prose, so a model that merely security instructions: a system prompt asking for best practices echoes the frame’s language gains nothing. Their validity moved their vulnerability rate by four points. The 500-output was established on hand-written fixtures before generation. result here is that a concrete specification helps, costs little, And Bandit, an instrument with no coupling to the frame, and never hurts. The measures differ and we do not compare corroborates the direction in every model. The recall of the

checkers is limited and documented, only registered classes are scored per task, and a class counts at most once per output, so the absolute counts understate the defects present. The paired design means this understatement applies to both arms equally. A further threat is that the frame arm adds 267 words and therefore more deliberation before code, so part of its effect could be reasoning rather than specification content [15]. Two facts argue that content matters here: the checkers score constraint-specific behavior (Decimal arithmetic, aware datetimes, idempotency keys) that a model has no reason to adopt from extra deliberation alone, and the two tasks that carried their own timezone constraint showed no frame effect at all, which is what a content account predicts and a deliberation account does not. But neither a length-matched neutral preamble nor a reason-first arm was run, so the confound is not excluded. Both are pre-registered arms of the follow-on study. Finally, the study did not measure functional correctness. Every output was required to be code, and no refusal or noncode output occurred, though 16 outputs were truncated in transport (Section VIII, item 9). Compile success and behavior against per-task tests were not scored, so a trade-off of the kind RealSec-bench reports for generic security prompts [18] cannot be ruled out here. A post hoc compile check (Python byte-compilation after fence stripping, exploratory, not preregistered) found 247 of 250 bare-arm and 236 of 250 framearm outputs compile. All 14 frame-arm failures and 2 of the 3 bare-arm failures are the truncated files, so among complete outputs compile success was 247 of 248 bare and 236 of 236 frame. Dai et al. [19] show the risk is real for the class of intervention: evaluated on functionality and security together, several secure-generation techniques degraded the base model by more than half, often by deleting the vulnerable lines or emitting code unrelated to the task. Per-task functional tests, scored jointly with the defect classes, are part of the follow-on design. Internal validity. The two arms differ only by the frame, verified by prompt checksums. Decoding parameters were held constant within each model. Two vendor endpoints did not accept the temperature parameter and it was omitted for them (Section VIII), so two models ran at their provider default. This affects both arms of each model equally and does not touch the paired comparison, but it does mean the pre-registered phrase “temperature 0” describes what was requested, not what every provider applied. The one-shot rule and the frozen analysis remove selection of outputs and of analyses. Resume after transient provider failures was pre-registered and does not regenerate any successful output. The author executed the study alone. The pre-registration, unit-tested scorers, per-model manifests with checksums, and full publication of raw outputs are the mitigation. A reader does not have to trust any of this: the dataset ships the scoring and analysis code that re-derives every figure in this paper from the raw bytes. External validity. Fifty tasks in four domains, written by one practitioner, in Python, through single-turn generation with no tools, do not represent all backend work. The tasks are realistic for finance, healthcare, and insurance backends but

are not sampled from a corpus. The models are five current flagships and will be superseded. The frame was instantiated once and not tuned per task. Results may differ for agentic multi-turn generation, other languages, and other domains. We make no claim beyond the population studied. Conclusion validity. Sample size is 50 pairs per model. The exact sign test on nonzero pairs discards ties (150 of 250 pairs tied), and the effect was still significant in every model after Holm adjustment. Confidence intervals come from a seeded bootstrap. The pooled result and the domain breakdown are descriptive and we draw no inferential claim from them. One run per arm means we did not measure within-model variance. Modern inference endpoints are not fully deterministic even at temperature zero, so an exact re-run will not reproduce every output byte for byte. Reproduction of the analysis is exact from the published outputs. Reproduction of the generation is approximate by nature and is stated as such. VIII. D EVIATIONS FROM THE P RE -R EGISTRATION Registered Report guidance at software engineering conferences asks that all deviations be documented in a section of the paper. We follow that practice although this paper is not itself a registered report. Every item below is also recorded in the dataset, either as a dated note in the provider source, a run-log entry, or a verifier-log entry. 1) Roster change before generation. One model on the planned roster was decommissioned by its provider before any call was made. It was replaced by a hosted openweight model from a different lineage under the openroster rule of Addendum C. No output was affected. 2) Flagship succession on run day. One provider’s flagship model changed on the day of the run. The successor flagship was used. The paper therefore tests the successor, and says so. 3) Temperature not accepted by two endpoints. The preregistration requests temperature 0 from every provider and the per-model manifest records the requested value as a harness constant. Two vendor endpoints (Anthropic, OpenAI) do not accept the parameter for the model families used, and the provider layer omits it for them, so those two models ran at the provider default. The provider source records the omission in a dated note. The other three models received temperature 0. 4) Output caps adapted during the run. The Anthropic cap was raised from 4000 to 16000 after the first cap starved several frame-arm answers to empty, because the model’s reasoning tokens share the output budget. The OpenAI endpoint required max_completion_tokens at 16000. The Google endpoint required the v1beta path and a preview identifier at 8192. The OpenAI-compatible cap was raised to 65536 to complete a residue of DeepSeek tasks cut off at the earlier cap, then lowered to 16384 and to 6000 for the Qwen leg as the host’s request rejections and per-minute token accounting required. Retries went from four to

seven attempts with waits of at least twenty seconds. All IX. C ONCLUSION changes affect both arms of the affected model equally. We pre-registered and ran a paired evaluation of a 267-word 5) Resume after transient failures. Empty or errored calls specification frame across 50 regulated-domain backend tasks during the Anthropic, DeepSeek, and Qwen legs wrote and five frontier models from five vendor lineages, generating no file and were retried under the pre-registered resume each output exactly once. The frame reduced the named defect protocol until each task had exactly one output per arm. classes in every model, with Holm-adjusted significance, never No successful output was replaced. The final dataset has made any model worse in any domain, and was corroborated no exclusions: all 500 outputs are present, 16 of them by an independent security scanner. The effect was largest truncated as described in item 9. The run log preserves where models were weakest, which makes the frame a floor on the retry history. engineering discipline that a team can own regardless of which 6) Secondary scanner execution. Bandit did not execute model sits behind its tooling. The whole result is re-derivable on the first scoring pass because of an environment path from published bytes. issue. It was run in a second scoring pass, recorded The practical instruction is short. Before a model generates in the verifier log. This was a re-scoring of existing code that touches money, time, retries, or access, put the outputs, which the pre-registration permits. No output specification in front of it. The evidence says it helps, it costs was regenerated. almost nothing, and it does not hurt. 7) Protocol wording realized. The pre-registration’s “idenD ISCLOSURE tical system message” was realized as no system message in either arm. Its “both arms in the same session per task” This work was self-funded. The author wrote the book that was realized as two consecutive stateless calls within the introduced the frame [6] and has no other competing interests. same run. DATA AVAILABILITY 8) Metadata fossil in the task file. The header of All materials are published under a DOI on Zenodo: the tasks.json in the published dataset retains the pre50 tasks and 100 prompts with checksums, the filled frame, Addendum-D count field n: 40 while its task array all 500 model outputs, the nine checkers with their unit holds 50 tasks. The array is authoritative and the harness gauntlet and fifty task fixture pairs (dataset version 1.1), the reads the array. The field will be corrected in dataset scoring and statistics code with seeds, the Bandit results, the v1.1. adjudication records, the verifier log, the run log, the pre1) Truncated outputs found after deposit. A byte-level registration with its four addenda, so every figure in this audit of the published archive on 19 September found 16 paper can be re-derived from the raw files. Dataset DOI: of the 500 outputs cut off mid-statement with an unclosed 10.5281/zenodo.22850887 (version 1.1, 19 September 2026) code fence: 14 from Claude (12 frame-arm, 2 bare-arm), [10]. Version 1.0 is 10.5281/zenodo.22598205. The harness 1 from DeepSeek (frame arm), and 1 from Gemini (frame fingerprint (SHA-256 of the exact code that ran) is recorded arm). Their sizes (1,251 to 12,194 characters) rule out in the dataset’s environment file. the output cap, so the cause was transport truncation the harness did not detect. Because the fence stripper leaves an unclosed fence in place, the AST-based checkers could not parse these files and returned no findings for them, while the regex-based checkers still ran (one finding was recorded on a truncated file). This can only bias the result toward the frame arm, since 14 of the 16 are frame-arm files. Sensitivity analysis excluding every pair with a truncated file in either arm: Claude 38 pairs, mean difference 0.84 (bootstrap 95 percent CI 0.47 to 1.29), 19 of 20 differing pairs improved, Holm-adjusted sign test p = 4.0 x 10^-5. DeepSeek: 49 pairs, 0.41 (0.24 to 0.57), 17 of 17. Gemini: 49 pairs, 0.71 (0.41 to 1.10), 24 of 26. Qwen and GPT: unchanged. Every per-model conclusion holds and the direction remains five of five. The 16 files are listed in the dataset’s version 1.1 note and remain published exactly as received. Nothing was regenerated.

No deviation changed a hypothesis, an analysis, or a successful output.

R EFERENCES [1] Veracode, “2026 GenAI Code Security Report,” Jul. 28, 2026. https://www.veracode.com/resources/analyst-reports/2026-genai-codesecurity-report/ [2] H. Pearce, B. Ahmad, B. Tan, B. Dolan-Gavitt, and R. Karri, “Asleep at the Keyboard? Assessing the Security of GitHub Copilot’s Code Contributions,” in Proc. IEEE Symposium on Security and Privacy, 2022. [3] N. Tihanyi, T. Bisztray, M. A. Ferrag, R. Jain, and L. C. Cordeiro, “How Secure is AI-Generated Code: A Large-Scale Comparison of Large Language Models,” Empirical Software Engineering, 2025, doi: 10.1007/s10664-024-10590-1 (preprint arXiv:2404.18353). [4] T. Gloaguen, N. Mündler, M. Müller, V. Raychev, et al., “Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?” arXiv:2602.11988, Feb. 2026. [5] Stack Overflow, “2025 Developer Survey: AI,” 2025. https://survey.stackoverflow.co/2025/ai [6] S. Dhuri, Delta: Closing the Specification Gap. Acuity Press, 2026. DOI 10.5281/zenodo.21584309. [7] H. Li, L. Zhu, B. Zhang, R. Feng, J. Wang, Y. Pan, E. T. Barr, F. Sarro, Z. Chu, and H. Ye, “ContextBench: A Benchmark for Context Retrieval in Coding Agents,” arXiv:2602.05892, Feb. 2026. [8] S. E. Farrag, “The Productivity-Reliability Paradox: SpecificationDriven Governance for AI-Augmented Software Development,” arXiv:2605.01160, May 2026. [9] G. Rosa, D. Moreno-Lumbreras, G. Robles, and J. M. GonzálezBarahona, “Understanding Specification-Driven Code Generation with LLMs: An Empirical Study Design,” arXiv:2601.03878, Jan. 2026 (Stage 1 Registered Report).

[10] S. Dhuri, “Delta Frame Validation Study: complete results dataset (500 generations, 5 model families, pre-registered paired evaluation),” Zenodo, version 1.1, Sep. 19, 2026. DOI 10.5281/zenodo.22850887 (version 1.0: 10.5281/zenodo.22598205). [11] I. Stoica et al., “Specifications: The Missing Link to Making the Development of LLM Systems an Engineering Discipline,” arXiv:2412.05299, Dec. 2024. [12] M. S. Patil, G. Ung, and M. Nyberg, “Towards Specification-Driven LLMBased Generation of Embedded Automotive Software,” arXiv:2411.13269, Nov. 2024. [13] D. Blain and M. Noiseux, “Broken by Default: A Formal Verification Study of Security Vulnerabilities in AI-Generated Code,” arXiv:2604.05292v2, Apr. 2026. [14] S. R. Marri, “Constitutional Spec-Driven Development: Enforcing Security by Construction in AI-Assisted Code Generation,” arXiv:2602.02584, Jan. 2026. [15] N. Garg, “When Spec-Driven Development Pays Off,” InfoQ, Sep. 10, 2026, reporting a study accepted at GAISS 2026. [16] B. Wang, Y. Zhong, M. Wan, W. Yu, Y. Ouyang, Y. Huang, and H. Li, “Is Your Prompt Poisoning Code? Defect Induction Rates and Security Mitigation Strategies,” Empirical Software Engineering, 2026, doi: 10.1007/s10664-026-10866-8 (preprint arXiv:2510.22944). [17] A. Soltanian Fard Jahromi, A. Tahir, P. Liang, and F. Khomh, “On Fixing Insecure AI-Generated Code through Model Fine-Tuning and Prompting Strategies,” arXiv:2605.05867, May 2026. [18] Y. Wang, Z. Zhang, C. Wang, X. Xu, M. Liu, Y. Wang, J. Chen, and Z. Zheng, “RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories,” in Findings of the Association for Computational Linguistics: ACL 2026 (preprint arXiv:2601.22706). [19] S.-C. Dai, J. Xu, and G. Tao, “Rethinking the Evaluation of Secure Code Generation,” in Proc. IEEE/ACM 48th International Conference on Software Engineering (ICSE 2026), Research Track. [20] M. Vero, N. Mündler, V. Chibotaru, V. Raychev, M. Baader, N. Jovanović, J. He, and M. Vechev, “BaxBench: Can LLMs Generate Correct and Secure Backends?” in Proc. 42nd International Conference on Machine Learning (ICML 2025), PMLR 267, pp. 61344-61390.

Record · ID 1028617 · SHA-256 0b3a9896c7c6902a
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.