No Accidental Software Agent-First Canonical Code for Human-Code Entropy Reduction and 30×–500× Lower Frontier-Model Requirements Jepson Taylor NeverHuman Research
arXiv:2606.14357v1 [cs.SE] 12 Jun 2026
Abstract The number is the hook, but the denominator is the science: frontier coding models are paying to learn human-code entropy, not just code. Raw human repositories contain behavior, incidents, tests, edge cases, migrations, product judgment, and operational scar tissue worth preserving, but those signals are entangled with accidental representation: language fashion, framework churn, naming drift, folder folklore, duplicated contracts, generated-source confusion, dependency rituals, continuous-integration dialects, weak proof routes, and review customs built for humans rather than agents. A coding agent pays for this entropy four times: during training, during context gathering, during reasoning/tool/retry loops, and during human review. This paper proposes agent-first canonical code: a governed, proof-carrying substrate that ports routine product software into canonical profiles, behavior cells, generated truth, typed change algebra, proof lanes, constrained edit grammars, reasoning digests, semantic patch cells, runtime negative memory, and proof-carrying change objects before training and before agent operation. The claim is deliberately ambitious, but every range has a denominator. Detailed conservative, central, and aggressive bands are hypotheses reported in the tables, not measured frontier results, and they must not be multiplied into a miracle. The decisive endpoint is all-in cost per verified correct change: source tokens, context, reasoning, tools, verification, security, provenance, review, failed loops, downstream defects, and amortized foundry cost under the same acceptance oracle. The theoretical endpoint is the No-Accident Horizon: under a declared oracle and supported routine-product distribution, all removable accident can vanish only until the residual floor of novelty, evidence, governance, risk, and future optionality remains; a mature foundry therefore has a strongest defensible planning limit near 100× all-in verified-change cost reduction, not an unrestricted all-code guarantee. The theoretical spine is quotienting software by behavior equivalence under a declared oracle. Many behavior-equivalent human encodings should collapse to one governed representative plus evidence, while residual novelty, legal risk, security exposure, and domain mismatch remain explicit in a disposition ledger. The limit is M INIMUM F UNCTIONAL D ESCRIPTION L ENGTH: the shortest canonical behavior specification, evidence bundle, proof obligations, and renderer that produce a working system under behavior, security, migration, provenance, and review constraints. Preliminary QLoRA evidence on Qwen2.5-Coder-14B shows that 64,088 canonically translated trajectories are learnable and suppress tested forbidden-language markers; it does not establish behavior preservation, scaling economics, or verified-change cost. The contribution is a falsifiable research program whose decisive endpoint is amortized cost per verified correct change. Index Terms code language models, agent-first programming, canonical code, training data, corpus compression, behavior cells, software engineering, model efficiency
One-sentence thesis. Raw repositories force models to learn both product behavior and accidental implementation orbits; canonical proof-carrying behavior lets them spend training, inference, review, and verification budget on verified change instead of rediscovering local software folklore. What would make this undeniable. On paired raw versus canonical repositories from the same lineage, the same model must spend materially fewer source, context, reasoning, tool, retry, and reviewer tokens to produce an accepted proof-carrying change. Only after that same-model result holds should specialist scaling curves be allowed to justify the 30×–150× central training-token claim and the 8.3×–10× dense-active serving-speed claim. I. I NTRODUCTION The conservative claim is already large: 10×–30× fewer training tokens to a fixed accepted-change capability, 3×– 10× fewer context/reasoning/tool/retry tokens per verified change, and 3×–10× lower all-in cost per verified correct change after foundry amortization. The central claim is larger: 30×–150× training-token efficiency, 10×–100× fewer reason-
ing/action/retry tokens, and 10×–50× lower verified-change cost on supported canonical product work. The denominator comes first because the thesis is not “better style.” It is that frontier coding models are paying to learn the wrong object. Frontier coding models are trained on a software record built by humans for humans. That record contains valuable behavior, but it is not a canonical representation of software. It
TABLE I T HE H OOK : C ONSERVATIVE - TO -AGGRESSIVE PAYOFF R ANGES BY M EASUREMENT D ENOMINATOR Outcome denominator
Conservative
Central
Aggressive
Training tokens to fixed acceptedchange target Context, reasoning, tool, and retry tokens per verified change
10×–30×
30×–150×
150×–1,000×
3×–10×
10×–100×
100×–10,000×
Cost per verified correct change
3×–10×
10×–50×
50×–1,000×
Same-model verified-change wallclock
2×–5×
5×–20×
20×–100×
Dense-active inference speed scenario
3×–5×
8.3×–10×
10×–20×
Dense-active infrastructure cost/token Effective action/representation space
67%–80% lower
88%–90% lower
90%–95% lower
10×–40×
40×–150×
100×–300×
Foundry amortization and reuse
2×–5×
10×–100×
100×–10,000×
What must be measured Paired raw/canonical scaling curves to the same hiddentest accepted-change capability. Files opened, planning tokens, hidden reasoning budget where measurable, tool calls, failed repair loops, and validation reruns. End-to-end dollars including source/context/action tokens, serving, verification, security/provenance, review, downstream defects, and amortized foundry cost. End-to-end elapsed time from issue receipt to accepted proof-carrying patch under equal tool and reviewer budgets. Conditional only: matched-quality canonical specialist versus 1T-class dense-active baseline on supported canonical work. Active-parameter, memory, batching, utilization, MoE, context, and verification-overhead accounting. Entropy of valid behavior-equivalent encodings and legal edit choices after contracts, profiles, and proof lanes are fixed. Reuse of behavior cells, semantic patches, negative memory, generated projections, and proof receipts across independent product lineages.
TABLE II E XTERNAL FACTS THE C ANONICAL -C ODE T HESIS M UST E XPLAIN Observed fact
Why it matters
Canonical-code interpretation
Public software is massive and multi-language: The Stack v2 contains over 3B files in 600+ languages and StarCoder2 trains on 3.3T–4.3T tokens across 619 languages [1], [2]. Duplication is structural: DéjàVu found 85M unique files among 428M files across 4.5M nonfork projects, and Jupyter studies found more than 70% exact snippet copies [3], [4]. Agent benchmarks are brittle: SWE-bench Verified improves curation, but UTBoost found insufficient tests and erroneous pass labels in SWE-bench-family results [5], [6]. Agent cost matters: resourceconstrained evaluation shows why resolve rate alone hides token, time, and failure cost [7], [8]. CI/security sprawl is real: recent GitHub Actions studies report workflow heterogeneity and security-practice gaps, and GitGuardian reported 28.65M new public-GitHub secrets in 2025 [9]– [12].
Broad code models pay to learn the whole human archive.
The archive is behavior raw material, not necessarily the optimal training substrate.
Human code repeats behavior and representation.
Cells and semantic patches should capture repeated behavior directly.
Passing tests is not equivalent to preserving behavior.
Ports and changes need evidence tiers, hidden tests, fuzzing, replay, mutation tests, and proof receipts.
Accuracy without cost accounting can be economically misleading.
The denominator must be cost per verified correct change.
Repository operations are part of the software behavior surface.
CI, secrets, permissions, and release policy must become canonical substrate objects.
is a survival archive: hundreds of languages, millions of local conventions, duplicated frameworks, inconsistent naming, weak tests, dependency sprawl, build folklore, continuous-integration (CI) dialects, repository settings, access policies, and release
rituals that were never designed for autonomous agents. A coding agent pays for that archive four times: during training, during context gathering, during reasoning/tool/retry loops, and during human review.
TABLE III F RONT-D OOR M EASUREMENT L EDGER Question
Primary denominator
First test
Failure mode
Does the substrate help before new training?
Same-model cost per accepted change
Broad model on paired raw vs. canonical repositories with the same issue lineage, hidden tests, and reviewer rubric.
Does canonical data train more efficiently?
Tokens-to-target accepted-change capability
Paired scaling curves from the same raw/canonical lineage.
Does the serving denominator improve?
Dense-active infrastructure cost/token and token/sec at matched quality Preservation tier and hidden oracle pass rate
Compare a broad dense-active baseline against a canonical specialist on supported canonical work. Differential tests, replay, fuzzing, contracts, migration checks, security checks, and human review. Foundry amortization plus training, serving, tools, verification, failed loops, and review.
Canonical form looks cleaner but does not reduce files opened, context/reasoning tokens, tool calls, failed lanes, or review. Canonical data requires similar tokens, similar model scale, or larger models for the same accepted-change target. Throughput gains disappear after batching, memory, context, verification, or quality corrections. Important compatibility ghosts are lost or security/migration regressions rise.
Does behavior survive porting? Does the economics close?
All-in cost per verified correct change
Foundry and verification overhead consume the gains.
The end-of-road story is stark. In the conservative version, examples from the old distribution. Canonical porting changes canonicalization is an engineering discipline that makes coding the distribution and the action substrate. The standard constrains agents cheaper and more reliable by cutting routine repository language roles, file grammar, naming, dependency ownership, search and failed repair loops. In the central version, it becomes generated boundaries, database truth, continuous integration, a new data substrate: fewer legal encodings, fewer legal edit repository policy, build/test commands, migration policy, proof paths, and lower tokens-to-target capability. In the aggressive lanes, and agent-readable edit scope. Observable behavior is version, routine product software approaches a behavior- preserved, rejected, or tiered according to an explicit evidence genome regime in which most CRUD, auth, policy, migration, envelope; arbitrary human degrees of freedom collapse only deployment, observability, billing, notification, workflow, and when they are not part of that envelope. The external-reader integration behavior is assembled from proof-carrying cells, claim is precise: human repositories are improvable because while bespoke code is reserved for minimum viable novelty. they mix durable behavior with accidental representation, not The paper is ambitious because the upside is enormous; it is because human behavior, product judgment, operational lessons, defensible only because every large number is tied to a named or edge cases are disposable. denominator and a falsification test. This is a deliberately strong position: for every supported The scale is enormous. GitHub reported 395 million public product-software concern there should be one governed way and open-source repositories in Octoverse 2025, roughly 230 to express it, and for other software classes there should new repositories per minute, and a major shift in the most- be a secondary governed profile rather than local preference. used language [13]. The Stack v2 reports 67.5TB full, 32.1TB Humans do not merely choose different surface styles; they deduplicated, roughly 900B train-full tokens, over 3B files, create mutually incompatible local dialects for routing, testing, and 658 languages [1]. StarCoder2 was trained on 3.3T–4.3T continuous integration, authorization, secrets, migrations, featokens across 619 programming languages [2]. Kimi K2 reports ture flags, deployment, observability, and review. Those dialects a 1T-total-parameter mixture-of-experts model with 32B active may work locally, but they are training noise for agents and parameters and 15.5T pretraining tokens [14], [15]. Modern reasoning overhead at inference time. For supported classes, the code-model work shows the same field-level direction: more falsifiable claim is that nearly every nontrivial human-authored languages, longer context, more tokens, more data curation, repository contains reducible accidental degrees of freedom: and more agentic training rather than a smaller human-code duplicate behavior, noncanonical representation, generated-truth target [16], [17]. drift, unconstrained edit surface, missing proof receipts, or The redundancy is also enormous. DéjàVu analyzed 4.5 ungoverned policy. million non-fork GitHub projects and found that only 85 million The decisive experiment is not a better-looking repository. of 428 million files were unique—about 70% were clones It is a same-model, same-lineage comparison: original human of previously created files [3]. Jupyter notebook code shows repository versus canonical repository, same issue, same similar patterns: more than 70% of code snippets were exact behavior contract, same hidden tests, same reviewer rubric, and copies [4]. Software has long been known to be highly repetitive a complete ledger of files opened, context tokens, reasoning and predictable [18]. tokens, tool calls, invalid edits, failed proof lanes, review The thesis is stronger than filtering or deduplication: the burden, wall time, and accepted-patch rate. If a broad model available training code should be ported into an agent-first performs materially better on the canonical repository before canonical standard before training. Filtering selects cleaner a canonical specialist is trained, then substrate compression is
real. If it does not, the thesis fails early. This paper makes the case through nine explicit assumptions, each with falsification criteria, a compression frontier catalog, preliminary training evidence, and a paired evaluation design. The core falsifiable thesis is simple: after controlling for behavior and task, canonical proof-carrying software should lower same-model search cost now, lower tokens-to-target accepted-change capability under paired scaling curves later, and lower amortized cost per verified correct change in production field trials. If those three gates fail, no title, figure, or compression catalog rescues the claim. II. P OSITIONING : W HAT E XISTING W ORK P ROVES , AND W HAT I T D OES N OT
moves entropy into metadata, the thesis should be rejected or narrowed. III. T HE M ISSING T HEORETICAL M OVE : T RAIN ON THE Q UOTIENT, N OT THE O RBIT The deepest version of the thesis is not that code can be standardized. It is that software learning should be quotiented by behavior. Let P be the set of source-level programs, repositories, configurations, schemas, tests, and deployment artifacts. Let O be a declared behavior oracle: public tests, hidden tests, trace equivalence, API contracts, migration replay, security policies, performance envelopes, accessibility checks, runtime invariants, and accepted incompatibility dispositions. Define an equivalence relation
Adjacent work strongly supports the premise that software p1 ∼O p2 ⇐⇒ O(p1 ) = O(p2 ) (1) contains compressible regularity, but it does not yet prove the thesis in this paper. Naturalness results show that code up to the declared preservation tier. The human softis unusually repetitive and predictable [18]; clone studies ware corpus samples many points in the syntactic orbit show massive exact and near-exact duplication in public [p]O : Python/FastAPI, Java/Spring, Node/Express, Ruby/Rails, repositories [3], [4]; modern code LLMs show that trillion-token Go/Gin, C#/ASP.NET, hand-written DTOs, generated DTOs, corpora can learn broad programming competence [2], [14], imperative migrations, declarative migrations, bespoke CI, and [16], [17]; and repository-level benchmarks show that agents ad hoc review rules can all encode the same product behavior. can now operate over real projects, although contamination, A raw code model learns both the quotient P/∼O and the orbit weak tests, and distribution mismatch make static benchmarks noise inside each class. A canonical foundry instead learns a fragile [5], [6], [19], [20]. This paper’s stronger claim is versioned normal-form map different: the target distribution itself is wrong. The right move κ : [p]O → (s, e, r, d), (2) is not only higher-quality filtering, more synthetic tasks, or larger context windows. It is to transform the code archive into where s is the shortest accepted canonical specification, e a canonical behavior-and-change substrate before training. is the evidence bundle, r is the renderer/generator, and d is The strongest allies are not only code LLM papers. They the disposition ledger for non-preserved behavior, legal risk, are program-analysis and programming-language systems that security defects, or domain mismatch. already separate meaning from textual accident. OpenAPI and This distinction clarifies why the proposal is more ambitious Protocol Buffers show that one source of truth can generate than deduplication. Deduplication removes copied files. Nearmultiple projections of an interface [21], [22]. MLIR shows that deduplication improves pretraining quality by removing redomain-specific dialects can lower through reusable compiler peated text. Canonical quotienting removes degrees of freedom infrastructure rather than forcing every domain into one flat that should not be in the model’s target language at all. The representation [23]. Coccinelle shows that semantic patches can model is no longer asked to learn every valid implementation encode collateral evolutions more compactly than hand-written orbit; it is asked to learn canonical representatives plus the diffs [24]. CodeQL shows that code can become queryable proof obligations that make the representative behaviorally semantic data for large-scale vulnerability reasoning [25]. E- acceptable. graphs show how equivalence classes of programs can be The theory also explains why the proposed breakthrough represented compactly for rewrite search [26]. DreamCoder and can coexist with modern scaling laws. Scaling-law work library-learning systems show that repeated programs can be implies that data quality and token efficiency matter; Chinchillacompressed into reusable abstractions that improve search [27]. style compute-optimal training links model size and training Agent-first canonical code combines these threads into one tokens under a compute budget [28]. Canonical code does substrate-level bet: learn from human software history, but do not repeal scaling laws. It changes the data distribution so not force future agents to imitate its incidental coordinates. that a token is closer to irreducible behavior and farther from This positioning also defines the paper’s burden of proof. A accidental representation. If the canonical token stream has reviewer should not accept the claim because the standard is lower conditional entropy and fewer irrelevant action branches, elegant. The claim earns belief only if paired experiments show the same physical training and inference budget should buy lower cost per verified correct change, lower tokens-to-target more verified software capability. accepted-change capability, lower action branching, preserved A useful analogy is compiler IR. LLVM and MLIR do behavior at declared evidence tiers, and lower infrastructure not eliminate source languages; they provide structured incost/token where model-size claims are invoked. Conversely, termediate forms that make analysis, lowering, optimization, if canonicalization removes useful implementation diversity, and reuse tractable [23], [29]. The canonical-code claim is harms robustness, fails to amortize foundry cost, or merely that product software needs a higher IR: entities, permissions,
TABLE IV F ROM R AW-C ODE I MITATION TO B EHAVIOR -Q UOTIENT L EARNING Layer
Raw-code learning target
Canonical quotient target
Representation
All languages, frameworks, layouts, generated artifacts, dependency rituals, and local repair customs. Infer semantics through many surface encodings.
One governed representative per behavior class, plus explicit profile escape hatches. Learn behavior primitives, typed deltas, proof obligations, and legal edit forms. Route to cells, owned files, generated truth, and proof lanes. Consume structured proof receipts and behavior-oracle deltas. Compile successful and failed trajectories into reasoning digests, semantic patch cells, and negative memory. Proof-carrying change accepted under a cost ledger.
Generalization Search Verification Memory
Explore file trees and patch locations in a high-entropy orbit. Interpret arbitrary tests, logs, CI conventions, and reviewer expectations. Remember past trajectories as chat/log traces.
Economic unit
Patch accepted after ad hoc reasoning and review.
policies, effects, state machines, migrations, API contracts, observability, proof obligations, rollout envelopes, and runtime invariants. Source code is then a projection, not the source of truth. Coccinelle’s semantic patches show that many source changes are better represented as semantic transformations than text diffs [24]; e-graphs show that equivalence classes can be exploited computationally [26]. The missing leap is to make these ideas the training substrate for software agents. The falsifiable prediction: after controlling for behavior and task, canonical quotient learning should reduce entropy, search, and proof cost before any new specialist model is trained. If a broad frontier model does not become cheaper and more reliable when the same software lineage is presented through the canonical representative, the quotient thesis is wrong.
IV. D EFINITIONS AND C LAIM D ENOMINATORS
The paper separates public facts, preliminary measurements, near-term targets, central hypotheses, and moonshot targets. Human-code entropy is the residual representation uncertainty left after behavior, contracts, and execution environment are fixed:
Hhuman = H(representation | behavior, contracts, environment).
(3)
A canonical substrate is useful only if it reduces this conditional entropy without erasing behavior evidence, safety constraints, provenance, compatibility ghosts, or deliberate diversity.
We report the following distinct reductions: raw source tokens Rsource = , (4) canonical source tokens H(raw valid forms) Rentropy = , (5) H(canonical valid forms) raw attempted/legal actions , (6) Raction = canonical legal actions raw tokens to target accuracy Rtrain = , (7) canonical tokens to target accuracy raw wall-clock to target accuracy Rtrain-time = , (8) canonical wall-clock to target accuracy raw wall-clock per verified change Rspeed = , (9) canonical wall-clock per verified change raw dense-active infrastructure cost/token , Rcost/token = canonical dense-active infrastructure cost/token (10) raw Treason+tool+plan+retry /∆ Rreason = canon , (11) Treason+tool+plan+retry /∆ Lraw failed proof/repair /∆ , (12) Rretry = canon Lfailed proof/repair /∆ Araw edit /∆ , (13) Rinvalid = invalid Acanon invalid edit /∆ raw cost per verified correct change Rcost = . canonical cost per verified correct change (14) Here ∆ denotes an accepted change. Rcost is the all-in version of the earlier Rchange notation, and Rcost/token is only the denseactive serving denominator. At fixed accelerator throughput, Rtrain-time ≈ Rtrain . Effective raw-corpus-equivalent training throughput is Seff = Sphysical Rtrain , so a 1M-token/s training run with a 10× canonical tokens-to-target reduction is measured as 10M raw-corpus-equivalent tokens/s against the same target capability. The 40×–150× range belongs to effective token-space reduction (Rentropy ); the 1,000×–100,000× range belongs to action-space pruning (Raction ); the 30×–150× central range belongs to training-token efficiency (Rtrain ); reasoning-token, retry-loop, invalid-edit, verified-change cost, and dense-active cost/token reductions belong to separate denominators (Rreason , Rretry , Rinvalid , Rcost , Rcost/token ). These do not multiply cleanly.
TABLE V C ORE D EFINITIONS Term
Definition
Canonical profile
A governed language-role, layout, naming, dependency, generated-boundary, data-truth, and validation configuration for a software class. The primary app profile is a versioned limited-stack product profile: one maintained backend/core lane, one typed product-surface lane, one browser/product-shell lane, one durable-data lane, one schema/contract lane, and one build/deploy/proof envelope. Systems software uses approved secondary profiles or governed primitives. A behavior-preserving or behavior-dispositioned transformation from a human artifact into a canonical profile. A machine-readable bundle: provenance, license status, original source, port disposition, behavior evidence, canonical artifact, rejected invalid paths, and validation results. Conditional representational uncertainty after behavior is fixed: the languages, frameworks, layouts, names, CI rituals, dependency choices, repository policies, and proof routes that vary without changing supported behavior. The inference-time expansion of context reads, reasoning tokens, tool calls, architecture discovery, invalid edits, proof-loop failures, and review burden caused by unconstrained human repositories. A named, versioned, proof-carrying software primitive: stable interface, schema, policy model, generated code, fixtures, tests, migration obligations, security negatives, repair memories, and known failure modes. A typed, governed change archetype such as add field, migrate nullable to required, add permission edge, rotate secret, add idempotency key, split table, add audit trail, or upgrade dependency safely. A standard local or CI validation route: exact command, environment, required artifacts, log schema, and acceptance rule. The structured unit of accepted change: intent, affected cells, typed diff, schema diff, migration diff, test delta, security delta, proof obligations, lane receipts, rollback plan, and provenance/license metadata. A compact, versioned summary of known architecture, generated zones, proof routes, repair strategies, failure modes, rejected plans, compatibility ghosts, and prior proof objects for a repository, profile, or cell. The machine-checkable set of legal files, typed edits, schema operations, migration forms, dependency updates, proof lanes, and generated-zone protections available to an agent. A product-level IR for entities, state machines, policies, permissions, effects, contracts, migrations, observability, proof obligations, and runtime constraints; source code is one generated projection. The mined atlas of recurring behavior families, variants, invariants, tests, proofs, provenance, repair memories, and generated projections across raw software history. The irreducible product intent, domain fact, policy tradeoff, external constraint, or novel algorithm that remains after routine behavior, architecture, tests, proofs, dependencies, and deployment are generated or inherited. A governed configuration for CI, branch protection, review, secrets, releases, ownership, generated zones, dependency updates, and deployment permissions. The set of implementations that expose the same observable behavior, interfaces, security invariants, and migration guarantees. Minimum Functional Description Length: the shortest canonical specification plus proofs plus renderer needed to produce a working system. Total dollars, source/context/reasoning/action tokens, tool calls, wall time, failed loops, validation runs, and human review needed to produce an accepted proof-carrying behavior-preserving change. Accelerator-side serving cost per generated or processed token under a specified active-parameter regime, context length, batch size, memory system, precision, and quality target; dense and MoE baselines must be reported separately.
Canonical port Canonical training object Human-code entropy Agent sprawl Canonical behavior cell Semantic patch cell Proof lane Proof-carrying change object Reasoning digest Constrained edit grammar Behavior intermediate representation Software genome Minimum viable novelty Canonical repository policy Behavior-equivalence class MFDL
Cost per verified correct change Dense-active cost/token
infrastructure
Many overlap. The composite effect must be measured on paired raw/canonical corpora and reported as cost per verified correct change. V. C ORRECT-C HANGE I NFORMATION T HEORY The missing theoretical move is to stop treating code as the primary object. The primary object is a verified behavior change. A repository is only one historical encoding of the behavior, tests, policies, provenance, and operational memory needed to accept that change. Let x be a raw repository state, i an issue or requested change, a1:k an edit/tool trajectory, and ∆ the accepted change object. A change is correct only under an oracle that includes behavior, tests, security, migration safety, provenance, and reviewer acceptance: O(x, i, a1:k ) → {ACCEPT, REJECT, ESCALATE}.
(15)
The canonical map Φ is useful only when it preserves the information relevant to ∆ while deleting information irrelevant
to acceptance: I(Φ(x); ∆ | i, O) ≈ I(x; ∆ | i, O),
(16)
H(Φ(x) | ∆, i, O) ≪ H(x | ∆, i, O).
(17)
This is the paper’s central learning-theoretic claim. The model should see less accidental entropy without losing the facts needed to make the right change. A. Accidental Representation Tax For a behavior target y, define the behavior-equivalence class Ey,τ = {x : O(x, y) ≥ τ }.
(18)
Raw software often contains many members of this class: different languages, frameworks, layouts, naming schemes, migration styles, tests, CI dialects, dependency wrappers, and repository policies. Canonicalization selects a smaller governed set Cy,τ plus audited exceptions. The accidental representation tax is: ART(y, τ ) = log |Ey,τ | − log |Cy,τ |. (19)
TABLE VI C LAIM S TATUS AND R EQUIRED E VIDENCE Claim
Status
Current evidence
Required evidence
Canonical trajectories are learnable
M EASURED
Canonical repositories reduce samemodel search cost
N EAR - TERM TARGET
QLoRA convergence on 64,088 translated trajectories and zero measured forbiddenlanguage markers. Substrate design: file grammar, proof lanes, generated zones, and reasoning digests.
Behavior preservation is adequate for training
C ENTRAL HYPOTHESIS
Dossier and preservation-tier design.
Training-token reduction
C ENTRAL HYPOTHESIS
Behavior-genome economics
M OONSHOT
Compression theory plus adjacent dataquality evidence. No direct measurement yet; proposed substrate architecture.
Replication across model sizes, heldout repositories, stronger profile-validity checks, and task success. Paired raw/canonical tasks with identical model, issue lineage, hidden tests, and cost ledger. Large paired ports with tiered tests, traces, contracts, fuzzing, differential replay, human review, and security/migration evidence. Paired scaling curves to target acceptedchange performance. Cell coverage census, mature foundry, runtime negative memory, and verifiedchange cost curves after amortization.
TABLE VII W HAT T HIS PAPER D OES N OT C LAIM Not claimed
Boundary
Universal stack destiny Automatic equivalence Literal 100× source shrink
The primary stack is a versioned product-app profile, not a permanent answer for all domains. Tests, traces, and contracts are evidence; only proof-backed cells get proof-backed claims. Large ranges refer to specific denominators such as action space or routine-domain tokens-totarget. Foundry, provenance, review, and verification costs must be amortized and included. Human artifacts carry behavior, edge cases, incidents, and product judgment; only accidental representation is targeted.
Free canonicalization Human code lacks value
The exact cardinalities are not directly observable for realistic software, but the tax can be estimated by source-token counts, identifier entropy, AST-pattern entropy, file-path entropy, dependency entropy, proof-lane entropy, repository-policy entropy, and model perplexity under paired raw/canonical corpora. This reframes the largest claim. The model is not merely learning fewer bytes. It is learning fewer behavior-equivalent encodings. That is why clone studies, naturalness results, software product-line engineering, compiler IRs, e-graphs, proof-carrying code, and library-learning systems all matter: they show different historical paths toward the same idea that repeated behavior should be represented by reusable structure rather than re-authored text [3], [18], [23], [26], [27], [29]–[31].
B. Correct-Change Search Work A coding agent pays for branching. At each step t, it faces legal or perceived choices At : files to inspect, commands to run, edit locations, migration strategies, dependency changes, proof lanes, and repair paths. Raw repositories inflate |At | with local folklore. Canonical repositories should shrink the legal set and reject invalid moves before they become failed
trajectories. A crude lower bound for search work is: Wsearch ∝
k X t=1
log |At | +
m X
ρ j Fj ,
(20)
j=1
where Fj are failed verification or review loops and ρj are their token/tool/human penalties. Canonicalization wins when it reduces both the action entropy and the expected number of expensive failures. This is why the strongest near-term experiment is samemodel raw/canonical ablation. It does not require a new model or a full training run. It asks whether the substrate itself reduces search work: fewer files opened, fewer tool calls, fewer invalid edits, fewer failed proof lanes, fewer review comments, and lower dollars per accepted change. C. Minimum Functional Description Length Minimum Functional Description Length is the productsoftware analogue of MDL: not the shortest text file, but the shortest behavior description plus evidence plus renderer that satisfies the declared oracle. The practical objective is: MFDLS,τ (y) = min |z| + λ|Π(z)| + µ|RS | + ν|G(z)| z∈ZS
s.t.
O(RS (z), y) ≥ τ.
(21)
Here z is the canonical behavior object, Π(z) its proof/evidence bundle, RS the renderer into source, tests, migrations, CI, documentation, observability, and deployment artifacts, and G(z) the governance/provenance burden. This extra governance term matters: compression that drops attribution, opt-out state, security evidence, or compatibility ghosts is not a substrate win. Falsification: this information theory fails if canonical artifacts preserve no more mutual information about accepted changes than raw repositories, or if they reduce entropy only by deleting behavior needed for hidden tests, product acceptance, security, migration safety, or provenance.
TABLE VIII F ROM C ODE I MITATION TO C ORRECT-C HANGE I NFORMATION Layer
Raw-model burden
Canonical information object
Text/source Repository
Change
Learn many strings for same behavior. Infer local architecture, generated zones, and proof commands. Re-implement auth, lifecycle, billing, upload, search, jobs, webhooks, audit, observability. Invent raw diffs and hope tests catch mistakes.
Evidence
Reviewer reconstructs intent from a diff and logs.
Memory
Incidents and failed patches disappear into history.
One governed rendering plus audited exceptions. File grammar, generated-zone manifest, proof lanes, repository-policy object. Versioned behavior cells with parameters, tests, proofs, and negative cases. Typed semantic patch cell with preconditions, generated deltas, proof obligations, rollback. Proof-carrying change object with receipts, provenance, security delta, migration replay, review rubric. Runtime negative memory: forbidden plans, compatibility ghosts, regression generators, patched cells.
Behavior
VI. T HE H UMAN -C ODE E NTROPY P ROBLEM
familiar projects, while METR’s 2026 update emphasizes that adoption and task-selection effects make real productivity Human code is chaotic because human software production measurement difficult [7], [8]. is chaotic. Teams choose languages for hiring markets, deThese findings do not prove every human program is broken. ployment constraints, fashion, deadlines, legacy compatibility, They prove the stronger substrate point: human software framework momentum, and personal preference. They choose lacks a single enforced shape. The code may run, but the continuous-integration systems and workflow YAML by copy- representation is filled with optionality that does not encode paste. They choose names under time pressure. They create product behavior. For an agent, optionality is not freedom; it is folders that reflect organization charts and local arguments. branching factor. If there are 20 plausible files, five plausible They copy data transfer objects instead of generating contracts. CI fixes, three dependency-update conventions, four migration They hand-edit generated artifacts. They split logic across styles, and several local naming dialects, the agent must spend services because the team split. They preserve test suites that tokens and tool calls discovering local folklore before changing prove mocks, not behavior. They keep dependencies because behavior. This is agent sprawl: the model becomes a repository updating them breaks a release. cartographer, build engineer, migration reviewer, dependency This is not an argument that human software history lacks archaeologist, security analyst, and code reviewer before it can value. The public and private code record contains product safely edit behavior. A canonical substrate removes that burden judgment, edge cases, incident responses, migration scars, by deleting opinions from the supported path. security repairs, tests, failure histories, and operational lessons This variation is not free. Every equivalent way to express a that a canonical substrate must preserve or disposition. The behavior becomes probability mass the model must learn. Every target is narrower and more aggressive: remove optional framework convention becomes a routing burden. Every naming representation that remains after behavior is fixed, while style becomes lexical entropy. Every CI dialect becomes an carrying forward the evidence that explains why the behavior operational trap. Every ungoverned repository setting becomes matters. hidden state. The frontier model pays for all of it. None of this is rare. GitGuardian found 28.65 million new Agent sprawl is not solved by a larger context window hardcoded secrets added to public GitHub commits in 2025, alone. A larger window lets the model carry more accidental a 34% year-over-year increase [12]. Synopsys reported that state; it does not remove the accidental state. The canonical 84% of assessed codebases contained open-source vulner- countermeasure is to move repeated reasoning into substrate abilities, 74% contained high-risk vulnerabilities, and 91% law: file grammar tells the model where behavior lives, used components ten or more versions behind the current generated-zone manifests tell it where not to edit, proof lanes version [32]. GitHub Actions studies find exactly the same tell it what proves completion, behavior cells identify the problem in automation: workflow complexity, heterogeneity, recurring primitive, semantic patch cells define legal change and compliance drift are now research subjects [9]; workflow operations, reasoning digests summarize known architecture files change continuously across large repository samples [10]; and failure modes, and proof-carrying change objects make security-practice adoption remains uneven even in a dominant review evidence explicit. CI/CD platform [11]. Travis CI evidence is similar: 3.7 million jobs across 1,276 projects showed noisy and heterogeneous VII. A SSUMPTION A: V ERSIONED P RODUCT P ROFILES AND C ANONICAL P ROGRAMMING build outcomes, including misleading pass/fail signals [33], and 9,312 Travis-using projects exhibited detectable CI feature Assumption A: For each supported software class, there misuse [34]. OpenSSF Scorecard exists because open-source should be one governed way to express, test, secure, projects do not consistently use secure practices [35], [36]. deploy, and repair each concern. The primary agent-first METR’s 2025 randomized trial found experienced open-source canonical profile targets web, product, SaaS, backend, and developers were 19% slower with early-2025 AI tools on mature data applications—not all software. We denote the current
TABLE IX H UMAN -C ODE VARIATION T HAT B LOATS THE T RAINING TARGET Variation class
Human-code reality
Canonical target
Language sprawl
Same backend behavior appears across many service, scripting, and systems languages. Express, FastAPI, Spring, Django, Rails, Laravel, Next, Remix, custom RPC, custom queues. User/account/customer/member; manager/service/handler/repo/dao; vague helpers and utils. Arbitrary folders, mixed generated and source files, stale migrations, implicit ownership. GitHub Actions, GitLab CI, Travis, Circle, custom runners, local env assumptions, undocumented gates, skipped or flaky tests. Unpinned libraries, abandoned packages, vulnerable transitive dependencies, risky workflows. Data-transfer-object forks, direct queries, app-owned durable truth, duplicated schema definitions. Branch protection, code owners, required checks, release permissions, environments, secrets, and deployment credentials vary per team. Model guesses where to edit and what proves completion.
One primary language per role, with secondary profiles for true infrastructure primitives. One canonical service grammar and one canonical UI/product grammar per profile. Controlled vocabulary tied to role, boundary, and domain.
Framework drift Naming drift Layout drift CI/build drift Dependency drift Data truth drift Repository-policy drift Agent ambiguity
2026 profile Papp as a versioned limited-stack product profile: one maintained backend/core lane, one typed product-surface lane, one browser/product-shell lane, one durable-data lane, one schema/contract lane, and one build/deploy/proof envelope, plus repository policy and observability. Systems code, kernels, compilers, databases, browsers, GPU kernels, embedded, native mobile, and high-performance computing use secondary canonical profiles or governed primitives. The thesis is not implementation monoculture. The thesis is that human preference should not decide the shape of routine product software. The profile is replaceable. The invariant is governed role ownership, generated truth, proof lanes, constrained edits, provenance, and observability: t+1 t Papp = Govern(Papp , Esecurity , Eagent ,
Eperformance , Eecosystem , Emigration ).
(22)
Enforced file grammar, generated boundaries, and ownership maps. One proof-lane grammar with reproducible commands, artifacts, permissions, and failure classes. Governed dependency closure, upgrade lanes, and scored exceptions. Governed durable-data truth, generated contracts, typed adapters, migration policy. One repository-policy object with generated GitHub/GitLab settings and auditable exceptions. Agent-readable scope, canonical commands, valid repair paths, and no-edit generated zones.
routine changes occur against explicit role ownership rather than local framework folklore. Second, file grammar is enforced. Source paths, generated paths, tests, migrations, contracts, adapters, and product surfaces live in predictable places. Third, naming is canonical. Variable and module names encode stable roles: actor, resource, command, event, policy, adapter, projection, migration, contract, and view. Fourth, generated boundaries are sacred. Contracts generate clients, schemas, test fixtures, and adapters. Generated outputs are not hand-edited. Fifth, continuous integration, build, and proof commands are standard. Every change maps to known local commands, CI jobs, permissions, logs, artifacts, and failure classes.
Sixth, repository policy, dependency policy, and data Profile components change only when evidence shows lower truth are governed. One durable-data lane owns persistent accepted-change cost, stronger verification reliability, better truth; dependencies are pinned, owned, scored, and upgraded security, or better ecosystem support after transition cost. through canonical lanes; branch protection, required checks, Canonical does not mean a single unchecked implementation. code owners, release gates, environment secrets, and deploy It means one governed interface, multiple independently permissions are generated from one repository-policy object. validated implementations where correlated failure would be The canonical standard is therefore a compression language catastrophic. Critical cells require independent implementations for software. It compresses by removing representational where feasible, adversarial review, conformance suites, canary degrees of freedom that do not change the product. If two rollout, version pinning, rollback, deprecation protocols, proverepositories implement the same account-management route nance records, and profile-specific exceptions. Unlimited drift with eight framework stacks, four CI systems, three naming is rejected; audited diversity is retained. conventions, and incompatible deploy policies, a broad model The canonical standard has six hard properties: pays for all of them. A canonical model pays for the behavior First, profile roles are fixed. The profile assigns one once, plus the governed exceptions that matter. supported lane to backend/core/tooling, one lane to typed product surfaces, one lane to the browser product shell, one Falsification: Assumption A fails if paired raw/canonical lane to durable data, one lane to deployment envelopes, and corpora do not reduce source tokens, identifier entropy, ASTone lane to generated contracts. The concrete technologies are pattern entropy, path grammar, perplexity, or action branching versioned profile choices, not the thesis. The invariant is that on target software classes.
TABLE X C ANONICAL S TANDARD AS A C OMPRESSION L ANGUAGE Standard element
Human-code space removed
Training effect
Limited language roles
Equivalent business behavior expressed across many backend, frontend, script, and migration languages. Local folder myths, mixed generated/source files, arbitrary helpers, unclear ownership. Synonym drift, vague service names, local abbreviations, duplicated domain labels. Data-transfer-object forks, hand-edited clients, duplicated schema truth, stale fixtures. README folklore, local scripts, skipped tests, inconsistent CI gates. Unpinned packages, stale transitive libraries, unsafe workflow permissions. Direct writes, app-owned durable truth, duplicate validation logic, unsafe migrations. Hand-configured GitHub/GitLab settings, inconsistent branch protection, unscoped workflow permissions, unclear owners.
Fewer grammars and lower cross-language transfer burden.
Canonical file grammar Controlled vocabulary Generated boundaries Strict build/test commands Governed dependencies Data truth policy Repository policy
VIII. A SSUMPTION B: G OVERNED F OUNDRY AND P ORT D ISPOSITIONS Assumption B: Every available artifact in the training set can be assigned a governed port disposition, and the cost of porting amortizes across repeated training runs, serving savings, and downstream reuse. Full coverage means every artifact enters the foundry. The claim is not that every repository becomes one product-app implementation. The claim is that every artifact is either transformed into a canonical training object or explicitly excluded with a reason that remains useful for governance, evaluation, or negative training. The maturity ladder is staged: first disposition coverage, then behavior-fixture extraction, then full canonical porting only where value, license status, and preservation evidence justify the cost. Every canonical training object is an auditable data product, not a loose rewrite. Its minimum schema includes source identity, license, profile, disposition, transformation trace, behavior evidence, security evidence, canonical artifact, known loss, and training-use policy. Source identity includes repository URL, commit, content hash or Software Heritage identifier when available, opt-out/removal state, and original artifact digest. License and supply-chain evidence should use existing machine-readable standards where possible: SPDX for license/provenance and software-bill-of-material information, CycloneDX for broader bill-of-materials and cyberrisk metadata, SLSA/in-toto-style attestations for build and transformation steps, OpenTelemetry semantic conventions for runtime evidence, and policy-as-code systems such as Open Policy Agent for inclusion and release rules [37]–[42]. Falsification: Assumption B fails if a representative corpus cannot be assigned governed dispositions, if license/provenance metadata cannot be preserved through transformation and removal propagation, if behavior fixtures cannot be extracted for high-value classes, or if porting cost does not amortize across at least 3× the initial foundry investment.
Repository layout becomes predictable context instead of hidden state. Lower lexical entropy and easier long-context retrieval. One source of truth replaces many invalid patch paths. Model learns standard completion actions instead of guessing rituals. Dependency updates become canonical maintenance tasks. Database and migration repairs become standard forms. Review, merge, release, and secret-handling become generated infrastructure.
IX. A SSUMPTION C: B EHAVIOR P RESERVATION U NDER D ECLARED C ONTRACTS Assumption C: Canonical ports can preserve, disposition, or explicitly reject behavior at a confidence level high enough for training. This is the hardest assumption in agentfirst canonical code. A canonical rewrite that silently deletes edge cases is not compression; it is corruption. Fully automatic behavior equivalence for arbitrary programs is not available in general; Rice’s theorem is the theoretical warning label on any claim that a tool can decide nontrivial semantic properties for all programs [43]. The foundry therefore treats behavior evidence as part of the training object, not a best-effort annotation. Every canonical training object carries a preservation and divergence dossier: original tests, canonical tests, differential traces, API compatibility checks, UI replay where applicable, migration replay, security-negative cases, performance budgets for behavior that depends on latency or resource use, provenance/license records, accepted incompatibilities, compatibility ghosts, and rejected port paths. A canonical port is never behavior-preserving merely because tests pass; it is behaviorpreserving only within a declared evidence envelope. The output is not binary pass/fail. It is graded behavior confidence for original artifact o and canonical artifact z: B(o, z) =
X i∈E
wi ei (o, z),
0 ≤ ei ≤ 1,
X
wi = 1.
i
(23) Here E includes test agreement, differential execution coverage, migration and data-invariant replay, security-negative coverage, UI/API replay coverage, human acceptance for high-value cases, performance/SLO agreement where applicable, and provenance/license completeness. Compatibility ghosts are first-class. If downstream users depend on behavior that looks accidental—ordering, timing, error text, weak validation, legacy API quirks, migration side effects, or undocumented defaults—the foundry must either preserve it, encode it as a compatibility fixture, or record an
TABLE XI G OVERNED P ORT D ISPOSITIONS Disposition
Meaning
Training use
transform-permitted
License, provenance, security, and behavior evidence permit a canonical artifact or full port. Artifact belongs to systems, mobile, embedded, scientific, or other non-product-app profile. Foundational component is retained behind stable wrappers, tests, versions, and ownership. Source cannot be transformed, but public behavior/specification can be represented independently. Tests, traces, issues, API examples, or migration fixtures are usable without source projection. Artifact teaches a failure mode, vulnerability, bad migration, or rejected pattern. Only aggregate metadata, disposition statistics, or provenance facts can be retained. Artifact appears unsafe, malicious, poisoned, undispositioned, or provenance-risky. License, attribution, opt-out, or derivative-work risk blocks inclusion.
Canonical source, behavior object, transformation trace, accepted patch examples. Profile-specific training and cross-profile interface examples.
secondary-profile governed-primitive spec-only fixture-only negative-only metadata-only quarantine license-excluded
Available codetraining universe
Classify role, behavior, language, license
Dependency, wrapper, compatibility, upgrade, and conformance examples. Abstract contracts, interface facts, and conformance targets. Behavior fixtures and weak-oracle warnings. Negative examples, refusal cases, and repair memories. Governance metrics and exclusion statistics. No positive training; security analysis and exclusion receipt only. No positive training artifact; removal propagation and audit receipt.
Disposition, profile, provenance policy
Minimize valid token space
Canonical training corpus
Fig. 1. Governed full-corpus foundry. The corpus is not filtered down to clean examples; every artifact receives provenance-aware disposition, and only qualified artifacts become positive canonical training objects. TABLE XII B EHAVIOR P RESERVATION D OSSIER Evidence lane
What it catches
Canonical training use
Original tests Differential traces Migration replay Security negatives UI/API replay Human acceptance Provenance/license
Known intended behavior and regression expectations. Runtime behavior across original and ported artifacts. Data-shape, rollback, lock-budget, and invariant failures. Auth bypass, injection, secret leakage, unsafe defaults. Browser, CLI, and API compatibility surfaces. High-value cases where automated tests are insufficient. Legal and origin constraints.
Positive evidence and weak-oracle metadata. Equivalence scoring and edge-case discovery. SQL/migration proof examples. Labeled bad paths for refusal and repair training. Product behavior fixtures. Gold labels and accepted-incompatibility manifests. Inclusion, exclusion, and attribution metadata.
accepted incompatibility signed by the relevant owner. Calling it accidental is not enough to delete it. The output is a preservation tier, not a blanket declaration of equivalence: The weak-oracle problem is explicit. Tests are evidence, not proof. SWE-bench-style patch evaluation can be distorted by insufficient tests and plausible-but-wrong patches; UTBoost found insufficient test cases and erroneous patches previously labeled as passed in SWE-bench-family evaluation [6]. Propertybased testing, differential testing, compiler fuzzing, mutation testing, and symbolic execution are stronger evidence lanes, but they remain scoped evidence rather than universal equivalence [44]–[47]. CompCert shows that semantic preservation can be a compiler contract rather than a slogan [48]. seL4 shows that proof-backed systems components are possible for security-critical kernels, while also making clear how much engineering work such proofs require [49]. The canonical foundry should use these methods where their cost is justified and reject overclaiming elsewhere. This section is intentionally early because behavior preservation is the trust anchor. The authorial claim that all human code is improvable does not mean all human code is disposable. It means every human artifact can be made more canonical if the
behavior worth keeping is first identified, tested, replayed, and carried forward. Falsification: Assumption C fails if behavior confidence remains low, human review rejects core ports, security/migration regressions rise, or important edge behavior is routinely lost during canonicalization. X. A SSUMPTION D: C ANONICAL B EHAVIOR C ELLS Assumption D: The highest-reuse software behaviors can be collapsed into certified, reusable behavior cells, and these cells can eventually absorb 70%–90% of routine product code if a behavior-cell census validates coverage. This is the breakthrough beyond canonical stack alignment. If the model is still writing 80 functions for create/read/update/delete resource behavior, authentication, forms, validation, errors, pagination, tests, clients, fixtures, and migrations, we are only standardizing the mess. If the model emits a cell declaration and parameters, we are removing entire classes of source tokens, legal edits, proof obligations, and review questions from the task. A canonical behavior cell is not a library. It is a reusable behavior unit with a stable interface, schema, policy model, generated projections across the full stack (service, typed client,
TABLE XIII C ANONICAL P RESERVATION T IERS Tier
Name
Meaning
Training use
P0
Rejected / metadata-only
P1
Syntactic
Exclusion, quarantine metadata, negative examples only. Low-weight syntax/profile examples only.
P2
Interface/test
P3
Trace/contract
P4
Property/fuzz/differential
P5
Proof-backed
Unsafe, unlicensed, malicious, undispositioned, or behavior evidence is too weak. Builds, formats, typechecks, or parses, but no behavior claim is made. Public interfaces are represented and original/canonical tests or fixtures pass. API, schema, migration, policy, compatibility ghosts, and negative cases match declared contracts and replay traces. Generated, fuzzed, property-based, symbolic, mutation-tested, or production-derived inputs exercise critical paths. A formal or mechanized proof discharges specified invariants or semantic preservation obligations.
Weak positive training with explicit weak-oracle label. Standard product-app training target. Higher-confidence port and edge-case curriculum. Certified cell, compiler/profile primitive, or critical dependency.
product UI, durable-data migration, tests, fixtures, observability, cover less than 50% of routine product-app behavior, or if cell and docs), known failure modes, security negatives, repair composition introduces more complexity than it removes. memories, compatibility ghosts, and proof obligations. Change XI. P RELIMINARY L EARNABILITY AND itself also becomes cell-shaped: common edits such as add P ROFILE -A DHERENCE E VIDENCE field, migrate nullable to required, add permission edge, rotate Claim scoped to this section: canonical data is learnable secret, add idempotency key, upgrade dependency safely, split table, and add audit trail become semantic patch cells rather and can enforce target-profile adherence. There is prior evidence that code model performance benefits heavily from than bespoke diff inventions. For routine domains, behavior cells are the main path from data quality: Arctic-SnowCoder-1.3B used 555B tokens in canonical code to MFDL. The agent should not decide from staged data refinement and beat or matched models trained on scratch how authorization, audit trails, pagination, idempotency, much larger token budgets [50]. The phi-1 work showed that upload signing, or migration rollback work in every repository. high-quality “textbook” data can produce remarkably capable It should select a certified cell, supply parameters, apply a small models [51]. To test whether canonical-translated data produces healthy semantic patch cell when behavior changes, and return a proofconvergence, we ran a quantized low-rank adaptation (QLoRA) carrying change object with receipts. The 70%–90% number is a target, not a measured result in fine-tuning experiment on Qwen2.5-Coder-14B-Instruct [17], this paper. It must be earned by a behavior-cell census that [52] using 64,088 canonically translated agentic coding trameasures coverage by accepted changes, source tokens, AST jectories (57,688 train, 6,400 validation). The dataset consists nodes, runtime traces, issue tickets, security-sensitive paths, of verified multi-turn coding sessions where an agent reads a review burden, and production incident classes. A cell is useful real GitHub issue, explores the codebase, applies a fix, and only if it reduces verified-change cost without hiding required passes automated tests, then translates the trajectory through the canonical porting pipeline to align with the target limited-stack variation. The compression ladder shows how gains compound across profile. Training ran for 7,106 steps on a single NVIDIA RTX 3090 the standardization levels: (24 GB) over approximately three days. Key hyperparameters: The endgame is rung 7: models manipulate canonical LoRA rank r=16, α=32, learning rate 2×10−5 with constantsoftware behavior graphs, and source code becomes a generated artifact. The training target moves from source text toward with-warmup schedule, effective batch size 1×8 gradient graph operations, semantic patches, proof obligations, and accumulation, maximum sequence length 2,048 tokens, 4-bit receipts. This is analogous to how LLVM created a reusable IR NF4 quantization with double quantization. Figure 2 shows the convergence curve generated from for program analysis and transformation [29], and how MLIR raw training telemetry. After restart de-duplication, the run explicitly aims to reduce software fragmentation [23]. Softcontains 7,106 training steps and 711 validation evaluations. ware product-line engineering has a related history: explicitly Raw training loss drops from 2.707 at step 1 to 0.419 at step modeling commonality and variability across product variants 7,106. Validation loss drops from 2.261 at step 10 to 0.646 at to build reusable core assets [30]. Equality-saturation systems step 7,106, closely tracking the raw training curve throughout. such as egg show how rewrite spaces can be represented and Two observations are significant: searched compactly when the intermediate representation is 1) No visible train/validation divergence. Raw training explicit [26]. and validation losses fall together through the run, with Cells have a lifecycle, not just a registry entry: Falsification: Assumption D fails if the top 500–2,000 cells validation ending at 0.646 after consuming 100% of the
TABLE XIV B EHAVIOR C ELLS : F ROM H UMAN -C ODE M ESS TO C ANONICAL P RIMITIVES Human-code mess
Canonical cell
Why it compresses
Rails users controller, Express users route, FastAPI users service, Spring user repository Custom login/session/token/password-reset/multi-factor code Hand-written role- or attribute-based permission checks scattered across layers Custom list endpoints with ad hoc pagination/sorting/filtering Ad hoc form validation duplicated across frontend/API/database/tests Recurring billing glue, Stripe webhook wrappers, proration logic Background retries, idempotency, dead-letter queues Object-storage upload variants, signed URLs, virus scanning Custom webhook verification, replay, deduplication Audit trails, event history, compliance logging Expand/contract migrations, backfills, rollback plans Custom notification channels, delivery tracking, templates
resource<User> auth.session / auth.oidc
One verified cell replaces thousands of reimplementations. Repeated in every app with dangerous variation.
policy.check
Security-critical code becomes a proven primitive.
list.page
Massive API/UI/DB repetition becomes one cell.
schema.form.validation
Four validation layers generated from one schema.
billing.subscription
Common SaaS surface with dangerous edge cases.
job.retry.idempotent blob.upload.signed_url webhook.verify.dispatch audit.append_only migration.expand notification.send
Same reliability logic reimplemented everywhere. Security and storage complexity repeated constantly. Integration-heavy apps repeat this badly. Required for compliance-heavy systems. High failure cost and high repetition. Repeated boilerplate across channels.
TABLE XV T HE C OMPRESSION L ADDER : F ROM R AW R EPOSITORY TO C ANONICAL B EHAVIOR G RAPH Rung
Level
What changes
Measurement
0 1 2
Raw human repo Canonical layout Canonical naming
Baseline File-search entropy, edit target count Identifier entropy
3
Single primary profile
4
Contract-first generation
5 6
Proof lanes Behavior cells
7
Behavior graph
Original code Fixed paths, ownership, generated zones Actor/resource/command/event/policy vocabulary Backend/core, typed surface, product shell, durable data, contracts, build/deploy/proof Schema generates clients, fixtures, validators, docs Fixed build/test/migrate/security commands Authentication, resource lifecycle, billing, search, uploads, jobs as primitives Model emits behavior graph operations plus proof obligations; source code is compiled
Language/framework entropy Duplicate data-transfer-object/client drift Failed-loop count Source-token and AST-pattern reduction Task-token and reasoning-token reduction
TABLE XVI B EHAVIOR C ELL L IFECYCLE
TABLE XVII P ILOT C ONFIGURATION
Stage
Required evidence
Field
Value
Propose
Behavior family, interface, parameters, variants, provenance, and expected coverage. Preservation tier, conformance suite, security negatives, fuzz/property lanes, rollout risk. Version pinning, generated projections, proof receipts, canary plan, rollback path. Runtime traces, incidents, performance budgets, compatibility ghosts, exploit reports. Migration path, old-version risk, user impact, and negative-memory update. Revocation, patched cell release, downstream blast-radius scan, and postmortem fixtures.
Base model Adaptation
Qwen2.5-Coder-14B-Instruct QLoRA, rank 16, alpha 32, 4-bit NF4, double quantization 64,088 translated trajectories; 57,688 train / 6,400 validation 7,106 training steps; 711 validation evaluations 2,048 tokens Single NVIDIA RTX 3090, 24 GB Average presence of off-profile markers such as Python def, Go func, and Java public class under target-profile service prompts
Certify Deploy Monitor Deprecate Incident response
training data. By contrast, raw open-source coding data typically requires aggressive deduplication: The Stack v2 reduced from 67.5 TB to 32.1 TB (52% deduplication) and then to 2.4 TB after quality filtering—a 96.4% reduction. 2) Zero forbidden-language violations. Throughout train-
Data Steps Sequence length Hardware Forbidden marker metric
ing, evaluation checkpoints measured off-profile markers (Python def, Go func, Java public class, etc.) in model outputs prompted with target-profile service tasks. All 711 evaluation checkpoints recorded zero markers, indicating that canonical fine-tuning successfully steered the model toward the target profile.
QLoRA Convergence on Canonical Coding Trajectories train loss, raw validation loss forbidden-language markers
2.0
markers
cross-entropy loss
2.5
1
1.5 1.0 0.5
0 0
1000
2000
3000
4000
training step
5000
6000
7000
Fig. 2. Raw training and validation loss convergence for QLoRA fine-tuning of Qwen2.5-Coder-14B on 64,088 canonically translated trajectories. The figure shows the real per-step training log after restart de-duplication: 7,106 training steps and 711 validation evaluations. Training loss starts at 2.707 and ends at a raw final value of 0.419. Validation loss drops from 2.261 at step 10 to 0.646 at step 7,106, with a minimum of 0.644 at step 6,780. Zero forbidden-language markers were recorded across all evaluation checkpoints. The vector figure is generated from raw per-step training logs and evaluation telemetry only.
TABLE XVIII W HAT THE P ILOT P ROVES AND D OES N OT P ROVE Status
Claim
Shows
Canonical translated trajectories are learnable under this parameter-efficient adaptation setup. The measured forbidden-language markers can be suppressed in the evaluation prompts. Behavior preservation, full stack correctness, held-out accepted-change success, or lower cost per verified change. Tokens-to-target reduction versus raw trajectories; that requires paired scaling curves from the same task lineage.
Shows Does not show
future edits through schemas, proof lanes, and semantic patch cells. Let X be raw human-code representations and Z be canonical behavior representations. Let Y be observable software behavior. The canonical standard is useful only if the canonical map preserves behavior while reducing irrelevant representation entropy: H(Z | Y ) ≪ H(X | Y ). (24)
This is the formal version of “one governed way per supported software class.” For a supported concern, if two implementations expose the same behavior and invariants, they should not remain two equally privileged training targets. One interface becomes canonical; independent implementations, port traces, These results provide early evidence that canonical data negative paths, compatibility fixtures, or excluded artifacts is learnable under this parameter-efficient adaptation setup. remain only when they carry evidence, diversity, governance, They do not yet prove behavior preservation, cost-per-correct- or security value. change reduction, or 100B replacing 1T. Those require the full A. Compression Accounting Protocol evaluation protocol described in Section XVII. All ranges in this paper are priors or targets until promoted Does not show
XII. A SSUMPTION E: C OMPRESSION T HEORY AND M INIMUM F UNCTIONAL D ESCRIPTION L ENGTH Assumption E: Removing accidental human degrees of freedom reduces effective representation entropy, legal action-space, and training-token demand. Agent-first canonical code is not a source minifier. It is a canonical coding theory for software behavior. The raw corpus contains many strings for the same behavior because humans choose different languages, frameworks, folders, continuous-integration workflows, repository settings, dependency wrappers, migration rituals, and naming schemes. The canonical standard maps that equivalence class to a smaller canonical space, then constrains
by measurement. A range becomes measured only when the numerator, denominator, task distribution, model, evidence tier, and confidence interval are reported. For task distribution τ and target quality q: Traw source (τ ) Rsource (τ ) = , (25) Tcanonical source (τ ) Draw (q, τ ) Rtrain (q, τ ) = , (26) Dcanonical (q, τ ) Treason+tool,raw (τ ) Rreason (τ ) = , (27) Treason+tool,canonical (τ ) Cverified change,raw (τ ) Rchange (τ ) = . (28) Cverified change,canonical (τ )
Composite ratios cannot be multiplied unless a fitted causal model validates separability on held-out paired tasks. This is especially important for language collapse, framework collapse, generated truth, behavior cells, and proof lanes, which often remove the same tokens, edits, or reasoning loops. The strongest effect is not literal token shrinkage. The strongest effect is action-space and reasoning-space collapse. A raw repository asks an agent to infer where code lives, what is generated, which tests matter, whether a migration is safe, how the CI system encodes permissions, what dependency path is acceptable, and what local naming convention should be preserved. A canonical repository turns those choices into explicit law. That law may add some metadata tokens, but it removes entire families of invalid edits, failed repair loops, and rediscovery reasoning. The final block in Figure 3 deliberately separates the raw independence bound from the claimable combined range. The conservative envelope is not the product of every row. It is anchored by constrained legal edits plus a small number of independent reductions in representation, retry, and reasoning work; stronger bands require evidence that those reductions remain separable on paired tasks. Contract-first generation is the first large step toward that limit: OpenAPI and Protocol Buffers are existing examples of schema-owned surfaces that generate clients, validators, and transport bindings instead of hand-maintaining parallel representations [21], [22]. Semantic patch systems such as Coccinelle and constrained generation systems such as XGrammar show complementary directions: patches and outputs can be represented as structured grammars rather than unconstrained text [24], [53]. The theoretical limit is Minimum Functional Description Length: a frontier over canonical specifications, evidence bundles, and renderers, not a magical globally shortest program. For substrate S, preservation threshold τ , behavior target y, proof/evidence bundle Π, renderer RS , and observable-behavior oracle O: MFDLS,τ (y) = min (|z| + λ|Π(z)| + µ|RS |) z∈ZS
s.t.
O(RS (z), y) ≥ τ. (29)
This is not code golf and not syntax preference. It is the end state where bespoke code appears only when the behavior is novel. Routine product behavior becomes behavior graph operations, cell composition, generated surfaces, policies, schemas, migrations, semantic patches, and proof receipts. In that limit, the model is no longer trained to imitate the long tail of human implementation choices. It is trained to operate a constrained software machine. Falsification: Assumption E fails if paired raw/canonical corpora do not reduce measured source tokens, identifier entropy, AST-pattern entropy, path grammar, proof-loop count, perplexity, or action branching on target software classes.
XIII. B EHAVIOR IR AND T YPED C HANGE A LGEBRA
A canonical substrate should not merely prescribe file layouts. It needs an application-level behavior intermediate representation. The behavior IR represents entities, state machines, authorization edges, external contracts, effects, idempotency requirements, migration semantics, observability, privacy zones, rollout constraints, performance envelopes, and proof obligations. The IR lowers into governed service code, typed product surfaces, durable-data operations, OpenAPI/Protocol Buffers schemas, tests, dashboards, deployment manifests, and documentation. OpenAPI and Protocol Buffers already show the value of contract-first generated projections [21], [22]; the canonical substrate generalizes this from API messages to product behavior. Routine edits should then be expressed as a typed change algebra rather than raw diffs:
∆ = (op, params, pre, post, proof s, rollback, provenance). (30) The operation determines the required proof lanes. For example, AddField(User.timezone, default=UTC) expands into schema diff, migration, API projection, form state, validation, fixtures, compatibility tests, rollout, and rollback. AddPermissionEdge(Manager, ApproveInvoice) expands into policy graph update, negative tests, audit event, UI affordance, threat-model delta, and authorization proof. AddIdempotencyKey(WebhookDispatch) expands into persistence, duplicate-replay tests, observability, retry semantics, and incident-runbook update. This is where the paper becomes more defensible and more ambitious. A raw diff asks the reviewer to reconstruct the operation. A typed change object states the operation first and makes generated diffs subordinate evidence. Proof-carrying code established the idea that producers can ship code with checkable evidence of safety [31]; CompCert and seL4 show that semantic preservation and machine-checked proofs can be practical in constrained high-value domains [48], [49]; SLSA and in-toto show that provenance and supply-chain integrity can be standardized as artifacts [41], [42]. Proof-carrying change applies the same philosophy to everyday product software: the unit of work is not a patch, but a patch plus typed intent, generated obligations, proof receipts, rollback, and provenance. The research bet is that most product changes are not semantically arbitrary. They are drawn from a finite and learnable algebra: add field, add resource, add role, add permission, add workflow state, add integration, add notification, add idempotency, split entity, migrate invariant, upgrade dependency, add audit, expose report, tighten validation, change rollout, repair flaky proof lane. If true, agentic software engineering should focus less on unconstrained patch generation and more on recognizing, composing, proving, and amortizing these typed operations.
TABLE XIX C ANONICAL R EDUCTION D ENOMINATORS Axis
Moonshot range
Literal source-token Effective representation-space Agent action-space
Meaning Physical source shrinkage from generation, cells, and fewer duplicate surfaces.
2×–8× conservative 40×–150× central
Reasoning-token space
1,000×–100,000× end 3×–50× hypothesis
Retry/review loop
3×–50× hypothesis
Training-token budget
30×–150× hypothesis
Cost per verified correct change
3×–100× hypothesis
high-
Fewer valid encodings for the same supported behavior. Fewer legal files, edits, workflows, proof lanes, migration paths, dependency choices, and generated-zone mistakes. Fewer planning, tool, architecture-discovery, retry, and repair tokens per accepted change. Fewer failed tests, invalid edits, ambiguous reviews, rollback surprises, and proof reruns. Fewer examples needed to reach target canonical accuracy, conditional on paired scaling curves. End-to-end cost after foundry, verification, serving, reasoning, failed-loop, and maintenance amortization.
Cumulative substrate stack, sorted by local gain
local band: conservative / central / aggressive 1.0x / 1.0x / 1.0x
Raw human software universe
baseline 2x / 5x / 15x
+ Corpus hygiene: dedup, provenance, generated/vendor/secret/malware filtering
source/corpus reduction 3x / 10x / 50x
+ Reasoning digests and proof caches: known plans, failure modes, prior proof objects
reasoning-token reduction 3x / 15x / 50x
+ Proof lanes and reusable receipts: verification, rollback, migration replay, policy checks
retry/review-loop reduction
+ Canonical profiles: language roles, layout, naming, repository policy, dependency law, proof lanes
effective representation 8x / 40x / 120x
+ Contract-first generated truth: schemas, clients, docs, fixtures, validators, migrations
effective representation
+ Behavior cells: auth, lifecycle, search, forms, billing, uploads, jobs, webhooks, audit, observability + Constrained edit grammar and semantic patch cells
4x / 20x / 60x
15x / 75x / 250x
routine-product representation 100x / 10,000x / 100,000x
legal-action reduction
Minimum Functional Description: irreducible product intent, external contracts, domain invariants, novelty, and evidence. All-layer combined envelope (effective correct-change search work, not source size or dollar cost). Gross local-band product if fully independent: 8.64 × 105 x / 4.5 × 1011 x / 6.75 × 1015 x (orientation only; not claimed). Overlap-adjusted claimable envelope: 102 –103 x conservative / 104 –106 x central / 107 –109 x aggressive. The conservative band is capped because source hygiene, canonical representation, cells, legal edits, proofs, and reasoning caches share causal mechanisms.
Fig. 3. Canonical Compression Cascade: From Human-Code Entropy to Minimum Functional Description. The stack starts at the 1.0× raw baseline and then adds layers sorted by conservative local gain, with ties ordered by central gain. The right column gives the local hypothesis band for each layer. The final block combines all row multipliers two ways: a gross independent-axis product for orientation, and a deliberately smaller overlap-adjusted envelope for effective correct-change search work. The local ranges are measured on different denominators and are not naively multiplied. Training time, inference cost, and verified-change cost remain separate outcome denominators.
XIV. A SSUMPTION I: R EASONING -T OKEN C OMPRESSION AND AGENT-S PRAWL C OLLAPSE Assumption I: Canonical software compresses the reasoning process required to produce a verified correct change. Agent sprawl is a substrate failure: when a repository does not state its own laws, teams compensate with planner agents, search agents, reviewer agents, test-fixing agents, security agents, and release agents that repeatedly rediscover the same facts. Modern coding agents do not spend only source tokens. They spend planning tokens, tool tokens, hidden or explicit reasoning tokens, file-inspection tokens, failed test loops, repair attempts, and human review attention. Reasoningmodel documentation and best-practice guidance already treat reasoning effort as a controllable inference resource [54], [55]. Chain-of-thought, self-consistency, ReAct, and test-time planning work show that intermediate reasoning and tool use
can improve outcomes, but also make inference cost a first-class denominator [56]–[59]. The correct-change cost model is: Cchange = Tcontext + Treason + Taction + Tverify + Trepair + Hreview + Afoundry .
(31)
Here Tcontext is repository reading, Treason is planning and architecture inference, Taction is patch generation and tool use, Tverify is proof-lane execution and log interpretation, Trepair is failed-loop recovery, Hreview is human review burden, and Afoundry is amortized canonical-foundry cost. Canonical repositories reduce these terms by compiling repeated reasoning into durable artifacts. Behavior cells encode architecture. Generated-zone manifests tell the agent where not to edit. Semantic patch cells encode common repairs. Proof lanes encode the validation sequence. Repository policy
TABLE XX T YPED C HANGES S HOULD C OMPILE TO P ROOF O BLIGATIONS , N OT M ERELY F ILES Change operation
Canonical parameters
Compiled obligations
AddField
entity, name, type, nullability, default/backfill actor, resource, action, conditions source, target, key, backfill, dualwrite window provider, event, signature, retry, idempotency package, version, risk class, migration notes provider, scope, rollout window
Schema diff, migration replay, generated clients, UI states, fixtures, compatibility, rollback. Policy graph proof, negative tests, audit log, UI affordance, threat-model delta. Data migration, read/write compatibility, rollback, observability, latency budget.
AddPermissionEdge SplitTable AddWebhook UpgradeDependency RotateSecret
Contract fixture, signature verifier, replay test, dead-letter path, alerting. SBOM/provenance check, changelog constraints, compatibility tests, security scan, rollback. Secret inventory, dual-read window, audit receipt, rollback, leak scan.
encodes review, secret, dependency, and deployment law. Reasoning digests summarize known plans, failure modes, rejected plans, compatibility ghosts, and prior proof objects. A proof-carrying change object then binds the intent, affected cells, typed diff, schema diff, migration diff, test delta, security delta, proof obligations, receipts, rollback plan, and provenance metadata. The moonshot is therefore not that agents “think less” in the abstract. It is that agents stop spending tokens on repository cartography, local policy discovery, command guessing, generated-zone detection, migration-law inference, and repairloop archaeology when the substrate can state those facts once and verify them mechanically. Structured outputs and constrained decoding are useful precedents: they move part of correctness from prompt convention into machine-checkable shape [53], [60]. Proof-carrying code established the broader principle that executable artifacts can carry evidence checked by a consumer rather than accepted on trust [31]. The measurement unit is deliberately concrete: files opened, context tokens, planning tokens, hidden reasoning budget where observable, tool calls, invalid edits, generated-zone violations, failed proof loops, validation reruns, reviewer comments, and retries per accepted change. Closed models may hide some reasoning tokens, so field trials should report direct reasoning-token counters where available and proxy them with planning text, tool loops, wall-clock, and failed proof-lane traces where not. The conservative target is 3×–10× fewer context/reasoning/tool/retry tokens per verified change. The central target is 10×–100×. The aggressive 100×–10,000× band is only for routine supported work after behavior cells, semantic patch cells, proof lanes, and negative memory mature. It is not a universal inference-speed claim. Reasoning compression is not separate from economics; it is one term in the verified-change cost equation in Assumption F. A reasoning digest only matters if it reduces measured files opened, context tokens, reasoning tokens, tool calls, invalid edits, repair attempts, reviewer comments, or wall-clock at the same accepted-change standard. Falsification: Assumption I fails if, on paired raw/canonical tasks with the same broad model, canonical repositories do not reduce files opened, tool calls, reasoning tokens, invalid edits, failed verification loops, wall time, reviewer comments,
or cost per accepted proof-carrying change. XV. A SSUMPTION F: T RAINING AND I NFERENCE E CONOMICS Assumption F: Canonical code reduces required training data, active inference compute, failed repairs, and review burden enough to lower cost per verified correct change. Training compute for dense transformers is commonly approximated as Ctrain ≈ 6N D, (32) where N is parameter count and D is training tokens [28], [61]. The canonical program may attack both terms, but that remains a scaling-curve hypothesis. It should be stated as paired tokensto-target-accuracy, not as a naked training-speed claim. At fixed hardware tokens/s, tokens-to-target reduction is also trainingtime reduction; equivalently, physical tokens/s become rawcorpus-equivalent tokens/s multiplied by Rtrain . A defensible central target is 30×–150× lower tokens-to-target capability under successful paired scaling curves, with additional savings from lower active inference, fewer reasoning/tool tokens, fewer failed repair loops, and lower review burden. For inference, the dense-active speed claim is deliberately narrower than the verified-change claim. If a 100B–120B dense canonical specialist reaches the same accepted-change quality as a 1T-class dense-active baseline on supported canonical work, then first-order active-parameter accounting implies roughly 8.3×–10× more same-hardware aggregate token/sec and roughly 88%–90% lower dense-active infrastructure cost/token before memory, batching, context, and utilization corrections. That claim does not apply unchanged to sparse mixture-ofexperts systems: a 1T-total MoE with 32B active parameters may already be cheap per token, so the canonical MoE claim must be measured through router entropy, expert specialization, and end-to-end cost per verified change rather than nominal total parameter count. The all-in economics of canonical software intelligence are: Ccorrect = Afoundry + Ctrain + Cinfer + Ctools + Cverify + Creview + Cfailed + Cmaint .
(33)
The foundry is a capital expenditure on the training substrate. Its cost must be amortized across repeated model runs, served tokens, downstream agent tasks, security repairs, and reusable
certified cells. Active inference compute includes context tokens, reasoning tokens, action/tool tokens, verification-log interpretation, and repair attempts. The strongest business metric is therefore cost per verified correct change, not raw token price. The central specialist scenario is a matched-quality conditional, not a proven empirical result: if the 100B–120B model fails to match broad-model accepted-change quality on supported canonical tasks, then the dense-active speed and cost/token scenario does not activate. The simplest break-even condition is: k∗ =
Cfoundry . Craw/change − Ccanon/change
(34)
Here k ∗ is the number of accepted changes required to repay the foundry cost under a fixed task distribution. If the canonical path does not reduce same-quality change cost, or if verification/governance cost erases the savings, the economic claim fails even when source and action-space metrics improve. Falsification: Assumption F fails if empirical scaling curves show no reduction in tokens to target canonical accuracy, if active inference and reasoning tokens do not fall on paired change tasks, if the 8.3×–10× dense-active token/sec scenario disappears after memory, batching, context, utilization, verification, or quality corrections, if a canonical specialist needs more than 50% of broad dense-active parameters to match broad-model accuracy on held-out canonical tasks, or if foundry cost does not amortize within 3× the initial investment. XVI. A SSUMPTION G: C ANONICAL ROLE M IXTURE - OF -E XPERTS Assumption G: Canonical code gives mixture-of-experts models cleaner expert domains, reducing router entropy and improving expert utilization. Kimi K2 shows why dense and mixture-of-experts baselines must be separated: it reports 1T total parameters but only about 32B active parameters per token [14], [15]. A dense 100B canonical model is not automatically faster than a 32Bactive mixture-of-experts model on raw per-token compute. The canonical opportunity for mixture-of-experts is different. The human-code mixture-of-experts model must route across many languages, frameworks, and repository idioms. A canonical-code mixture-of-experts model routes across fewer, cleaner domains: product surface, service/core, durable data and migration, contracts/generated artifacts, security policy, dependency governance, verification, repair, and planning. Experts specialize around canonical software roles rather than arbitrary language/framework combinations. This matters because routing uncertainty is another form of human-code entropy. In a raw repository, a billing change might require a controller, browser form, secret, CI permission, SQL migration, queue worker, and webhook verifier, each following local conventions. In the canonical substrate, those concerns have named roles, owned files, generated boundaries, and proof lanes. The router can learn the role graph directly: billing cell, policy cell, SQL migration expert, UI resource-state expert,
verifier/repair expert. The model spends less capacity deciding what kind of software world it is inside. The canonical mixture-of-experts hypothesis is therefore not merely “use sparse models.” It is that canonicalization gives sparse models better expert boundaries. A role expert can specialize in migration invariants rather than every dialect of migrations; a security expert can specialize in policy edges and negative cases rather than every framework’s middleware idiom; a verifier expert can specialize in proof receipts and failed lanes rather than repository-specific CI folklore. The metrics are explicit. Router entropy should fall for a fixed task distribution: X Hrouter (x) = − p(e | x) log p(e | x). (35) e
Success also requires expert load balance, low dropped-token or overflow rate, specialization purity by software role, low cross-expert repair churn, and equal or better task success at fixed active parameters. A canonical role mixture-of-experts should be compared against dense models and broad mixtureof-experts baselines by accepted-patch rate and cost per verified correct change, not by nominal total parameter count. Falsification: Assumption G fails if router entropy, expert utilization, or task success do not improve versus language/framework-oriented mixture-of-experts routing. XVII. A SSUMPTION H: E VALUATION AND FALSIFICATION Assumption H: The canonical substrate reduces invalid actions and end-to-end change cost, not merely produces cleaner-looking code. The evaluation target is not HumanEvalstyle function synthesis. The target is whether a model operating inside the canonical universe can implement, modify, test, migrate, secure, and deploy ported software at lower cost than a broad model operating inside the human-chaotic universe. SWE-bench Verified is a useful precedent [5], [19], but static issue benchmarks are vulnerable to weak tests, contamination, and benchmark aging [6], [20], [62]. This paper’s claim needs paired raw/canonical tasks from the same software lineage. The benchmark object is the tuple: (raw_repo, canonical_repo, issue, behavior_contract, preservation_tier, proof_lanes, hidden_tests, reviewer_rubric, cost_ledger)
The proof program has three gates. Gate 1 is the samemodel substrate test: run the same broad frontier model on raw and canonical versions of the same issue lineage. This is the first falsification point because it isolates substrate value before specialist training. Gate 2 trains or adapts a canonical specialist and compares it on canonical tasks against the same broad baseline and the raw/canonical same-model result. Gate 3 measures paired scaling curves and verified-change economics: tokens-to-target, model size, serving cost, foundry amortization, and all-in cost per accepted proof-carrying change. The key ablation is direct: if the same broad frontier model performs much better on canonical repos than on the original raw repos, before canonical specialist training, then substrate compression is real. That proof
TABLE XXI C OST C OMPONENTS H IDDEN BY R AW T OKEN P RICING Component
Raw repository burden
Canonical compression route
Training compute
Verification cost
Long-tail languages, frameworks, layouts, generated drift, and weak labels. Long context reads, architecture rediscovery, exploratory tool use. Search, test discovery, command guessing, dependency archaeology. Ad hoc commands, flaky gates, unclear migration safety.
Review burden
Reviewers reconstruct intent and proof from a raw diff.
Failed repairs
Invalid edits, generated-file modifications, unsafe migrations, repeated test failures.
Paired scaling curves on canonical training objects and behavior graphs. Reasoning digests, file grammar, generated-zone manifests, proof-carrying change objects. Proof lanes, repository policy, dependency closure, standard receipts. Deterministic proof lanes, migration replay, rollback receipts, policy checks. Intent, typed diff, security delta, test delta, receipts, and rollback plan in one object. Constrained edit grammar, semantic patch cells, negativepath curriculum.
Active inference compute Tool calls
TABLE XXII C ANONICAL M ODEL T RAINING L ADDER Stage
Model/data configuration
Purpose
1B–3B scout
Canonical syntax, file grammar, naming, schema, and build command data. Messy-to-canonical traces, generated-boundary repairs, migrations, dependency updates. Full canonical repositories with long-context tasks, UI/product changes, DB changes, and security fixes. Full canonical corpus, behavior fixtures, profile routing, and long-horizon planning data. Experts for product surface, service/core, durable data and migration, contracts/generated artifacts, security/policy, infra/release, verifier/repair, and porting.
Verify that the standard is learnable under small-model budgets. Measure how much porting logic fits in small specialists.
7B–14B repairer 32B–70B system model 100B–120B central model Mixture-of-experts canonical frontier
TABLE XXIII C ANONICAL ROLE M IXTURE - OF -E XPERTS L AYOUT Expert family
Canonical role
Product surface
UI state, accessibility, visual behavior, and typed user-facing flows. Domain logic, APIs, adapters, and service invariants. Durable truth, schema changes, migration safety, and rollback. Schemas, generated clients, fixtures, adapters, and docs. Authentication, authorization, input boundaries, secrets, dependency risk. Deployment envelopes, release policy, observability, and operational permissions. Validation lanes, failed traces, blocked changes. Messy-to-canonical translation and disposition assignment.
Service/core Durable data/migration Contracts/generated Security/policy Infra/release Verifier/repair Porter
Establish capability curve before 100B-class runs. Central target for matching broad dense-active coding models on canonical work. Reduce active compute and improve router precision.
privacy, and consent obligations for code data [1], [64]. Software Heritage’s 2025 activity report is a reminder that the public software archive is enormous and must be treated as infrastructure, not scrapeable exhaust [65]. Compressionfactor overlap is handled by paired-corpus measurement; the paper never multiplies language collapse, framework collapse, cell collapse, action-space collapse, reasoning-token collapse, and repair-path collapse naively. Representational compression is not automatically model compression; that requires scaling curves. Falsification: Assumption H fails if cost per verified correct change is not at least 3× lower after including foundry, verification, serving, reasoning, failed-loop, review, and maintenance cost amortization. XVIII. C OST L EDGER AND F IELD T RIAL D ESIGN
does not require waiting for a 100B-class specialist. The paired task must measure accepted patch rate and the full trajectory: files opened, context tokens, reasoning tokens, tool calls, invalid edits, generated-zone violations, failed tests, failed migrations, wall time, reviewer comments, cost ledger entries, and rollback/proof receipts. Major risks are first-class evaluation targets and are expanded in Section XIX. Licensing and provenance must travel with every source artifact; The Stack v2 and the BigCode Governance Card both emphasize governance, provenance,
The paper should be judged by one primary endpoint: amortized cost per verified correct change. A result that reduces source tokens but increases proof failures is a loss. A result that improves benchmark pass rate but increases reviewer burden is incomplete. A result that lowers token cost while increasing incident remediation is not a win. The ledger must include source, context, reasoning, and action tokens; serving infrastructure; verification; security and provenance work; tool runtime; wall time; human review; failed attempts; rollback work; downstream defect cost where measurable; and foundry amortization.
TABLE XXIV C ANONICAL P ORT E VALUATION A RMS Arm
Model
Repository form
Purpose
A
Broad frontier model
Original human repo
B C D E
Broad frontier model Canonical specialist Canonical specialist + cells Porter model
Canonical repo Canonical repo Cellized canonical repo Human repo to canonical object
Current-world baseline: capability and cost in the messy corpus. Measures substrate benefit without a new model. Measures model plus substrate effect. Measures behavior-cell compression benefit. Measures foundry automation throughput.
TABLE XXV K ILLER A BLATION : R AW R EPO V ERSUS C ANONICAL R EPO WITH THE S AME M ODEL Metric
Raw repository expectation
Canonical repository target
Files opened Tool calls Reasoning tokens
Broad search across local architecture. Discover build, tests, migrations, policy, dependencies. Rediscover architecture, ownership, generated zones, repair strategy. Generated-file edits, wrong layer, unsafe migration, policy bypass. Flaky or misselected commands and repeated repair loops. Baseline for current-world agents.
Fewer predictable file reads from canonical path grammar. Invoke known proof lanes and typed patch tools. Use reasoning digest, cell contracts, and proof obligations.
Invalid edits Failed verification Accepted patch rate
Rejected by constrained edit grammar before proof lanes. Standard receipts and classified failures. Higher acceptance at equal or lower total cost.
TABLE XXVI C ONTAMINATION -AWARE B ENCHMARK T IERS Tier
Use
Risk controlled
Public paired
Reproducible raw/canonical ablations and open leaderboard tasks. Partner repositories, undisclosed issues, hidden tests, and blinded review. New issues evaluated before public disclosure and paired canonicalization. SWE-smith-style generated task families and preservation traps [63].
Transparent but contamination-prone; not sufficient for frontier claims. Reduces training contamination and patch memorization.
Private held-out Live / post-cutoff Generated stress tasks
raw + Araw C∆ foundry /N Rchange = canon , C∆ + Acanon foundry /N
(36)
Tests current repository reasoning under fresh task distributions. Scales coverage and tests false-green, generated-zone, and compatibility-ghost failures.
corpus with honest rejected dispositions is more valuable than a clean-looking corpus that silently drops behavior.
XIX. R ISKS AND T HREAT M ODEL where Afoundry is foundry construction and porting cost The canonical substrate is not risk-free. It deliberately amortized across N future changes. The aggressive case is concentrates software structure, so its failure modes must be impossible if the foundry never amortizes; the conservative named as design constraints rather than left as caveats. case can be true even while the foundry is still expensive if The central controlled-diversity rule is: canonicalize behavior repeated verified changes become much cheaper. interfaces, proof receipts, and edit grammars; diversify critical The first field trial should avoid the common benchmark implementations and evaluators where common-mode failure trap of measuring only task success. It should choose 20– would be catastrophic. N-version programming is a warning 50 production-like repositories with known issue lineage and as much as a precedent: diversity is useful only when failures hidden reviewer rubrics, port each into the canonical substrate, are sufficiently independent [66]. Unlimited local dialects are and run arms A–D from Table XXIV. The decisive result is not useful diversity. Independently validated lowerings, fuzzers, not that the canonical specialist wins. The decisive result is that the same broad model in arm B becomes cheaper and policy engines, monitors, and rollback paths are. more reliable than arm A. That isolates substrate value from XX. C OMPRESSION F RONTIER S UMMARY model value. Then arm C tests specialist training, and arm D The full 30-domain compression catalog is supplemental tests behavior-cell amortization. The paper should publish every failure. Failed ports reveal source in compression_catalog.tex. The main paper weak behavior oracles. Compatibility ghosts reveal hidden keeps only the proof spine: every domain is a prior or target contracts. Non-preserved behavior reveals product decisions. until paired measurement promotes it. The catalog is useful Negative examples become training material. The foundry for search-space accounting, but it is not the evidence base by should not hide these cases; it should classify them. A canonical itself.
TABLE XXVII C ANONICAL S TANDARD A BLATION L ADDER Ablation
What changes
What it proves
Naming only
Controlled vocabulary and role names; no layout or dependency changes. Canonical folders, ownership, test locations, generated/source paths. Contracts generate clients/adapters/fixtures; hand edits forbidden. Branch protection, code owners, required checks, workflow permissions, secrets, release gates. Pinned, owned, scored dependencies and upgrade lanes. Standard build/test/migration/security commands and acceptance artifacts. Top 50 cells with generated expansion.
Separates lexical entropy from structural entropy.
Path grammar only Generated boundaries only Repository policy only Dependency governance only Proof lanes only Behavior cells only Full standard
Language roles, file grammar, naming, generated boundaries, repository policy, dependencies, data truth, cells, validation.
Measures repository-navigation and edit-location reduction. Measures duplicate-truth and invalid generated-edit reduction. Measures hidden-state and CI/action-space reduction. Measures supply-chain and repair-loop reduction. Measures failed-loop reduction without full language/profile collapse. Measures source-token and action-space reduction from cells alone. Measures total canonical substrate effect.
TABLE XXVIII C ORE A SSUMPTIONS , S TRENGTH A SSESSMENTS , AND FALSIFICATION C RITERIA ID
Assumption
Strength
Falsified if
A
One governed way can cover each supported product-software concern. Full-corpus foundry can assign governed dispositions and amortize cost. Behavior can be preserved or dispositioned with graded confidence. Behavior cells absorb high-reuse routine product code. Canonicalization reduces representation entropy and legal action-space. Canonical data reduces required model size and training tokens. Canonical mixture-of-experts routes better by software role than by language accident. Canonical substrate lowers cost per verified correct change. Canonical substrate compresses reasoning, tool, planning, and retry tokens.
Strong
Hardest
Primary profile covers less than 60% of economically valuable product/web/backend/data tasks. Large fractions remain undispositioned or foundry cost exceeds downstream savings. Human review rejects core ports or security/migration regressions rise.
Very plausible
Top 500–2,000 cells cover less than 50% of product-app behavior.
Central claim
Plausible
Paired corpora show no entropy, path, AST, perplexity, or action-branching reduction. Scaling curves show no lower token or parameter need for target canonical accuracy. Router entropy, utilization, and task success do not improve.
Business claim
End-to-end change cost is not at least 3× lower after amortization.
Moonshot claim
Same-model canonical tasks do not reduce files opened, tool calls, reasoning tokens, invalid edits, failed tests, wall time, or reviewer burden.
B C D E F G H I
Hard
Hypothesis
The frontier should be read as an accounting program. Literal source-token reduction is expected to be modest compared with legal-action and reasoning-space reduction. Training-token reduction is the hardest claim and remains conditional on paired scaling curves. Cost per verified correct change is the final denominator because it includes foundry, verification, serving, reasoning, failed-loop, review, and maintenance cost. XXI. F UTURE V ISION : M INIMUM V IABLE N OVELTY The long-term prize is not a smaller corpus of cleaner code. It is a post-source-code substrate in which routine software is compiled from compact intent, behavior cells, invariants, policies, proofs, provenance, runtime evidence, and reusable reasoning. Source code remains useful, inspectable, and deployable, but it becomes one generated projection of a deeper object. This is the extreme form of MFDL: the system should contain only minimum viable novelty. Humans still decide product
goals, legal obligations, risk tolerance, taste, and genuinely new domain behavior. But all routine coding decisions—language, framework, folder, generated boundary, migration idiom, dependency wrapper, proof lane, review evidence, rollback form, and repair strategy—move out of human discretion and into governed substrate law. The defensible path starts with facts already visible in the field. Public software archives are massive and structured; Software Heritage reported over 27 billion unique source files from 421 million projects in its 2025 activity report [65]. Raw code is repetitive and predictable [18]; large clone studies and notebook studies show heavy duplication [3], [4]. Product-line engineering has long treated software families as managed commonality plus variability [30]. LLVM, MLIR, e-graphs, OpenAPI, Protocol Buffers, and Coccinelle are all partial precedents for moving from raw text toward intermediate representations, generated projections, equivalence classes, contracts, and semantic transformations [21]–[24], [26], [29].
TABLE XXIX M INIMUM L EDGER FOR A D EFENSIBLE F IELD T RIAL Ledger item
Raw repository measurement
Canonical repository measurement
Discovery
Files opened, grep/search calls, dependency reads, architecture notes. Input tokens, cache hit rate, long-context latency.
Cell lookups, profile docs read, reasoning digest reads.
Context/pre-fill Planning/reasoning Generation/action
Planning tokens, hidden reasoning budget where measurable, self-repair tokens, tool-selection deliberation. Output tokens, patch size, tool calls, action tokens.
Review Reliability
Dense-active or MoE cost/token, token/sec, memory, batching, utilization. Test commands, failed lanes, log interpretation, flaky reruns. SBOM, license/provenance review, secret scan, policy checks. Reviewer comments, review minutes, rework cycles. Escaped defects, incident tickets, rollback time.
Amortization
None or local scripts only.
Serving Verification Security/provenance
Reduced context tokens, generated-zone manifests, cached profile/cell context. Typed change recognition, bounded proof-plan tokens, noreason lanes. Typed operation, generated diff, proof object, rollback object. Matched-quality canonical specialist serving ledger plus verification overhead. Standard proof-lane receipts, classified failures, deterministic replay. Carried metadata, transformation receipts, governed policy gates. Intent audit, proof receipt audit, residual novelty audit. Runtime invariant violations, negative-memory updates, rejected cells. Foundry cost spread over future changes, repos, profiles, and cells.
TABLE XXX C ANONICAL S UBSTRATE T HREAT M ODEL Risk
Failure mode
Required mitigation
Behavior loss
Port drops hidden edge behavior, migration side effects, timing assumptions, or compatibility ghosts. Tests pass while behavior diverges.
Declared contracts, preservation tiers, differential replay, accepted-incompatibility manifests, and human escalation. Treat tests as evidence only; add fuzzing, property tests, mutation tests, generated tests, hidden review, and trace replay. SPDX/CycloneDX metadata, source hashes, transformation traces, removal propagation, and policy gates. One governed interface with independent implementations for critical cells, conformance suites, canaries, rollback, and cell advisories. Secondary profiles, governed primitives, explicit exception process, and profile lifecycle evidence.
Weak oracles Licensing/provenance Canonical monoculture Useful-diversity loss Benchmark contamination Foundry cost Governance capture Generated-code opacity Scaling-curve uncertainty
Canonical artifact loses attribution, opt-out state, source lineage, or derivative-risk status. One flawed auth, migration, billing, or policy cell creates correlated failure at scale. Profile excludes a domain-specific implementation choice that carries performance, safety, accessibility, regulatory, or ecosystem value. Public paired tasks leak into training or static tests age out. Porting, proof, review, and governance costs exceed downstream savings. Canonical rules favor one vendor, stack, or implementation without evidence. Source becomes generated but reviewers cannot understand obligations or failures. Representation compression does not translate to trainingtoken or model-size reduction.
Public, private held-out, and live/post-cutoff benchmark tiers. Break-even accounting, staged dispositions, high-reuse prioritization, and amortization receipts. Versioned public specs, independent conformance suites, audit logs, and appeal paths. Proof-carrying change objects, generated-zone manifests, readable diffs, and reviewer rubrics. Paired raw/canonical scaling curves and fixed-target accepted-change benchmarks.
Proof-carrying code, CompCert, seL4, and SLSA show that evidence and provenance can become first-class artifacts rather than after-the-fact paperwork [31], [41], [48], [49].
learning supports the general idea that learned abstractions can shorten future program search [27]; the canonical foundry applies that logic to production software behavior.
Software genome. The foundry should mine raw repositories, issues, tests, traces, incidents, dependency histories, and repair patches into a global atlas of behavior families: authentication, tenant boundaries, resource lifecycle, billing, uploads, jobs, webhooks, audit, observability, migrations, notifications, permissions, retries, cache invalidation, and product workflows. The output is not a snippet library. It is a behavior gene bank: canonical intent, variants, anti-examples, contracts, tests, proof templates, repair memories, provenance, legal transformations, and generated implementations. DreamCoder-style library
Behavior IR. The substrate needs an application-level intermediate representation above source code. A behavior IR represents entities, state machines, policies, permissions, effects, external contracts, migration semantics, observability, runtime SLOs, and proof obligations. It lowers into governed service code, typed product surfaces, durable-data operations, UI, docs, tests, dashboards, deployment manifests, and formal/specification artifacts. This is the product-software analogue of compiler IR: agents edit semantic deltas, not arbitrary file trees.
TABLE XXXI M AIN -PAPER C OMPRESSION F RONTIER S UMMARY Layer
Current status
Main mechanism
Required measurement
Corpus hygiene
Partly measured in prior corpora
Raw/canonical token counts, removal receipts, license disposition rates.
Canonical profile
Near-term target
Contract-first generation
Central hypothesis
Behavior cells
Central hypothesis
Semantic patch cells
Central hypothesis
Runtime and negative memory
Moonshot
Deduplication, generated/vendor/secret/malware filtering, provenance assignment. Fixed roles, path grammar, generated zones, proof lanes, dependency law, repository policy. Schemas generate clients, fixtures, validators, docs, migrations, and adapters. Auth, resource lifecycle, search, billing, uploads, jobs, webhooks, audit, observability as certified cells. Typed edits such as add field, split table, rotate secret, add idempotency key, and add permission edge. Incidents, reverted patches, production traces, and rejected plans become invariants and forbidden paths.
Same-model raw/canonical files opened, invalid edits, proof-lane failures, accepted patch rate. Duplicate-truth removal, generated-edit violations, schema drift rate, migration failures. Behavior-cell census by accepted changes, AST nodes, traces, issues, review burden, and security paths. Legal action count, invalid edit count, proof receipt completeness, repair-loop reduction. Incident recurrence, compatibility-ghost capture, negative-test reuse, post-deployment regression rate.
TABLE XXXII T HREE B UILD -O UT S TAGES FOR THE C ANONICAL S UBSTRATE Stage
Substrate built
Main compression denominator
1. Canonical repo
Primary/secondary profiles, file grammar, generated zones, proof lanes, repository policy.
Training tokens, active context, tool discovery, invalid edits.
2. Correct-change substrate
Behavior cells, semantic patch cells, proofcarrying change objects, reasoning digests, negative corpus. Behavior IR, software genome, runtime-derived invariants, agent-native OS, verification markets, evolutionary foundry.
Action space, repair loops, review burden, repeated reasoning.
3. Behavior genome
Routine-domain training, action/reasoning/search space, institutional memory, non-novel implementation.
TABLE XXXIII T YPED C HANGE A LGEBRA E XAMPLES Operation
Parameters
AddField
User, timezone, nullable=false, backfill=UTC actor=Manager, action=ApproveInvoice source=Events, target=AuditEvents endpoint=WebhookDispatch provider=Stripe
AddPermissionEdge SplitTable AddIdempotencyKey RotateSecret
Typed change algebra. Routine changes should become typed operations with known proof obligations: Each operation expands into schema diff, migration plan, generated client updates, UI state changes, security review, observability deltas, tests, rollback, and receipts. Reasoning compiler. The most important future gain is cognitive amortization. Successful and failed trajectories should compile into reusable plans: issue type to affected cells, repository profile to legal actions, failure signature to repair strategy, migration class to proof obligations, dependency update to compatibility checks. Current reasoning and acting methods show why intermediate reasoning helps [56]–[58]; the canonical future is to spend that reasoning once, store it as substrate law, and execute known work through no-reason or
Target range 3×–10× training; 2×–8× inference/reasoning/tool; 2×–5× verified-change cost. 30×–150× training; 10×–100× inference/reasoning/tool; 10×–50× verifiedchange cost. 150×–1,000× routine-domain training; 100×–10,000× inference/action/reasoning; 50×–1,000× verified-change cost.
bounded-reason lanes. Proof-carrying changes. Future agents should not submit raw diffs. They should submit typed change objects containing intent, affected behavior cells, semantic patch, generated source diff, schema diff, migration diff, tests, security delta, proof obligations, receipts, rollback plan, and provenance/license metadata. Human review shifts from reconstructing intent from text to auditing changed obligations, assumptions, and residual novelty. Evolutionary foundry. Once behavior cells have evaluators, cells can improve continuously. Agents propose variants; tests, proofs, fuzzers, benchmarks, security scanners, and runtime monitors select survivors. AlphaDev and FunSearch are narrow but important evidence that evaluator-guided program search can discover useful algorithms beyond ordinary human implementation [67], [68]. The canonical version extends the loop to product behavior: better retries, indexes, migrations, policy encodings, UI state machines, rollback paths, and proof envelopes. Runtime and negative memory. The substrate should learn from production traces, incidents, rejected patches, reverted commits, review objections, failed migrations, security bugs, flaky tests, and rollbacks. Runtime behavior reveals implicit compatibility contracts; negative examples become forbidden plans, red-team tests, proof obligations, and repair memories.
This is how the system avoids destroying hidden behavior while still compressing historical accident. Verification markets and agent-native software OS. A mature substrate needs independent evaluators: proof-lane providers, fuzzing services, conformance suites, cell auditors, provenance validators, and runtime monitors competing on evidence quality. The agent-facing operating system is then not a terminal plus repository checkout. It is a governed environment exposing typed changes, behavior cells, proof lanes, policy decisions, provenance receipts, runtime traces, and negative memory as native system calls. The extreme endpoint is not zero code and not zero engineering judgment. It is no accidental software: no accidental architecture, no accidental dependency choice, no accidental CI, no accidental migration ritual, no accidental security policy, no accidental review burden, and no repeated reasoning where the substrate has already learned the law. The human and agent frontier moves to irreducible novelty.
invariants would turn production traces, rollbacks, failed migrations, and reverted patches into negative memory and proof obligations. Verification markets would let independent proof-lane providers, fuzzers, conformance suites, cell auditors, provenance validators, and runtime monitors compete on evidence quality. Agent-native operating systems would expose typed changes, behavior cells, proof lanes, policy decisions, provenance receipts, runtime traces, and negative memory as native system calls rather than leaving agents inside an unstructured terminal and repository checkout. A mature canonical foundry should therefore publish two ledgers. The corpus ledger reports what was ingested, rejected, ported, preserved, proven, license-cleared, or quarantined. The change ledger reports every accepted and failed task: context tokens, reasoning tokens, tool calls, invalid actions, proof failures, repair loops, wall-clock, review comments, and dollars. The paper’s central numbers should rise or fall with those ledgers.
XXII. T HE F IRST S IX E XPERIMENTS T HAT D ECIDE THE T HESIS
XXIV. T HE N O -ACCIDENT H ORIZON
The paper becomes stronger if it stops asking readers to believe a grand end state and instead offers a near-term kill chain. The first decisive program should be small enough to run before a full foundry exists and strong enough that a negative result hurts. These experiments also protect the paper from its own ambition. The 150×–1,000× story is not the starting claim; it is the long-horizon envelope after cells, proof lanes, negative memory, and behavior IR mature. The starting claim is harsher and cleaner: if same-model raw/canonical ablation does not produce obvious cost and search-work savings, the trainingtoken moonshot should not be believed.
The strongest form of this paper is not that all future software becomes free. That would be false for arbitrary future programs and misleading for reviewers. The stronger and defensible claim is narrower: once accidental representation, repeated architecture, duplicated contracts, routine behavior families, proof-route discovery, generated-surface drift, dependency rituals, migration folklore, and repair loops have been quotiented away, software converges to a residual floor. That floor is the No-Accident Horizon. A. The Limit Question
Let At be the admissible software evidence available at time t: repositories, issues, pull requests, tests, traces, incidents, vulnerabilities, reviews, schemas, deployment histories, rejected XXIII. F RONTIER E XTENSIONS T HAT M AKE THE T HESIS patches, documentation, provenance, license metadata, and H ARDER TO I GNORE rollback evidence. Let Mt = Φ(At ) be the best governed The current paper should be read as a substrate thesis plus a canonical memory produced from that evidence. The memory research agenda. The strongest next ideas are not cosmetic; they is admissible only if it preserves provenance, weak-oracle turn the claim from a compression manifesto into a defensible labels, rejected dispositions, security constraints, migration law, and behavior evidence. experimental program. Let Y ∼ Pt (Y | At ) be a future software demand drawn The most important addition is the behavior-cell coverage from a declared workload distribution, and let Oτ,H be an census. If 70–90% of routine product-software changes fall into acceptance oracle at evidence tier τ and future horizon H. The reusable cells and semantic patch cells, the aggressive thesis oracle includes behavior, hidden compatibility, security, migrabecomes plausible. If coverage stalls at 20–30%, canonical code tion safety, provenance, review policy, and future adaptability. may still be valuable, but the paper becomes an engineeringA foundry F maps (M , Y ) to a proof-carrying change object t efficiency paper rather than a training-substrate revolution. The or to a refusal. Its all-in accepted-change cost is second most important addition is the paired repository arena, because it directly tests the economic question: can the same CF (Y ) = Cintent + Ccontext + Creason + Ctool + Cverify or smaller model produce accepted proof-carrying changes + Creview + Crisk + Camortized foundry + Cdefect . with fewer total tokens, tool calls, failed loops, and reviewer (37) interventions? Four frontier extensions should remain explicitly future work, The raw baseline Craw (Y ) is the same ledger for the original not current evidence. Behavior genomes would mine raw human repository and process under the same oracle. The repositories, traces, incidents, and repairs into a global atlas denominator is therefore cost per verified correct change, not of behavior families with variants, invariants, tests, proofs, lines of code, source bytes, prompt tokens, or benchmark solve provenance, and generated projections. Runtime-derived rate.
TABLE XXXIV S IX E XPERIMENTS T HAT C ONVERT THE T HESIS INTO E VIDENCE Experiment
Pass condition
Fails the thesis if
Same-model substrate ablation
A broad frontier or strong open model solves paired canonical tasks with materially fewer files, tokens, tool calls, failed loops, and reviewer comments than raw tasks. Same model family reaches fixed accepted-change quality with fewer canonical tokens than raw tokens, with confidence intervals. Paired canonical tasks show fewer planning tokens, hidden reasoning tokens where measurable, tool calls, invalid edits, proof failures, validation reruns, and retries. Ports survive hidden tests, fuzzing, mutation tests, replay, migration checks, security negatives, and human adjudication. Porting, proof, governance, serving, verification, review, and maintenance costs are repaid by repeated accepted-change savings. A canonical specialist matches or beats a broad model on supported canonical work at lower all-in cost.
Canonical repositories look cleaner but do not reduce search work or accepted-change cost.
Training tokens to acceptedchange target Reasoning/tool/retry reduction Foundry behavior preservation Cost per verified correct change Specialist model
versus
broad
Representation compression does not become sampleefficiency gain. The model spends the same search budget despite the canonical substrate. Canonicalization passes easy tests while losing compatibility ghosts or security behavior. Foundry cost or verification overhead erases downstream economics. Supported-domain specialization fails to improve acceptedchange quality, speed, or cost after corrections.
TABLE XXXV H IGHEST-L EVERAGE A DDITIONS FOR A S TRONGER C ANONICAL -C ODE P ROGRAM Addition
Why it matters
Concrete artifact
Behavior-cell coverage census
Converts moonshot language into a measured market map: what fraction of product software is resource lifecycle, auth, policy, workflow, billing, notification, observability, integration, or migration routine? Prevents the foundry from optimizing prettiness instead of preserved behavior and lower change cost.
A labeled corpus with per-file and per-change coverage by cell family, plus residual novelty estimates.
Canonicalization loss function Paired repository arena
Makes the decisive claim testable without waiting for fullcorpus porting.
Negative memory bank
Turns incidents, reverted patches, flaky tests, security bugs, and failed migrations into reusable anti-examples. Keeps the profile ambitious without becoming brittle or cultish.
Renderer pluralism Foundry amortization ledger
Makes the economics credible.
Adversarial equivalence audit
Protects against canonical ports that pass easy tests while losing edge behavior.
For a distribution and oracle, define the removable-work fraction as ΛNA (Pt , Oτ,H ) = 1 −
inf F ∈Fadm EY ∼Pt [CF (Y )] . EY ∼Pt [Craw (Y )]
(38)
The corresponding ideal multiplier is MNA =
1 . 1 − ΛNA
(39)
This normalization is severe. A 10× result means 90% removable work, a 100× result means 99% removable work, and a 1,000× result means 99.9% removable work. B. No Universal Software Inverse No universal software inverse. For arbitrary computable programs or adversarial future software demands, no computable
Multi-objective score: behavior preservation, proof strength, source entropy, action entropy, security posture, provenance quality, and amortized cost. Raw/canonical repository pairs, identical issues, hidden tests, proof lanes, reviewer rubrics, and full token/tool/dollar traces. Versioned forbidden-plan cells and regression generators attached to behavior cells and proof lanes. One behavior IR rendered to the primary product profile and to secondary governed profiles where domain constraints require it. Per-port cost, review cost, proof cost, reuse count, servedtoken savings, accepted-change savings, and break-even k∗ . Differential replay, fuzzing, mutation testing, symbolic checks where feasible, security review, and human productowner adjudication.
foundry has a positive guaranteed No-Accident Horizon. In the distribution-free case, inf ΛNA (P, Oτ,H ) = 0. P
(40)
Proof sketch. If a system could always construct, prove, or correctly reject every arbitrary future behavior at finite bounded residual cost, it could decide nontrivial semantic properties of arbitrary programs by encoding those properties as software intents. That contradicts the Turing–Rice boundary [43], [69], [70]. If the future distribution is unconstrained, no-freelunch reasoning gives the learning-theoretic version of the same warning: an optimizer wins only by exploiting nonuniform structure in the task distribution [71]. The shortest adequate description of an arbitrary target is also not generally computable in the Kolmogorov sense [72]. A foundry can
dominate useful regions of software space, but it cannot own the shortest proof-carrying description of every possible future program. This negative result is the boundary that makes the positive theory scientific. The claim is not that a model learns all possible future programs. The claim is that economically important software is concentrated in repeated behavior families, organizational patterns, integration rituals, schemas, policies, migrations, operational failures, and product conventions. That concentration is exactly what canonical code exploits [3], [4], [18], [27], [30]. C. The Horizon Equation Normalize today’s raw all-in cost mass to one. Decompose it into an irreducible floor η and disjoint reducible strata q1 , . . . , qJ : η+
J X
qj = 1,
qj ≥ 0,
η ≥ 0.
(41)
j=1
Each qj is a removable source of accidental work: representation search, repository discovery, duplicate truth, prooflane discovery, invalid edit attempts, routine behavior implementation, generated surface maintenance, retry loops, local architecture reconstruction, dependency rituals, migration folklore, or repeated review reasoning. Let rj ≥ 1 be the asymptotic reduction achieved on stratum j by behavior cells, semantic patch cells, proof-carrying changes, canonical profiles, negative memory, renderers, and reusable evidence. The residual cost fraction is J X qj 1 sNA = η + , MNA = . (42) r sNA j=1 j Equation 42 is the Amdahl law of inverse software [73]. If two techniques reduce the same failed-search loop, they do not multiply; they compete for the same qj . The exhausted-options limit is 1 lim MNA = . (43) r1 ,...,rJ →∞ η The foundry does not escape the denominator. It drives every reducible term toward zero until only η remains. D. Future-Adaptive Minimum Functional Description Length The irreducible floor is not merely source length. It is the cost of new information, acceptance, evidence, governance, and adaptability. An oracle-relative form is E[CF (Y )] ≥ E αKO (Y | Mt ) + βLO (Y ) + γEτ (Y ) + δG(Y ) + ζVH (Y ) . (44) Here KO (Y | Mt ) is the target behavior not already implied by canonical memory, LO (Y ) is the description length of the acceptance oracle, Eτ (Y ) is the minimum evidence burden, G(Y ) is governance and provenance burden, and VH (Y ) is the option-value burden across horizon H. This is the MFDL boundary in future-adaptive form: canonicalization can shorten descriptions and make proofs reusable, but it cannot make required new information disappear [72], [74].
E. Numerical Limit Estimate The strongest defensible numerical answer is a regime table, not a single slogan. Table XXXVI states the removable-work fraction ΛNA and reciprocal multiplier MNA for increasingly favorable assumptions. The table says three things. First, the universal mathematical problem has no positive guaranteed compression. Second, broad commercial software is compressible but still contains enough hardware, scientific, regulatory, adversarial, organizational, and algorithmic novelty that a universal 100× all-in claim is not defensible. Third, the paper’s real frontier is supported routine product software after behavior-genome maturity. In that regime, the strongest defensible final-limit planning number is Λ⋆NA,product ≈ 0.990
⇐⇒
⋆ MNA,product ≈ 100 × . (45)
A cautious theoretical band is Λ5%–95% NA,product ∈ [0.967, 0.998], 5%–95% MNA,product ∈ [30×, 500×].
(46)
This is a theoretical limit estimate, not a measured result. It is stronger than the near-term central target because it assumes mature behavior cells, semantic patch cells, reusable proof receipts, negative memory, renderers, and broad amortization. It is weaker than an unqualified 1,000× claim because Equation 43 respects the residual floor. F. Falsification Rules The theory is designed to fail cleanly. Estimate η, qj , and rj on paired raw/canonical repositories from the same lineage under the same model, issue distribution, hidden tests, proof lanes, reviewer rubric, and cost ledger. The No-Accident Horizon collapses into ordinary engineering efficiency if any of the following hold: 1) the irreducible floor η remains above 3% for routine product changes after mature canonicalization; 2) behavior-cell and semantic-patch coverage fail to exceed roughly 85% of accepted routine change cost mass; 3) covered strata do not reach at least 100× reduction in context, reasoning, invalid-action, retry, and proofdiscovery cost; 4) proof, security, provenance, migration, or review overhead grows enough to erase the saved search cost; 5) canonical ports lose important compatibility ghosts or increase downstream defects at equal review standards; 6) foundry amortization does not compound across independent repositories. Conversely, the 100× exhausted-options limit becomes conservative if η < 0.5%, covered routine-change mass exceeds 95%, cell-covered strata reach thousands-fold reductions, and proof receipts reuse across many unrelated product lineages without loss of behavior or provenance. The final claim is therefore precise enough to defend and large enough to matter: the universe of economically
Improvement multiplier (larger reductions lower on chart)
1x
1x current
3x
10x
Conservative
30x
No-Accident Horizon 100x planning limit
100x
300x
Aggressive
500x 0
0.25
0.50
0.75
1.0
Canonical substrate maturity Fig. 4. Convergence from current all-in verified-change cost toward the No-Accident Horizon. The conservative curve approaches a 30× mature-foundry limit, the aggressive curve approaches a 500× upper-tail limit for supported routine-product work, and the dashed line marks the 100× theoretical planning horizon. The band is a theoretical exhausted-options estimate, not a measured result and not a distribution-free guarantee.
important code can be built increasingly close to the NoAccident Horizon, but never beyond it. For arbitrary future programs, the guaranteed horizon is zero. For routine product/application software under a mature behavior-genome foundry, the strongest defensible final-limit estimate is about 100× lower all-in cost per verified correct change, with a cautious 30×–500× theoretical band. Values above 1,000× belong to closed behavior-cell lanes where the irreducible floor is below 0.1%, not to unconstrained software as a whole. XXV. L IMITS AND C ONCLUSION This paper does not prove the full economic thesis. It does not prove behavior preservation at corpus scale, does not prove paired training-token reduction, does not prove that a 100B-class canonical specialist replaces all broad frontier coding systems, does not apply to all novel systems or unsupported domains, and does not prove that behavior cells cover 70%–90% of routine product behavior. The QLoRA pilot is deliberately narrower: canonical translated trajectories are learnable under a parameter-efficient adaptation setup, and the measured forbidden-language markers remain at zero. The rest of the thesis must be earned through the claim ledger in Table VI, beginning with the same-model raw/canonical substrate test. The research agenda is therefore concrete. Build paired raw/canonical repositories from the same lineage. Assign every source artifact a legal/provenance disposition. Port
behavior only within declared contracts. Report preservation tiers, weak-oracle failures, compatibility ghosts, and accepted incompatibilities. Run the same broad model on raw and canonical tasks before training a specialist. Measure files opened, context tokens, reasoning tokens, tool calls, invalid edits, failed proof lanes, reviewer comments, wall-clock, and cost per accepted proof-carrying change. Then run paired scaling curves before claiming training-token compression. The thesis remains bold because the target is not cleaner code. The target is a canonical mirror of the software universe: behavior preserved or dispositioned, routine product work collapsed into certified cells, changes expressed as typed operations, proof lanes generated and checked, provenance carried through every transformation, and runtime failures compiled into negative memory. Human code contains the behavior, edge cases, incidents, product judgment, and hardwon lessons that matter. Raw human representation is the wrong thing to imitate when it encodes local accident rather than durable behavior. The theoretical limit is the No-Accident Horizon developed in Section XXIV. In that limit, agents do not invent local architecture, rediscover repository folklore, guess migration law, or re-learn the same auth, billing, upload, policy, and audit patterns in every codebase. They operate a constrained software machine and spend compute only where novelty, evidence, governance, risk, and future optionality remain. The end state is no accidental software: no accidental
TABLE XXXVI N O -ACCIDENT H ORIZON : F INAL -L IMIT E STIMATES BY S COPE Scope
Interpretation
Arbitrary programs as mathematical objects
Distribution-free future software, including adversarial and incompressible targets. Product apps plus systems, embedded, data, scientific, security, mobile, infrastructure, creative tools, and genuinely new algorithms. SaaS, internal tools, resource lifecycles, auth, policy, billing, uploads, jobs, webhooks, audit, migrations, observability, deployment, and integrations. Narrow lanes where the future change is almost entirely cell selection, parameterization, renderer output, and reusable proof receipts.
Broad economically requested software
Routine product and application software
Closed behavior-cell lanes
Conservative
Central exhausted-options limit
Aggressive upper tail
0% / 1×
0% / 1×
No positive universal bound
75% / 4×
88% / 8.3×
96% / 25×
96.7% / 30×
99.0% / 100×
99.8% / 500×
99.0% / 100×
99.9% / 1,000×
99.98% / 5,000×
architecture, dependency choice, CI ritual, migration path, security policy, review burden, or repeated reasoning once the substrate has learned the law. This paper is a falsifiable program for replacing accidental software representation with a canonical, proof-carrying substrate for verified correct change. In the strongest form, every remaining unit of engineering buys novelty, judgment, evidence, governance, or safety. A PPENDIX A S UPPLEMENTAL C OMPRESSION C ATALOG : C ODE -S PACE G AIN P RIORS This supplemental catalog lists the major coding-work gains available once agent-first canonical code is treated as a substrate rather than a style guide. Each domain states what is rebuilt, what becomes generated, what agent action paths disappear, and conservative, central, and aggressive gain estimates. The gains do not multiply cleanly—many overlap, and several domains share the same underlying token mass. The composite effect must be measured on paired raw/canonical corpora; the estimates presented here are prior ranges and target hypotheses, not measured results. Where a per-domain source range such as 3×–8× is stated, it denotes the ratio of raw-corpus tokens to canonical-corpus tokens required to represent equivalent behavior. Composite ranges use the explicit denominator named in the row: representation space, action space, training tokens, reasoning/tool/retry tokens, or cost per verified correct change.
1) Language Role Collapse: Approach. Rebuild service, UI, migration, and scripting roles into fixed profile lanes; generate adapters at profile boundaries; remove languagechoice, runtime-choice, and cross-language repair paths. The same backend behavior—HTTP routing, database access, queue consumption, serialization—appears in Python/FastAPI, Java/Spring, Node/Express, Ruby/Rails, PHP/Laravel, Go/Gin, and C#/ASP.NET, among others. A canonical standard collapses these into one service/core lane and one typed product-surface lane. The model no longer spends capacity learning seven syntactically distinct encodings of identical semantics. Estimates. Conservative: 4×–6×. Central: 6×–12×. Aggressive: 12×–20×. 2) Framework Collapse: Approach. Rebuild routing, middleware, state, and persistence into one service/UI grammar; generate framework glue; remove local framework DSL choices and handler-shape variants. Within a single language, framework conventions— middleware chains, routing DSLs, ORM query builders, template engines, configuration idioms—introduce a second layer of representational divergence. Express, Django, Rails, Spring, and Laravel each impose a distinct grammar over the same behavioral primitives. Canonical conversion replaces all framework-specific grammars with one canonical service grammar, eliminating the combinatorial surface the model must memorize.
A. Infrastructure Compression
Estimates. Conservative: 3×–5×. Central: 5×–10×. Aggressive: 10×–15×.
Infrastructure compression eliminates the representational cost of decisions that carry no behavioral consequence: which language encodes a REST handler, which framework provides routing, which folder tree organizes source files. These decisions dominate open-source corpora yet contribute nothing to the space of behaviors a model must learn to produce. Domains 1–6 target this stratum.
3) Generated Truth Collapse: Approach. Rebuild schemas as the source of truth; generate data transfer objects, clients, validators, fixtures, mocks, and docs; remove hand-edited projections and drift repairs. Hand-written data transfer objects, API client stubs, request/response validators, serialization fixtures, and test factories are all downstream projections of a single schema contract.
In raw corpora these projections diverge, drift, and duplicate. Canonical form replaces them with contract-first generation: one source of truth (an API schema or cell contract) produces all projections deterministically. The model learns to emit the contract, not its many manual echoes. Estimates. Conservative: 2×–3×. Central: 3×–6×. Aggressive: 6×–10×. 4) Layout and Naming Collapse: Approach. Rebuild repositories into canonical paths and controlled vocabulary; generate owner maps and generated-zone manifests; remove file-hunt, synonym, and helper-placement choices. Arbitrary folder structures (src/controllers/ vs. app/handlers/ vs. lib/api/), file-naming conventions (userService.ts vs. user_service.py vs. UserService.java), and organizational myths consume representational capacity without encoding behavior. A canonical repository grammar fixes vocabulary and topology, converting layout noise into a stable, predictable structure the model can exploit as prior knowledge rather than re-learn per repository. Estimates. Conservative: 2×–3×. Central: 3×–5×. Aggressive: 5×–8×. 5) Dependency Collapse: Approach. Rebuild integrations behind governed adapters; generate configuration, retry, error, and observability wrappers; remove bespoke vendor glue and unsafe upgrade routes. Every production codebase wraps third-party services— Stripe, S3, GitHub, Slack, OpenAI, Redis, email providers, observability backends—in locally invented adapter layers with inconsistent error handling, retry policies, and configuration surfaces. Canonical adapters provide governed wrappers with standard interfaces. The model learns one adapter contract per external service rather than hundreds of bespoke wrappers, and local glue code vanishes entirely. Estimates. Conservative: 2×–3×. Central: 3×–5×. Aggressive: 5×–8×. 6) Build and Test Collapse: Approach. Rebuild local and CI validation as proof lanes; generate GitHub/GitLab workflow settings, artifacts, and required checks; remove build folklore and permission guesswork. Custom shell scripts, undocumented CI gate configurations, Makefile folklore, local environment assumptions, and ad-hoc test harnesses constitute a significant fraction of repository tokens. Canonical form replaces them with standard proof lanes: a fixed set of commands that produce expected artifacts (type-checked binary, migration proof, test evidence, security scan). The model learns the proof protocol, not the archaeology of each project’s build system. Estimates. Conservative: 2×–3×. Central: 2×–4×. Aggressive: 4×–6×. B. Behavior Cell Compression Behavior cells are the canonical unit of application logic: a named, typed, proof-carrying declaration of a behavioral intent
that the canonical runtime can instantiate, compose, and verify. Where infrastructure compression removes representational divergence, cell compression removes behavioral divergence— the phenomenon whereby identical application semantics (pagination, authentication, billing) are re-implemented from scratch in every codebase with local variation that encodes no new information. Domains 7–20 target this stratum. To keep the catalog compact, several high-win cells are folded into adjacent domains: tenant/org/user belongs to identity and policy, feature flags and experiments belong to configuration policy, cache/invalidation belongs to governed dependency and observability contracts, and UI resource state belongs to list/search cells. Patch intermediate representation, reasoning digests, supply-chain closure, and notebook-to-pipeline compression are treated separately in the meta-compression layer because their denominators are action, reasoning, governance, and workflow entropy rather than source tokens alone. 1) Resource Lifecycle Cells: Approach. Rebuild resource lifecycle work as typed cell declarations; generate routes, SQL, UI state, tests, and fixtures; remove handwritten create/read/update/delete surfaces and repeated controller edits. Create, read, update, delete, list, archive, restore, and status transition constitute the single most repeated pattern in product software. In raw corpora, each resource lifecycle surface is hand-written with per-project validation, authorization, serialization, and error handling. Canonical resource lifecycle cells reduce this to a typed resource declaration: the model emits cell parameters (fields, transitions, access policies) rather than writing route handlers, SQL queries, and test fixtures. Estimates. Conservative: 3×–8×. Central: 8×–20×. Aggressive: 20×–50×. 2) Authentication / Identity / Policy Cells: Approach. Rebuild identity, tenant, organization, session, and policy behavior as governed cells; generate guards, recovery flows, fixtures, and negative tests; remove bespoke auth plumbing. Login flows, session management, token issuance and validation, external identity-provider integration, role-based access control, attribute-based access control, permission checks, password reset, multi-factor authentication, and account recovery are repeated in nearly every production application— with dangerous variation. Security-critical code is the worst candidate for bespoke re-implementation and the best candidate for cell compression. A canonical identity cell encapsulates the full policy surface; the model configures it rather than writing cryptographic plumbing. Estimates. Conservative: 3×–10×. Central: 10×–30×. Aggressive: 30×–100×. 3) Form and Input Validation Cells: Approach. Rebuild validation as schema-owned constraints; generate frontend, API, database, and test projections; remove four-layer drift and hand-maintained form logic. Validation logic is duplicated across at least four layers in typical applications: frontend form validators, API request schemas, database constraints, and test assertions. Each layer uses a different validation DSL, and drift between layers is a
TABLE XXXVII I NFRASTRUCTURE C OMPRESSION D OMAINS (1–6): G AIN E STIMATES #
Domain
1 2 3 4 5 6
Language Role Collapse Framework Collapse Generated Truth Collapse Layout/Naming Collapse Dependency Collapse Build/Test Collapse
Conservative
Central
Aggressive
4×–6× 3×–5× 2×–3× 2×–3× 2×–3× 2×–3×
6×–12× 5×–10× 3×–6× 3×–5× 3×–5× 2×–4×
12×–20× 10×–15× 6×–10× 5×–8× 5×–8× 4×–6×
persistent source of bugs. Canonical validation cells generate all four projections from a single schema, eliminating crosslayer inconsistency and the representational cost of maintaining parallel validation grammars.
retry semantics and idempotency checks. Canonical job cells declare the reliability contract; the runtime enforces it.
Estimates. Conservative: 2×–5×. Central: 5×–15×. Aggressive: 15×–30×.
7) File Upload and Media Cells: Approach. Rebuild uploads as storage and policy declarations; generate signed URLs, scanners, transforms, fixtures, and audit events; remove ad hoc storage/security code. Signed URL generation, virus scanning, image processing pipelines, storage backend abstraction, content-type validation, and access control for uploaded media combine security, storage, and processing complexity that is repeated in nearly every user-facing application. Canonical media cells declare upload constraints, processing steps, and access policies; infrastructure details are resolved by the runtime.
4) Pagination / Search / Filter Cells: Approach. Rebuild list/search behavior as a resource-state cell; generate durable-data queries, API cursors, product-surface state, empty/error/loading views, and tests; remove custom list implementations. Paginated lists with sorting, filtering, full-text search, cursorbased or offset pagination, and search index synchronization represent massive UI/API/database repetition. The same behavioral contract—“return a filtered, sorted, paginated view of a resource collection”—is implemented independently at every layer with per-project conventions. A canonical pagination cell parameterizes the contract once; all layers are generated. Estimates. Conservative: 3×–8×. Central: 8×–20×. Aggressive: 20×–50×.
Estimates. Conservative: 2×–5×. Central: 5×–15×. Aggressive: 15×–30×.
Estimates. Conservative: 3×–8×. Central: 8×–20×. Aggressive: 20×–40×. 8) Billing / Subscription / Payment Cells: Approach. Rebuild payments as subscription and ledger state machines; generate webhook handlers, idempotency, reconciliation, fixtures, and alerts; remove brittle payment glue. Subscription lifecycle management, invoice preview and proration, payment method lifecycle, webhook handling for Stripe and other payment providers, dunning flows, and tax calculation represent a common SaaS surface with dangerous edge cases in currency handling, idempotency, and state reconciliation. Canonical billing cells parameterize the subscription model and payment provider; the full webhook surface, idempotency layer, and state machine are generated.
5) SQL Migration and Data Invariant Cells: Approach. Rebuild data change as declared invariants, lock budgets, and expand/contract plans; generate SQL and replay checks; remove opaque migration scripts and unsafe rollbacks. Expand/contract migrations, data backfills, rollback plans, data-shape invariants, and lock-budget management carry high failure cost and high repetition. In raw corpora, migration files are opaque SQL scripts with no machine-readable invariant or proof of safety. Canonical migration cells declare the schema transition, invariant constraints, and acceptable lock budgets; Estimates. Conservative: 3×–10×. Central: 10×–30×. the runtime generates the SQL and verifies safety properties Aggressive: 30×–100×. before execution. 9) Notification / Email / Template Cells: Approach. ReEstimates. Conservative: 2×–5×. Central: 5×–10×. Ag- build outbound messaging as channel and template policy; gressive: 10×–20×. generate provider adapters, preferences, retries, and observabil6) Background Job / Queue / Scheduler Cells: Approach. ity; remove local email/SMS/push wrappers. Multi-channel notification dispatch (email, SMS, push, inRebuild async work as idempotent job contracts; generate app), delivery tracking, template rendering with variable retries, dead-letter routing, schedules, metrics, and tests; remove substitution, user preference management, and retry logic for queue-specific reliability rewrites. Retry policies, idempotency guarantees, exponential backoff, transient delivery failures. Canonical notification cells declare dead-letter queue routing, scheduled task management, and con- channels, templates, and delivery policies; the runtime handles currency limits constitute the reliability layer of asynchronous provider integration and observability. processing. The same reliability logic is re-implemented in every codebase that uses background jobs, with subtle bugs in
Estimates. Conservative: 2×–5×. Central: 5×–15×. Aggressive: 15×–30×.
10) Webhook and Integration Cells: Approach. Rebuild integrations as signed event contracts; generate verification, replay, deduplication, dispatch, and fixtures; remove providerspecific webhook boilerplate. Inbound webhook verification, signature validation, idempotent event processing, replay handling, deduplication, and provider-specific payload parsing. Integration-heavy applications repeat this pattern per external provider, each with a different signing scheme and payload format. Canonical integration cells declare the provider contract and verification method; the idempotency and processing pipeline are standard.
14) Error Handling and Result Envelope Cells: Approach. Rebuild failure handling as typed result envelopes; generate error codes, retry classifications, UI messages, and tests; remove stringly exceptions and null-path ambiguity. Typed error envelopes, machine-readable error codes, retry classification (transient vs. permanent), user-facing vs. internal error separation, and structured error metadata. Raw codebases exhibit a chaotic mixture of exceptions, HTTP status codes, null returns, and string error messages. Canonical error cells impose a typed result algebra; the model learns to classify and route errors rather than invent ad-hoc handling per call site.
Estimates. Conservative: 2×–5×. Central: 5×–15×. Aggressive: 15×–30×.
Estimates. Conservative: 2×–4×. Central: 4×–8×. Aggressive: 8×–12×.
11) Audit Log and Event History Cells: Approach. Rebuild audit as append-only event policy; generate event writers, retention, queries, and tamper checks; remove scattered compliance logging. Append-only event streams, compliance audit trails, actor/action/resource/timestamp recording, and tamper-evident log integrity are required in compliance-heavy systems (healthcare, finance, government). In raw corpora, audit logging is either absent, ad-hoc, or inconsistently applied. Canonical audit cells declare the event schema and retention policy; the runtime guarantees append-only semantics and query access. Estimates. Conservative: 2×–5×. Central: 5×–10×. Aggressive: 10×–20×. 12) Observability / Logging / Tracing Cells: Approach. Rebuild telemetry as a cross-cutting contract; generate logs, spans, metrics, health checks, and correlation IDs; remove inconsistent instrumentation edits. Structured logging, distributed tracing with context propagation, metrics collection and export, request-ID propagation across service boundaries, and health-check endpoints. These cross-cutting concerns are woven inconsistently through raw codebases, often with conflicting log formats and missing trace context. Canonical observability cells declare the instrumentation contract; the runtime injects it uniformly. Estimates. Conservative: 2×–3×. Central: 3×–8×. Aggressive: 8×–15×.
C. Meta-Compression Meta-compression operates not on source tokens but on the action space and reasoning space of the agent itself. Where infrastructure and cell compression shrink the representation a model must learn to read and write, meta-compression shrinks the space of trajectories a model must explore during generation. The gains here are measured not in token-count ratios but in the reduction of invalid action paths, wasted inference steps, unproductive exploration, review burden, supply-chain ambiguity, and workflow reinvention. Domains 21–30 target this stratum. 1) Repair-Path Collapse: Approach. Rebuild patching as constrained edit policy; generate legal edit maps, proof routes, and no-edit zones; remove arbitrary file, command, migration, and dependency choices. In a canonical environment, the agent no longer selects among arbitrary files to edit, frameworks to invoke, proof commands to run, generated artifacts to modify, or migration strategies to attempt. The canonical grammar constrains every dimension of the action space simultaneously: which files are mutable, which cells accept parameters, which proof lanes must pass, which artifacts are generated. Invalid edit paths— the overwhelming majority of the raw action space—vanish entirely. This is the single largest compression domain by estimated magnitude, because the space of wrong actions in unconstrained codebases is combinatorially vast.
13) Configuration / Secrets / Environment Cells: Approach. Estimates. Conservative: 10×–100×. Central: 100×– Rebuild configuration, secrets, feature flags, and experiments 10,000×. Aggressive: 10,000×–100,000×. as typed policy; generate env validation, rotation checks, flag 2) Negative-Path Curriculum: Approach. Rebuild misgates, and rollout tests; remove .env folklore. takes as labeled curriculum; generate invalid edits, missing Secret rotation, environment-specific configuration, feature proofs, unsafe migrations, and generated-zone violations; flag management, configuration validation at startup, and secret- remove ambiguous definitions of wrongness. leakage prevention. In raw corpora, .env files, hardcoded Canonical form enables a training curriculum that includes secrets, and undocumented environment variables constitute a not only correct canonical code but also labeled negative persistent security and reliability risk. Canonical configuration examples: editing generated files, applying patches to the cells declare the configuration schema with types, defaults, and wrong architectural layer, writing unsafe migrations without secret classifications; the runtime validates at boot and enforces lock-budget declarations, duplicating generated data-transferrotation policies. object definitions, and skipping proof lanes. These negatives Estimates. Conservative: 2×–4×. Central: 4×–8×. Ag- are cheap to produce in a canonical environment (any violation gressive: 8×–15×. of the grammar is a negative) and expensive to produce in raw
TABLE XXXVIII B EHAVIOR C ELL C OMPRESSION D OMAINS (7–20): G AIN E STIMATES #
Domain
7 8 9 10 11 12 13 14 15 16 17 18 19 20
Resource Lifecycle Authentication / Identity / Policy Form / Input Validation Pagination / Search / Filter SQL Migration / Data Invariants Background Job / Queue / Scheduler File Upload / Media Billing / Subscription / Payment Notification / Email / Template Webhook / Integration Audit Log / Event History Observability / Logging / Tracing Config / Secrets / Environment Error Handling / Result Envelope
corpora (where “wrong” is ambiguous). The model learns the boundary of correct behavior, not just its interior. Estimates. Conservative: 2×–5×. Central: 5×–15×. Aggressive: 15×–50×.
Conservative
Central
Aggressive
3×–8× 3×–10× 2×–5× 3×–8× 2×–5× 2×–5× 3×–8× 3×–10× 2×–5× 2×–5× 2×–5× 2×–3× 2×–4× 2×–4×
8×–20× 10×–30× 5×–15× 8×–20× 5×–10× 5×–15× 8×–20× 10×–30× 5×–15× 5×–15× 5×–10× 3×–8× 4×–8× 4×–8×
20×–50× 30×–100× 15×–30× 20×–50× 10×–20× 15×–30× 20×–40× 30×–100× 15×–30× 15×–30× 10×–20× 8×–15× 8×–15× 8×–12×
Many changes recur across repositories: add field, migrate nullable to required, add permission edge, rotate secret, add idempotency key, upgrade dependency safely, split table, add audit trail, deprecate API field, or convert sync work to an idempotent job. Semantic patch cells make those changes firstclass. The agent binds parameters and discharges obligations instead of inventing edits file by file.
3) Proof-Carrying Change Objects: Approach. Rebuild diffs as structured change objects and patch IR; generate intent, affected cells, proofs, receipts, and rollback evidence; remove Estimates. Conservative: 10×–50×. Central: 50×– bare diff-only completion paths. 1,000×. Aggressive: 1,000×–10,000×. In canonical form, a patch is not a raw diff. It is a structured 6) Constrained Edit Grammar: Approach. Rebuild editchange object comprising: the declared intent, the affected cells, the generated delta, the migration proof, the security ing as legal typed operations over behavior IR, cells, schemas, proof, and the test evidence. The model learns to produce the and policies; reject illegal files, operations, and proof omissions proof route—the full trajectory from intent to verified change— before generation. rather than a bare textual patch whose correctness must be Repair-path collapse removes many wrong trajectories after inferred post-hoc. This transforms code generation from a the agent chooses a path. A constrained edit grammar removes single-shot prediction problem into a structured reasoning chain them before choice. Generated files are immutable; migrations with verifiable intermediate steps. must declare lock budgets and rollback; policy changes must Estimates. Conservative: 2×–5×. Central: 5×–10×. Ag- emit negative tests; dependency changes must route through governed upgrade lanes. The agent sees a small legal action gressive: 10×–20×. set instead of a whole filesystem. 4) Reasoning Digest and Plan-Cache Compression: Approach. Rebuild repeated agent cognition as versioned rea- Estimates. Conservative: 100×–1,000×. Central: 1,000×– soning digests; generate task routes, known failure modes, 100,000×. Aggressive: 100,000×+. proof obligations, and repair playbooks; remove repeated repo 7) Proof-Lane Receipt Reuse: Approach. Rebuild verifiarchaeology. cation as reusable receipts; generate migration replay, rollback Current agents spend inference tokens rediscovering archi- evidence, policy checks, and security receipts; remove repeated tecture, ownership, generated zones, tests, migration rituals, proof interpretation and review reconstruction. and safe repair strategies. Canonical repositories should store a Proof lanes define commands. Receipt reuse defines the compact reasoning map: task class to affected cells, issue type evidence object that survives the command: what was checked, to legal edit set, failure signature to repair strategy, policy delta which obligations changed, which logs matter, what failed to required test matrix, and dependency update to compatibility before repair, and what rollback means. Reviewers and agents checks. This turns repeated thought into substrate memory. inspect the receipt instead of reconstructing proof meaning Estimates. Conservative: 3×–8×. Central: 8×–50×. Ag- from raw logs. gressive: 50×–200×. 5) Semantic Patch Cells: Approach. Rebuild recurring changes as typed transformations; generate blast radius, preconditions, proofs, tests, and rollbacks; remove raw-diff invention for common edits.
Estimates. Conservative: 3×–8×. Central: 8×–30×. Aggressive: 30×–100×. 8) Supply-Chain and Dependency Closure: Approach. Rebuild package selection as governed capability resolution; generate adapters, provenance, license, bill-of-materials, vul-
nerability, upgrade, and rollback evidence; remove arbitrary package choice. Dependencies are hidden code expansion. A single import can carry transitive code, build scripts, licenses, vulnerabilities, and runtime behavior. Canonical form should resolve verified capabilities rather than package names: password hashing, email send, payment webhook, rate limit, or object storage. Each capability has approved providers, adapters, provenance, proof lanes, and upgrade routes. Estimates. Conservative: 2×–5×. Central: 5×–20×. Aggressive: 20×–100×. 9) Notebook-to-Pipeline Compression: Approach. Rebuild exploratory notebooks as typed data pipelines; generate environment locks, provenance, tests, parameter cells, and scheduled jobs; remove copy-pasted exploratory state. Notebook-heavy software often stores executable history rather than reproducible behavior. Canonical conversion separates exploration from pipeline truth: cells become typed data steps, parameters become schema, plots and reports become generated projections, and environments become locked artifacts. The agent learns the pipeline graph rather than the accidental order of notebook execution. Estimates. Conservative: 3×–10×. Central: 10×–50×. Aggressive: 50×–200×. 10) Runtime and Negative Memory Compression: Approach. Rebuild incidents, reverts, failed patches, hidden compatibility, and production traces as living invariants and forbidden plans; remove repeated rediscovery of bad worlds. The raw corpus mostly preserves accepted code, but software expertise also lives in what failed: unsafe migrations, auth bypasses, flaky tests, reverted patches, support incidents, and production-only edge cases. Canonical memory turns those failures into anti-examples, negative tests, proof obligations, and reasoning warnings. Every failure becomes a compression artifact. Estimates. Conservative: 2×–5×. Central: 5×–25×. Aggressive: 25×–100×. 11) Composite Gain Estimates: The per-domain estimates above cannot be multiplied naı̈vely: language collapse and framework collapse share token mass; resource lifecycle cells and validation cells overlap on the same source files; repair-path collapse subsumes portions of every other domain. Table XL presents composite estimates across five measurement axes, derived from conservative overlap-discounting of the individual domains. These composites represent the full-stack prior: the expected gain when all canonical transformations are applied simultaneously to a representative product codebase. The distinction between literal source-token reduction and effective representation-space reduction is critical. Source tokens measure the physical size of the corpus; representation space measures the number of distinct behavioral encodings the model must learn to produce. A 5× source-token reduction that also collapses seven languages into one yields a far larger effective reduction, because the model no longer maintains
seven parallel decoders for the same semantic space. Agent action-space reduction is larger still, because it compounds source compression with the elimination of invalid edit trajectories (Domain 21). Reasoning-token reduction is a separate denominator: it measures how often the agent must rediscover architecture, tests, proof routes, repair strategy, dependency law, and review expectations that can instead be compiled into reasoning digests, proof receipts, and semantic patch cells. 12) Theoretical Limit: Minimum Functional Description Length: The compression frontier converges toward a theoretical limit we term the Minimum Functional Description Length: the shortest canonical specification, plus proofs, plus renderer needed to produce a fully working system. Minimum Functional Description Length is not a code-golf metric; it is the information-theoretic minimum over the space of behavioral specifications that can be mechanically verified and rendered into executable artifacts. For routine product software—the SaaS platforms, internal tools, and resource-heavy services that constitute the majority of commercial codebases—we estimate that 70–90% of behavior may eventually be expressible as cell composition, generated surfaces, and policy configuration. The remaining 10–30% is true domain-specific logic: novel algorithms, unique business rules, and irreducible problemspecific reasoning that no canonical grammar can absorb without becoming domain-specific itself. The theoretical limit is not shorter code. The theoretical limit is no bespoke code unless the behavior is novel. R EFERENCES [1] BigCode Project, “The Stack v2,” https://huggingface.co/datasets/bi gcode/the-stack-v2, 2024, dataset card: 67.5TB full, 32.1TB deduped, roughly 900B train-full tokens, over 3B files, 658 languages; accessed: 2026-05-14. [2] A. Lozhkov, R. Li, L. B. Allal, F. Cassano, J. Lamy-Poirier, N. Tazi, A. Tang, D. Pykhtar, J. Liu, Y. Wei, T. Liu, M. Tian, D. Kocetkov, A. Zucker, Y. Belkada, Z. Wang, Q. Liu, D. Abulkhanov, I. Paul, Z. Li, W.-D. Li, M. Risdal, J. Li, J. Zhu, T. Y. Zhuo, E. Zheltonozhskii, N. O. O. Dade, W. Yu, L. Krauss, N. Jain, Y. Su, X. He, M. Dey, E. Abati, Y. Chai, N. Muennighoff, X. Tang, M. Oblokulov, C. Akiki, M. Marone, C. Mou, M. Mishra, A. Gu, B. Hui, T. Dao, A. Zebaze, O. Dehaene, N. Patry, C. Xu, J. McAuley, H. Hu, T. Scholak, S. Paquet, J. Robinson, C. J. Anderson, N. Chapados, M. Patwary, N. Tajbakhsh, Y. Jernite, C. M. Ferrandis, L. Zhang, S. Hughes, T. Wolf, A. Guha, L. von Werra, and H. de Vries, “StarCoder 2 and The Stack v2: The next generation,” arXiv preprint arXiv:2402.19173, 2024. [Online]. Available: https://arxiv.org/abs/2402.19173 [3] C. V. Lopes, P. Maj, P. Martins, V. Saini, D. Yang, J. Zitny, H. Sajnani, and J. Vitek, “DéjàVu: A map of code duplicates on GitHub,” in Proceedings of the ACM on Programming Languages, vol. 1, no. OOPSLA, 2017, pp. 84:1–84:28, analyzed 4.5M non-fork GitHub projects: only 85M of 428M files were unique (about 70% clones). [4] M. H. Nguyen, B. Adams, and A. E. Hassan, “Jupyter notebooks on GitHub: Characteristics and code clones,” arXiv preprint arXiv:2007.10146, 2020, found more than 70% of all code snippets were exact copies and about half of notebooks had no unique snippet. [Online]. Available: https://arxiv.org/abs/2007.10146 [5] OpenAI and Princeton NLP, “SWE-bench Verified: Human-validated software engineering tasks,” https://www.swebench.com/verified.html, 2024, human-validated subset of 500 SWE-bench instances. [6] B. Yu, Y. Zhu, P. He, and D. Kang, “UTBoost: Rigorous evaluation of coding agents on SWE-Bench,” arXiv preprint arXiv:2506.09289, 2025, aCL 2025. [Online]. Available: https://arxiv.org/abs/2506.09289
TABLE XXXIX M ETA -C OMPRESSION D OMAINS (21–30): G AIN E STIMATES #
Domain
Conservative
Central
Aggressive
21 22 23 24 25 26 27 28 29 30
Repair-Path Collapse Negative-Path Curriculum Proof-Carrying Change Objects Reasoning Digests / Plan Cache Semantic Patch Cells Constrained Edit Grammar Proof-Lane Receipt Reuse Supply-Chain / Dependency Closure Notebook-to-Pipeline Runtime / Negative Memory
10×–100× 2×–5× 2×–5× 3×–8× 10×–50× 100×–1,000× 3×–8× 2×–5× 3×–10× 2×–5×
100×–10,000× 5×–15× 5×–10× 8×–50× 50×–1,000× 1,000×–100,000× 8×–30× 5×–20× 10×–50× 5×–25×
10,000×–100,000× 15×–50× 10×–20× 50×–200× 1,000×–10,000× 100,000×+ 30×–100× 20×–100× 50×–200× 25×–100×
TABLE XL C OMPOSITE G AIN E STIMATES ACROSS A LL 30 C OMPRESSION D OMAINS (D ENOMINATOR -S PECIFIC ; N OT M ULTIPLICATIVE ) Measurement Axis
Conservative
Central
Aggressive
Literal source-token reduction Effective representation-space reduction Agent action-space reduction Training-token efficiency Reasoning/tool/retry-token reduction Cost per verified correct change
2×–8× 20×–40× 100×–1,000× 10×–30× 3×–10× 3×–10×
3×–10× 40×–150× 1,000×–100,000× 30×–150× 10×–100× 10×–50×
10×–30× 100×–300× 100,000×+ 150×–1,000× 100×–10,000× 50×–1,000×
[7] J. Becker, A. Rush, B. Barnes, and N. Rush, “Measuring the impact of early-2025 ai on experienced open-source developer productivity,” arXiv preprint arXiv:2507.09089, 2025, randomized controlled trial reporting 19% longer completion time with AI tooling in the studied setting. [Online]. Available: https://arxiv.org/abs/2507.09089 [8] J. Becker, N. Rush, T. Cunningham, D. Rein, and K. Mahamud, “We are changing our developer productivity experiment design,” https://metr.org /blog/2026-02-24-uplift-update/, 2 2026, discusses selection effects and updated evidence after the early-2025 developer productivity experiment; accessed: 2026-05-28. [9] E. Abrokwah and T. A. Ghaleb, “An empirical study of complexity, heterogeneity, and compliance of GitHub Actions workflows,” arXiv preprint arXiv:2507.18062, 2025, registered report accepted at ICSME 2025. [Online]. Available: https://arxiv.org/abs/2507.18062 [10] R.-M. Karampatsis, M. Linares-Vásquez, O. Chaparro, and G. Bavota, “An empirical study of the evolution of GitHub Actions workflows,” https://arxiv.org/abs/2602.14572, 2026, reports analysis of 49K+ repositories, 267K+ workflow change histories, and 3.4M+ workflow file versions. [11] Y. Kubo, F. Kanei, M. Akiyama, T. Wakai, and T. Mori, “Action required: A mixed-methods study of security practices in GitHub Actions,” in Proceedings of the Network and Distributed System Security Symposium (NDSS), 2026, analyzed 338,812 public repositories and found low implementation rates across five GitHub Actions security practices. [Online]. Available: https://www.ndss-symposium.org/ndss-p aper/action-required-a-mixed-methods-study-of-security-practices-in-g ithub-actions/ [12] GitGuardian, “State of secrets sprawl 2026,” https://www.gitguardian. com/state-of-secrets-sprawl-report-2026, 2026, reports 28.65 million new hardcoded secrets in public GitHub commits in 2025 and 34% year-over-year growth; accessed: 2026-05-28. [13] GitHub, “Octoverse 2025: A new developer joins github every second as ai leads typescript to number one,” https://github.blog/news-insights/ octoverse/octoverse-a-new-developer-joins-github-every-second-as-ai-l eads-typescript-to-1/, 2025, reports more than 180 million developers, 395 million public/open-source repositories, about 230 new repositories per minute, and TypeScript overtaking Python and JavaScript in August 2025; accessed: 2026-05-28. [14] Moonshot AI, “Kimi K2: Open agentic intelligence,” arXiv preprint arXiv:2507.20534, 2025. [Online]. Available: https://arxiv.org/abs/2507 .20534 [15] ——, “Kimi K2 model documentation,” https://platform.kimi.ai/docs/mo
dels, 2025, documents 1T total parameters, 32B activated parameters, and long-context mixture-of-experts serving profile; accessed: 2026-05-14. [16] D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu, Y. K. Li, F. Luo, Y. Xiong, and W. Liang, “DeepSeek-Coder: When the large language model meets programming,” arXiv preprint arXiv:2401.14196, 2024. [Online]. Available: https://arxiv.org/abs/2401.14196 [17] Qwen Team, “Qwen2.5-Coder technical report,” arXiv preprint arXiv:2409.12186, 2024. [Online]. Available: https://arxiv.org/abs/2409.1 2186 [18] A. Hindle, E. T. Barr, Z. Su, M. Gabel, and P. Devanbu, “On the naturalness of software,” in Proceedings of the 34th International Conference on Software Engineering (ICSE), 2012, pp. 837–847, foundational work showing that software is highly repetitive and predictable, supporting the use of statistical language models for code. [19] C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan, “SWE-bench: Can language models resolve real-world github issues?” arXiv preprint arXiv:2310.06770, 2023. [Online]. Available: https://arxiv.org/abs/2310.06770 [20] I. Badertdinov, A. Golubev, M. Nekrashevich, A. Shevtsov, S. Karasik, A. Andriushchenko, M. Trofimova, D. Litvintseva, and B. Yangel, “SWE-rebench: An automated pipeline for task collection and decontaminated evaluation of software engineering agents,” arXiv preprint arXiv:2505.20411, 2025. [Online]. Available: https://arxiv.org/ abs/2505.20411 [21] OpenAPI Initiative, “OpenAPI Specification version 3.1.0,” https://spec .openapis.org/oas/v3.1.0.html, 2021, accessed: 2026-05-28. [22] Google, “Protocol buffers overview,” https://protobuf.dev/overview/, 2026, accessed: 2026-05-28. [23] C. Lattner, M. Amini, U. Bondhugula, A. Cohen, A. Davis, J. Pienaar, R. Riddle, T. Shpeisman, N. Vasilache, and O. Zinenko, “MLIR: Scaling compiler infrastructure for domain specific computation,” in Proceedings of the IEEE/ACM International Symposium on Code Generation and Optimization (CGO), 2021, pp. 2–14. [24] Linux Kernel Documentation, “Coccinelle,” https://docs.kernel.org/dev-t ools/coccinelle.html, 2026, accessed: 2026-05-28. [25] GitHub, “CodeQL documentation,” https://codeql.github.com/docs/, 2026, accessed: 2026-05-14. [26] M. Willsey, C. Nandi, Y. R. Wang, O. Flatt, Z. Tatlock, and P. Panchekha, “egg: Fast and extensible equality saturation,” Proceedings of the ACM on Programming Languages, vol. 5, no. POPL, pp. 1–29, 2021. [Online]. Available: https://arxiv.org/abs/2004.03082
[27] K. Ellis, C. Wong, M. Nye, M. Sable-Meyer, L. Cary, L. Morales, L. Hewitt, A. Solar-Lezama, and J. B. Tenenbaum, “DreamCoder: Growing generalizable, interpretable knowledge with wake-sleep bayesian program learning,” arXiv preprint arXiv:2006.08381, 2020. [Online]. Available: https://arxiv.org/abs/2006.08381 [28] J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre, “Training compute-optimal large language models,” arXiv preprint arXiv:2203.15556, 2022. [Online]. Available: https://arxiv.org/abs/2203.15556 [29] C. Lattner and V. Adve, “LLVM: A compilation framework for lifelong program analysis and transformation,” in Proceedings of the International Symposium on Code Generation and Optimization (CGO), 2004, pp. 75– 86. [30] P. Clements and L. Northrop, Software Product Lines: Practices and Patterns. Addison-Wesley, 2002. [31] G. C. Necula, “Proof-carrying code,” in Proceedings of the 24th ACM SIGPLAN-SIGACT Symposium on Principles of Programming Languages, 1997, pp. 106–119. [Online]. Available: https://people.eecs. berkeley.edu/∼necula/papers.html [32] Synopsys, “New synopsys report finds 74% of codebases contained highrisk open source vulnerabilities,” https://news.synopsys.com/2024-02-2 7-New-Synopsys-Report-Finds-74-of-Codebases-Contained-High-Ris k-Open-Source-Vulnerabilities%2C-Surging-54-Since-Last-Year, 2024, reports 84% of assessed codebases contained open-source vulnerabilities, 74% contained high-risk vulnerabilities, and 91% used components ten or more versions behind; accessed: 2026-05-28. [33] K. Gallaba, C. Macho, M. Pinzger, and S. McIntosh, “Noise and heterogeneity in historical build data: An empirical study of Travis CI,” in Proceedings of the International Conference on Automated Software Engineering (ASE), 2018, pp. 87–97, analyzed 3.7 million build jobs across 1,276 open-source projects. [34] K. Gallaba and S. McIntosh, “Use and misuse of continuous integration features: An empirical study of projects that (mis)use Travis CI,” IEEE Transactions on Software Engineering, vol. 46, no. 1, pp. 33–50, 2020, study of CI feature use and misuse across 9,312 open-source systems. [35] OpenSSF, “OpenSSF Scorecard,” https://github.com/ossf/scorecard, 2026, accessed: 2026-05-14. [36] L. Simon, A. Birgisson, A. Ghodsi, T. Lauinger, B. Callaway, Z. Durumeric, M. Hicks, D. Stefan, and S. Torres-Arias, “On the path toward ecosystem-wide automated security metrics for open source software,” in Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security, 2022, pp. 1607–1620. [37] SPDX Project, “SPDX overview,” https://spdx.dev/about/overview/, 2026, open standard for communicating software bill-of-material information, provenance, license, and security metadata; accessed: 2026-05-28. [38] OWASP Foundation, “CycloneDX: Bill of materials standard,” https: //cyclonedx.org/, 2026, eCMA-424 bill-of-materials standard; accessed: 2026-05-28. [39] OpenTelemetry, “Semantic conventions,” https://opentelemetry.io/docs/ concepts/semantic-conventions/, 2026, accessed: 2026-05-28. [40] Open Policy Agent, “Open Policy Agent documentation,” https://ww w.openpolicyagent.org/docs, 2026, general-purpose policy engine and Rego policy language; accessed: 2026-05-28. [41] OpenSSF, “SLSA: Supply-chain levels for software artifacts,” https: //slsa.dev/, 2024, accessed: 2026-05-14. [42] in-toto Project, “in-toto: A framework to secure the integrity of software supply chains,” https://in-toto.io/, 2026, accessed: 2026-05-14. [43] MIT OpenCourseWare, “Rice’s theorem,” https://ocw.mit.edu/course s/6-045j-automata-computability-and-complexity-spring-2011/reso urces/mit6 045js11 lec09/, 2011, lecture notes for 6.045J Automata, Computability, and Complexity; accessed: 2026-05-28. [44] K. Claessen and J. Hughes, “QuickCheck: A lightweight tool for random testing of Haskell programs,” in Proceedings of the ACM SIGPLAN International Conference on Functional Programming (ICFP), 2000, pp. 268–279. [45] X. Yang, Y. Chen, E. Eide, and J. Regehr, “Finding and understanding bugs in C compilers,” in Proceedings of the ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI), 2011, pp. 283–294.
[46] Y. Jia and M. Harman, “An analysis and survey of the development of mutation testing,” IEEE Transactions on Software Engineering, vol. 37, no. 5, pp. 649–678, 2011. [47] C. Cadar, D. Dunbar, and D. Engler, “KLEE: Unassisted and automatic generation of high-coverage tests for complex systems programs,” in Proceedings of the 8th USENIX Symposium on Operating Systems Design and Implementation (OSDI), 2008, pp. 209–224. [Online]. Available: https://www.usenix.org/event/osdi08/tech/full papers/cadar/cadar.pdf [48] X. Leroy, S. Blazy, D. Kästner, B. Schommer, M. Pister, and C. Ferdinand, “The CompCert C verified compiler: Documentation and user’s manual,” https://compcert.org/man/manual.pdf, 2024, accessed: 2026-05-28. [49] seL4 Foundation, “seL4 verification proofs,” https://sel4.systems/Verific ation/proofs.html, 2026, accessed: 2026-05-28. [50] Y. Wei, F. Cassano, J. Liu, Y. Ding, N. Jain, H. de Vries, L. von Werra, A. Guha, and L. Zhang, “Arctic-SnowCoder: Demystifying high-quality data in code pretraining,” arXiv preprint arXiv:2409.02326, 2024, 1.3B model using 555B tokens in staged data refinement beat or matched models trained on much larger token budgets. [Online]. Available: https://arxiv.org/abs/2409.02326 [51] S. Gunasekar, Y. Zhang, J. Aneja, C. C. T. Mendes, A. D. Giorno, S. Gopi, M. Javaheripi, P. Kauffmann, G. de Rosa, O. Saarikivi, A. Salim, S. Shah, H. S. Behl, X. Wang, S. Bubeck, R. Eldan, A. T. Kalai, Y. T. Lee, and Y. Li, “Textbooks are all you need,” arXiv preprint arXiv:2306.11644, 2023. [Online]. Available: https://arxiv.org/abs/2306.11644 [52] T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “QLoRA: Efficient finetuning of quantized LLMs,” arXiv preprint arXiv:2305.14314, 2023. [Online]. Available: https://arxiv.org/abs/2305.1 4314 [53] Y. Dong, C. F. Ruan, Y. Cai, R. Lai, Z. Xu, Y. Zhao, and T. Chen, “XGrammar: Flexible and efficient structured generation engine for large language models,” arXiv preprint arXiv:2411.15100, 2024, mLSys 2025. [Online]. Available: https://arxiv.org/abs/2411.15100 [54] OpenAI, “Reasoning models,” https://platform.openai.com/docs/guides/ reasoning, 2026, openAI API documentation; accessed: 2026-05-28. [55] ——, “Reasoning best practices,” https://platform.openai.com/docs/guide s/reasoning-best-practices, 2026, openAI API documentation; accessed: 2026-05-28. [56] J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” arXiv preprint arXiv:2201.11903, 2022. [Online]. Available: https://arxiv.org/abs/2201.11903 [57] X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” arXiv preprint arXiv:2203.11171, 2022. [Online]. Available: https://arxiv.org/abs/2203.11171 [58] S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, “ReAct: Synergizing reasoning and acting in language models,” arXiv preprint arXiv:2210.03629, 2022. [Online]. Available: https://arxiv.org/abs/2210.03629 [59] J. Lin, X. Zeng, J. Zhu, S. Wang, J. Shun, J. Wu, and D. Zhou, “Plan and budget: Effective and efficient test-time scaling on reasoning large language models,” arXiv preprint arXiv:2505.16122, 2025, revised 2026; accepted to ICLR 2026. [Online]. Available: https://arxiv.org/abs/2505.16122 [60] OpenAI, “Structured model outputs,” https://platform.openai.com/docs/g uides/structured-outputs, 2026, openAI API documentation; accessed: 2026-05-28. [61] J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, “Scaling laws for neural language models,” arXiv preprint arXiv:2001.08361, 2020. [Online]. Available: https://arxiv.org/abs/2001.08361 [62] OpenAI, “Why SWE-bench Verified no longer measures frontier coding capabilities,” https://openai.com/index/why-we-no-longer-evaluate-swe -bench-verified/, 2026, accessed: 2026-05-28. [63] J. Yang, K. Lieret, C. E. Jimenez, A. Wettig, K. Khandpur, Y. Zhang, B. Hui, O. Press, L. Schmidt, and D. Yang, “SWE-smith: Scaling data for software engineering agents,” arXiv preprint arXiv:2504.21798, 2025. [Online]. Available: https://arxiv.org/abs/2504.21798 [64] BigCode Collaboration, S. Hughes, H. de Vries, J. Robinson, C. M. Ferrandis, L. B. Allal, L. von Werra, J. Ding, S. Paquet, and Y. Jernite, “The BigCode project governance card,” arXiv preprint arXiv:2312.03872, 2023. [Online]. Available: https://arxiv.org/abs/2312.03872
[65] Software Heritage, “Software heritage activity report: 2025,” https://ww w.softwareheritage.org/2026/01/16/software-heritage-activity-report-2 025/, 2026, reports 2025 archive activity; accessed: 2026-05-28. [66] A. Avizienis, “The N-Version approach to fault-tolerant software,” IEEE Transactions on Software Engineering, vol. SE-11, no. 12, pp. 1491–1501, 1985. [67] D. J. Mankowitz, A. Michi, A. Zhernov, M. Gelmi, M. Selvi, C. Paduraru, E. Leurent, S. Iqbal, J.-B. Lespiau, A. Ahern, T. Köppe, K. Millikin, S. Gaffney, S. Elster, J. Broshear, C. Gamble, K. Milan, R. Tung, M. Jiang, H. Wang, E. Özcan, D. Silver, D. Hassabis, P. Kohli, M. Riedmiller, O. Vinyals, and D. H. Silver, “Faster sorting algorithms discovered using deep reinforcement learning,” Nature, vol. 618, pp. 257–263, 2023. [Online]. Available: https://www.nature.com/articles/s41586-023-06004-9 [68] B. Romera-Paredes, M. Barekatain, A. Novikov, M. Balog, M. P. Kumar, E. Dupont, F. J. R. Ruiz, J. S. Ellenberg, P. Wang, O. Fawzi, P. Kohli, and A. Fawzi, “Mathematical discoveries from program search with large language models,” Nature, vol. 625, pp. 468–475, 2024. [Online]. Available: https://www.nature.com/articles/s41586-023-06924-6 [69] A. M. Turing, “On computable numbers, with an application to the Entscheidungsproblem,” Proceedings of the London Mathematical Society, vol. s2-42, no. 1, pp. 230–265, 1936. [70] H. G. Rice, “Classes of recursively enumerable sets and their decision problems,” Transactions of the American Mathematical Society, vol. 74, no. 2, pp. 358–366, 1953. [71] D. H. Wolpert and W. G. Macready, “No free lunch theorems for optimization,” IEEE Transactions on Evolutionary Computation, vol. 1, no. 1, pp. 67–82, 1997. [72] M. Li and P. M. B. Vitányi, An Introduction to Kolmogorov Complexity and Its Applications, 3rd ed. Springer, 2008. [73] G. M. Amdahl, “Validity of the single processor approach to achieving large scale computing capabilities,” in Proceedings of the AFIPS Spring Joint Computer Conference, 1967, pp. 483–485. [74] P. D. Grünwald, The Minimum Description Length Principle. MIT Press, 2007.