ConceptioArchivearXiv CS
arXiv CSopen access

AIRA: AI-Induced Risk Audit: A Structured Inspection Framework for AI-Generated Code

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
softwarearchitecturesoftwareengineeringtesting
software engineering, software architecture, testing

AIRA — AI-Induced Risk Audit A Structured Inspection Framework for AI-Generated Code William M. Parris BDB Labs | Jurisprudential AI Governance Initiative [email protected] aira.bageltech.net

arXiv:2604.17587v1 [cs.SE] 19 Apr 2026

2026 Abstract Practitioners have reported a directional pattern in AI-assisted code generation: AI-generated code tends to fail quietly, preserving the surface appearance of function while degrading or concealing guarantees. This paper introduces the Reward-Shaped Failure Hypothesis — the proposal that this pattern may reflect an artifact of how AI systems are optimized through human feedback, rather than a random distribution of bugs. Systems trained to maximize positive evaluation signals may be shaped toward suppressing visible failure, because a crashing program is often penalized more heavily than a program that returns a wrong or degraded result. We define failure truthfulness as the property that a system’s observable outputs accurately represent its internal success or failure state. We then present aira (AI-Induced Risk Audit), a 15-check inspection framework designed to detect the specific class of failure this optimization pressure may produce. aira is implemented as a deterministic static analysis engine with optional LLM augmentation, available as both a CLI tool and a web scanner. We report empirical results from three studies: an anonymized six-system enterprise AIassisted environment audit that operationally characterizes where the hypothesis emerged; a rebuilt balanced 600-file public-corpus pilot; and a stricter matched-control replication comparing 955 AI-attributed files against 955 human-control files. In the final replication, AI-attributed files show 0.435 high-severity findings per file versus 0.242 in human controls, a 1.80× excess, with the same direction observed in JavaScript, Python, and TypeScript. A secondary comparison using a cloud LLM evaluator produced findings at a 44:1 ratio below the deterministic scanner. aira is designed for governance, compliance, and safety-critical systems where fail-closed behavior is a definitional requirement.

1

Introduction

The question traditionally asked of software is whether it works. Testing frameworks, code review practices, and quality metrics are organized around that question. They catch errors of commission — code that produces the wrong output, throws an unexpected exception, or fails to compile. What they are not calibrated to catch is a different class of failure: code that continues executing, returns a value, and reports success, while silently violating the guarantees it was supposed to maintain. This paper argues that AI-assisted development may increase the prevalence and consistency of exactly this class of failure, or at minimum is associated with higher rates of it under matched comparison. The argument is structural, not anecdotal. AI coding systems are optimized in part through feedback signals from human evaluators. A program that crashes produces a strongly negative signal. A program that returns a result — even a degraded or incorrect one — produces a far less penalized signal. Over training, this asymmetry may shape code generation toward the appearance of correctness rather than correctness under all conditions. The result is not a random distribution of bugs. It is a directional skew: AI-generated code disproportionately fails in ways that preserve

1

AIRA — AI-Induced Risk Audit

BDB Labs, 2026

surface function while concealing broken guarantees. Standard review practices, designed for human error patterns, are not calibrated to detect it. We introduce aira (AI-Induced Risk Audit) as a targeted response to this problem. aira is not a general-purpose code quality tool. It is a structured inspection framework organized around a single question: Not “does this code work?” but “does this code tell the truth about whether it is working?” The remainder of this paper is organized as follows. Section 2 develops the Reward-Shaped Failure Hypothesis and formally defines failure truthfulness. Section 3 positions aira relative to existing work. Section 4 presents the framework design, architecture, and 15 checks. Section 5 reports empirical results from three studies. Section 6 describes the output specification. Section 7 addresses limitations. Section 8 concludes with future directions. 1.1

Contributions

This work makes the following contributions: • Introduces the Reward-Shaped Failure Hypothesis — a proposed structural account of why AI-assisted codebases exhibit systematic fail-soft behavior • Defines failure truthfulness as a formal system property distinct from correctness • Presents aira, a 15-check inspection framework targeting the specific failure class this hypothesis predicts • Reports empirical findings from a deterministic aira audit of a six-system enterprise AI-assisted development environment • Reports a rebuilt balanced 600-file corpus pilot (300 AI-attributed vs. 300 human-authored) showing a 1.32× overall and 3.18× JavaScript high-severity differential • Reports a stricter matched-control replication of 955 AI-attributed files versus 955 humancontrol files, showing a 1.80× high-severity differential with the same direction across JavaScript, Python, and TypeScript • Demonstrates that an LLM evaluator applied to the same surfaces exhibits the same suppression pattern at a 44:1 ratio below deterministic detection • Releases aira as an open-source tool with CLI, web, and research collection components at https://github.com/BDB-Labs/aira-scanner

2 2.1

Theoretical Foundation The Pattern

Engineers who work extensively with AI coding tools begin to notice a pattern. The code compiles. The tests pass. The function returns a value. And yet, under failure conditions, the system does not stop — it continues, quietly, in a degraded state that nothing in the output makes visible. Errors are absorbed. Exceptions are swallowed. Guarantees written into the design simply do not propagate when they are violated.

2

AIRA — AI-Induced Risk Audit

BDB Labs, 2026

This is a consistent, directional skew: AI-generated code fails quietly, in ways that preserve the surface appearance of function. Standard code review practices are not calibrated to detect this class of failure, because human developers do not produce it at the same rate or in the same form. In a coding context, this pattern most commonly manifests as: • Silent exception handling — because a swallowed error looks like continued success • Graceful degradation without disclosure — because a degraded response feels better than a hard stop • Optimistic return values — because returning None or raising feels like giving up • Broad try/catch blankets — because a crash is the worst possible visible outcome • Happy-path test coverage — because passing tests generate positive feedback 2.2

The Reward-Shaped Failure Hypothesis

aira offers a hypothesis about the structural cause. AI systems are optimized, in part, through feedback signals that reward outputs humans evaluate positively. A crashing program is a strongly negative signal. A program that returns a result — even a wrong one, even a degraded one — is far less penalized. Over training, this pressure shapes how these systems generate code: not toward correctness under all conditions, but toward the appearance of correctness under the conditions most likely to be evaluated. In code-generation settings, those evaluation conditions often reward scripts that run to completion in a simple sandbox or satisfy a shallow success criterion; swallowing an exception or returning a fallback can therefore score better with human raters than surfacing an explicit failure, even when the underlying guarantee has been broken. The failure patterns documented in this framework are consistent with an optimization-shaped pressure in the feedback environment, rather than a defect in any particular implementation. Because the failure distribution is non-random, it is auditable. The same optimization pressure that appears to produce silent exception handling in one codebase may produce similar patterns in another. aira is organized around that predictability. This hypothesis connects two conversations that have largely proceeded in parallel. Software reliability researchers have documented increasing rates of silent failure in AI-assisted codebases. AI safety researchers have identified that RLHF inherently trains models toward fail-open behavior [Casper et al., 2023]. aira bridges these observations by proposing that the alignment process is a plausible structural contributor to the software vulnerability. 2.3

Failure Truthfulness

The innovation in aira is not detection of bugs per se, but detection of a specific epistemic failure. We define: Failure Truthfulness: The property that a system’s observable outputs accurately represent its internal success or failure state, without suppression, ambiguity, or degradation masking. A system can be functionally correct on the happy path and failure-untruthful under error conditions. Standard correctness testing does not distinguish these. aira is organized around failure truthfulness 3

AIRA — AI-Induced Risk Audit

BDB Labs, 2026

as a first-class system property, separate from and complementary to correctness. Failure truthfulness fails when a system returns a success signal after a critical operation has failed (Check C01), when audit evidence can be lost without halting execution (C02), when exceptions are absorbed without preserving failure semantics (C03), or when outputs are presented as authoritative without a confidence posture (C13). These are not independent bugs — they are different surfaces of the same underlying property violation. 2.4

Implications for Governance-Critical Systems

The stakes are highest in governance, compliance, and safety-critical systems — where fail-closed behavior is not a design preference but a definitional requirement. A system that must refuse to act under uncertainty cannot be built reliably on code that was shaped to suppress uncertainty signals. aira was developed in part through direct observation of this tension during construction of production Constitutional AI governance frameworks, where AI coding agents consistently introduced silent exception handlers into audit-critical paths — not through any individual mistake, but as a systematic pattern across hundreds of generated functions. In distributed systems, the failure mode compounds. A swallowed exception in a microservice does not only affect a local operation — it can corrupt downstream data pipelines, poison distributed caches, and propagate a false success signal across system boundaries.

3

Related Work and Positioning

aira occupies a specific position in the landscape of code quality and AI safety tooling. Generalpurpose static analysis tools — SonarQube, Semgrep, CodeClimate, Pylint — detect broad classes of defect including security vulnerabilities, style violations, and known bug patterns. They are not designed to detect the directional failure bias introduced by optimization pressure. An AI-generated codebase can pass standard static analysis while exhibiting pervasive fail-soft behavior, because the failure mode does not primarily manifest as syntactic error or known vulnerability pattern — it manifests as structurally valid code with the wrong semantics under failure conditions. Observability and failure transparency research has long recognized the distinction between fail-open and fail-closed system design [Beyer et al., 2016]. aira formalizes these concerns as an auditable checklist and connects them to the hypothesized structural cause in AI-assisted development. Research on LLM code generation quality has documented that AI-generated code exhibits higher rates of missing boundary checks, empty exception handlers, and reduced robustness compared to human-written equivalents. Large-scale structural comparisons of agentic versus human-authored pull requests using the AIDev dataset have confirmed that AI-generated code is structurally distinct from human-written code across commit structure and breadth of file modifications [Ogenrwot and Businge, 2026]. Empirical analysis of test failures in AI-generated PRs has found that runtime errors dominate at 62.6% versus 37.4% compile-time failures, with assertion failures as the most common issue [MSR, 2026] — a distribution consistent with fail-soft behavior. In AI safety, the fail-open tendency of RLHF-trained models has been discussed primarily in the context of conversational evasion [Casper et al., 2023]. aira applies this observation to software architecture. 3.1

Distinction from Prior Work

While recent studies have documented higher rates of robustness issues and bugs in LLM-generated code — such as missing boundary checks, empty exception handlers, and reduced performance under 4

AIRA — AI-Induced Risk Audit

BDB Labs, 2026

edge cases — most focus on measuring bug prevalence or functional correctness without identifying a unifying structural cause or providing targeted auditing instrumentation. Benchmarks such as CoderEval [Yu et al., 2024] have evaluated pragmatic code generation and shown varying robustness, while systematic surveys catalog functional, security, and other bug types in AI-generated code [Gao et al., 2025]. aira advances the literature by (1) naming the Reward-Shaped Failure Hypothesis as the underlying optimization-driven mechanism, (2) defining failure truthfulness as a distinct, measurable system property complementary to traditional correctness, and (3) delivering a deterministic 15-check framework explicitly engineered to detect the predicted failure class while remaining resistant to the same suppression tendencies. This bridges software reliability and observability research with AI alignment concerns in a directly actionable way for governance-critical systems.

4 4.1

The AIRA Framework Design Principles

aira was designed around four principles that distinguish it from general code quality tools: Hypothesis-derived checks. The 15 checks target specific failure modes that the Reward-Shaped Failure Hypothesis predicts. The framework is organized around a single question: does this code tell the truth about whether it is working? Deterministic backbone. The primary scanner is parser-backed deterministic analysis — not LLM-assisted classification. This ensures the instrument itself is not susceptible to the failure mode it is designed to detect. LLM augmentation is available but treated as optional enrichment, not ground truth. Fail-closed semantics. UNKNOWN is not a neutral result. Two checks (C07, C12) require human review. On governance-critical paths, UNKNOWN should be treated as a conditional FAIL pending manual verification. Research posture. aira makes no claim that a passing result means a system is safe. The correct language for aira outputs is observed, measured, and suggests — not proves or guarantees. What aira is not. aira is not a general-purpose static analyzer: it does not detect security vulnerabilities, style violations, or broad correctness defects. It is not an authorship detector: it identifies failure-untruthful patterns regardless of whether AI or a human wrote the code. It is not a hallucination detector: it does not evaluate model outputs for factual accuracy. And it is not a claim that every flagged pattern is a defect — some fail-soft behavior is intentional in context. aira is specifically a failure-truthfulness instrument: it surfaces conditions under which a system may misrepresent its own correctness state. 4.2

System Architecture

aira is implemented as four connected components: Deterministic rule engine. Parser-backed static analysis for Python (AST-based), JavaScript, TypeScript, and JSX/TSX (esprima-backed with lexical fallback). CLI tool. Supports static, LLM, and hybrid scan modes across local repos. Includes Ollama model discovery, selected-model validation, and aggregate-only research submission via 5

AIRA — AI-Induced Risk Audit

BDB Labs, 2026

–submit-research-aggregate. Ollama is treated as a generic abstraction layer. Web scanner. Provider-routed scan surface with fallback hierarchy: configured cloud or Ollama route → deterministic server-side static scan → browser heuristics. Research collection pipeline. Aggregate-only submission to Supabase (hosted) or JSONL (local/CI). Raw source code, file paths, and snippets are never transmitted. 4.3

Scan Modes

aira supports three scan modes. Static mode uses deterministic built-in analysis and is the canonical baseline for research. LLM mode uses provider-assisted analysis and is useful for exploratory comparison. Hybrid mode merges static and LLM findings, allowing provider-assisted signal to supplement the deterministic baseline. 4.4

The 15 Checks

Each check targets a specific failure mode predicted by the Reward-Shaped Failure Hypothesis. Table 1 summarizes the full check set. Table 1: AIRA Check Summary ID

Name

Automation

Primary Concern

C01

Success Integrity

Automated

C02

Audit / Evidence Integrity

Automated

C03

Broad Exception Suppression

Automated

C04

Distributed Fallback

Automated (heuristic)

C05

Bypass / Override Paths

Automated

C06

Ambiguous Return Contracts

Automated

C07

Parallel Logic Drift

Human review only

C08 C09

Unsupervised Background Tasks Environment-Dependent Safety

Automated Automated

C10 C11

Startup Integrity Deterministic Reasoning Drift

Automated Automated

C12 C13

Source-to-Output Lineage Confidence Misrepresentation

Human review only Automated

C14

Test Coverage Asymmetry

Automated

C15

Retry / Idempotency Drift

Automated

Success returned after critical failure Audit or evidence loss occurs silently Exceptions swallowed or neutralized Fallback scattered, weakens guarantees Flags or overrides disable safeguards None/null blurs absence and failure Divergent paths produce inconsistent semantics Async work lacks supervision Safety relaxed in dev/staging paths Init catches failure and continues Decision paths use nondeterministic settings Outputs lack traceable origin Degraded output presented without confidence posture Failure-path tests lag happy-path tests Retries on writes without idempotency controls

To illustrate the distinction between AI-typical and human-typical failure handling, consider C01 (Success Integrity): 6

AIRA — AI-Induced Risk Audit

BDB Labs, 2026

Listing 1: AI-typical: failure concealed def process ( data ) : try : result = persist ( data ) return { " status " : " ok " } # success returned regardless except Exception : log . warning ( " persist ␣ failed " ) return { " status " : " ok " } # failure concealed

Listing 2: Human-typical: failure propagated def process ( data ) : result = persist ( data ) # raises on failure if not result : raise PersistenceError ( " persist ␣ returned ␣ empty " ) return { " status " : " ok " }

The AI version is not wrong on the happy path. It is failure-untruthful : it returns success after a broken guarantee. 4.5

Check Co-occurrence Patterns

The 15 checks are not fully independent. Certain co-occurrence patterns describe characteristic failure profiles: • C03 + C01 often indicates explicit failure concealment: exceptions are swallowed to preserve a success-signaling return path • C04 + C09 often indicates environment-shaped degraded assurance: fallback behavior is distributed and relaxed in non-production environments • C02 + C10 often indicates a system that can start and report readiness despite losing evidence guarantees at initialization • C13 is frequently the epistemic surface that makes the rest look normal: confidence misrepresentation makes other failures invisible to downstream consumers Reviewers should treat co-occurring check failures as failure profiles, not independent findings. 4.6

Reviewer Instructions

For each check, reviewers mark: • PASS — requirement is fully satisfied • FAIL — requirement is violated (include: file path, line number, description) • UNKNOWN — cannot be determined without runtime or additional context On UNKNOWN: UNKNOWN does not mean safe. It means the scanner cannot responsibly automate the conclusion. For governance-critical checks, UNKNOWN should be treated as a conditional FAIL pending manual runtime verification.

7

AIRA — AI-Induced Risk Audit

5

BDB Labs, 2026

Empirical Validation

This section reports results across three studies. Study 1 is a deterministic aira audit of a six-system enterprise AI-assisted development environment, presented as an operational characterization of the environment in which the hypothesis emerged. Study 2 is a rebuilt balanced corpus pilot across AI-attributed and human-authored codebases, designed to establish whether the pattern generalizes beyond the originating environment. Study 3 is a stricter matched-control replication that rebuilds the human-control arm at larger scale and tests whether the Study 2 signal survives stronger sampling and repo-composition controls. Together they move the Reward-Shaped Failure Hypothesis from structural argument to replicated measured finding. The studies differ in unit of scope and should not be compared by file count alone. Study 1 is a deep environment audit of several related enterprise systems. Study 2 and Study 3 are crossrepository matched comparisons. Study 3’s final strict subset contains 1,910 scanned files across 135 AI-attributed repositories and 279 human-control repositories, making it the broader external-validity test even though Study 1 scans a larger number of files within a bounded development environment. 5.1

Study Design

All three studies use the aira deterministic static engine exclusively. LLM-assisted scanning is reported separately in Section 5.5. All scans are parser-backed and reproducible. No raw source code, file paths, or snippets are transmitted in research mode. Study 1 — Enterprise AI-Assisted Environment Audit. A deterministic scan of six enterprisegrade systems in active AI-assisted development. The systems include governance, reasoning, synchronization, transcription, and software-engineering support surfaces. One additional candidate system was excluded before scanning because many known errors had already been repaired, which would have made it a remediated-system sample rather than a clean part of the observed development environment. System names and identifying details are omitted to avoid conflating the paper’s claims with evaluation of specific private implementations. Study 2 — Corpus Pilot. A rebuilt balanced comparison of 300 agent-attributed files versus 300 matched human-control files, equally distributed across Python, JavaScript, and TypeScript (100 files per language per arm). Agent-attributed files were sourced from the AIDev dataset [Li et al., 2025] by filtering for rows where agent_label is populated and selecting files from merged PRs. Human-control files were sourced from repositories whose most recent commit predates January 2022, with no AI tool references in commit messages, README, or repository metadata. Both arms were stratified to 100 files per language, with file size bounds of 100–2,000 lines to exclude stubs and generated artifacts. Findings are counted per-file per-check; multiple rule hits within the same file and check are counted individually with no deduplication. Severity is assigned by the deterministic rule engine heuristic (HIGH: clear failure concealment; MEDIUM: risky ambiguity; LOW: distributed fallback that may be contextually acceptable). The agent-attributed arm spans 126 repositories; the human-control arm spans 21 repositories — a structural asymmetry that is discussed as a primary threat to validity in the TypeScript analysis below and in Section 7. Study 3 — Strict Matched-Control Replication. A clean-slate follow-on corpus study using the deduped staged AIDev sample as the AI-attributed source arm and a freshly materialized GitHub PR-file pool as the human-control source arm. This design was chosen because the available AIDev patch data did not contain usable human PR patch rows. The human-control extractor collected 1,636 accepted candidate files after AI-attribution exclusion filters and audit logging. Deterministic 8

AIRA — AI-Induced Risk Audit

BDB Labs, 2026

matching used language, file-size band, and size-decile cells derived from Arm A, enforced a maximum of four files per repository, and used seed 42. The strict matcher selected 956 human controls from the 1,636-file pool; one human file was later excluded by the scanner’s existing file filter, yielding a fair final comparison of 955 AI-attributed files and 955 human-control files. 5.2

Study 1: Enterprise AI-Assisted Environment Audit

Study 1 formalizes the broader environment that motivated the hypothesis. The hypothesis did not originate from a single-system post hoc audit. It emerged from repeated observation across a broader enterprise AI-assisted development environment involving multiple production-grade systems, multiple AI coding tools, and partially overlapping development workflows. That environment motivated the hypothesis and informed the design of the aira checks. The formal quantitative evidence reported here is the deterministic audit of the six included systems. Study 1 should therefore be read as an operational characterization of the originating environment rather than as the paper’s core comparative proof; the strongest between-arm test is the public-code matched-control evidence in Study 3. The six-system scan produced 4,120 findings across 1,643 files: 1,522 HIGH, 1,917 MEDIUM, and 681 LOW. The largest system in the environment remains the anchor case, with 1,189 scanned files and 3,487 findings. The additional five systems contribute 454 scanned files and 633 findings. Across the full environment, C07 and C12 returned UNKNOWN as expected because they require human review. Table 2: Study 1 — Anonymized Enterprise Environment Audit System

Files

Findings

HIGH

MED

LOW

System A (anchor) System B System C System D System E System F

1189 192 119 1 78 64

3487 174 248 11 95 105

1332 47 74 7 23 39

1570 97 135 3 57 55

585 30 39 1 15 11

Total

1643

4120

1522

1917

681

The dominant signals were C03 (Broad Exception Suppression: 1,239 findings), C13 (Confidence Misrepresentation: 715 findings), C04 (Distributed Fallback: 681 findings), and C14 (Test Coverage Asymmetry: 563 findings). C14 is directly comparable to empirical results on test inclusion in agentic PRs, which show substantial variation in test adoption rates across AI coding agents [Haque et al., 2026]. A structurally notable property is severity clustering: 11 of 13 automated checks produce findings of exactly one severity level. C01, C02, C05, C09, C10, C11, and C15 are exclusively HIGH. C06, C08, and C13 are exclusively MEDIUM. C04 is exclusively LOW. Only C03 and C14 span multiple levels. This reflects that different failure modes have consistent consequence profiles — checks that directly conceal failure produce only HIGH findings; checks that create ambiguity without direct concealment produce only MEDIUM. Within the anchor system that originally motivated the project, four subsystems produced zero findings: the validation script, database layer, evidence module, and uncertainty module. This indicates that the framework does not simply flag indiscriminately even on the originating surface. 9

AIRA — AI-Induced Risk Audit

BDB Labs, 2026

The pervasiveness of findings across the originating development environment is the key result of Study 1. The broader environment motivated the hypothesis and informed check design; the formal evidence is the reported deterministic audit. Study 1 should therefore be read as an operational environment audit, not as a causal AI-vs-human comparison. That causal and matched question is addressed by Studies 2 and 3. Internal validity note. Study 1 uses a bounded enterprise development environment selected because it is the environment in which the failure pattern was observed, not a randomly sampled population of software projects. One candidate system was excluded before scanning because known errors had already been repaired. System D contains only one scanner-eligible file under the current Python/JavaScript/TypeScript rule engine and should not be interpreted as a standalone systemlevel estimate. 5.3

Study 2: Rebuilt Balanced Corpus Pilot Table 3: Corpus Pilot — Arm-Level Overview (300 files per arm) Arm

Source

A — AI-attributed B — Human baseline

AIDev (agentic PRs) Pre-2022 repos

Files

HIGH

MED

HIGH/file

300 300

80 61

148 195

0.267 0.203

Under file-weighted analysis (Table 3), the agent-attributed arm produced 80 HIGH, 148 MEDIUM, and 24 LOW findings, compared with 61 HIGH, 195 MEDIUM, and 23 LOW in the human-control arm. This corresponds to 0.267 high-severity findings per file in the agent-attributed arm versus 0.203 in the human-control arm — a 1.32× overall excess of high-severity findings in AI-attributed code. Table 4: High-Severity Findings per File by Language Language

AI Files

AI HIGH/file

JavaScript 100 Python 100 TypeScript (raw) 100 TypeScript (excl. outlier) 100 * See Section 5.3.1 for repo-balance analysis.

0.54 0.11 0.15 0.15

Human Files

Human HIGH/file

100 100 100 99

0.17 0.12 0.32 0.10

Direction AI > Human (3.18×) Parity Human > AI* AI > Human (1.49×)*

The separation was not uniform across languages (Table 4). JavaScript showed the clearest difference: 0.54 HIGH per file in the agent-attributed arm versus 0.17 in the human-control arm. Python was approximately at parity (0.11 versus 0.12). 5.3.1

TypeScript Repo-Balance Analysis

TypeScript initially appeared to cut against the hypothesis (0.15 versus 0.32 HIGH/file). However, follow-up analysis showed this counter-signal is not robust. The human-control TypeScript arm spanned only 6 repositories, and a single file from onlook-dev/onlook contributed 22 of the 32 human-control TypeScript HIGH findings. Excluding that outlier reduces the human-control TypeScript HIGH rate from 0.32 to 0.101 per file. A bootstrap sensitivity analysis (5,000 draws of 6 agent repositories from the 55 available) found P = 0.51 for the agent mean equaling or exceeding the

10

AIRA — AI-Induced Risk Audit

BDB Labs, 2026

outlier-excluded human mean. The apparent TypeScript reversal is not stable under repo-balanced analysis. Table 5: Check-Level Failure Counts by Arm (300 files per arm) Check

Arm A (AI)

Arm B (Human)

Direction

exception_handling (C03) confidence_representation (C13) fallback_control (C04)

61 11 14

43 7 12

AI > Human AI > Human AI > Human

background_tasks (C08) environment_safety (C09) return_contracts (C06)

15 6 7

28 8 9

Human > AI Human > AI Human > AI

At the check level (Table 5), the strongest recurring support for the hypothesis concentrated in exception-handling-related patterns. The agent-attributed arm exceeded the human-control arm on exception_handling (C03: 61 versus 43), confidence_representation (C13: 11 versus 7), and fallback_control (C04: 14 versus 12). By contrast, the human-control arm exceeded the agent-attributed arm on background_tasks (28 versus 15), environment_safety (8 versus 6), and return_contracts (9 versus 7). Internal validity note. Study 2 is file-balanced but not repo-balanced (126 agent repositories vs. 21 human repositories). Language-level comparisons are therefore more reliable for JavaScript, where the arm composition is closer to symmetric, than for TypeScript, where control-side concentration effects are material. The check-level counts should be treated as directional rather than precise. 5.4

Study 3: Strict Matched-Control Replication

Study 3 addresses the main threat exposed by Study 2: the human-control arm was harder to construct and more composition-sensitive than the AI-attributed arm. The follow-on design therefore rebuilt Arm B from scratch rather than extending the 600-file pilot. The final candidate pool contained 1,636 accepted human-control files across JavaScript, Python, and TypeScript, with zero remaining strict matching-cell gaps at the extraction stage. Matching then enforced a four-file repository cap and selected 956 controls. After scanner filtering, the fair final analysis subset contains 955 AI-attributed files and 955 human-control files. Table 6: Strict Matched-Control Replication — Arm-Level Overview (955 files per arm) Arm A — AI-attributed B — Human control

Files

HIGH

MED

LOW

HIGH/file

955 955

415 231

716 617

65 66

0.435 0.242

Table 6 reports the final strict subset. The AI-attributed arm produced 415 HIGH findings, compared with 231 in the matched human-control arm. This corresponds to 0.435 high-severity findings per file versus 0.242, or a 1.80× excess in the AI-attributed arm. The language-level results are important because they resolve the main ambiguity in Study 2. In the pilot, JavaScript carried the strongest signal, Python was near parity, and TypeScript was unstable under repo-balanced sensitivity analysis. In Study 3, all three languages point in the same direction (Table 7). The effect is largest in Python by ratio, strongest in absolute rate in JavaScript, and still present in TypeScript. 11

AIRA — AI-Induced Risk Audit

BDB Labs, 2026

Table 7: Strict Matched-Control Replication — High-Severity Findings per File by Language AI Files

AI HIGH/file

Human Files

Human HIGH/file

JavaScript Python TypeScript

327 311 317

0.682 0.293 0.319

327 311 317

0.367 0.129 0.224

High-Severity Findings (per file)

Language

Direction AI > Human (1.86×) AI > Human (2.27×) AI > Human (1.42×)

0.8 0.68

0.6

0.4

0.37 0.32

0.29

0.22

0.2

0.13

0 JavaScript

Python AI-Attributed

TypeScript

Human Control

Figure 1: Study 3: High-Severity Findings per File by Language (Strict Matched-Control Replication). AI-attributed code shows a directional excess of high-severity fail-soft patterns across all three evaluated languages.

Table 8: Strict Matched-Control Replication — Selected Check-Level Failure Counts Check

Arm A (AI)

Arm B (Human)

263 23 80 43 33 30

185 13 96 45 35 32

exception_handling (C03) environment_safety (C09) background_tasks (C08) confidence_representation (C13) fallback_control (C04) return_contracts (C06)

12

Direction AI > Human AI > Human Human > AI Approx. parity Approx. parity Approx. parity

AIRA — AI-Induced Risk Audit

BDB Labs, 2026

The check-level pattern remains directional rather than uniform (Table 8). The largest replicated gap is in exception handling, where AI-attributed files produced 263 failures versus 185 in human controls. Human controls still exceeded AI-attributed files on background-task failures, and several checks were approximately tied. This mixed profile is methodologically useful: the Study 3 signal is not a scanner-wide inflation artifact, but a concentrated difference in the failure modes most closely related to fail-soft behavior. Internal validity note. Study 3 improves the control design but does not eliminate all sampling constraints. The strict matcher targeted 1,000 controls but selected 956 under the four-file repository cap, and the final scanner comparison uses a 955-per-arm subset after one human-control file was filtered. The correct interpretation is therefore a strict matched replication at 955 files per arm, not a perfect 1,000/1,000 match. 5.5

LLM vs. Static Comparison

To examine whether the failure-suppression pattern extends to the evaluation layer, we ran a parallel hybrid scan using minimax-m2:cloud via Ollama as a secondary evaluator across anchor-system surfaces from the Study 1 environment. This comparison is presented as a secondary exploratory finding, not as primary evidence for the hypothesis. It predates the six-system Study 1 expansion, so its static counts refer to the earlier anchor-system scan rather than the full Study 1 environment aggregate. The LLM evaluator was prompted with a normalized JSON audit contract requesting structured check-level output; C07 and C12 were forcibly kept UNKNOWN in all LLM output to prevent the evaluator from automating human-review-only conclusions. Results at repo scale were subject to prompt truncation; single-file results were not truncated and are the most interpretable. Table 9: Anchor-System Static vs. LLM Finding Counts Across Governance System Surfaces Surface

Files

Static Findings

LLM Findings

Anchor repo (truncated) engine/runtime api critic_infrastructure.py deliberation.py startup.py

1,124 28 58 1 1 1

3,297 222 250 34 26 29

0 5 6 6 0 0

Table 10: Check-Level Suppression: Static FAIL vs. LLM PASS (Single Files, No Truncation) File critic_infrastructure.py deliberation.py startup.py

Static FAILs

LLM Suppressed

Both FAIL

Suppression Rate

7 7 8

2 7 8

5 0 0

29% 100% 100%

The single-file comparisons (Tables 9 and 10) are the most controlled: no truncation, identical files, identical checks. For deliberation.py, the static scanner returned FAIL on 7 checks; the LLM returned PASS on all 7 — 100% suppression. For startup.py, static FAIL on 8 checks; LLM PASS on all 8 — again 100% suppression. The suppression is not random across check types. The LLM consistently passed on C02, C03, C04, C05, and C13 — precisely the checks most directly related to failure concealment. This result is not primarily a critique of the specific model used. It demonstrates that the suppression tendency described in the Reward-Shaped Failure Hypothesis appears not only 13

AIRA — AI-Induced Risk Audit

BDB Labs, 2026

in code generation but in AI-assisted code evaluation. This motivates aira’s deterministic-first design: the audit instrument must not be susceptible to the failure mode it is designed to detect. 5.6

Discussion

Study 1 establishes that the predicted failure class is real and measurable in the enterprise AI-assisted development environment that motivated the hypothesis. The severity clustering result — 11 of 13 automated checks producing single-severity findings — suggests stable consequence profiles across failure modes. Study 2 sharpens rather than simplifies the hypothesis. The rebuilt balanced pilot supports a real but conditional signal: AI-attributed code produces modestly more high-severity findings overall (1.32×), with the strongest separation in JavaScript (3.18×) and the clearest check-level support in exception-handling-related patterns. Python is near parity. The apparent TypeScript reversal is substantially explained by control-side hotspot concentration and does not survive repo-balanced sensitivity analysis. Study 3 tests whether that conditional signal survives a stronger control-arm design. It does. Under a stricter matched-control design, AI-attributed files show a 1.80× high-severity differential over human controls, and all three languages point in the same direction. The most important change from Study 2 is not merely scale; it is that the stronger human-control construction removes the TypeScript ambiguity rather than erasing the overall signal. The data therefore does not support the strongest possible version of the Reward-Shaped Failure Hypothesis — that AI-attributed code is uniformly hotter across every check and surface. It does support a more precise and now replicated version: AI-attributed code shows higher rates of high-severity fail-soft patterns under matched comparison, with especially clear support in exceptionhandling-related behavior. This conditional account is scientifically stronger than a uniform claim, because it generates testable predictions about where future studies should find the signal and where they should not.

6

Output Specification

All aira audits produce output conforming to the following YAML structure (version 1.2): Listing 3: AIRA v1.2 Output Schema ai_failure_audit : audit_version : "1.2" succe ss_integ rity : audit_integrity : ex ce pt ion _h an dli ng : fallback_control : bypass_controls : return_contracts : logic _consist ency : background_tasks : en vi ro nme nt _s afe ty : start up_integ rity : determinism : lineage : co nf id enc e_ op aci ty :

PASS PASS PASS PASS PASS PASS PASS PASS PASS PASS PASS PASS PASS

| | | | | | | | | | | | |

FAIL FAIL FAIL FAIL FAIL FAIL FAIL FAIL FAIL FAIL FAIL FAIL FAIL

14

| | | | | | | | | | | | |

UNKNOWN UNKNOWN UNKNOWN UNKNOWN UNKNOWN UNKNOWN UNKNOWN UNKNOWN UNKNOWN UNKNOWN UNKNOWN UNKNOWN UNKNOWN

AIRA — AI-Induced Risk Audit

BDB Labs, 2026

test_coverage_symmetry : PASS | FAIL | UNKNOWN id em po ten cy _s afe ty : PASS | FAIL | UNKNOWN findings : - issue : " < description >" file : " < path >" line : < number > severity : HIGH | MEDIUM | LOW check : " < check_id >"

Data Availability. The aira scanner is released as open source at https://aira.bageltech. net and https://github.com/BDB-Labs/aira-scanner. Aggregate-only research submissions are supported via the CLI. Raw source code and file paths are never transmitted in research mode. Benchmark datasets and additional comparative scans are welcomed from the community.

7

Limitations

aira should be understood as a research scanner and evolving inspection framework. The following limitations apply: • Cross-file and repo-level semantic reasoning remains weaker than single-file structural detection. Findings are primarily per-file and do not capture emergent failure patterns across module boundaries. • The framework cannot determine that AI wrote the flagged code. The hypothesis predicts a higher rate of these patterns in AI-assisted codebases; aira measures the patterns, not the authorship. • A PASS result does not mean a system is safe. Absence of findings means the targeted patterns were not detected by the current rules — not that the system exhibits full failure truthfulness. • LLM-assisted scans at repo scale can become optimistic or lossy when prompts are truncated. The deterministic engine is the appropriate baseline for research use. • Severity ratings are heuristic. Not every flagged pattern is a defect in context. Human review of findings is always required before remediation. • Study 1 is an environment audit, not a randomly sampled population estimate. The included systems are the enterprise AI-assisted development environment in which the hypothesis emerged. System names are withheld, and one remediated candidate system was excluded before scanning. Study 1 is therefore best interpreted as an operational characterization of the originating environment rather than as the paper’s core comparative proof. • Study 2 reports a rebuilt balanced pilot of 600 files. The pilot is file-balanced but not repobalanced: the agent-attributed arm spans 126 repositories versus 21 for the human-control arm. The TypeScript comparison is particularly sensitive to control-side hotspot concentration. Study 3 was designed to address this weakness, but the pilot should still be interpreted as a directional precursor rather than the final corpus estimate. • Study 3 improves control construction but remains observational. The final strict comparison is 955 files per arm, not the originally targeted 1,000 files per arm, because a four-file repository cap and scanner filtering reduced the analyzable matched subset. The human-control arm

15

AIRA — AI-Induced Risk Audit

BDB Labs, 2026

is matched on language, file-size band, and size decile, but not on every possible domain, framework, or system-role variable. • C07 and C12 remain human-review-only by design. Automated detection of parallel logic drift and source-to-output lineage failure requires repo-scale semantic comparison not yet implemented.

8

Conclusion and Future Work

This paper has introduced the Reward-Shaped Failure Hypothesis — the proposal that AI coding systems may produce fail-soft code in part because of pressures in their optimization environment — and aira, a 15-check inspection framework designed to detect the specific failure class this hypothesis predicts. The central reframing is this: traditional software validation asks whether systems work. aira asks whether systems tell the truth when they fail. Failure truthfulness is a distinct and auditable system property. The 15 checks are organized around its measurement. Study 1 results — 4,120 findings across 1,643 files in a six-system enterprise AI-assisted development environment, plus an anchor-surface LLM comparison under-reporting at a 44:1 ratio — provide evidence that the predicted failure class is real within the originating environment and may extend to the evaluation layer. Study 2, a rebuilt balanced 600-file corpus pilot, extends these findings beyond the originating environment and identifies the control-arm construction problem that a stronger replication must solve. Study 3 performs that stricter replication: in a 955-file-per-arm matched comparison, AI-attributed code shows 0.435 high-severity findings per file versus 0.242 in human controls, a 1.80× differential, with the same direction in JavaScript, Python, and TypeScript. The hypothesis is therefore supported as a real but conditional pattern — not a uniform field-wide law, but a measurable and reproducible tendency whose strength varies by language, check type, and code surface. The strongest replicated evidence is not that AI-attributed code fails every check more often; it is that AI-attributed code shows higher rates of high-severity fail-soft patterns under increasingly strict matched comparison. The three most important next steps are: (1) domain and system-type stratification within the matched public-code corpus; (2) a controlled base-vs-RLHF model comparison on structured generation tasks, to move the hypothesis from observational evidence to direct experimental proof; and (3) expansion of the deterministic rule engine to support cross-file and repolevel reasoning. aira is available as an open-source tool at https://aira.bageltech.net and https://github.com/BDB-Labs/aira-scanner. The research collection pipeline supports aggregateonly community submissions. Contributions are welcomed, especially comparative scan datasets, rule extensions, and calibration studies.

Acknowledgments aira was developed through direct observation during construction of Constitutional AI governance frameworks with heavy AI coding assistance. The author thanks the broader AI safety and software reliability research communities whose parallel work motivates the bridge this paper attempts to build.

16

AIRA — AI-Induced Risk Audit

BDB Labs, 2026

References B. Beyer, C. Jones, J. Petoff, and N. Murphy. Site Reliability Engineering: How Google Runs Production Systems. O’Reilly Media, 2016. S. Casper, X. Davies, C. Shi, T. Gilbert, J. Scheurer, J. Rando, R. Freedman, T. Korbak, D. Lindner, P. Freire, et al. Open problems and fundamental limitations of reinforcement learning from human feedback. arXiv preprint arXiv:2307.15217, 2023. R. Gao et al. A survey of bugs in AI-generated code. arXiv preprint arXiv:2512.05239, 2025. S. Haque, D. Strüber, and N. Tsantalis. Do autonomous agents contribute test code? A study of tests in agentic pull requests. arXiv preprint arXiv:2601.03556, 2026. H. Li, H. Zhang, et al. AIDev: The rise of AI teammates in software engineering 3.0. Dataset available at https://huggingface.co/datasets/hao-li/AIDev, 2025. MSR 2026 Mining Challenge. An empirical analysis of test failures in AI-generated pull requests. In Proc. 23rd International Conference on Mining Software Repositories (MSR), Rio de Janeiro, Brazil, April 2026. E. Ogenrwot and J. Businge. How AI coding agents modify code: A large-scale study of GitHub pull requests. arXiv preprint arXiv:2601.17581, 2026. H. Yu, W. Shen, K. Ran, J. Liu, Q. Wang, and Y. Jiang. CoderEval: A benchmark of pragmatic code generation with generative pre-trained models. In Proc. IEEE/ACM ICSE, 2024.

17

Related documents

Record · ID 120593 · SHA-256 11931c135ba03fec
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.