ConceptioArchivearXiv CS
arXiv CSopen access

Is US Defense Acquisition Ready to Acquire AI-Enabled Capabilities? Assessing the DoD Software Acquisition Pathway Through a Scenario-Based Policy Analysis

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
softwarearchitecturesoftwareengineeringtesting
software engineering, software architecture, testing

Is US Defense Acquisition Ready to Acquire AI-Enabled Capabilities? Assessing the DoD Software Acquisition Pathway Through a Scenario-Based Policy Analysis

DANIEL LUGO, Purdue University, USA JAMES C. DAVIS, Purdue University, USA As AI systems transition from experimental prototypes to mission-critical tools, their dependence on dynamic data, evolving models, and governance raises questions about whether existing acquisition pathways can keep pace. The U.S. Department of Defense has modernized its acquisition processes through the Adaptive Acquisition Framework, with the Software Acquisition Pathway (SWP)

arXiv:2606.07393v1 [cs.SE] 5 Jun 2026

serving as the primary mechanism for acquiring software-intensive capabilities. This paper evaluates whether SWP is sufficient to address the unique demands of AI acquisition. In this work, we perform a scenario-based evaluation that traces a notional AI-enabled program through key SWP planning activities to assess how policy translates into program artifacts and decisions. We use Policy Scenario Analysis to examine whether the SWP-centered governance stack provides sufficient actionable support for AI acquisition. The governance stack provides a viable foundation for iterative delivery and AI testing. However, we identify a recurring actionability problem in the core guidance. AI-specific controls for data provenance, lifecycle management, and human oversight remain distributed across supplemental documents rather than embedded in the program-facing mechanisms through which SWP is executed. This disconnect leaves program offices reliant on inconsistent local interpretation. We conclude by recommending an AI-supporting sub-path and targeted artifact refinements to better bridge this policy-to-artifact gap. CCS Concepts: • Software and its engineering → Software development process management; • Social and professional topics → Government technology policy. Additional Key Words and Phrases: Software acquisition, AI acquisition, Adaptive Acquisition Framework (AAF), Software Acquisition Pathway (SWP), Policy Scenario Analysis (PSA), Test and evaluation, Responsible AI, Government procurement ACM Reference Format: Daniel Lugo and James C. Davis. 2026. Is US Defense Acquisition Ready to Acquire AI-Enabled Capabilities?: Assessing the DoD Software Acquisition Pathway Through a Scenario-Based Policy Analysis. 1, 1 (June 2026), 30 pages. https://doi.org/XXXXXXX.XXXXXXX

1

Introduction

Governments are integrating AI capabilities to improve operational efficiency in domains such as healthcare, education, transportation, finance, and national defense [32, 52, 77]. Public-sector delivery is constrained by formal systems of policy, oversight, and budgeting that shape what programs can require, document, approve, and sustain [28]. As a result, some governments are developing AI-specific procurement and governance frameworks that make expectations for testing, transparency, risk management, and post-deployment oversight more explicit [16, 18]. For defense acquisition, this raises a broader question about whether existing guidance can accommodate the distinctive demands of AI Authors’ Contact Information: Daniel Lugo, [email protected], Purdue University, West Lafayette, Indiana, USA; James C. Davis, [email protected], Purdue University, West Lafayette, Indiana, USA. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. Manuscript submitted to ACM Manuscript submitted to ACM

1

2

Lugo & Davis

acquisition. This tension is particularly acute for AI-enabled capabilities whose performance depends on data quality, model updates, validation practices, and continued monitoring over time [4, 27, 58]. In the U.S. Department of Defense (DoD), capability acquisition is organized through the Adaptive Acquisition Framework (AAF) [42]. Within AAF, the Software Acquisition Pathway (SWP) was introduced to support rapid, iterative delivery through continuous integration and responsiveness to user needs [43]. SWP was designed to move software programs away from rigid, hardware-oriented acquisition approaches and toward processes more consistent with modern software practice. By mandating SWP for software-intensive efforts, DoD has positioned it as the principal organizing framework for software acquisition [29, 38]. Acquisition teams must use it within a broader governance environment that includes statutory, regulatory, departmental, and AI-specific guidance [11, 42]. In this study, we consider whether AAF’s SWP-centered governance stack provides DoD program offices with sufficiently actionable guidance for the acquisition of AI capabilities. To answer that question, we conduct a scenariobased policy assessment [8] of the DoD acquisition governance stack. The assessment compares a synthetic AI-enabled acquisition scenario with a conventional software acquisition scenario. We decompose both scenarios into discrete acquisition episodes and trace them through the Planning Phase artifacts, decision points, and review structures that organize software acquisition in practice. We then evaluate whether the governing corpus provides explicit, partial, or absent support for AI-relevant acquisition properties at the points where programs must produce artifacts and make decisions. Our results show that SWP can accommodate the high-level needs of software-intensive AI programs. Its actionability becomes uneven, however, when program success depends on lifecycle controls that are more salient for AI than for conventional software. Across our scenario-based assessment, SWP consistently supports modular execution and general cybersecurity expectations, but offers weaker or absent procedural direction for training-data governance, model traceability, retraining triggers, and performance-drift management. Although supplemental guidance improves coverage in several of these areas, important gaps remain. Overall, our findings indicate that the broader SWP-centered governance stack is unevenly specified for AI-specific acquisition needs. In summary, our main contributions are: • A scenario-based assessment of AI acquisition under SWP. We apply an established policy-assessment method to the Software Acquisition Pathway and identify where the current governance stack provides explicit, partial, or absent support for AI-specific acquisition demands. • Recommendations to make SWP more AI-ready. Where SWP provides partial or absent guidance, we propose targeted enhancements to current guidance, framed as improvements to the existing acquisition-process rather than a wholly separate pathway. Significance: As AI-enabled capabilities become central to defense missions, acquisition policy must address lifecycle demands that extend beyond conventional software delivery, including data governance, model assurance, ongoing monitoring, and responsible-use oversight. This study contributes both a practical framework for navigating existing guidance when planning AI-enabled acquisitions and an evidence-based basis for refining AAF/SWP guidance where needed. 2

Background and Related Work

This section provides the background for our policy assessment. §2.1 explains the origins and intent of the Adaptive Acquisition Framework (AAF) and the Software Acquisition Pathway (SWP). §2.2 identifies the key properties that Manuscript submitted to ACM

Is US Defense Acquisition Ready to Acquire AI-Enabled Capabilities?

3

Fig. 1. The Software Acquisition Pathway (Reused from DoDI 5000.87 [43]). This figure illustrates the SWP; it has two main phases (Planning and Execution) and proposes active user engagements and iterative development. There are several plan-code-build-test iterations where software releases can be delivered.

distinguish AI from conventional software and defines the dimensions that we later use in the methodology to assess policy support. §2.3 then situates SWP within the broader literature on acquisition reform and AI governance. 2.1

The Software Acquisition Pathway (SWP)

The Adaptive Acquisition Framework (AAF) was introduced in 2020 as part of the Department of Defense’s broader acquisition reform effort [42], replacing the long-standing DoDI 5000.02-centered structure that dated back to 2008 and was later revised in 2015 [64, 65]. DoDI 5000.02 established the policy framework for that change and responded to longstanding criticism that earlier DoD 5000-series processes were too rigid, too slow, and poorly matched to modern software and other emerging technologies [3, 42]. The AAF established six acquisition pathways, including Urgent Capability, Middle Tier, Major Capability Acquisition, and Software Acquisition (SWP), each intended to tailor governance and oversight to the characteristics of different acquisition efforts [42]. AAF did not change acquisition practice through a single top-level policy alone. Its introduction triggered a broader realignment of acquisition guidance across the Department, including pathway-specific instructions and updates to functional and Service-level policies for engineering, testing, cybersecurity, sustainment, and related activities. As a result, acquisition personnel do not execute AAF through a single self-contained instruction. They work instead within a distributed acquisition-governance environment that combines pathway guidance with other authoritative policy instruments. Manuscript submitted to ACM

4

Lugo & Davis Within that acquisition environment, DoDI 5000.87 serves as the core instruction for the SWP [43]. It organizes the

software lifecycle into two primary phases: the Planning Phase and the Execution Phase (see Figure 1). The Planning Phase is designed to establish the program’s technical and governance foundations, requiring the development of a Functional Strategy and a Value Assessment that replace traditional, static requirements with a more dynamic, user-centered approach [9]. During this phase, acquisition teams must also produce a suite of foundational strategies, including the Acquisition Strategy, Test Strategy, Cybersecurity Strategy, and Product Support Strategy, which define the operational envelope for the capability [12]. Upon successful completion of these planning milestones, the program enters the Execution Phase, which is characterized by rapid, iterative cycles of development, integration, and delivery [9]. This phase leverages DevSecOps principles to maintain continuous user engagement and shortened delivery timelines, ideally resulting in a Minimum Viable Product (MVP) and subsequent software increments [43]. By institutionalizing Continuous Integration/Continuous Delivery (CI/CD) and continuous Authority to Operate (cATO), the SWP provides the procedural framework intended to keep defense software responsive to changing operational needs [3, 43]. 2.1.1 Documents Used by Acquisition Personnel Under SWP. Acquisition personnel do not execute SWP through the pathway instruction alone. They must navigate a broader policy environment that includes the pathway instruction, the functional policies it invokes, and supplemental guidance for engineering, testing, cybersecurity, and sustainment [10, 39, 41, 42, 44]. This is especially relevant for AI-enabled capabilities, which are commonly acquired under SWP when custom software is the primary means of delivering the capability [9, 43]. Some requirements are addressed through standard software structures, while others rely on cross-cutting mandates from related acquisition functions [39, 41, 44]. Although these materials vary in formal procedural status, ranging from mandatory instructions to advisory or explanatory guidance, they collectively constitute the operational guidance environment for a DoD program acquiring software capability [10, 42]. Table 1 summarizes the principal documents that structure the planning and execution of software-intensive programs. Together, they represent the distributed governance stack that defines the compliance and delivery boundaries for acquisition practitioners.

2.2

AI-Relevant Acquisition Properties

Software now underpins capability delivery and sustainment across the DoD, with a growing share of that investment directed toward AI-enabled capabilities. For FY2025, the DoD requested approximately $1.8 billion specifically for artificial intelligence initiatives [46]; by FY2026, this expanded into a dedicated $13.4 billion portfolio for autonomy and AI systems [47]. To put this in perspective, this dedicated AI and autonomy funding is now equivalent in scale to approximately 20% of the Department’s entire $66 billion IT budget. This investment supports a wide range of technologies that extend beyond traditional deterministic software. According to published FY2026 program acquisition cost estimates, these investments are concentrated in several high-priority categories, including aerial autonomy and drones, maritime and underwater systems, and AI-enabled infrastructure and integration [48]. AI-enabled capabilities differ from conventional software because they require the continuous interaction of code, data, models, and operations [4, 54]. As a result, acquisition success depends not only on delivering executable code, but also on governing data quality, evaluating probabilistic behavior, and maintaining oversight of model evolution. Table 2 summarizes the AI-relevant properties discussed in this section. These properties were identified through our Manuscript submitted to ACM

Is US Defense Acquisition Ready to Acquire AI-Enabled Capabilities?

5

Table 1. Documents that acquisition personnel may use when planning and executing software-intensive programs under the Adaptive Acquisition Framework and the Software Acquisition Pathway. Identified via the DAU AAF Acquisition Policies repository [10], the DoD Publications repository [72], and the Federal CIO Council policy index [74]. (Snapshot Nov 2025). Layer

Title / Identifier

Role in acquisition environment

Date

Overarching

DoDI 5000.02, Operation of the Adaptive Acquisition Framework

AAF structure and governance

Jun 2022

Pathway

DoDI 5000.87, Operation of the Software Acquisition Pathway

SWP procedural anchor

Oct 2020

Functional

DoDI 5000.73, Cost Analysis Guidance and Procedures DoDI 5000.82, Acquisition of Digital Capabilities DoDI 5000.83, Technology and Program Protection to Maintain Technological Advantage DoDI 5000.88, Engineering of Defense Systems DoDI 5000.89, Test and Evaluation DoDI 5000.90, Cybersecurity for Acquisition Decision Authorities and Program Managers DoDI 5000.91, Product Support Management for the Adaptive Acquisition Framework DoDI 5000.95, Human Systems Integration in Defense Acquisition DoDI 5000.97, Digital Engineering DoDI 5010.44, Intellectual Property (IP) Acquisition and Licensing

Cost analysis expectations Digital capability planning Program protection expectations

Oct 2024 Jun 2023 May 2021

Systems engineering expectations T&E expectations Cybersecurity expectations

Nov 2020 Nov 2020 Dec 2020

Sustainment expectations

Nov 2021

Human factors

Apr 2022

Digital engineering practices IP and data/software rights strategy

Dec 2023 Oct 2019

DAFI 63-101/20-101, Integrated Life Cycle Management SECNAVINST 5000.2G, Defense Acquisition System and Joint Capabilities Integration and Development System Implementation Army Regulation 70–1, Army Acquisition Policy

Air Force implementation overlay

Feb 2024

Navy implementation overlay

Apr 2022

Army implementation overlay

Nov 2023

DoD Responsible AI Strategy and Implementation Pathway CDAO Responsible AI Toolkit CDAO Test and Evaluation Strategy Framework CDAO AI Test and Evaluation Guidance

DoD AI governance expectations

2022

Operational RAI support AI TEVV framing AI T&E implementation detail

2023 2024 2024

federal AI acquisition guidance

Apr 2025

Service

Supplemental Department

Supplemental Government-wide

OMB Memorandum M-25-22, Driving Efficient Acquisition of Artificial Intelligence in Government

literature review on how AI-enabled systems differ from conventional software and were selected to highlight distinct acquisition implications for planning, oversight, testing, sustainment, and governance. • Data as a first-class deliverable. Conventional software often treats data primarily as an input or output. In AI systems, training, tuning, and evaluation data materially shape capability performance, reuse, and sustainment. Data quality, lineage, representativeness, and licensing become acquisition concerns alongside code and interfaces [19, 36]. • Models evolve and can drift. Conventional software logic usually changes through managed releases. AI models may degrade as environments, missions, or data distributions change, creating ongoing needs for monitoring, retraining, and revalidation [2, 4]. Manuscript submitted to ACM

6

Lugo & Davis

Table 2. AI-relevant properties used as evaluation dimensions in the policy assessment, identified through our literature review on how AI-enabled systems differ from conventional software and selected for their acquisition implications. AI-relevant property

Acquisition implication relative to conventional software

Data as a first-class deliverable

Performance and sustainment depend on training, tuning, and evaluation data, making data quality, lineage, and licensing acquisition concerns alongside code and interfaces [19, 36]. Performance may change as missions, environments, or data distributions shift, creating continuing needs for monitoring, retraining, and revalidation [2, 4]. Behavior may vary across inputs and operating conditions, making exhaustive testing infeasible and increasing the importance of robustness evaluation and context-aware assurance [4, 54]. Limited interpretability increases the need for traceability, documentation, and provenance linking model behavior to data, design choices, and operational use [79, 80]. Fairness, privacy, lawful use, and human oversight can determine whether a capability is acceptable for deployment and continued use [33, 76]. Dependence on upstream models, datasets, and toolchains increases provenance, security, and licensing risks across the lifecycle [25, 55, 57, 60]. Fast-moving models, tools, and deployment stacks increase the need for continuous validation, monitoring, and adaptation [21, 31, 37]. Performance, latency, throughput, and cost often depend on accelerator-rich environments, making deployment context more consequential [26, 56].

Models evolve and drift Non-determinism and distribution sensitivity Explainability and accountability Ethics and governance constraints Supply-chain and transitive risk Rapid ecosystem churn Compute and specialized hardware

• Assurance is harder under non-determinism and distribution sensitivity. Many AI systems are probabilistic, datadependent, and difficult to evaluate exhaustively. AI assurance must address robustness, distribution shift, monitoring, and performance across changing operational contexts rather than static test accuracy alone [4, 54]. • Explainability and accountability matter for oversight. Many AI components are only partially interpretable, which complicates assurance, traceability, certification, and human accountability. This increases the importance of documentation and provenance-aware mechanisms that connect model behavior to underlying data, design choices, and operational use [79, 80]. • Ethics and governance become operational requirements. Fairness, bias mitigation, privacy, lawful use, and human oversight are not optional enhancements; they can determine whether an AI capability is acceptable for deployment and continued use [33, 76]. • Supply-chain security and transitive risk increase. AI systems depend on upstream artifacts such as pre-trained models, datasets, and specialized toolchains. Unlike traditional software, these dependencies include model weights that can propagate backdoor vulnerabilities or data-poisoning risks into fielded capabilities. This necessitates provenance, integrity verification, and a more granular AI Bill of Materials (AIBOM) that extends traditional software provenance to include dataset lineage and model training configurations [25, 37, 55, 57, 60]. • The ecosystem changes quickly. AI models, tools, and deployment stacks evolve rapidly, resetting baselines and increasing the need for continuous validation, monitoring, and adaptation across the lifecycle [21, 31, 37]. • Compute and hardware dependencies are often mission-relevant. Many AI workloads depend on accelerator-rich compute environments such as GPUs, TPUs, or NPUs. As a result, performance, latency, throughput, and cost may vary significantly across hardware targets, including edge and cloud deployments [23, 26, 56]. Manuscript submitted to ACM

Is US Defense Acquisition Ready to Acquire AI-Enabled Capabilities?

7

Summary. These properties motivate the later scenario design and coding framework. In particular, they highlight why AI acquisition places greater pressure on data governance, probabilistic TEVV, lifecycle monitoring, provenance, and human-governance mechanisms than conventional software-intensive acquisition. 2.3

Prior Work on SWP and AI Acquisition

The DoD’s transition to AAF and SWP reflects a broader effort to support iterative delivery for mission-tailored software. Existing literature, however, remains limited and divided. One stream focuses on the institutionalization of Agile and DevSecOps, while another treats AI as a distinct governance and contracting problem. The first stream of literature treats the SWP primarily as a reform mechanism for improving software-delivery speed. GAO found that, despite policy updates intended to support modern software practices, implementation remained inconsistent across programs [75]. Dunlap likewise emphasizes entrenched business practices, legacy certification processes, and workforce constraints as persistent barriers to software-acquisition reform [17]. Tate and Bailey further argue that SWP suitability depends on program conditions and organizational context rather than serving as a universal solution [59]. While these studies identify constraints on SWP adoption for general software, they offer limited insight into whether these artifacts can accommodate AI acquisition. The second stream of literature shifts attention to AI as a distinct acquisition and contracting problem driven by data dependence and continuous model evolution. The Center for Strategic and International Studies (CSIS) [53] argues that procurement mechanisms designed for static software deliverables are ill-suited for AI systems requiring ongoing, data-driven retraining. Complementing this, DAU’s analysis [49] advocates for a responsible AI framework that embeds ethics, transparency, and bias mitigation into oversight. These analyses provide a vision for what AI acquisition should entail, yet they remain largely decoupled from the concrete execution mechanisms of the SWP. Recent DoD fielding efforts further complicate this landscape by introducing two diverging acquisition patterns. On one hand, generative AI is increasingly delivered through enterprise-scale commercial integration—such as the GenAI.mil rollout and large-scale Palantir awards—using license- and platform-centered contracting vehicles [20, 61]. Conversely, certain mission contexts, including forward operating positions, space systems, embedded systems, and contested environments, will require specialized, edge-based AI resident on mission systems and designed for real-time operation under tactical constraints [62, 63, 73]. In those settings, AI acquisition aligns more closely with traditional program responsibilities, with the program office responsible for integration, sustainment, verification, and upgrade planning for AI components deployed on mission systems. Prior research therefore, supports both SWP-driven modernization and AI-specific governance reform, but it says less about how those two agendas meet in practice. Despite rising AI funding and DoD’s reliance on SWP, relatively little work has examined whether SWP’s planning artifacts provide the actionable specificity required for AI-enabled systems, especially in tactical edge environments [6]. This paper addresses that gap through a scenario-based policy assessment that maps SWP planning artifacts to AI technical requirements, identifying key points of tension and proposing refinements for an AI-ready acquisition pathway. 3

Research Question, Methodology, and Corpus Selection

AI-enabled capabilities inherit many features of software-intensive systems, but they also introduce distinctive dependencies. Given these differences, it is crucial to understand whether the Software Acquisition Pathway (SWP) provides sufficiently specific guidance for acquiring such capabilities. We therefore investigate the following research question: Manuscript submitted to ACM

8

Lugo & Davis

RQ: To what extent does the Software Acquisition Pathway (SWP), as implemented through the broader SWP-centered governance stack, accommodate AI’s distinctive technical and operational characteristics, and where does it remain underspecified?

3.1

Methodology Overview

Evaluating acquisition policy through observed program outcomes is difficult because such studies require long time horizons and are strongly shaped by local tailoring, organizational capacity, and operational context. They also depend on access to program artifacts, internal documentation, and personnel across multiple programs, which may be limited or unavailable in practice. In this setting, a lighter-weight pre-implementation assessment can provide an initial evidence base for whether a policy appears actionable and whether a more resource-intensive evaluation is warranted. We therefore employ Policy Scenario Analysis (PSA) [8]. PSA evaluates policy through scenarios that combine relevant circumstances with the policy actions available to decision-makers, enabling analysts to examine likely implications before real-world execution [8]. In this study, the documents described in Section 2.1.1 constitute the policy action set. These documents define artifacts, roles, responsibilities, decision points, and review structures through which an SWP program is expected to plan and govern its activities. We operationalize PSA through a structured scenario-based assessment of a synthetic but grounded DoD-like AI acquisition case, alongside a conventional software comparator. In this way, the analysis helps surface weaknesses in the governing policy stack before they emerge in execution or sustainment.

3.2

Scenario Design and Scope

This subsection defines the baseline and comparator scenarios, the study boundary, and the stakeholder perspective used to evaluate policy actionability. Scenario set design (baseline + comparator). PSA typically employs a small number of scenarios, usually beginning with a baseline case and comparing it with a single alternative; additional scenarios are introduced only when analytically necessary to clarify policy implications [8]. In our assessment, we use two scenarios: (i) a baseline AI acquisition under SWP, and (ii) a comparator conventional software-intensive acquisition under SWP. The comparator serves as a methodological control that holds the acquisition framework constant while varying the technical character of the capability. This design allows us to distinguish AI-specific underspecification from broader SWP reliance on local tailoring. Baseline scenario design: AI-enabled program (Targeting AI). Our primary evaluation case is a hypothetical capability termed the Targeting AI program. The scenario is intentionally synthetic but grounded: rather than reproducing any single real-world program, it represents a recognizable class of DoD AI acquisition efforts in which machine-learning models are integrated into mission workflows under operational and governance constraints. The notional system fuses multi-source sensor data, including optical imagery, radar, and signals intelligence (SIGINT) reports, to generate targeting recommendations for human-in-the-loop decision support (see Figure 2a). We assume that the capability must operate at the tactical edge, requiring on-premises execution in bandwidth-constrained and contested environments. This assumption is consistent with current DoD demand signals indicating that, alongside enterprise AI platforms, some Manuscript submitted to ACM

Is US Defense Acquisition Ready to Acquire AI-Enabled Capabilities?

9

(a) Notional AI-enabled targeting capability.

(b) Notional flight planning optimizer. Fig. 2. Comparison of the two notional scenarios used in this study. (a) The AI-enabled targeting capability integrates diverse data streams—including signals intelligence (INTEL), optical sensors, radar, and mission parameters—into an AI inference engine that generates targeting recommendations for a weapon system, while a human operator remains in the loop for final decision authority and supervision. (b) The flight planning optimizer uses authoritative data sets such as aircraft performance specifications, airspace restrictions, and weather products within an optimization algorithm to produce a fuel-efficient, constraint-satisfying route for review by mission planners before transmission to aircrews.

mission contexts will require specialized, edge-based AI resident on mission systems1 and operating under real-time tactical constraints [62, 63, 73]. We further ground the scenario by mapping its principal features to publicly documented DoD program patterns, each of which contributes an acquisition-relevant burden preserved in the baseline. • Analytic function: Informed by Project Maven, the scenario retains the requirement to process heterogeneous intelligence inputs using computer vision and machine learning in support of faster analysis and human decisionmaking [66]. • Operational integration: Informed by the PM IS&A portfolio, the system is situated within existing expeditionary intelligence-processing workflows, ensuring that the software must interoperate with mission systems and support tactical information requirements [50]. • Edge constraints: Informed by current demand signals associated with CJADC2 and contested operations, the scenario assumes disconnected or intermittent connectivity, making update timing, local model validation, and sustainment planning acquisition-relevant burdens rather than merely architectural details [69]. 1 See DoD FY2026 Budget Estimates, SOCOM (2025): “Investments in ’AI at the Edge’ are critical for autonomous ISR platforms operating in bandwidth-

constrained and GPS-denied environments, where centralized AI processing is not a viable COA [Course of Action].”

Manuscript submitted to ACM

10

Lugo & Davis We design the Targeting AI program scenario to preserve several of the AI-relevant acquisition properties summarized

in Table 2 while remaining sufficiently bounded for structured policy analysis. It therefore serves as a demanding but analytically tractable test case: multi-source fusion makes data a first-class deliverable; environmental and mission drift create a need for monitoring and retraining to manage model evolution; the non-deterministic character of AI complicates conventional test logic; and reliance on specialized hardware at the edge makes deployment context an acquisition-facing concern. Comparator scenario design: conventional software program (Flight Planning Optimization). To distinguish AI-specific underspecification from broader SWP reliance on local tailoring, we define a conventional software-intensive comparator with a similar operational context but without learning-based behavior: a notional flight planning optimization application that generates fuel-efficient flight plans by integrating deterministic data sources and operational constraints such as aircraft performance models, routing constraints, and weather services, with mission-planning users reviewing outputs before dissemination to aircrews (see Figure 2b). As with the Targeting AI baseline, this scenario is synthetic but grounded in publicly described DoD mission-planning software with comparable operational context, delivery pressures, and user interaction requirements [45]. We trace it through the same SWP decision points and planning artifacts as the AI scenario so that the acquisition framework remains constant while the technical character of the capability varies. The comparator therefore serves as a methodological control, helping us distinguish AI-specific acquisition burdens from broader forms of SWP underspecification that may also affect conventional software. Scope and time horizon. We bound the study to the SWP Planning Phase, beginning after AAF pathway selection and continuing through completion of the planning activities required under SWP (§2.1). We assume that SWP has been formally designated as the program’s primary governance mechanism, consistent with Department memoranda directing its use for software-dominant capabilities [29, 38]. Within this boundary, we evaluate the guidance used to develop the required planning artifacts. We focus on the Planning Phase because it is the point at which written policy is converted into executable program governance. This makes it possible to assess whether the SWP-centered governance stack provides acquisition teams with sufficient direction before the program enters resource-intensive development and deployment. Stakeholders. The PSA is conducted from the perspective of the Integrated Product Team (IPT), which is responsible for translating SWP policy into executable plans, artifacts, and reviewable evidence. Core IPT roles include the Program Manager and Lead Engineer, supported by functional experts in TEVV, cybersecurity, and contracting. The Decision Authority provides pathway and gate approvals and accepts the program’s risk posture, while functional authorities influence the minimum evidence required for release decisions. We adopt this perspective because SWP’s actionability for AI depends on whether program teams are given sufficiently clear guidance to execute consistently across iterations. 3.3

Evaluation Procedure

This subsection describes the procedure used to evaluate the study scenarios, including the episode structure, assessment criteria, coding scheme, and resulting analytic outputs. Scenario Episodes (Circumstances). To instantiate the circumstances side of the PSA framework, we organize our scenarios into five episodes: (1) data sourcing, labeling, and rights negotiation; (2) TEVV design; (3) cybersecurity and release provenance for model and software artifacts; (4) lifecycle management; and (5) human oversight constraints for operational use. We initially identified ten candidate episodes from the intersection of three elements in the study design: Manuscript submitted to ACM

Is US Defense Acquisition Ready to Acquire AI-Enabled Capabilities?

11

Fig. 3. Acquisition Policy Scenario Analysis methodology. We apply Policy Scenario Analysis to evaluate SWP actionability for AI by (Phase 1) defining study parameters, selecting the governing policy corpus, and constructing a baseline AI scenario plus a conventional-software comparator; (Phase 2) walking each scenario through SWP Planning-phase activities to map required artifacts to policy guidance and identify gaps; and (Phase 3) producing comparative outputs—a coverage scorecard and a summary of key differentiators and recommendations for decision makers.

the AI-relevant acquisition properties identified in Table 2, the mandatory SWP Planning Phase artifacts and review structures described in the governing corpus, and the recurring acquisition burdens identified in the AI-governance and defense-acquisition literature reviewed in §2.3. From that broader set, we selected five because they provided the strongest non-redundant coverage of AI-relevant stress points within the SWP Planning Phase. We do not claim that these five episodes exhaust all possible AI-acquisition circumstances; instead, they were chosen as representative and analytically useful cases for evaluating the SWP-centered governance stack. Evaluation concerns. Within each episode, we assess whether SWP guidance provides actionable policy support for five acquisition concerns that are especially salient for AI and that can also be examined in the conventional software comparator: • Data governance and data rights: whether policy supports ownership and accountability, provenance and labeling expectations, and rights sufficient for training, test, operation, and sustainment. • Assurance and T&E (TEVV): whether policy supports the evidence and evaluation methods needed to justify performance across operational conditions, including deterministic software verification as well as more conditionsensitive or non-deterministic AI behavior. • Cybersecurity and traceability: whether policy supports model and pipeline security (including supply chain) and traceability across data, model, software, and system releases. Manuscript submitted to ACM

12

Lugo & Davis • Lifecycle management: whether policy supports monitoring, update and change-control responsibilities, and triggers for retraining, drift detection, software revision, or re-authorization following material changes in data, models, code, or operational conditions. • Human oversight: whether policy supports human oversight and responsible or ethical compliance mechanisms appropriate to the use context. These same concern areas are applied to the conventional software comparator, but they arise in different forms

and with different levels of acquisition burden. For example, data governance in the comparator primarily concerns authoritative input data and interface control rather than training-data rights; lifecycle management centers on software versioning and rule updates rather than retraining or model drift; and assurance focuses on deterministic verification and test coverage rather than stochastic model behavior. Applying a common concern structure across both scenarios allows observed differences in policy support to be attributed more clearly to AI-specific properties rather than to differences in the evaluation framework. By selecting these five areas, we operationalize acquisition actionability for AI in a manner consistent with the DoD Responsible AI Strategy’s tenets of being Traceable, Reliable, and Governable [71]. Episode Procedure and Coding. To implement PSA in a repeatable manner, we developed a structured episode worksheet [34]. This instrument records the specific context of an episode, including: circumstances, assumptions, policy actions, governing corpus evidence, coverage judgments, gap types, and downstream consequences (Figure 4). It operationalizes the scenario-based logic of PSA and supports consistent cross-episode comparison. The unit of analysis is the episode-by-concern assessment of whether the applicable policy action set provides actionable support through required artifacts, minimum content expectations, or decision gates. Each episode is analyzed in three steps: (1) specify the circumstances and assumptions; (2) identify the SWP policy actions invoked and extract the associated artifact content from the governing corpus; and (3) evaluate the actionability of that content for the concern under review. Following Cunningham’s observation that policy failure often stems from underspecification rather than a total absence of guidance [8], we evaluate actionability using a ternary coding scheme: Explicit, Partial, and Absent. This scheme distinguishes between guidance that is directly actionable, guidance that is relevant but insufficiently specified for repeatable execution, and guidance that is missing at the point of decision. It also aligns with qualitative approaches that distinguish degrees of support or maturity rather than treating adequacy as purely binary [22, 51]. We define these codes as follows: • Explicit: the policy and/or required artifact includes a direct mandate, role assignment, decision trigger, or specific content requirement addressing the property. • Partial: the policy provides a relevant artifact or general software guidance, but lacks the technical specificity, standards, or triggers needed to manage the risk in a repeatable way. • Absent: the policy corpus provides no meaningful guidance for the property at that decision point, leaving resolution to local tailoring or undocumented practice. Codes are assigned conservatively based on the most specific applicable policy language in the governing corpus for the episode and concern under review. In this paper, an Explicit rating does not require that every element of support appear in core SWP guidance alone; rather, it requires that the applicable SWP-centered governance stack provide direct and actionable support for the property at issue. When coverage is Partial or Absent, we additionally classify the gap type (missing artifact content, unclear responsibility assignment, missing decision gate or trigger, or missing contracting or rights mechanism) and record a brief downstream consequence describing the likely execution or sustainment risk if the gap persists. Manuscript submitted to ACM

Is US Defense Acquisition Ready to Acquire AI-Enabled Capabilities?

13

Example Episode Worksheet Episode / concern: Data governance and data rights for the Targeting AI program / whether policy supports ownership and accountability, provenance and labeling expectations, and rights sufficient for training, test, operation, and sustainment Coverage judgment: Partial Gap type: Missing artifact content Circumstances and assumptions: The program must acquire, label, protect, and sustain multi-source data for training, validation, and operation, with future retraining and sustainment expected. Policy actions invoked: Planning Phase artifacts include the Program Protection Plan, IP Strategy, and Data Strategy within the Acquisition Strategy.

• DoDI 5000.83 supports protection of controlled technical information and data assets, helping the program define how sensitive training and operational data should be safeguarded and governed. • DoDI 5010.44 supports negotiation of data and IP rights, helping the program address ownership, access, and reuse rights for the data needed for training, testing, operation, and later retraining. • DoDI 5000.82 and the DoD Data, Analytics, and AI Adoption Strategy support enterprise data governance, helping frame expectations for stewardship, accountability, and management of data across the AI lifecycle. Coverage is partial because these sources support protection, rights-planning, and governance at a general level, but do not specify program-level requirements for data provenance, labeling quality, quality-assurance thresholds, or retraining-ready data deliverables. Downstream consequence: Data may be protected and rights partially secured, yet provenance, labeling quality assurance, versioning, and retrainingready deliverables remain insufficiently specified.

Fig. 4. Episode worksheet developed for this study to operationalize Policy Scenario Analysis in the software acquisition context. The example shows how the data-governance and data-rights episode was coded.

Answering the research question and outputs. We answer the research question by using the episode worksheet to determine whether the governing corpus provides program-actionable guidance for AI-enabled acquisition and to identify where the SWP-centered governance stack remains underspecified. We aggregate episode-level judgments into two outputs: (i) a scenario-to-policy mapping linking SWP steps to governing provisions, required artifacts, and responsible roles or authorities; and (ii) a coverage-and-gaps matrix summarizing Explicit, Partial, and Absent judgments and associated gap types across the artifact-producing guidance in the corpus. Together, these outputs show how well SWP accommodates AI’s distinctive characteristics and where additional specificity is required for consistent execution across program offices. Completeness and stopping criteria. Consistent with PSA, we treat scenarios as decision-support instruments rather than literal predictions and therefore use a small number of representative scenarios designed to reveal hidden weaknesses and their consequences [8]. We consider the assessment complete when (i) all mandatory SWP Planning Phase artifacts, decision points, and phase entry or exit criteria within scope have been enumerated; (ii) each in-scope AI-relevant property is exercised by at least one episode such that it must map to a policy-backed artifact, role, or decision gate; and (iii) further passes through the governing corpus do not introduce additional artifacts, responsibilities, or decision mechanisms that would change the existing mappings. Following scenario-based evaluation guidance, we do not attempt to enumerate all possible operational states. Instead, we select episodes to cover representative conditions the program is likely to encounter and to expose where guidance is technology-agnostic or absent [30]. Method tailoring. Our implementation adapts PSA to the constraints and goals of this study. Rather than constructing alternative futures or using expert elicitation to assign future probabilities, we use a baseline/comparator scenario design and a structured episode worksheet to trace how the applicable policy guidance is translated into roles, artifacts, and Manuscript submitted to ACM

14

Lugo & Davis

gates. This tailoring preserves PSA’s core purpose of surfacing hidden weaknesses and anticipating their consequences, while producing audit-ready, program-actionable findings for acquisition practice [8]. 3.4

Limitations and Threats to Validity

We discuss construct, internal, and external threats to validity as they apply to this scenario-based policy assessment. Because PSA is a decision-support method rather than a predictive evaluation of observed program outcomes, these threats primarily concern how we operationalize acquisition actionability, how we draw inferences from scenario analysis, and how far the findings can reasonably be generalized beyond the study setting. Construct validity. Our central construct is acquisition actionability: whether the SWP-centered governance stack gives a program office sufficiently clear roles, artifacts, content expectations, and decision triggers to manage AI-relevant acquisition burdens. We operationalize this construct through five evaluation concerns and an episode-based PSA worksheet focused on Planning Phase artifacts and gates. This framing is appropriate to our research question, but it does not capture every possible dimension of AI governance or acquisition success. In particular, our study evaluates policy support for planning and governance rather than downstream operational effectiveness, contractor performance, or organizational adoption in practice. Internal validity. Our analysis relies on qualitative judgments when mapping policy language to scenario episodes and when assigning Explicit, Partial, or Absent codes. The coding was performed by a single analyst, and we therefore do not report inter-rater agreement. This reflects the specialized nature of the assessment: it required familiarity with DoD acquisition policy, pathway artifacts, and military acquisition practice, and one author contributed substantially more domain expertise in those areas than the other. We mitigate this limitation through a structured worksheet, explicit code definitions, a bounded Planning Phase scope, and a baseline/comparator design that holds the acquisition framework constant across scenarios. Even so, the findings should be read as structured analytic judgments rather than coder-independent measurements. We discuss this division of expertise further in the positionality statement below (§3.5). External validity. Our results are derived from a synthetic but grounded DoD-like AI baseline, a conventional software comparator, and a governing corpus centered on SWP and associated guidance. The findings are therefore most applicable to software-intensive defense programs in which AI-enabled capabilities have similar deployment requirements. They may be less applicable to simpler enterprise AI deployments such as GenAI.mil (§2.3), autonomous systems outside the scope of our baseline, or acquisition environments outside the DoD. More broadly, PSA does not aim to predict the outcome of any single real program [8]; rather, it provides analytically useful insight into where the current SWP-centered governance stack is likely to be actionable, underspecified, or dependent on undocumented local practice. 3.5

Statement of Positionality

The research team comprises an acquisition professional and a software engineering researcher. This dual perspective informed the study’s design by balancing methodological rigor with operational realism. Our combined expertise led us to prioritize software-centric guidance as the primary analytical lens and ensured that the selected scenarios reflect the practical constraints of the defense acquisition environment. Manuscript submitted to ACM

Is US Defense Acquisition Ready to Acquire AI-Enabled Capabilities?

15

Because the acquisition-policy coding required familiarity with DoD acquisition pathways, artifacts, and practical application, the primary coding judgments were made solely by the author with direct defense-acquisition experience. The findings and interpretations presented herein are strictly those of the authors and do not represent the official positions of any government agency, military department, or commercial vendor. 4

Results

This section presents the episode-level results of applying PSA to the baseline Targeting AI scenario through the structured episode worksheet developed for this study (Figure 4). For each episode, we assess how the SWP-centered governance stack addresses the relevant acquisition concern, identify remaining gaps and their likely consequences, and use the conventional software comparator to distinguish AI-specific underspecification from broader limitations in SWP guidance. 4.1

Summary of Baseline Scenario Findings

Table 3 summarizes the PSA results for the baseline Targeting AI scenario. Overall, the SWP-centered governance stack provides a usable acquisition structure but uneven AI-specific actionability. The strongest support appears in AI assurance and TEVV, where SWP artifacts and supplemental CDAO guidance provide relatively direct direction for planning evaluation activities. In the remaining episodes, the guidance provides a workable structure but leaves important elements of artifact content, review triggers, and decision criteria to local interpretation. 4.2

Data Governance and Data Rights

Finding: Coverage is Partial because the governance stack fails to translate enterprise quality standards into programlevel expectations for training data provenance, labeling QA, or retraining-ready data deliverables. 4.2.1

Baseline Scenario: Targeting AI.

Circumstance. In this episode, the Targeting AI program must acquire, prepare, and sustain the data needed to train, validate, and operate a targeting model that fuses multiple sensor sources and provides targeting recommendations. The scenario assumes that the data requires labeling and quality assurance before use, and must remain available for later retraining, test, and sustainment as mission conditions evolve. These circumstances force the program to address acquisition-time questions of dataset provenance, labeling integrity, version control, protection of sensitive data products, and government rights to use, modify, and sustain data and associated artifacts over time. Policy Actions Invoked. Under the SWP-centered governance stack, the IPT is required to develop a set of planning artifacts that translate pathway guidance into executable program decisions. In this episode, the circumstance brings three required artifacts directly into play: the Program Protection Plan (PPP), the Intellectual Property (IP) Strategy, and the program’s Data Strategy within the Acquisition Strategy. These artifacts are therefore the principal policy actions through which the IPT must address the episode’s data-governance and data-rights issues. The governing documents implicated by this episode include DoDI 5000.83 for program protection, DoDI 5010.44 for intellectual property and data rights, and DoDI 5000.82 together with the DoD Data, Analytics, and AI Adoption Strategy for enterprise data governance [6, 40, 67, 70]. Manuscript submitted to ACM

16

Lugo & Davis

Table 3. Summary of PSA baseline findings for the Targeting AI scenario, assessed at the level of the SWP-centered governance stack. Episode / concern

Coverage

Primary artifact(s)

Key finding

Data governance and data rights (Illustrated Fig. 4)

Partial

Program Protection Plan; IP Strategy; Data Strategy within Acquisition Strategy

Governance stack provides explicit mechanisms for data protection and rights-planning, but does not specify minimum content for provenance, labeling quality, versioning, or the data deliverables needed for retraining and sustainment. (Gap type: missing artifact content)

AI assurance and TEVV

Explicit

Test Strategy; T&E planning artifacts; CDAO AI T&E guidance

Governance stack provides explicit support for AI-specific TEVV when the required Test Strategy is combined with CDAO guidance on stochastic performance criteria, operationalenvelope characterization, and other AIspecific evaluation factors.

Cybersecurity and traceability

Partial

Cybersecurity Strategy within Acquisition Strategy; Program Protection Plan; Cybersecurity Strategy Annex

Governance stack is strong on lifecycle cybersecurity, supply-chain risk management, and recurring assessment, but remains underspecified for end-to-end provenance linking datasets, model artifacts, software builds, dependencies, and deployed system configurations. (Gap type: missing artifact content)

Lifecycle management

Partial

Product Support Strategy; lifecyclesupport content within the Acquisition Strategy; OMB M-25-22

Governance stack provides meaningful support for post-award monitoring and update governance, but drift detection, retraining triggers, and AI-specific performance thresholds remain only partially specified. (Gap type: missing decision gate/trigger)

Human oversight

Partial

Acquisition Strategy; DoD Responsible AI Strategy and Implementation Pathway; human-systems planning artifacts

Governance stack provides meaningful guidance on accountability, intended-use constraints, and governability, but does not yet embed these expectations in pathway-specific artifact sections, review gates, or standard decision criteria. (Gap type: missing decision gate/trigger)

Assessment. At the episode-by-concern level, coverage for data governance and data rights is Partial. When evaluated against our concern whether policy supports ownership and accountability, provenance and labeling expectations, and rights sufficient for training, test, operation, and sustainment, the governance stack exhibits uneven actionability. The PPP, governed primarily by DoDI 5000.83, provides an explicit mechanism for identifying and protecting controlled technical information and mission-critical program information across the lifecycle [70]. In the Targeting AI case, this mechanism can be applied to training datasets, model weights, and feature pipelines as sensitive program assets. However, the guidance is oriented toward preventing exfiltration or compromise rather than ensuring the usability, integrity, and technical reliability of AI data. It therefore provides an explicit basis for protection, but only indirect support for broader AI data governance. The IP Strategy provides more direct support. DoDI 5010.44 explicitly directs programs to plan early for the intellectual property and data rights needed to keep systems functional, affordable, and supportable over time [67]. For the Targeting Manuscript submitted to ACM

Is US Defense Acquisition Ready to Acquire AI-Enabled Capabilities?

17

AI program, this gives the IPT a clear policy basis to identify and negotiate rights for training datasets, documentation, model-supporting technical data, and related deliverables needed for sustainment. In that sense, coverage is explicit: the policy clearly requires early planning for the rights the government will need to use, modify, and support the capability over time. The program’s Data Strategy, developed as part of the Acquisition Strategy, draws on DoDI 5000.82 and the DoD Data, Analytics, and AI Adoption Strategy to provide a meaningful enterprise-level basis for data governance. In particular, this guidance treats data as a product, emphasizes lifecycle assessment using VAULTIS (Visible, Accessible, Understandable, Linked, Trustworthy, Interoperable, and Secure) [6, 40] and related data-quality dimensions, and assigns responsibility to data domain owners and data product teams. Despite these high-level principles, coverage is coded Partial for program-actionable acquisition guidance because provenance and labeling expectations are left to local implementation and are not translated into standard minimum content for SWP planning artifacts. Consequently, the pathway currently lacks the technical specificity needed to manage these AI-relevant properties in a repeatable, actionable way. Gap and Downstream Consequence. The gap in this episode is missing artifact content. The governing documents do not clearly require planning artifacts to define training data provenance, labeling quality-assurance thresholds, or dataset versioning rules. As a result, a program may secure enough protection and rights to field an initial capability while still lacking the metadata and governance needed to sustain model performance. Without clear responsibility for producing and maintaining labels and dataset documentation, retraining may become infeasible in practice. 4.2.2

Comparator Scenario: Flight Planning Optimizer.

Episode Summary and Assessment. The comparator episode involves a similar acquisition problem: ensuring that the flight-planning system has continuing access to the data it depends on, including aircraft performance data, weather products, airspace restrictions, and related mission-planning inputs. The same general planning artifacts are implicated, particularly the Program Protection Plan, the Intellectual Property Strategy, and the data-related portions of the Acquisition Strategy. In this case, however, the relevant inputs are established operational data products with designated owners, stewardship mechanisms, and update processes. As a result, the SWP-centered governance stack is generally sufficient to document access, protection, and sustainment arrangements without introducing the additional governance burden associated with training and evaluation datasets. Comparative Interpretation. The comparator shows that SWP is generally adequate when a program depends on established operational data with known stewardship and update processes. The burden becomes more demanding in the Targeting AI baseline because training and evaluation datasets must themselves be governed as enduring elements of the capability, including their provenance, labeling, reuse rights, and retraining suitability. 4.3

AI Assurance and TEVV

Finding: Coverage is Explicit because the required SWP testing structure, when combined with CDAO AI T&E guidance, directly supports operational-envelope characterization and AI-specific TEVV. 4.3.1

Baseline Scenario: Targeting AI.

Circumstance. In this episode, the Targeting AI program must define how the system will be tested, evaluated, verified, and validated before and during iterative delivery. The program must assess a model-based targeting capability whose Manuscript submitted to ACM

18

Lugo & Davis

behavior may vary across operational and environmental conditions, including cases in which performance degrades or fails unpredictably. These circumstances raise acquisition-time questions about what constitutes acceptable model performance, how to characterize behavior across an operational envelope, and what forms of evidence are sufficient to support release and continued operational use. Policy Actions Invoked. Within the SWP-centered governance stack, the IPT is required to produce a Test Strategy during the Planning Phase. In this episode, the Test Strategy and related T&E planning artifacts are therefore the principal documents through which the program must translate AI-assurance needs into executable test planning. The governing policy documents implicated by this episode are anchored in DoDI 5000.87, which establishes SWP testing expectations, and DoDI 5000.89, which provides the Department’s functional policy for test and evaluation and the framework used to develop the Test Strategy. They also include supplemental AI-specific T&E guidance from CDAO that bears directly on how those expectations can be operationalized for AI-enabled systems [7, 39, 43]. Assessment. At the episode-by-concern level, coverage for AI assurance and TEVV is Explicit. DoDI 5000.87 and DoDI 5000.89 provide the required planning structure and baseline software T&E expectations, while supplemental CDAO AI T&E guidance supplies the more specific evaluation content needed to assess non-deterministic behavior across operational conditions. Taken together, these documents provide direct support for defining acceptable model performance, characterizing behavior across an operational envelope, and specifying evidence appropriate for release and continued operational use. DoDI 5000.87 describes the Test Strategy as the program’s blueprint for how capabilities, features, and user stories will be tested to satisfy developmental test and evaluation criteria and to demonstrate operational effectiveness, suitability, interoperability, survivability, and cyber survivability throughout the software lifecycle [43]. Likewise, DoDI 5000.89 emphasizes continuous integration, automated data collection, and reusable test scripts, tools, and libraries to support iterative evaluation [39]. However, these core SWP documents are less specific on AI-tailored acceptance logic, such as stochastic performance criteria, operational-envelope characterization, and evaluation under distributional or environmental variation. Those more specific expectations are supplied by the supplemental CDAO Test and Evaluation Strategy Framework and companion AI T&E guidance, which explicitly address AI-specific metrics, operational-condition characterization, and related evaluation factors in acquisition-relevant terms [7]. As a result, coverage for AI TEVV is explicit when this supplemental guidance is incorporated in our governance stack. Gap and Downstream Consequence. Because the SWP-centered governance stack provides explicit support for AIspecific TEVV when read as a whole, the weakness in this episode is incomplete integration of that guidance into core SWP testing expectations. The AI-specific content needed for strong TEVV planning is not fully embedded in the core SWP testing guidance itself. As a result, program offices that do not identify and apply the relevant supplemental guidance may still default to narrower or overly deterministic test logic, creating variability in implementation quality across programs. 4.3.2

Comparator Scenario: Flight Planning Optimizer.

Episode Summary and Assessment. In the comparator episode, the program team must define how the flight-planning capability will be tested before and during iterative delivery. The program must verify and validate a deterministic Manuscript submitted to ACM

Is US Defense Acquisition Ready to Acquire AI-Enabled Capabilities?

19

optimization application operating over known software logic, authoritative data inputs, and representative missionplanning conditions. The same general T&E artifacts are implicated, particularly the Test Strategy and related planning mechanisms under DoDI 5000.87 and DoDI 5000.89. Because the comparator system does not depend on learned model behavior, drift-sensitive inference, or stochastic outputs, the core SWP and functional T&E guidance are generally sufficient to support regression testing, scenario-based evaluation, and system-level verification. Comparative Interpretation. The comparator shows that SWP is generally adequate when the capability can be evaluated through established software-verification and mission-scenario methods. In the Targeting AI baseline, acceptable performance must also be defined across stochastic behavior and varying operational conditions, which makes supplemental AI-specific TEVV guidance more important to consistent program execution. 4.4

Cybersecurity and Traceability

Finding: Coverage is Partial; while lifecycle security is strong, the governance stack does not require release-level provenance linking datasets, models, dependencies, software builds, and deployed configurations. 4.4.1

Baseline Scenario: Targeting AI.

Circumstance. In this episode, the Targeting AI program must prepare to release and sustain an AI-enabled targeting capability in a contested environment while maintaining confidence in both its cybersecurity posture and the provenance of each deployed release. The delivered capability depends on an evolving combination of code, model artifacts, data products, dependencies, configuration settings, and deployment pipelines. These conditions raise acquisition-time questions about how to secure development and deployment environments, manage supply-chain risk, and preserve traceability from a fielded release back to the specific data, model, weights, software, and configuration elements used to produce it. Policy Actions Invoked. Within the SWP-centered governance stack, the IPT is required to address cybersecurity during planning through the Cybersecurity Strategy section of the Acquisition Strategy and through the Program Protection Plan (PPP) and associated Cybersecurity Strategy Annex. DoDI 5000.87 establishes cybersecurity as a continuous lifecycle obligation and requires a risk-based approach to be embedded across design, infrastructure, development, test, integration, delivery, and operations [43]. It further requires recurring assessment of the supply chain, development environment, processes, and tools, together with continuous automated cybersecurity testing and operational evaluation. DoDI 5000.90 strengthens these requirements by assigning explicit PM accountability for cybersecurity across acquisition stages, requiring identification and documentation of breach consequences, and enforcing cybersecurity through RMF and supply-chain risk management (SCRM) mechanisms [41]. These artifacts and responsibilities are therefore the principal policy actions through which the IPT must address cybersecurity and release provenance in this episode. Assessment. Coverage for cybersecurity and traceability is Partial. When evaluated against our concern whether policy supports model and pipeline security (including supply chain) and traceability across data, model, software, and system releases, the governance stack exhibits uneven actionability. For cybersecurity, the guidance is comparatively strong. DoDI 5000.87 and DoDI 5000.90 clearly require programs to plan for cybersecurity across the lifecycle, assess development and supply-chain risks, and maintain continuous security monitoring and testing [41, 43]. In the Targeting AI case, this gives the IPT a clear policy basis to secure the Manuscript submitted to ACM

20

Lugo & Davis

development pipeline, assess tools and dependencies, document breach consequences, and integrate cybersecurity into release planning and operational use. The weaker area is traceability, particularly for AI supply-chain provenance. While DoDI 5000.87 mandates continuous cybersecurity practices, it does not translate these into AI-specific requirements for dataset integrity, model artifact protection, or vetting of third-party foundation models. Nor does it clearly require a release-level provenance artifact linking data versions, training configurations, model weights, dependencies, and deployed system configurations. In the Targeting AI case, that omission creates a critical visibility gap: the IPT may be unable to determine whether a targeting failure stems from a code exploit, a data-poisoning attack, model compromise, or an untrusted upstream dependency. Coverage for AI-specific traceability and supply-chain provenance is coded as Partial. Gap and Downstream Consequence. The gap in this episode is missing artifact content. The governing documents do not clearly require planning artifacts to define how dataset versions, training configurations, model artifacts, dependencies, and deployed system configurations will be linked and preserved across releases. If these linkages are not established during planning, the program may be able to secure and deploy the capability while still lacking the traceability needed to investigate failures, support forensic analysis, manage retraining-related supply-chain risk, or confidently roll back to a known-good configuration. 4.4.2

Comparator Scenario: Flight Planning Optimizer.

Episode Summary and Assessment. In the comparator episode, the program office must secure the flight-planning capability and maintain confidence in released software versions. The program must protect conventional software builds, interfaces, dependencies, and operational data feeds and ensure that released versions can be traced through ordinary software configuration-management practices. The same policy actions are implicated, particularly the Cybersecurity Strategy, PPP, and associated RMF artifacts. Because the comparator system relies on authoritative data products rather than trained model artifacts, conventional configuration control and software-assurance practices are generally sufficient to preserve release provenance. Comparative Interpretation. The comparator shows that SWP is generally adequate when release provenance can be maintained through conventional software configuration management and established cybersecurity controls. In the Targeting AI baseline, the provenance problem extends across data, models, weights, dependencies, and deployment configurations, which makes the current lack of explicit AI-specific traceability expectations more consequential. 4.5

Lifecycle Management

Finding: Coverage for this episode is coded as Partial because the governance stack does not embed clear drift-detection, retraining, and re-authorization triggers. 4.5.1

Baseline Scenario: Targeting AI.

Circumstance. In this episode, the program team must plan for how the capability will be sustained, monitored, and updated after initial fielding. Operational performance may change over time as data distributions shift, environments evolve, and models are retrained or replaced. These conditions raise acquisition-time questions about who is responsible for post-deployment monitoring, what evidence should trigger retraining or update decisions, and how material model changes should be reviewed, approved, and re-authorized during iterative execution. Manuscript submitted to ACM

Is US Defense Acquisition Ready to Acquire AI-Enabled Capabilities?

21

Policy Actions Invoked. Within the SWP-centered governance stack, the IPT is required to develop a Product Support Strategy during the Planning Phase to define how the capability will remain functional, supportable, and affordable across the lifecycle. In software acquisition practice, this artifact may appear as a stand-alone Product Support Strategy or as lifecycle-support content within the Acquisition Strategy [13, 14]. The governing policy documents implicated by this episode therefore include both the product-support guidance associated with SWP and the supplemental AI-relevant acquisition guidance that informs post-award AI governance. In particular, OMB Memorandum M-25-22 is included because it provides acquisition-relevant direction on post-award AI monitoring, evaluation, and update governance [78]. Assessment. Overall coverage for lifecycle management is Partial. When evaluated against our concern whether policy supports monitoring, update and change-control responsibilities, and triggers for retraining, drift detection, or re-authorization following material changes, the governance stack exhibits uneven actionability. The Product Support Strategy provides explicit mechanisms for the lifecycle sustainment of software systems, aligning with modern iterative practice by requiring plans for continuous delivery, software versioning, and long-term sustainment [13, 14]. Furthermore, the governance stack is materially strengthened by OMB Memorandum M-2522, which provides acquisition-relevant mechanisms for governing AI systems after award. It explicitly supports contractual terms for ongoing testing and monitoring, recommends use of agency-controlled validation or testing datasets representative of deployed conditions, and provides policy-backed hooks for vendor access, independent evaluation, update governance, rollback expectations, and performance standards that new versions must satisfy before deployment [78]. For the Targeting AI scenario, these provisions make post-deployment monitoring and update governance sufficiently direct and actionable to justify an Explicit sub-assessment within the broader episode. The primary limitation for the episode is the lack of drift detection as a standardized technical control. While OMB M-25-22 mandates ongoing monitoring, it provides no baseline for identifying distributional shift or determining specific retraining thresholds. As such, the SWP-centered governance stack facilitates lifecycle administration but provides only Partial support for the automated or statistical detection of performance degradation. Gap and Downstream Consequence. The primary shortfall in this episode is missing decision gates and triggers. Because AI-specific monitoring and re-authorization expectations are not embedded in core SWP guidance, they must be assembled from a fragmented collection of supplemental documents. This creates a risk of procedural flattening, where the IPT may treat material model changes as ordinary software updates, bypassing the rigorous performance-integrity checks required for non-deterministic systems. Even with the full governance stack, the lack of standardized driftdetection criteria leaves program offices with persistent uncertainty regarding when declining operational performance necessitates formal intervention. 4.5.2

Comparator Scenario: Flight Planning Optimizer.

Episode Summary and Assessment. The comparator episode involves a similar acquisition problem: planning for longterm sustainment and iterative updates of the flight-planning capability. In that case, lifecycle management primarily concerns software maintenance, backlog refinement, and interface stability for deterministic functionality. While this includes managing external data providers, the focus remains on ensuring API uptime and schema consistency rather than monitoring the underlying distribution of the data itself. Integration efforts are centered on building interfaces to known providers, where update cadences are predictable and rarely require full re-validation of the system’s core logic. Consequently, the same general product-support artifacts are invoked, and the SWP-centered governance stack is sufficient to govern these activities. Manuscript submitted to ACM

22

Lugo & Davis Comparative Interpretation. The comparator shows that SWP is generally adequate when sustainment centers

on conventional software maintenance, versioning, and interface stability. In the Targeting AI baseline, lifecycle management must also govern monitoring, retraining, and re-authorization as performance conditions evolve, making the absence of clear decision gates and drift-detection triggers harder for program offices to manage consistently. 4.6

Human Oversight

Finding: Coverage for human oversight is coded as Partial due to a structural integration gap; while high-level Responsible AI (RAI) guidance is extensive, the pathway fails to operationalize these principles into mandatory Planning Phase artifacts or review gates. 4.6.1

Baseline Scenario: Targeting AI.

Circumstance. In this episode, the Targeting AI program must plan for the human oversight, use constraints, and governance safeguards needed to deploy an AI-enabled targeting capability in an operationally consequential setting. The system’s outputs may shape targeting decisions under time pressure and uncertain conditions. These conditions raise acquisition-time questions about who is accountable for operational use, how intended use and limitations will be defined, what evidence is needed to support responsible deployment, and what mechanisms must exist for human supervision, override, or disengagement when the system behaves unexpectedly. Policy Actions Invoked. The core SWP guidance does not include a dedicated pathway artifact specifically for AI human oversight. This requirement instead comes from responsible-AI guidance within the broader SWP-centered governance stack, particularly the DoD AI Ethical Principles, the 2021 implementation memorandum, and the 2022 Responsible AI (RAI) Strategy and Implementation Pathway [15, 68, 71]. Together, these documents provide the principal policy basis for translating high-level commitments into acquisition-relevant governance expectations. Assessment. At the episode-by-concern level, coverage for human oversight is Partial. When evaluated against our concern whether policy supports human oversight and responsible or ethical compliance mechanisms appropriate to the use context including accountability, override, and disengagement, the governance stack exhibits a significant translation failure. The SWP-centered governance stack provides substantial direction through the DoD AI Ethical Principles and the RAI Strategy and Implementation Pathway. These documents establish operational expectations for accountability, use constraints, and human-in-the-loop mechanisms [15, 68, 71]. However, support for human oversight within the core pathway is coded as Partial because the SWP does not translate these high-level commitments into a dedicated artifact requirement or standard decision criterion. In the Targeting AI case, this omission means that while the Program Manager is subject to RAI policy, there is no fixed mechanism within DoDI 5000.87 to verify that meaningful human oversight is technically enabled before entering the Execution Phase. The principal weakness is structural: the pathway lacks the connective tissue required to move RAI from a statement of intent to an actionable acquisition control. Gap and Downstream Consequence. The gap in this episode is the lack of structural integration of human-oversight expectations into core pathway artifacts and review gates. Because SWP does not mandate an Oversight Plan or specific RAI review gates, program actionability depends entirely on local tailoring and supplemental guidance. The downstream consequence for high-consequence systems like Targeting AI is the risk of deploying a capability that lacks documented Manuscript submitted to ACM

Is US Defense Acquisition Ready to Acquire AI-Enabled Capabilities?

23

use limits or explicit technical mechanisms for human override, potentially leading to unintended system consequences in a mission environment. 4.6.2

Comparator Scenario: Flight Planning Optimizer.

Episode Summary and Assessment. In the comparator, the program office must ensure that the flight-planning capability remains subject to operator review. Human review in this case is centered on functional verification: confirming that deterministic software has correctly calculated a route using fixed inputs such as fuel, weight, and weather. Because the logic is transparent and the outputs are reproducible, the core SWP guidance is generally sufficient to support effective human oversight through ordinary usability, verification, and correctness-oriented planning artifacts. Comparative Interpretation. For the flight-planning comparator, operator review is mainly a matter of checking correct execution against transparent logic and known inputs. In Targeting AI, operator involvement carries a broader governance function: users must supervise probabilistic recommendations, recognize when the system is operating outside expected conditions, and retain meaningful authority to override or disengage. Those demands are not clearly embedded in standard SWP artifacts. 5

Discussion

The discussion proceeds in three parts. In §5.1, we identify the core actionability problem in the SWP-centered governance stack. In §5.2, we consider what DoD could do about that problem within SWP, including both an AIsupporting sub-path and targeted artifact refinements. Finally, in §5.3, we consider broader implications beyond SWP reform, including institutionalizing lessons learned and drawing comparative insights from external AI procurement frameworks.. 5.1

Actionability Problem in the SWP-Centered Governance Stack

The episode findings in sections 4.2 to 4.6 show that the central problem in the SWP-centered governance stack is not the total absence of AI-relevant policy, but the inconsistent translation of that policy into core pathway guidance. Across the baseline episodes, acquisition teams can often locate relevant guidance somewhere in the broader corpus, yet still lack sufficiently explicit direction on what must be planned, documented, or reviewed at the program level. The result is a recurring actionability problem: AI-relevant controls are often present in principle, but not consistently embedded in the program-facing mechanisms through which SWP is executed. For example, while enterprise-level data strategies exist, they are not translated into specific SWP deliverables for dataset labeling or versioning (§4.2). This matters because tailoring is not equally sufficient across capability types. In the conventional software comparator, established software artifacts and review practices provided a sufficient governance floor. In the Targeting AI baseline, by contrast, acquisition success depended on AI-specific lifecycle controls that were not yet fully embedded in core SWP artifacts. A concrete example is model drift (see §4.5): in principle, a program could address it through local tailoring by specifying monitoring metrics, retraining triggers, and rollback criteria. In practice, however, if pathway guidance does not make those expectations visible, the government must already know to ask for them; otherwise, programs are more likely to rely on contractor-defined approaches or local expertise. Because AI practices are still maturing, they necessitate more prescriptive guidance than conventional software to effectively institutionalize emerging capabilities [5]. Without a solid governance floor, effective controls remain siloed in individual program offices rather than becoming repeatable acquisition norms. Manuscript submitted to ACM

24

Lugo & Davis

5.2

How SWP Could Better Support the Acquisition of AI-Enabled Capabilities

The findings of our Policy Scenario Assessment (§5.1) suggest that strengthening SWP support for AI acquisition requires more explicit translation of AI-relevant expectations into pathway structure and program-facing artifacts. Within SWP itself, two responses appear especially important: providing a clearer institutional home for AI-enabled programs and refining the planning artifacts through which AI-relevant controls are implemented. The following subsections consider each response in turn. 5.2.1 An AI-Supporting Sub-Path Within SWP. DoDI 5000.87 currently organizes SWP around applications and embedded software sub-paths, distinguishing software primarily by where and how it is delivered. The actionability problem identified in §5.1, however, suggests that some AI-enabled programs are not fully served by either sub-path as currently structured. Their acquisition burdens arise not only from deployment context, but from the need to plan, document, and sustain AI-specific controls that conventional software programs may not require. One possible institutional response would be an AI-supporting sub-path within SWP that makes these recurring expectations explicit through defined artifacts, content requirements, and review logic. Such a sub-path would provide a clearer structural home for the controls this study found to be underspecified and reduce the burden on program offices to assemble and interpret guidance across multiple documents. This issue becomes more significant as DoD investment in AI-enabled capabilities grows. As discussed in §2.2, recent budget materials show a substantial increase from earlier AI-specific funding to a much larger autonomy and AI portfolio concentrated in mission areas such as aerial autonomy, maritime systems, and AI-enabled infrastructure. For some of these acquisitions, the kinds of data, assurance, and lifecycle controls examined in this study are likely to be central to acquisition success. A more explicit pathway for AI-enabled programs would help ensure that those expectations are treated as standard acquisition concerns rather than as local tailoring. 5.2.2 Targeted Artifact Refinements for AI Acquisition. If SWP were to establish an AI-supporting sub-path, that subpath would need to define more explicit expectations for the planning artifacts through which AI-relevant controls are translated into program practice. The baseline findings in sections 4.2 and 4.4 to 4.6 identify several areas where current SWP guidance remains underspecified for AI-enabled systems. The following refinements illustrate the kinds of artifact-level expectations that an AI-supporting sub-path could formalize. AI-Enhanced Data Strategy: The data-governance findings show that current planning artifacts do not specify provenance, labeling quality, versioning, or sustainment-oriented data deliverables (§4.2). The Data Strategy should therefore move beyond general references to stewardship by requiring an AI Data Annex that specifies expectations for training and evaluation datasets. Model-Based Product Support Strategy: The lifecycle-management findings show that monitoring, retraining, and re-authorization triggers remain only partially embedded in current sustainment guidance (§4.5). For AI-capability acquisition programs, the Product Support Strategy (PSS) guidance should therefore be updated to require that planning artifacts explicitly define the triggers for monitoring, retraining, rollback, and re-authorization. Integrated Cybersecurity and Human Oversight: The findings for cybersecurity and human oversight reveal that AI release provenance and responsible-AI (RAI) expectations are not yet effectively operationalized in standard artifacts or review gates (sections 4.4 and 4.6). For instance, the lack of a mandatory provenance chain linking specific data versions to model weights creates a critical visibility gap (§4.4). To address this, these expectations should be codified as mandatory content within the AI sub-path. Manuscript submitted to ACM

Is US Defense Acquisition Ready to Acquire AI-Enabled Capabilities?

25

These refinements illustrate the kinds of expectations that an AI-supporting sub-path within SWP could formalize. By specifying where AI-relevant controls belong in planning artifacts and how they should be reviewed, such a sub-path would make AI-enabled acquisition more consistent within the existing SWP structure. 5.3

Implications Beyond SWP Reform

The recommendations above focus on reforms that could be implemented within SWP itself. The findings also suggest broader implications for how AI acquisition governance may need to evolve over time. In particular, they raise questions about how DoD institutionalizes lessons from program experience and how developments in external AI procurement frameworks may inform future acquisition practice. 5.3.1 Institutionalizing Lessons Learned for AI Acquisition. As discussed in §5.1, local tailoring is not inherently undesirable; it is one way SWP accommodates variation across programs. The concern is that when AI-enabled acquisition depends too heavily on program-level expertise and interpretation, effective solutions may remain siloed within individual offices rather than becoming visible and reusable across the Department. For that reason, improving SWP for AI-enabled acquisition should involve more than a lessons-learned repository or reporting requirement. Recent evidence from high-reliability software organizations suggests that when lessons learned remain informal, ad hoc, or weakly tied to operational decision points, failures recur and experience remains fragmented [1]. A more effective approach would be to institutionalize lessons learned at defined pathway review points so that emerging practices are assessed and incorporated into future acquisition decisions. Such a function could be housed within organizations such as CDAO or DAU to help ensure that program experience informs future acquisition norms rather than remaining isolated within individual efforts. 5.3.2 Comparative Notes from External AI Procurement Frameworks. As discussed in §1, governments are developing AI procurement guidance that makes expectations for testing, transparency, and lifecycle oversight more explicit [16, 18]. This trend is reflected in binding regulations, procurement-oriented guidance, and statutory frameworks that regulate vendor obligations, including the EU AI Act and South Korea’s AI statutory frameworks [18, 35]. For the DoD, these developments offer useful examples of how broad AI-governance principles can be translated into procurement-facing requirements. By observing how other public-sector buyers formalize those expectations, the Department can refine the technical gates and oversight mechanisms within SWP and make AI-enabled acquisition more explicit and repeatable. These comparative developments may also carry implications beyond the Department, because U.S. defense acquisition approaches often influence allied reform efforts. Recent partner-nation modernization work has used U.S. acquisition practice as a reference point for structuring software and rapid capability acquisition [24]. As a result, a clearer SWP framework for AI-enabled capabilities could matter not only for DoD programs, but also as a more legible and actionable model for partner nations seeking to modernize their own defense acquisition processes. 6

Conclusion

This paper asked: To what extent does the Software Acquisition Pathway (SWP), as implemented through the broader SWP-centered governance stack, accommodate AI’s distinctive technical and operational characteristics, and where does it remain underspecified? Our analysis indicates that the current governance structure provides a workable foundation for acquiring AI-enabled capabilities, but remains uneven in how it translates high-level AI policy into program-facing acquisition practice. The core weakness identified in this study is a policy-to-artifact gap: although AI-relevant controls exist within the broader governance stack, they are not consistently translated into artifacts, content expectations, Manuscript submitted to ACM

26

Lugo & Davis

and review logic needed to govern programs effectively under SWP. This gap can be bridged by establishing a clear institutional home for AI acquisition programs within SWP, together with explicit artifact-level requirements for the AI-relevant controls that acquisition teams must plan, document, and review. Future Work. This study is intentionally bounded to the SWP’s Planning Phase governance and two representative scenarios. Future work should examine how these planning-phase gaps manifest during execution, including model updates, release decisions, and operational monitoring under continuous delivery conditions. It should also compare formal pathway expectations with real acquisition programs, contracts, and solicitations to assess how program offices instantiate AI governance in practice. Acknowledgments The views expressed in this article are those of the authors and do not reflect the official policy or position of the United States Space Force, Department of the Air Force, Department of Defense, or the U.S. Government.

Manuscript submitted to ACM

Is US Defense Acquisition Ready to Acquire AI-Enabled Capabilities?

27

References [1] Dharun Anandayuvaraj, Tanmay Singla, Zain A. H. Hammadeh, Andreas Lund, Alexandra Holloway, and James C. Davis. 2026. Learning From Software Failures: A Case Study at a National Space Research Center. In Proceedings of the 2026 IEEE/ACM 48th International Conference on Software Engineering (ICSE ’26). ACM, Rio de Janeiro, Brazil. doi:10.1145/3744916.3773149 [2] Denis Baylor, Eric Breck, Heng-Tze Cheng, Noah Fiedel, Chuan Yu Foo, Zakaria Haque, Salem Haykal, Mustafa Ispir, Vihan Jain, Levent Koc, Chiu Yuen Koo, Lukasz Lew, Clemens Mewald, Akshay Naresh Modi, Neoklis Polyzotis, Sukriti Ramesh, Sudip Roy, Steven Euijong Whang, Martin Wicke, Jarek Wilkiewicz, Xin Zhang, and Martin Zinkevich. 2017. TFX: A TensorFlow-Based Production-Scale Machine Learning Platform. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 1387–1395. doi:10.1145/3097983.3098021 [3] Defense Innovation Board. 2019. Software Is Never Done: Refactoring the Acquisition Code for Competitive Advantage. Technical Report. U.S. Department of Defense. https://media.defense.gov/2019/Apr/30/2002124828/-1/-1/0/SOFTWAREISNEVERDONE_ REFACTORINGTHEACQUISITIONCODEFORCOMPETITIVEADVANTAGE_FINAL.SWAP.REPORT.PDF [4] Eric Breck, Shanqing Cai, Eric Nielsen, Michael Salib, and D. Sculley. 2017. The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction. In 2017 IEEE International Conference on Big Data (Big Data). 1123–1132. doi:10.1109/BigData.2017.8258038 [5] Nils Brunsson and Bengt Jacobsson. 2000. A World of Standards. Oxford University Press. [6] Chief Digital and Artificial Intelligence Office. 2023. 2023 Data, Analytics, and Artificial Intelligence Adoption Strategy. https://media.defense.gov/ 2023/Nov/02/2003333300/-1/-1/1/dod_data_analytics_ai_adoption_strategy.pdf. [7] Chief Digital and Artificial Intelligence Office (CDAO). 2024. Systems Integration Test and Evaluation of Artificial Intelligence-Enabled Capabilities: What to Consider in a Test & Evaluation Strategy. CDAO Test and Evaluation Strategy Framework (Guidance and Best Practices). U.S. Department of Defense. https://www.ai.mil/Portals/137/Documents/Resources%20Page/CDAO_TE_Framework_SI_TES_RELEASED_APRIL_2024-compressed.pdf [8] Stephen Cunningham. 2016. Policy Scenario Analysis: A Methodological Overview. Technical Report BOBP/WB/OPP/REP 20. Bay of Bengal Programme Inter-Governmental Organisation, Chennai, Tamil Nadu, India. [9] Defense Acquisition University. [n. d.]. Software Acquisition. https://aaf.dau.edu/aaf/software/. [10] Defense Acquisition University. 2025. Acquisition Policies | Adaptive Acquisition Framework. https://aaf.dau.edu/policy/ [11] Defense Acquisition University. 2026. Adaptive Acquisition Framework (AAF) Overview. https://aaf.dau.edu/ [12] Defense Acquisition University. n.d.. Develop Strategies (Software Acquisition Pathway). https://aaf.dau.edu/aaf/software/develop-strategies/ [13] Defense Acquisition University (DAU). 2022. Software Acquisition Strategy: Agile Guidance. https://aaf.dau.edu/storage/2022/06/AcquisitionStrategy-Guidance.pdf. [14] Department of the Air Force. 2024. Department of the Air Force Instruction 63-101/20-101, Integrated Life Cycle Management. https://static.epublishing.af.mil/production/1/saf_aq/publication/dafi63-101_20-101/dafi63-101_20-101.pdf. Implementing Responsible Artificial Intelligence in the Department of Defense. Memoran[15] Deputy Secretary of Defense. 2021. dum. https://media.defense.gov/2021/May/27/2002730593/-1/-1/0/IMPLEMENTING-RESPONSIBLE-ARTIFICIAL-INTELLIGENCE-IN-THEDEPARTMENT-OF-DEFENSE.PDF [16] Digital Agency of Japan. 2025. Guidelines for Japanese Government’s Procurement and Utilization of Generative AI (Provisional Translation). https://www.digital.go.jp/assets/contents/node/basic_page/field_ref_resources/e2a06143-ed29-4f1d-9c31-0f06fca67afc/6e45a64f/20250527_ resources_standard_guidelines_guideline_04.pdf [17] Jeffrey Dunlap. 2024. Innovation in Software Acquisition: The Good, Bad, and Ugly. In Proceedings of the Twenty-First Annual Acquisition Research Symposium. Monterey, CA. https://dair.nps.edu/bitstream/123456789/5136/1/SYM-AM-24-073.pdf [18] European Parliament and Council of the European Union. 2024. Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union (July 2024). https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng [19] Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2021. Datasheets for Datasets. Commun. ACM 64, 12 (2021), 86–92. doi:10.1145/3458723 [20] Google Cloud. 2025. Chief Digital and Artificial Intelligence Office Selects Google Cloud’s AI to Power GenAI.mil. https://www. googlecloudpresscorner.com/2025-12-09-Chief-Digital-and-Artificial-Intelligence-Office-Selects-Google-Clouds-AI-to-Power-GenAI-mil. [21] Google Cloud Architecture Center. 2024. MLOps: Continuous Delivery and Automation Pipelines in Machine Learning. https://cloud.google.com/ architecture/mlops-continuous-delivery-and-automation-pipelines-in-machine-learning. [22] ISACA. 2024. CMMI Levels of Capability and Performance. https://cmmiinstitute.com/learning/appraisals/levels. [23] Purvish Jajal, Wenxin Jiang, Arav Tewari, Erik Kocinare, Joseph Woo, Anusha Sarraf, Yung-Hsiang Lu, George K. Thiruvathukal, and James C. Davis. 2024. Interoperability in Deep Learning: A User Survey and Failure Analysis of ONNX Model Converters. In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (Vienna, Austria) (ISSTA 2024). Association for Computing Machinery, New York, NY, USA, 1466–1478. doi:10.1145/3650212.3680374 [24] Won-Joon Jang and Hea Ji Park. 2024. Closing the Gap: Modernizing South Korea’s Defense Acquisition Framework. Research Papers 24/6. Korea Institute for Industrial Economics and Trade. https://EconPapers.repec.org/RePEc:ris:kietrp:2024_006 [25] Wenxin Jiang, Nicholas Synovic, Matt Hyatt, Taylor R. Schorlemmer, Rohan Sethi, Yung-Hsiang Lu, George K. Thiruvathukal, and James C. Davis. 2023. An Empirical Study of Pre-Trained Model Reuse in the Hugging Face Deep Learning Model Registry. In Proceedings of the 45th International Manuscript submitted to ACM

28

Lugo & Davis

Conference on Software Engineering (Melbourne, Victoria, Australia) (ICSE ’23). IEEE Press, 2463–2475. doi:10.1109/ICSE48619.2023.00206 [26] Norman P. Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, and Rick et al. Boyle. 2017. In-Datacenter Performance Analysis of a Tensor Processing Unit. In Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA ’17). ACM/IEEE, 1–12. doi:10.1145/3079856.3080246 [27] Andrej Karpathy. 2017. Software 2.0. https://karpathy.medium.com/software-2-0-a64152b37c35. [28] Rainer Kattel and Wolfgang Drechsler. 2022. Public Procurement and Innovation: Theory and Practice. Standardization and Public Policy Review 14, 2 (2022), 125–143. doi:10.1007/978-3-642-40258-6 [29] William A. LaPlante. 2022. Use of the Software Acquisition Pathway for Defense Business Systems. Memorandum USA000671-22. Office of the Under Secretary of Defense for Acquisition and Sustainment, 3010 Defense Pentagon, Washington, DC 20301-3010. https://aaf.dau.edu/storage/2022/09/ USDA-S-signed-Memo-Use-of-the-SWP-for-DBS_20220824.pdf [30] Nik Looker, David Webster, Duncan Russell, and Jie Xu. 2008. Scenario Based Evaluation. In 2008 11th IEEE International Symposium on Object and Component-Oriented Real-Time Distributed Computing (ISORC). IEEE, 148–154. doi:10.1109/ISORC.2008.56 [31] Neal Maslej et al. 2025. Artificial Intelligence Index Report 2025. https://hai.stanford.edu/assets/files/hai_ai_index_report_2025.pdf [32] Sehl Mellouli, Marijn Janssen, and Adegboyega Ojo. 2024. Introduction to the Issue on Artificial Intelligence in the Public Sector: Risks and Benefits of AI for Governments. Digit. Gov.: Res. Pract. 5, 1, Article 1 (March 2024), 6 pages. doi:10.1145/3636550 [33] Microsoft Worldwide Public Sector. 2023. Advancing AI Procurement and Adoption in the Public Sector: Considerations, Use Cases and Practical Approaches. Technical Report. Microsoft Worldwide Public Sector, Redmond, WA. https://wwps.microsoft.com/wp-content/uploads/2023/12/ Microsoft-AI-Procurement-Paper_Final.pdf [34] Matthew B. Miles, A. Michael Huberman, and Johnny Saldaña. 2014. Qualitative Data Analysis: A Methods Sourcebook (3rd ed.). SAGE Publications, Thousand Oaks, CA. [35] Ministry of Government Legislation, Republic of Korea. 2025. Framework Act on the Development of Artificial Intelligence. Technical Report. Ministry of Government Legislation, Republic of Korea. https://cset.georgetown.edu/publication/south-korea-ai-law-2025/ [36] Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model Cards for Model Reporting. In FAT* ’19: Conference on Fairness, Accountability, and Transparency. 220–229. doi:10.1145/ 3287560.3287596 [37] National Institute of Standards and Technology. 2023. AI Risk Management Framework (AI RMF 1.0). Technical Report NIST AI 100-1. National Institute of Standards and Technology. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf [38] Secretary of Defense. 2025. Directing Modern Software Acquisition to Maximize Lethality. Memorandum. U.S. Department of Defense, Washington, DC. https://media.defense.gov/2025/Mar/07/2003662943/-1/-1/1/DIRECTING-MODERN-SOFTWARE-ACQUISITION-TO-MAXIMIZE-LETHALITY. PDF [39] U.S. Department of Defense. 2020. DoD Instruction 5000.89: Test and Evaluation. Technical Report. Office of the Under Secretary of Defense for Research and Engineering and Office of the Director, Operational Test and Evaluation. https://www.esd.whs.mil/Portals/54/Documents/DD/ issuances/dodi/500089p.PDF [40] U.S. Department of Defense. 2023. DoD Instruction 5000.82: Requirements for the Acquisition of Digital Capabilities. https://www.esd.whs.mil/ Portals/54/Documents/DD/issuances/dodi/500082p.pdf. [41] Office of the Under Secretary of Defense for Acquisition and Sustainment. 2020. Cybersecurity for Acquisition Decision Authorities and Program Managers. DoD Instruction 5000.90. U.S. Department of Defense. https://www.esd.whs.mil/Portals/54/Documents/DD/issuances/dodi/500090p.PDF [42] Office of the Under Secretary of Defense for Acquisition and Sustainment. 2020. Operation of the Adaptive Acquisition Framework. DoD Instruction 5000.02. U.S. Department of Defense. https://www.esd.whs.mil/Portals/54/Documents/DD/issuances/dodi/500002p.pdf?ver=2020-01-23-144114-093 [43] Office of the Under Secretary of Defense for Acquisition and Sustainment. 2020. Operation of the Software Acquisition Pathway. DoD Instruction 5000.87. U.S. Department of Defense. https://www.esd.whs.mil/Portals/54/Documents/DD/issuances/dodi/500087p.pdf [44] Office of the Under Secretary of Defense for Research and Engineering. 2020. Engineering of Defense Systems. DoD Instruction 5000.88. U.S. Department of Defense. https://www.esd.whs.mil/Portals/54/Documents/DD/issuances/dodi/500088p.PDF [45] Office of the Director, Operational Test and Evaluation. 2018. Mission Planning System (MPS) / Joint Mission Planning System–Air Force (JMPS-AF). In DOT&E FY 2018 Annual Report. U.S. Department of Defense, 197–198. https://www.dote.osd.mil/Portals/97/pub/reports/FY2018/af/2018afmps. pdf?ver=2019-08-21-155843-447 [46] Office of the Under Secretary of Defense (Comptroller). 2024. Defense Budget Overview, Fiscal Year 2025. Technical Report. Office of the Under Secretary of Defense (Comptroller), Washington, DC. https://comptroller.defense.gov/Budget-Materials/ [47] Office of the Under Secretary of Defense (Comptroller). 2025. Defense Budget Overview, Fiscal Year 2026. Technical Report. Office of the Under Secretary of Defense (Comptroller), Washington, DC. https://comptroller.defense.gov/Budget-Materials/ [48] Office of the Under Secretary of Defense (Comptroller). 2025. FY 2026 Program Acquisition Costs by Weapon System. Technical Report. U.S. Department of Defense. https://comptroller.war.gov/Portals/45/Documents/defbudget/FY2026/FY2026_Weapons.pdf [49] Cheryl Pellerin, Stephen Wood, and Mark Allen. 2022. Artificial Intelligence (AI) and Machine Learning (ML) Acquisition and Policy Implications. Technical Report. Defense Acquisition University (DAU), Fort Belvoir, VA. https://www.dau.edu/library/arj/p/AI-ML-Acquisition-Implications [50] Program Executive Office Intelligence, Electronic Warfare & Sensors. 2026. Project Manager Intelligence Systems & Analytics. https://cpeisw.army. mil/pm-isa/. Manuscript submitted to ACM

Is US Defense Acquisition Ready to Acquire AI-Enabled Capabilities?

29

[51] Charles C. Ragin. 2008. Redesigning Social Inquiry: Fuzzy Sets and Beyond. University of Chicago Press, Chicago. doi:10.7208/chicago/9780226702797. 001.0001 [52] Soumyadeep Roy, Sowmya S. Sundaram, Dominik Wolff, and Niloy Ganguly. 2025. Building Trustworthy AI Models for Medicine: From Theory to Applications. In Proceedings of the Eighteenth ACM International Conference on Web Search and Data Mining (WSDM ’25). Association for Computing Machinery, Hannover, Germany, 1012–1015. doi:10.1145/3701551.3703477 [53] Kelley Sayler, Andrew Hunter, and Paul Scharre. 2023. A New Path Forward: An Analysis of Current AI Software Acquisition Procedures. Technical Report. Center for Strategic and International Studies (CSIS), Washington, D.C. https://www.csis.org/analysis/new-path-forward-ai-software-acquisition [54] D. Sculley, G. Holt, D. Golovin, E. Davydov, T. Phillips, D. Ebner, V. Chaudhary, M. Young, J.-F. Crespo, and D. Dennison. 2015. Hidden Technical Debt in Machine Learning Systems. Advances in Neural Information Processing Systems 28 (2015). [55] Sigstore. 2021. sigstore: Software Signing for Everyone. https://www.sigstore.dev [56] Cristina Silvano, Daniele Ielmini, Fabrizio Ferrandi, Leandro Fiorin, Serena Curzel, Luca Benini, Francesco Conti, Angelo Garofalo, Cristian Zambelli, Enrico Calore, Sebastiano Schifano, Maurizio Palesi, Giuseppe Ascia, Davide Patti, Nicola Petra, Davide De Caro, Luciano Lavagno, Teodoro Urso, Valeria Cardellini, Gian Carlo Cardarilli, Robert Birke, and Stefania Perri. 2025. A Survey on Deep Learning Hardware Accelerators for Heterogeneous HPC Platforms. ACM Comput. Surv. 57, 11, Article 286 (June 2025), 39 pages. doi:10.1145/3729215 [57] SLSA Community. 2023. Supply-chain Levels for Software Artifacts (SLSA) v1.0. https://slsa.dev/spec/v1.0 [58] Monika Steidl, Michael Felderer, and Rudolf Ramler. 2023. The Pipeline for the Continuous Development of Artificial Intelligence Models: Current State of Research and Practice. Journal of Systems and Software 199 (2023), 111615. doi:10.1016/j.jss.2023.111615 [59] David M. Tate and John Bailey. 2022. When is it Feasible (or Desirable) to Use the Software Acquisition Pathway? Technical Report. Acquisition Research Program, Naval Postgraduate School, Monterey, CA. https://www.dair.nps.edu/handle/123456789/4569 [60] Santiago Torres-Arias, Hammad Afzali, Trishank Karthik Kuppusamy, Reza Curtmola, and Justin Cappos. 2019. in-toto: Providing farm-totable guarantees for bits and bytes. In 28th USENIX Security Symposium (USENIX Security 19). USENIX Association, Santa Clara, CA, 1393–1410. https://www.usenix.org/system/files/sec19-torres-arias.pdf [61] U.S. Army Public Affairs. 2025. U.S. Army Awards Enterprise Service Agreement to Enhance Military Readiness and Drive Operational Efficiency. https://www.army.mil/article/287506/u_s_army_awards_enterprise_service_agreement_to_enhance_military_readiness_and_drive_ operational_efficiency. [62] U.S. Army xTech Program. 2025. PEO Demand Signal Slides – Master Deck. https://xtech.army.mil/wp-content/uploads/2025/03/PEO-DemandSignal-Slides-Master-Deck.pdf. [63] U.S. Army xTech Program. 2025. xTechOverwatch Informational Session Presentation Deck. https://xtech.army.mil/wp-content/uploads/2025/04/ xTechOverwatch-Informational-Session-Presentation-Deck.pdf. [64] U.S. Department of Defense. 2008. DoD Instruction 5000.02: Operation of the Defense Acquisition System. Department of Defense Instruction 5000.02. U.S. Department of Defense. https://www.dami.army.pentagon.mil/site/artpc/docs/DoDI%205000_02p.pdf [65] U.S. Department of Defense. 2015. DoD Instruction 5000.02: Operation of the Defense Acquisition System. Department of Defense Instruction 5000.02. U.S. Department of Defense. [66] U.S. Department of Defense. 2017. Establishment of an Algorithmic Warfare Cross-Functional Team (Project Maven). https://dodcio.defense.gov/ Portals/0/Documents/Project%20Maven%20DSD%20Memo%2020170425.pdf [67] U.S. Department of Defense. 2019. DoD Instruction 5010.44: Intellectual Property (IP) Acquisition and Licensing. Technical Report. Office of the Under Secretary of Defense for Acquisition and Sustainment. https://www.esd.whs.mil/Portals/54/Documents/DD/issuances/dodi/501044p.pdf [68] U.S. Department of Defense. 2020. DoD Adopts Ethical Principles for Artificial Intelligence. Defense.gov Release. https://www.defense.gov/ Newsroom/Releases/Release/Article/2091996/dod-adopts-ethical-principles-for-artificial-intelligence/ [69] U.S. Department of Defense. 2020. DoD C3 Modernization Strategy. https://dodcio.defense.gov/Portals/0/Documents/DoD-C3-Strategy.pdf. [70] U.S. Department of Defense. 2020. DoD Instruction 5000.83: Technology and Program Protection to Maintain Technological Advantage. https: //www.esd.whs.mil/Portals/54/Documents/DD/issuances/dodi/500083p.pdf. [71] U.S. Department of Defense. 2022. Responsible Artificial Intelligence Strategy and Implementation Pathway. Strategy and Implementation Pathway AD1215042. Chief Digital and Artificial Intelligence Office (CDAO), Washington, DC. https://media.defense.gov/2022/Jun/22/2003022604/-1/1/0/Department-of-Defense-Responsible-Artificial-Intelligence-Strategy-and-Implementation-Pathway.PDF [72] U.S. Department of Defense. 2025. Department of Defense — Publications (Search: “Software Pathway”, since Jan 2021). https://www.war.gov/News/ Publications/StartDate/2020-01-01/?Search=software+pathway [73] U.S. Department of Defense, United States Special Operations Command. 2025. Fiscal Year 2026 Budget Estimates: Research, Development, Test & Evaluation, Defense-Wide, United States Special Operations Command. https://comptroller.defense.gov/Portals/45/Documents/defbudget/FY2026/ budget_justification/pdfs/03_RDT_and_E/RDTE_SOCOM_PB_2026.pdf. [74] U.S. Federal Chief Information Officers Council. 2025. Resources. https://www.councils.gov/resources/ [75] U.S. Government Accountability Office. 2021. DOD Software Acquisition: Status of and Challenges Related to Reform Efforts. Technical Report GAO-21-105298. U.S. Government Accountability Office, Washington, DC. https://www.gao.gov/products/gao-21-105298 [76] U.S. Government Accountability Office. 2023. Artificial Intelligence: DoD Needs Department-Wide Guidance to Inform Acquisitions. https://www.gao. gov/products/gao-23-105850 Manuscript submitted to ACM

30

Lugo & Davis

[77] Lydia Velazquez-Garcia, Antonio Cedillo-Hernandez, Maria Del Pilar Longar-Blanco, and Eduardo Bustos-Farias. 2025. Enhancing Educational Gamification through AI in Higher Education. In Proceedings of the 2024 16th International Conference on Education Technology and Computers (ICETC ’24). Association for Computing Machinery, New York, NY, USA, 213–218. doi:10.1145/3702163.3702416 [78] Russell T. Vought. 2025. Driving Efficient Acquisition of Artificial Intelligence in Government. Memorandum M-25-22. Office of Management and Budget, Executive Office of the President. https://www.whitehouse.gov/wp-content/uploads/2025/02/M-25-22-Driving-Efficient-Acquisition-ofArtificial-Intelligence-in-Government.pdf [79] Michael Winikoff, John Thangarajah, and Sebastian Rodriguez. 2025. A Scoresheet for Explainable AI. In Proceedings of the 24th International Conference on Autonomous Agents and Multiagent Systems (Detroit, MI, USA) (AAMAS ’25). International Foundation for Autonomous Agents and Multiagent Systems, Richland, SC, 2171–2180. [80] Jiachi Zhang, Wenchao Zhou, and Benjamin E. Ujcich. 2024. Provenance-Enabled Explainable AI. Proc. ACM Manag. Data 2, 6, Article 250 (Dec. 2024), 27 pages. doi:10.1145/3698826

Received 2026; revised 2026; accepted 2026

Manuscript submitted to ACM

Record · ID 266237 · SHA-256 7c23592cd4e61465
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.