E DUCATING THE AGENTIC E NGINEER : C URRICULA , C OLLABORATION , AND C ONTINUOUS L EARNING IN THE AI E RA
Mamdouh Alenezi Saudi Data and Artificial Intelligence (SDAIA) Riyadh, Saudi Arabia
arXiv:2607.29610v1 [cs.SE] 31 Jul 2026
August 3, 2026
A BSTRACT Generative and agentic artificial intelligence (AI) are reconfiguring software and systems engineering from a discipline centered on the human authorship of artifacts toward one centered on the direction, verification, and governance of autonomous systems. This transition demands a new professional archetype—the agentic engineer—whose durable value lies in intent specification, orchestration of multi-agent workflows, critical evaluation of machine-generated outputs, and accountable ethical judgment. This article presents an integrative conceptual synthesis of research spanning engineering education, computing education, human–AI interaction, human factors, and the learning sciences to derive an evidence-grounded educational architecture for this archetype. We introduce the ACCEL framework (Agentic Competencies through Curricula, Collaboration, and Enduring Learning), which organizes five competency pillars—intent specification and problem framing; orchestration and delegation; verification, validation, and critical evaluation; ethical governance and accountability; and adaptive self-directed learning—and maps them onto three institutional delivery vectors: curricula, collaboration structures, and continuous learning pathways. Drawing on agency theory, trust-in-automation research, and empirical studies of AI-assisted programming— including randomized and observational evidence that AI benefits are unevenly realized and systematically misperceived—we specify a scaffolded curricular progression, a delegation–verification pedagogical loop for human–AI teaming, redesigned assessment regimes, and governance-literate ethics integration, and we position the framework against current curricular guidelines and international AI-competency frameworks. We articulate the risks the framework is designed to mitigate— automation bias, deskilling, superficial engagement, and diffuse accountability—and close with a research agenda and institutional implications. The analysis argues that educating the agentic engineer requires systemic transformation rather than incremental curricular addition: the unit of instruction must shift from the production of artifacts to the exercise of judgment over increasingly autonomous socio-technical systems. Keywords Agentic engineering · Generative AI · Engineering education · Human-AI collaboration · Curriculum design · Lifelong learning · AI literacy · Assessment
1
Introduction
For most of its modern history, engineering education has rested on a stable implicit contract: universities teach students to produce technical artifacts—code, designs, analyses—and professional value accrues to those who produce them well. The rapid maturation of generative and agentic AI systems has destabilized this contract. Contemporary AI agents plan multi-step workflows, write and refactor code, generate architectural alternatives, run analyses, and iterate on their own outputs with diminishing human intervention [1, 2]; systematic reviews now document large language models and LLM-based agents operating across every phase of the software lifecycle, from requirements elicitation through design, implementation, testing, repair, and maintenance [3, 4]. Field and laboratory evidence
A PREPRINT - AUGUST 3, 2026
indicates that these systems can deliver substantial productivity gains on routine knowledge work [5–7], while simultaneously introducing new failure modes—plausible-but-wrong outputs, opaque reasoning, and uneven reliability across task types—that demand vigilant human oversight [8, 9]. The gains, however, are neither automatic nor universally realized. In a large field deployment, generative AI raised customer-support productivity by 14% on average but concentrated its benefits among less-experienced workers, acting as a skill-leveler rather than a uniform multiplier [6]. More starkly, a randomized controlled trial with experienced open-source developers found that early-2025 AI tooling slowed task completion by 19% even as participants believed it had accelerated them by 20% [10]—a misperception consistent with large-scale observational evidence that developers’ felt productivity with AI assistants tracks suggestion-acceptance behavior rather than measured throughput [11]. That divergence between perceived and measured benefit is precisely a failure of calibrated judgment—and it foreshadows the educational argument of this article. The consequence for the profession is a shift in the locus of engineering value: from the volume and speed of manual production toward the quality of judgment exercised over autonomous production [12, 13]. We term the professional archetype adequate to this shift the agentic engineer: an engineer who specifies intent precisely, decomposes problems for delegation to human and machine collaborators, orchestrates and supervises multi-agent workflows, verifies and validates machine-generated artifacts with the rigor applied to human colleagues’ work, and remains ethically and legally accountable for outcomes. The term is deliberately double-voiced. It refers, first, to engineers who work with AI agents; and second, in [14]’s [14] sense of human agency, to engineers who exercise intentionality, forethought, self-reactiveness, and self-reflectiveness rather than being passively carried by technological change. Educational institutions are not yet organized to produce this archetype. Curricula remain heavily weighted toward individual artifact production; assessment regimes largely measure unassisted output; ethics is commonly quarantined in standalone modules; and degree structures presume that professional preparation is substantially complete at graduation [15–17]. Even the most recent curricular guidelines, which elevate AI to a core knowledge area and adopt competency-based framing [18], and the international AI-competency frameworks issued for students and teachers [19, 20], stop short of specifying the orchestration, verification, and accountability competencies that supervising agentic systems demands. Meanwhile, empirical computing-education research shows that generative AI both disrupts existing pedagogy—undermining traditional assessments and enabling superficial completion of learning tasks—and opens genuinely new pedagogical affordances when integration is deliberate [21–24]. The field therefore faces a design problem: not whether to integrate AI into engineering education, but how to re-architect education so that graduates can direct AI rather than be displaced or deskilled by it. This article addresses that design problem through an integrative conceptual synthesis. Our contributions are fourfold: 1. Conceptualization. We provide a theoretically grounded definition of the agentic engineer that couples the socio-cognitive construct of human agency [14] with empirical findings on AI-native software engineering practice [1, 25]. 2. Framework. We introduce ACCEL (Agentic Competencies through Curricula, Collaboration, and Enduring Learning), a framework that organizes five competency pillars and maps them onto three institutional delivery vectors (Section 4). 3. Design specification. We translate the framework into actionable curricular scaffolds, a delegation– verification pedagogical loop for human–AI teaming, assessment redesign principles, and a governanceliterate approach to ethics integration (Sections 5–9). 4. Research agenda. We identify open empirical questions concerning deskilling, trust calibration, assessment validity, and institutional change, and state the limitations of a conceptual synthesis (Sections 10–11). The remainder of the article proceeds as follows. Section 2 describes the synthesis approach. Section 3 develops the conceptual foundations. Section 4 presents the ACCEL framework. Sections 5–7 elaborate the three delivery vectors, with ethics and assessment treated as cross-cutting threads in Sections 8 and 9. Section 10 discusses implications and a research agenda; Section 11 states limitations; Section 12 concludes.
2
Approach: An Integrative Conceptual Synthesis
This article is an integrative conceptual synthesis in the tradition of [26] and [27]: its purpose is not the exhaustive, protocol-driven coverage of a systematic review, but the construction of a new conceptual architecture from mature yet disconnected literatures. We are explicit about the methodological commitments this entails, and about their limits. The argument was constructed in three stages. 2
A PREPRINT - AUGUST 3, 2026
Stage 1: Corpus assembly. Candidate sources were identified through keyword searches of Scopus, Web of Science, the ACM Digital Library, and IEEE Xplore (search window: 1983–2026; core search terms combined generative AI, agentic AI, engineering education, computing education, human–AI collaboration, trust in automation, and lifelong learning), supplemented by backward and forward citation chaining from anchor works in each community. From this pool we assembled literature across five research communities whose intersection defines the problem space: (i) engineering education research on curricular change, interdisciplinarity, and lifelong learning [16, 28, 29]; (ii) computing education research on generative AI in programming instruction [15, 21–24, 30–32]; (iii) human–AI interaction and human factors research on trust, reliance, and automation [33–40]; (iv) learning sciences research on active learning, cognitive load, self-regulation, and reflective practice [41–47]; and (v) emerging scholarship and field evidence on AI-native and agentic software engineering [1–4, 6, 7, 10–13, 25, 48, 49]. Where the argument required theoretical machinery for delegation itself, we drew additionally on agency theory in its canonical principal–agent form [50] and its recent extension to delegation to and from agentic information-system artifacts [51]. Sources were selected for peerreviewed standing, citational influence, and direct relevance to the education of engineers who supervise autonomous systems; recent preprints and working papers were admitted only where they report emerging practice or field evidence not yet available in the archival literature, and every claim they support in this article is additionally anchored to at least one peer-reviewed source wherever possible. Stage 2: Thematic integration. Following the logic of integrative conceptual review [26], we iteratively coded the corpus for competency claims (what agentic engineers must be able to do), pedagogical claims (how such competencies are developed), and risk claims (what failure modes education must mitigate). Convergent claims across at least two research communities were promoted to framework elements; claims supported within a single community were retained as design suggestions and flagged as such. Coding was performed by the author; the absence of independent coding and inter-rater checks is acknowledged as a limitation in Section 11, and is partially offset by the convergence rule just stated, which requires cross-community corroboration before any claim enters the framework. Stage 3: Framework construction and stress-testing. The resulting elements were organized into the ACCEL framework and stress-tested against three adversarial questions: Does each element survive plausible near-term advances in AI capability? Does each element have at least one concrete curricular or assessment instantiation? And does the framework as a whole mitigate—rather than merely acknowledge—the documented risks of automation bias, deskilling, and superficial engagement [24, 35, 37]? The limitations inherent in this approach—notably the absence of primary empirical validation—are addressed in Section 11.
3
Conceptual Foundations
3.1
From tools to agents: the changing object of engineering work
The distinction between an AI tool and an AI agent is consequential for education. A tool augments a human-executed task: a compiler, a static analyzer, an autocomplete engine. An agent, by contrast, accepts an intent, plans a course of action, executes multi-step workflows—often invoking other tools and agents—and returns candidate outcomes for human review [1, 2]. The distinction is not merely rhetorical: LLM-based agents combining planning, memory, tool invocation, and iterative self-correction are now documented across the entire software lifecycle, including multiagent architectures in which specialized planner, coder, tester, and reviewer agents coordinate under human governance [4,49], and systematic review of the broader LLM-for-software-engineering literature confirms the shift of engineering effort from authoring artifacts to prompting, reviewing, and validating machine-generated ones [3]. Classic humanfactors theory supplies the vocabulary for reasoning about this shift: [36] model automation as operating at graded levels across four cognitive stages—information acquisition, analysis, decision selection, and action execution—and show that the design question is never whether to automate but which functions, to what level, with what human oversight. Agentic systems push automation to high levels across all four stages simultaneously, which is exactly the configuration in which their model predicts the greatest supervisory burden. Studies of AI-assisted development document the transitional phase: developers using code-generation assistants report substantial acceleration on routine tasks alongside recurring difficulties in understanding, debugging, and trusting generated code [8, 9, 25]. Grounded-theory observation adds interaction-level texture: programmers engage codegenerating models in two distinct modes—acceleration, in which the programmer already knows the intended code and vets suggestions in seconds against a mental template, and exploration, in which the model supplies candidate approaches that demand deliberate prompting, comparison, and validation—and effective use of either mode is a learned regulatory meta-skill with documented expert–novice differences, not a byproduct of language knowledge [48]. Controlled evidence quantifies both edges of the capability boundary: large productivity gains on tasks within the technology’s competence frontier, and degraded performance when workers rely on AI beyond that frontier [5, 7]. [7] 3
A PREPRINT - AUGUST 3, 2026
memorably describe this uneven boundary as a “jagged frontier,” and field evidence shows the realized benefit is further moderated by worker experience, with the largest gains accruing to novices on routine work [5, 6]. Subsequent randomized evidence sharpens the point: experienced developers using early-2025 tools on mature, high-context codebases were measurably slower with AI assistance yet perceived themselves faster, a roughly forty-percentage-point gap between forecast and measured effect [10]. The contrast with the 56% speedup observed on bounded, well-specified greenfield tasks [5] is instructive: the same class of tool produces opposite effects in different task regimes, with contextual complexity, quality standards, and verification cost as the moderating variables. Self-perception, meanwhile, is an unreliable instrument for locating the frontier—perceived productivity tracks suggestion-acceptance rates rather than measured output [11]—and the central educational implication of the present article is that mapping, probing, and respecting that frontier is a learnable, teachable professional competency rather than a matter of intuition. As autonomy increases, the engineer’s relationship to the system shifts from operator to supervisor and from author to editor-of-record. Classic human factors research anticipated the hazards of exactly this shift: [37] showed that automating routine work paradoxically raises the skill demands on the human supervisor, who must intervene precisely when the automation fails in unfamiliar ways; [35] catalogued the misuse (over-reliance) and disuse (under-reliance) patterns that follow poorly calibrated trust; and [34] demonstrated that appropriate reliance depends on the human’s ability to assess automation performance across contexts. Engineering education has largely ignored this literature because its graduates were artifact producers, not automation supervisors. That exemption has ended. 3.2
Agency as the organizing construct
We ground the agentic engineer in [14]’s socio-cognitive theory of human agency, which identifies four core properties: intentionality (forming action plans and strategies), forethought (anticipating outcomes and setting goals), self-reactiveness (motivating and regulating execution), and self-reflectiveness (examining one’s own functioning and correcting course). Each property maps directly onto a demand of AI-native practice: intentionality onto precise intent specification and problem framing; forethought onto anticipating agent failure modes and designing guardrails before delegation; self-reactiveness onto monitoring, intervening in, and steering agentic workflows in flight; and selfreflectiveness onto post-hoc evaluation of both the artifact and the human–AI process that produced it, including one’s own reliance behavior. The socio-cognitive account is complemented by a second theoretical lens: agency theory in its principal–agent form. The canonical analysis of delegation identifies its enduring costs—information asymmetry between principal and agent, divergence between delegated intent and executed behavior, and the monitoring expenditure required to contain both [50]. [51] extend this apparatus to information systems that are themselves agentic, arguing that as artifacts acquire autonomy the research question shifts from technology use to bidirectional delegation: humans delegate tasks to agentic artifacts, artifacts delegate sub-decisions and exceptions back to humans, and the appropriateness of each delegation depends on task characteristics, artifact capability, and the allocation of accountability. Transposed to engineering education, this framing converts orchestration from a vague soft skill into a structured decision problem—what to delegate, under what monitoring regime, with what reversibility—whose costs and failure modes are theoretically characterized in advance. Together the two lenses do theoretical work that a purely skills-based account cannot. They explain why the agentic engineer is not reducible to a checklist of tool proficiencies: agency is a disposition exercised over whatever tools exist, which is precisely what makes it durable across technology generations, and delegation is a principal’s problem that persists regardless of which agent technology occupies the other side of the contract. The grounding also connects the professional construct to well-developed educational machinery—self-regulated learning [44], reflective practice [45], and active learning [42]—giving curriculum designers established levers rather than requiring pedagogical invention from scratch. 3.3
What the empirical record already shows
Four findings from the recent empirical record constrain any credible educational response. First, generative AI collapses traditional novice tasks. Code-generation models solve typical introductory programming assessments at or above student level [31], undermining the signaling value of unassisted-production assessment and forcing a re-examination of what introductory courses are for [21, 23]. Second, unstructured access produces divergent learning outcomes. Novices given AI code generators complete tasks faster without uniform learning loss, but exhibit distinct usage profiles: some use generation as scaffolding for understanding, while others develop dependency patterns associated with weaker subsequent unassisted performance [24]. Observational replication work confirms and extends the divide: generative AI accelerates already-capable 4
A PREPRINT - AUGUST 3, 2026
novices while compounding the metacognitive difficulties of struggling ones, who frequently finish with an “illusion of competence”—believing they performed better than they did [30]. A recent systematic review of 58 studies of AI agents in programming education reports over-reliance producing superficial learning in roughly two-thirds of the studies examined [32]. The workplace analogue is now documented at scale: AI assistance functions as a skill-leveler that lifts weaker performers most [5,6]—an equity opportunity, but also a mechanism by which novices can experience an artificial inflation of capability precisely when their independent verification skill is weakest. Instructional structure, not access per se, determines whether AI assistance builds or borrows competence—an instance of the general finding that AI interventions unmoored from cognitive theory and coherent instructional design underperform [43, 46, 47, 52]. Third, intent specification is a distinct, isolable bottleneck. A causal-intervention study of beginning students’ failed text-to-code prompts tested whether the deficit lies in style (lacking technical vocabulary) or substance (not understanding how much information the model needs): surgically replacing informal vocabulary with correct terminology largely failed to rescue performance, while the information content of the prompt—whether it specified inputs, outputs, edge cases, and behavioral constraints—predicted success [53]. The same study documents a characteristic “stuck” revision pattern in which students respond to failed generations with cosmetic rewording rather than added specification, iterating in place without converging. Specification competence is thus conceptual rather than lexical; it will not be produced by tool exposure, prompt templates, or vocabulary drills, and it must be taught and assessed in its own right. Fourth, professional practice is reorganizing around review and orchestration. Studies of AI-assisted developers show effort shifting from writing code to specifying, prompting, evaluating, and integrating it [9, 25, 48]; systematic reviews document LLMs and LLM-based agents deployed across requirements, design, implementation, testing, debugging, repair, and maintenance [3,4]; and accounts of agentic DevOps document pipelines in which agents propose, test, and deploy changes under human governance [2, 13]. Notably, the open challenges catalogued by the agentic-SE research community read as an inverted competency specification for its human supervisors: agents remain unreliable at requirement disambiguation, their end-to-end autonomy amplifies error propagation across pipeline stages, and orchestration quality often dominates raw model capability in determining outcomes [4]. Education that continues to optimize for unassisted production is therefore optimizing for a shrinking slice of professional activity. Together these findings define the design constraints: preserve foundational understanding (because supervision without comprehension is vacuous), teach calibrated reliance (because the frontier is jagged), teach specification explicitly (because underspecification, not phrasing, is the documented novice failure), restructure assessment (because unassisted production no longer discriminates), and institutionalize continuous adaptation (because the frontier moves).
4
The ACCEL Framework
We now integrate these foundations into a single architecture. ACCEL—Agentic Competencies through Curricula, Collaboration, and Enduring Learning—comprises (i) a competency core of five pillars that specify what the agentic engineer must be able to do, and (ii) three institutional delivery vectors that specify how educational systems develop those pillars. Figure 1 depicts the architecture; Table 1 operationalizes the pillars as learning outcomes and assessment evidence. 4.1
The five competency pillars
P1: Intent specification and problem framing. The upstream competency on which all delegation depends. It encompasses eliciting and formalizing requirements, decomposing ill-structured problems into delegable units, and expressing intent as executable specifications with explicit constraints, acceptance criteria, and guardrails [1]. Prompt engineering is the currently visible surface of this pillar, but the durable competency is specification discipline—a lineage engineering education already possesses through requirements engineering [54] and which now requires generalization to machine addressees. The pillar’s status as a genuine, isolable skill is empirically established: novice failure at AI-mediated code generation is driven by underspecification of the machine interlocutor’s information requirements, not by deficient phrasing, and it persists under vocabulary correction [53]. Requirement disambiguation is likewise the documented weak point of current agentic systems themselves [4], making human specification competence the binding input to hybrid-team performance. P2: Orchestration and delegation. The capacity to allocate work across a hybrid team of humans and agents: selecting appropriate agents for sub-tasks, sequencing and parallelizing workflows, managing hand-offs and shared context, and adapting the allocation as evidence about agent performance accumulates [2, 13]. The pillar is theoretically structured by delegation theory—what to delegate, under what monitoring regime, with what accountability allocation [50, 51]—and by the levels-of-automation model, which frames each delegation as a choice of automation 5
A PREPRINT - AUGUST 3, 2026
P2 Orchestration & delegation P1 Intent specification & problem framing
P3 Verification, validation & critical evaluation
The Agentic Engineer intentionality · forethought · self-reactiveness · self-reflectiveness
P5 Adaptive selfdirected learning
P4 Ethical governance & accountability
Curricula
Collaboration
Continuous learning
scaffolded integration; AI literacy; PBL
human–AI teaming; interdisciplinarity; industry
CEE; micro-credentials; reflective portfolios
Institutional delivery vectors
Figure 1: The ACCEL framework. Five competency pillars (P1–P5) converge on the agentic engineer, whose core is defined by Bandura’s four properties of human agency. Three institutional delivery vectors—curricula, collaboration structures, and continuous learning pathways—develop and sustain the pillars across the professional lifespan.
level per cognitive stage rather than a binary hand-off [36]. Its technical substrate is now well documented: multi-agent architectures coordinate specialized planner, coder, tester, and reviewer agents through structured communication, and their emergent interactions can be difficult to predict [4, 49], which is precisely why orchestration quality, not raw model capability, often dominates outcomes. This pillar recasts classical project management and software architecture skills for teams whose members include non-human contributors of variable and shifting reliability. P3: Verification, validation, and critical evaluation. The defining supervisory competency: reviewing machinegenerated artifacts with the rigor applied to human colleagues’ work, detecting plausible-but-wrong outputs, designing test regimes and acceptance gates for AI contributions, interpreting uncertainty, and calibrating reliance to demonstrated performance [8, 34]. Empirically, this is where AI-assisted practice most often fails [7]; pedagogically, it is where automation bias and deskilling must be confronted directly [35,37]. Importantly, overreliance is not immutable: experimental work shows it responds to the cost–benefit structure of engagement, declining when explanations make verification cheap relative to blind acceptance and when task stakes make errors salient [40]—which converts critical evaluation from an exhortation into a designable property of learning environments. P4: Ethical governance and accountability. Beyond conventional professional ethics, this pillar comprises literacy in algorithmic bias and fairness [55], transparency and interpretability demands for safety-relevant systems, intellectual-property and privacy obligations attached to generated artifacts, and working knowledge of the maturing regulatory apparatus—risk-management frameworks and statutory regimes—within which engineered AI systems must now operate [17, 56–58]. The pillar’s design rationale follows [59]: high-level ethical principles cannot by themselves guarantee ethical AI, because they lack the technical and institutional mechanisms of implementation; the engineer must therefore be able to translate principles into system constraints—traceable decision paths, audit trails, human-oversight points, and fail-safes—in the spirit of human-centered AI design, which treats reliability, safety, and trustworthiness as engineered properties supporting human control rather than aspirations appended to it [38, 39]. The pillar’s core commitment is non-delegable accountability: authority over decisions may be shared with machines; responsibility may not. 6
A PREPRINT - AUGUST 3, 2026
Table 1: The ACCEL competency pillars operationalized as exemplar learning outcomes and assessment evidence. Pillar
Exemplar learning outcomes
Exemplar assessment evidence
P1 Intent specification & problem framing
Decompose an ill-structured problem into delegable units; author executable specifications with acceptance criteria and guardrails; justify framing choices against stakeholder needs.
Specification documents graded for completeness and testability; comparative critique of alternative decompositions; traceability from requirements to delegated tasks.
P2 Orchestration & delegation
Allocate sub-tasks across humans and agents with explicit rationale; design multi-agent workflows with hand-off and context-management protocols; re-plan when agent performance deviates.
Orchestration logs and workflow diagrams; design reviews defending allocation decisions; documented re-planning episodes in project retrospectives.
P3 Verification, validation & critical evaluation
Detect seeded defects in AI-generated artifacts; design acceptance gates and test regimes for machine contributions; demonstrate calibrated reliance across task types.
Defect-detection exercises with known ground truth; review reports on AI contributions; reliancecalibration records comparing trust decisions with measured agent accuracy.
P4 Ethical governance & accountability
Conduct bias and impact assessments; map system decisions to accountable humans; apply riskmanagement and regulatory requirements to a concrete deployment.
Impact-assessment artifacts; accountability matrices; red-team exercise reports with reflective debriefs.
P5 Adaptive selfdirected learning
Systematically evaluate an unfamiliar AI tool against task requirements; maintain a reflective record of one’s own human–AI process; construct and defend a personal development plan.
Tool-evaluation reports with explicit criteria; reflective portfolios; evidence of self-initiated learning (micro-credentials, community contributions).
P5: Adaptive self-directed learning. The meta-competency that keeps the other four current as the capability frontier moves. It comprises self-regulated learning strategies [44], reflective practice on one’s own human–AI process [45], systematic evaluation of unfamiliar tools and models, and construction of a personal learning infrastructure— communities of practice, curated information flows, credentialing pathways—that sustains professional currency across a career [60]. The pillar rests on the consensus synthesis of the learning sciences: durable expertise develops through active engagement, deliberate practice, feedback, and metacognitive monitoring of one’s own understanding, and self-regulation is a trainable capability rather than a fixed trait [41]. 4.2
The three delivery vectors
The pillars are developed through three mutually reinforcing institutional vectors. Curricula (Section 5) provide vertically and horizontally integrated, scaffolded development of the pillars across a degree program. Collaboration structures (Section 6) supply the authentic hybrid-team contexts—human–AI, interdisciplinary, and industry-connected—in which the pillars are exercised under realistic conditions. Continuous learning pathways (Section 7) extend development beyond graduation through continuing engineering education, micro-credentials, and reflective professional practice. Ethics (Section 8) and assessment (Section 9) are deliberately treated as cross-cutting threads woven through all three vectors rather than as vectors of their own: the framework’s central design claim is that quarantining either one recreates precisely the fragmentation it seeks to repair.
5
Vector I — Curricula: Scaffolded Integration Rather Than Additive Modules
5.1
Integration, not addition
The dominant institutional reflex—adding an “AI tools” elective to an otherwise unchanged program—fails on both theoretical and empirical grounds. Isolated, tool-specific instruction does not produce transferable competence [52,55], and bolting AI onto unrevised pedagogy erects barriers to the active learning on which durable understanding depends [42]. The historical analysis of [28] shows that engineering education’s major advances came from systemic shifts— toward outcomes-based accreditation, learning-sciences-informed pedagogy, and context-rich instruction—not from course-level additions. ACCEL therefore prescribes vertical integration (pillar development sequenced across program years) and horizontal integration (pillar exercise embedded within disciplinary courses, from mechanics to databases), with AI literacy in [61]’s sense—knowing what AI is, what it can and cannot do, and how to critique its outputs— treated as a program-wide foundation rather than a course topic, consistent with its recent institutional codification 7
A PREPRINT - AUGUST 3, 2026
Table 2: Scaffolded curricular progression for agentic engineering education. AI autonomy is expanded stage-by-stage, contingent on demonstrated verification competence. Stage
AI role
Student focus
Dominant pillars
1. Foundations (Year 1)
Tutor and explainer; generation restricted in assessed work
Core disciplinary concepts; unassisted problem solving; AI literacy; first critique of AI outputs against known ground truth
P3 (nascent), P5
2. Assisted practice (Year 2)
Pair collaborator on bounded tasks
Specification writing; systematic review of AI contributions; seeded-defect detection; documentation of AI usage
P1, P3
3. Supervised delegation (Year 3)
Semi-autonomous multi-step tasks
on
Workflow orchestration; acceptance gates; reliance calibration; interdisciplinary team projects with AI members
P2, P3, P4
4. Orchestration (Year 4 / capstone)
Multi-agent teams across the full lifecycle
End-to-end intent-tooutcome responsibility; governance artifacts; reflective analysis of the human–AI process
All pillars
agent
in international competency frameworks [19] and in the CS2023 guidelines’ elevation of AI to a core, cross-cutting knowledge area [18]. 5.2
A scaffolded progression
Table 2 specifies a four-stage progression, each stage defined by the autonomy granted to AI collaborators and the corresponding supervisory demand placed on the student. The progression operationalizes two learning-sciences constraints. First, foundations precede supervision: students cannot verify what they cannot understand, so early stages deliberately restrict AI assistance in foundational skill-building while using AI tutoring to support—not substitute for—practice [24, 43]; this ordering follows directly from the consensus finding that expertise is built through active engagement and deliberate practice rather than exposure to solutions [41]. Second, difficulty is granted, not evaded: at each stage, AI autonomy increases only as verification competence is demonstrated, preventing the dependency profiles observed under unstructured access [23, 24]. The staging also responds to the broader analysis of LLMs in education: the technologies shift learners from producers of every artifact toward supervisors and evaluators of machine-generated work, and educational objectives must move correspondingly toward problem formulation, verification, judgment, and ethical reasoning [47]. 5.3
Project-based learning as the load-bearing pedagogy
Project-based learning (PBL) is the natural vehicle for the upper stages because agentic competence is exercised, not recited: it emerges from consequential decisions about delegation, trust, and integration under ambiguity [62], and the engagement dividends of well-designed experiential formats in software engineering education are durable rather than novelty effects [63]. In an ACCEL-aligned capstone, students define problem boundaries, select and configure AI collaborators, negotiate evolving goals, and remain accountable for the delivered system. Critically, assessment weight shifts from the final artifact—which agents can increasingly produce—to the orchestration record: the quality of specifications, the defensibility of delegation decisions, the rigor of verification, and the honesty of reflection [21, 22]. Section 9 develops this shift in full. 5.4
Adaptive platforms under pedagogical control
AI-powered adaptive learning platforms can personalize pacing, diagnose misconceptions, and simulate complex engineering scenarios at scale. Their integration, however, must be governed by the finding that educational technology unanchored in learning theory prioritizes novelty over effectiveness [46,52]: adaptive difficulty must respect cognitive8
A PREPRINT - AUGUST 3, 2026
1. Specify intent
2. Delegate
frame, decompose, set acceptance criteria (P1)
select agents, set guardrails (P2, P4)
3. Agent executes plans, generates, iterates
rejection
re-delegation
6. Reflect & capture rationale
5. Integrate or re-delegate
4. Verify & validate
document decisions, update trust model (P5)
accept, revise intent, or reassign (P2)
test, inspect, calibrate reliance (P3)
Figure 2: The delegation–verification loop. Students traverse the loop explicitly, producing assessable artifacts at every stage; accountability checks (P4) gate integration of any irreversible change. Dashed edges mark the rejection and re-delegation paths whose exercise distinguishes calibrated from credulous reliance.
load constraints [43], tutoring must scaffold rather than answer, and educators must retain design authority over the systems that mediate their students’ learning. The agentic-engineering program is itself an appropriate site for this scrutiny—students who dissect their own adaptive platform’s recommendation logic are simultaneously exercising P3 and P4 on a system with immediate personal stakes.
6
Vector II — Collaboration: Hybrid Teams as the Unit of Practice
6.1
The delegation–verification loop
The pedagogical core of human–AI teaming instruction is a disciplined loop that students traverse repeatedly, first slowly and explicitly, later fluently (Figure 2). The loop renders visible—and therefore teachable and assessable—the supervisory cycle that expert AI-native practitioners execute tacitly [9,25]. Its stages instantiate the pillars in sequence: intent specification and decomposition (P1), delegation with explicit guardrails (P2), agent execution, verification against acceptance criteria (P3), integration or re-delegation, and reflective rationale capture (P5), with accountability checks (P4) gating irreversible actions. Three design features guard against the loop degenerating into ritual. First, curricula must engineer encounters with failure: exercises seeded with plausible-but-wrong agent outputs, “AI coworker” simulations in which over-trust has visible consequences, and post-incident debriefs [34, 35]. Students who have never caught an agent being confidently wrong have not learned verification; they have learned deference. Crucially, calibration must be measured rather than self-reported: even experienced professionals systematically misjudge whether AI assistance is helping them [10, 11], so curricula should confront students with objective records of their own acceptance decisions against ground truth, not with reflection alone. Second, the cost structure of verification must be deliberately engineered. Overreliance is sensitive to the relative effort of engaging critically versus accepting blindly: when explanations and supporting evidence make verification cheap, and when stakes make errors salient, people verify more [40]. Loop-based exercises should therefore require agents (and students configuring them) to surface reasoning, intermediate artifacts, and confidence signals that lower the cost of inspection—simultaneously training students to demand inspectable output as a delegation precondition. Because verification depth should also vary by interaction mode—shallow pattern-matching suffices in acceleration mode, while exploration mode demands documentation lookup, comparison, and testing [48]—students must learn to diagnose which mode they are in and match their evaluation strategy to it. Third, guidance on human–AI interaction—making capability boundaries visible, supporting efficient correction, calibrating expectations—should be taught as design knowledge students apply when they build agentic systems for others, not merely experience as users [33, 38, 39]. 6.2
Interdisciplinary teams and AI as boundary object
Modern AI-enabled systems entangle computer science, data science, security, human factors, law, ethics, and domain expertise; interdisciplinary competence is accordingly a validated engineering learning outcome with a developed research base [29]. Agentic practice sharpens the demand in a specific way: AI agents function as boundary objects that silently encode assumptions from multiple fields—statistical assumptions from their training regimes, value assumptions from their alignment procedures, domain assumptions from their data. Interrogating those assumptions is intrinsically a multi-disciplinary act. ACCEL therefore prescribes team-based experiences in which engineering students collaborate with peers from data science, design, and the humanities to audit and adapt agent behavior for 9
A PREPRINT - AUGUST 3, 2026
domain-specific constraints—rehearsing the distributed cognition of professional practice while exercising P3 and P4 on live systems. 6.3
Industry partnership as curricular metabolism
AI capabilities evolve on cycles measured in months; university curriculum revision operates on cycles measured in years. Sustained industry partnership—mentored projects, internships in AI-native teams, challenge-based learning, co-designed learning outcomes—is therefore not enrichment but metabolic necessity, the mechanism by which programs sense and absorb changes at the practice frontier [15, 16]. National capability-building initiatives, in which academies co-locate curriculum design with government and industry demand signals, exemplify the institutional form this metabolism can take [12].
7
Vector III — Continuous Learning: The Career-Length Program
If the capability frontier moves continuously, terminal degrees cannot be terminal. ACCEL treats the degree as the first phase of a career-length program with three components. Self-directed learning capacity as a designed outcome. The disposition to identify one’s own knowledge gaps, locate resources, and regulate learning is a trainable competency, not a temperament [44, 60]. Programs build it deliberately: learning-to-learn instruction embedded in disciplinary courses, reflective portfolios that make students’ own learning processes objects of analysis [45], and repeated tool-evaluation exercises (P5) that rehearse the professional act of confronting an unfamiliar system and systematically establishing what it can be trusted to do. Structured continuing engineering education. Reskilling and upskilling pathways—stackable micro-credentials, modular certificates, executive programs—give the post-graduation phase institutional form. Taxonomic standardization of continuing engineering education enables benchmarking and portability across providers and borders, and alignment between credential frameworks and evolving job architectures allows both institutions and governments to forecast and provision workforce capability [12, 58]. The design principle carried over from Section 5 applies with full force: continuing education must develop pillar competencies, not perishable tool proficiencies. Personal learning infrastructure. Between formal episodes, professionals sustain currency through infrastructure they curate themselves: communities of practice, monitored research and standards flows, personal experimentation sandboxes, and—reflexively—AI-powered learning tools whose recommendations they evaluate with the same calibrated skepticism they apply to any agent (P3 applied to P5). Graduates who leave with such an infrastructure already operating, seeded by the reflective portfolio and tool-evaluation practices above, cross the education-to-practice boundary without interrupting their learning.
8
Cross-Cutting Thread I — Ethics and Governance
Standalone ethics modules demonstrably underperform: codes of conduct imparted in isolation do not transfer to technical decision-making [17,55]. The deeper diagnosis is [59]’s [59]: unlike medicine, AI practice lacks the common aims, professional norms, and accountability mechanisms that make principle-based ethics effective, so principles alone cannot guarantee ethical systems—they must be translated into concrete technical and institutional mechanisms. That translation is engineering work, and it belongs in the engineering curriculum. ACCEL therefore embeds ethical reasoning as a continuous thread with three strands. First, situated ethical reasoning within technical courses: bias audits inside machine-learning coursework, privacy analyses inside database and systems design, accountability mapping inside software architecture—so that ethical analysis is encountered as part of engineering judgment, not adjacent to it [55]. Second, adversarial and reflective exercises: structured red-team activities in which students deliberately probe AI systems for bias, unsafe behavior, and misuse potential, followed by debriefs that convert the experience into articulated professional obligation [17]. These exercises serve a dual function, simultaneously training P3 (finding failure) and P4 (owning its implications). Third, governance literacy: working fluency with the risk-management frameworks and statutory regimes that now bind deployed AI systems—including risk classification, documentation, human-oversight, and transparency obligations [56,57]. For the agentic engineer, regulation is not compliance overhead but design input: guardrails, audit trails, and oversight mechanisms are engineered artifacts, and their construction belongs in the curriculum alongside every other artifact class [39, 58]. 10
A PREPRINT - AUGUST 3, 2026
Table 3: Reorienting assessment for agentic engineering education. Dimension
Traditional paradigm
Agentic paradigm
Primary object
Final artifact produced without assistance
Judgment: specification quality, delegation rationale, verification rigor, reflective honesty
AI usage
Prohibited or ignored
Declared, documented, and itself assessed
Mode
Summative examination; isolated tasks
Portfolios, design reviews, orchestration logs, defect-detection exercises, viva-style defenses
Integrity model
Detection and prohibition
Transparency and responsible-use demonstration
Foundational knowledge
Assumed measured by artifact production
Measured directly in deliberately AI-restricted components
The thread terminates in the framework’s non-negotiable commitment: accountability does not delegate. Whatever autonomy is granted to agents, a named human remains answerable for outcomes. Curricula operationalize this through accountability matrices in projects, sign-off protocols at integration gates (Figure 2), and assessment that asks not only “does it work?” but “who answers for it, and on what evidence?”
9
Cross-Cutting Thread II — Assessment: From Artifact to Judgment
Assessment is where the paradigm shift becomes unavoidable. When AI systems solve standard assessments at student level [31], unassisted-artifact grading loses validity as a measure of professional readiness—and prohibition is neither enforceable nor aligned with the practice graduates will enter [21, 22]. Table 3 summarizes the required reorientation: the object of assessment shifts from the artifact to the judgment exercised in producing it. Three design principles govern the reorientation. First, assess the loop, not only its output: orchestration logs, specification documents, review reports, and reflective analyses generated by the delegation–verification loop (Figure 2) are first-class assessment evidence, graded against the pillar outcomes of Table 1. Process traces are not merely available but diagnostic: a student’s prompt-revision trajectory, for example, observably distinguishes reasoning about specification completeness from cosmetic thrashing [53], giving instructors an evidence-bearing signal of P1 competence that no final artifact can supply. The move toward evaluating reasoning processes, project work, and oral defense— with AI as a declared, legitimate tool—is likewise the convergent recommendation of the broader LLMs-in-education literature [21, 47]. Second, protect a foundational core: because supervision presupposes comprehension, programs retain deliberately AI-restricted assessment components—concept-focused examinations, whiteboard reasoning, live defect-finding—that certify the understanding on which everything else rests [23, 43]. This core is also the countermeasure to the documented “illusion of competence,” in which AI-assisted students sincerely overestimate what they have learned [30]: self-assessment cannot be trusted to detect the gap that AI-restricted assessment reveals. Third, make transparency the integrity mechanism: students declare and document AI usage as professionals must, and the quality of that documentation is itself graded, converting integrity from a policing problem into a competency [21,22]. AI-supported assessment infrastructure—analytics on collaboration patterns, automated formative feedback on openended work—can enrich this regime, provided it is designed under the same learning-sciences discipline demanded of adaptive platforms and is itself subject to the bias and validity scrutiny of Section 8 [46]. Reducing judgment competencies to convenient proxies would reproduce, inside assessment, exactly the automation credulity the curriculum exists to prevent.
10
Discussion
10.1
What ACCEL adds
ACCEL occupies ground that adjacent frameworks do not. AI-literacy frameworks [19,20,61] specify what any citizen, student, or teacher should understand about AI but stop short of the supervisory, orchestration, and accountability competencies specific to professional engineering. Established curriculum guidelines either predate agentic practice entirely [54,64] or—as in CS2023, which elevates AI to a core knowledge area and adopts competency-based framing— address AI as subject matter and assistant without yet specifying orchestration, verification, and governance of autonomous collaborators as graduate competencies [18]. Human-factors accounts of automation supervision [34–36] diagnose the reliance problem without prescribing an educational architecture for it. And the technical literature on 11
A PREPRINT - AUGUST 3, 2026
agentic software engineering establishes, comprehensively, what LLM-based agents can do across the lifecycle [3,4]— but leaves open the institutional question of who is educated to direct them, and how; ACCEL is addressed to precisely that residual. Relative to these prior treatments, the framework makes three moves that we regard as its principal contributions. First, it unifies literatures that have addressed the problem in isolation: computing-education findings on generative AI [21, 22], human-factors theory on automation supervision [34–37], socio-cognitive agency theory and principal–agent delegation theory [14,50,51], and emerging accounts of AI-native practice [1,25,48]—yielding an educational architecture in which pedagogical prescriptions inherit theoretical warrant. Second, it conditions autonomy on verification: the scaffolded progression (Table 2) grants AI autonomy stage-by-stage against demonstrated supervisory competence, directly operationalizing the empirical finding that instructional structure determines whether AI assistance builds or borrows competence [24]. Third, it relocates assessment: by making the delegation–verification loop itself the object of assessment, the framework restores validity to evaluation in an era when artifacts no longer evidence unassisted competence [31]. 10.2
Risks the framework is designed to mitigate
Four failure modes recur in the record, and each is met by a specific framework mechanism. Automation bias and over-reliance [34, 35]—whose contemporary signature is the measured gap between perceived and actual AI benefit [10, 11]—are countered by engineered failure encounters, verification-cost design [40], and reliance-calibration assessment (Section 6). Deskilling [37] is countered by the protected foundational core and staged autonomy (Sections 5, 9). Superficial engagement—completing tasks through AI without learning, now documented at scale and known to disproportionately harm weaker students [24, 30, 32, 42]—is countered by assessing the loop rather than the artifact. Diffuse accountability is countered by the non-delegation commitment and its curricular instruments (Section 8). We emphasize that these are design intentions with empirical warrant, not demonstrated effects; establishing the effects is the first item of the research agenda. 10.3
Institutional and policy implications
For universities, the analysis implies that faculty development is the binding constraint: educators cannot teach calibrated supervision of systems they have not themselves learned to supervise, and institutions must resource that learning explicitly [16,47,52]. The equity stakes deserve emphasis: because AI assistance levels skills on routine work [5,6] while effective use presupposes specification and verification competence that novices systematically lack [30, 53], unstructured access widens capability gaps even as it appears to democratize practice—scaffolded, explicitly taught progression is therefore an equity instrument, not only a pedagogical preference. Accreditation bodies face pressure to evolve outcome frameworks: even the most recent guidelines, though AI-aware and competency-based [18], do not yet name orchestration, verification, and governance of autonomous collaborators as explicit graduate outcomes, and older frameworks predate agentic practice entirely [54, 64]. For governments and national capability programs, ACCEL’s continuous-learning vector implies investment in credential architectures and continuing-education taxonomies that keep national workforces adaptive rather than episodically retrained [12, 58]; national AI academies that integrate degree-level, professional, and executive education within one competency architecture are a natural institutional vehicle. 10.4
Research agenda
Five empirical questions follow directly. (1) Longitudinal skill formation: do graduates of scaffolded-autonomy programs exhibit stronger unassisted foundations and better-calibrated reliance than graduates of unrestricted-access programs? (2) Verification pedagogy: which exercise designs most efficiently produce durable defect-detection skill on AI-generated artifacts, and how does that skill transfer across artifact types? (3) Assessment validity: do orchestrationlog and portfolio assessments predict professional performance in AI-native teams better than traditional artifact grading? (4) Trust calibration measurement: can reliance calibration be instrumented reliably enough—for example, by comparing students’ acceptance decisions against measured agent accuracy—to serve as a formal learning outcome? (5) Institutional change: which faculty-development and curriculum-governance models allow programs to track a capability frontier that moves faster than revision cycles? Each question is tractable with existing education-research methods; none is answerable from the armchair.
11
Limitations
Five limitations bound the claims of this article. First, it is a conceptual synthesis: ACCEL is derived from, and consistent with, the empirical record, but the framework itself has not been implemented and evaluated as a whole; 12
A PREPRINT - AUGUST 3, 2026
its mitigation claims are design hypotheses pending the studies outlined above. Second, the corpus is weighted toward software and computing education, where the empirical record on generative AI is deepest; extension to other engineering disciplines—where physical artifacts, safety certification, and licensure alter the delegation calculus— requires discipline-specific validation. Third, the underlying technology is moving: findings about current model failure modes [7, 8] may age quickly, though we have deliberately anchored the framework in competencies (specification, verification, governance, learning) chosen for robustness to capability shifts rather than in tool-specific skills. Fourth, the analysis largely presumes institutional contexts with the resources to execute systemic reform; adaptation to resource-constrained settings, where faculty capacity and infrastructure are limiting, is an open design problem of the first importance [58]. Fifth, the synthesis was conducted by a single author: corpus coding was not subject to independent inter-rater checks, and a portion of the sources characterizing AI-native practice are the author’s own recent work, some of it available only as preprints at the time of writing. We have mitigated both risks by the crosscommunity convergence rule of Section 2 and by anchoring every load-bearing claim to at least one independent, peer-reviewed source, but readers should weigh these provenance facts when assessing the argument.
12
Conclusion
The rise of agentic AI does not diminish the engineer; it relocates the engineer’s value. When autonomous systems can produce artifacts, the scarce and durable human contributions become the framing of intent, the orchestration of hybrid teams, the verification of machine work, the ownership of ethical consequence, and the discipline of continuous adaptation. Educating for these contributions is not accomplished by adding a course, licensing a tool, or prohibiting one. It requires the systemic re-architecture this article has specified: competency pillars grounded in agency theory and human-factors evidence; curricula that grant AI autonomy only as verification competence is demonstrated; collaboration structures that make the delegation–verification loop explicit, teachable, and assessable; assessment that measures judgment rather than artifacts; ethics practiced as governance engineering rather than recited as code; and learning designed to outlast the degree that begins it. The stakes exceed pedagogy. Societies are delegating consequential decisions to increasingly autonomous systems, and the engineers who supervise those systems constitute the profession’s—and the public’s—primary line of accountable oversight. Whether that oversight is exercised with calibrated judgment or credulous deference will be decided, in large part, by how the current generation of engineers is educated. The agentic engineer, in the full Bandurian sense, is therefore this era’s central educational objective: a professional whose learning never stops, whose collaboration spans human and artificial minds, and whose agency ensures that autonomous technology remains an instrument of human intention, dignity, and progress.
References [1] Mamdouh Alenezi. From determinism to delegation: AI-native software engineering and the evolution of the agentic engineer. arXiv preprint arXiv:2606.28791, 2026. [2] Mohammed Akour and Mamdouh Alenezi. Agentic AI in DevOps: Boosting software automation and collaboration. In 2025 International Conference on Artificial Intelligence Security and Applications. IEEE, 2025. [3] Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang. Large language models for software engineering: A systematic literature review. ACM Transactions on Software Engineering and Methodology, 33(8):Article 220, 2024. [4] Junwei Liu, Kaixin Wang, Yixuan Chen, Xin Peng, Zhenpeng Chen, Lingming Zhang, and Yiling Lou. Large language model-based agents for software engineering: A survey. ACM Transactions on Software Engineering and Methodology, 2026. https://doi.org/10.1145/3796507. [5] Sida Peng, Eirini Kalliamvakou, Peter Cihon, and Mert Demirer. The impact of AI on developer productivity: Evidence from GitHub Copilot. arXiv preprint arXiv:2302.06590, 2023. [6] Erik Brynjolfsson, Danielle Li, and Lindsey Raymond. Generative AI at work. The Quarterly Journal of Economics, 140(2):889–942, 2025. [7] Fabrizio Dell’Acqua, Edward McFowland III, Ethan R. Mollick, Hila Lifshitz-Assaf, Katherine Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, and Karim R. Lakhani. Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality. Technical Report Working Paper 24-013, Harvard Business School, 2023. 13
A PREPRINT - AUGUST 3, 2026
[8] Priyan Vaithilingam, Tianyi Zhang, and Elena L. Glassman. Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models. In Extended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems, pages 1–7. ACM, 2022. [9] Advait Sarkar, Andrew D. Gordon, Carina Negreanu, Christian Poelitz, Sruti Srinivasa Ragavan, and Ben Zorn. What is it like to program with artificial intelligence? In Proceedings of the 33rd Annual Conference of the Psychology of Programming Interest Group (PPIG 2022), 2022. [10] Joel Becker, Nate Rush, Beth Barnes, and David Rein. Measuring the impact of early-2025 AI on experienced open-source developer productivity. arXiv preprint arXiv:2507.09089, 2025. [11] Albert Ziegler, Eirini Kalliamvakou, X. Alice Li, Andrew Rice, Devon Rifkin, Shawn Simister, Ganesh Sittampalam, and Edward Aftandilian. Measuring GitHub Copilot’s impact on productivity. Communications of the ACM, 67(3):54–61, 2024. [12] Mamdouh Alenezi. The rise of AI-native software engineering: Implications for practice, education, and the future workforce. arXiv preprint arXiv:2606.12986, 2026. [13] Mamdouh Alenezi. Human–AI collaboration and the transformation of software engineering work. arXiv preprint arXiv:2606.03394, 2026. [14] Albert Bandura. Toward a psychology of human agency. Perspectives on Psychological Science, 1(2):164–180, 2006. [15] Cigdem Sengul, Rumyana Neykova, and Giuseppe Destefanis. Software engineering education in the era of conversational AI: Current trends and future directions. Frontiers in Artificial Intelligence, 7:1436350, 2024. [16] Aditya Johri, Andrew S. Katz, Junaid Qadir, and Ashish Kohli. Generative artificial intelligence and engineering education. Journal of Engineering Education, 112(3):572–577, 2023. [17] Jason Borenstein and Ayanna Howard. Emerging challenges in AI and the need for AI ethics education. AI and Ethics, 1:61–65, 2021. [18] ACM/IEEE-CS/AAAI Joint Task Force on Computing Curricula. Computer science curricula 2023 (CS2023): The final report. Technical report, ACM, IEEE Computer Society, and AAAI, 2024. [19] Fengchun Miao and Kelly Shiohira. AI Competency Framework for Students. UNESCO, Paris, 2024. [20] Fengchun Miao and Mutlu Cukurova. AI Competency Framework for Teachers. UNESCO, Paris, 2024. [21] Paul Denny, James Prather, Brett A. Becker, James Finnie-Ansley, Arto Hellas, Juho Leinonen, Andrew LuxtonReilly, Brent N. Reeves, Eddie Antonio Santos, and Sami Sarsa. Computing education in the era of generative AI. Communications of the ACM, 67(2):56–67, 2024. [22] James Prather, Paul Denny, Juho Leinonen, Brett A. Becker, Ibrahim Albluwi, Michelle Craig, Hieke Keuning, Natalie Kiesler, Tobias Kohn, Andrew Luxton-Reilly, Stephen MacNeil, Andrew Petersen, Raymond Pettit, Brent N. Reeves, and Jaromir Savelka. The robots are here: Navigating the generative AI revolution in computing education. In Proceedings of the 2023 Working Group Reports on Innovation and Technology in Computer Science Education (ITiCSE-WGR ’23), pages 108–159. ACM, 2023. [23] Brett A. Becker, Paul Denny, James Finnie-Ansley, Andrew Luxton-Reilly, James Prather, and Eddie Antonio Santos. Programming is hard—or at least it used to be: Educational opportunities and challenges of AI code generation. In Proceedings of the 54th ACM Technical Symposium on Computer Science Education (SIGCSE ’23), pages 500–506. ACM, 2023. [24] Majeed Kazemitabaar, Justin Chow, Carl Ka To Ma, Barbara J. Ericson, David Weintrop, and Tovi Grossman. Studying the effect of AI code generators on supporting novice learners in introductory programming. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–23. ACM, 2023. [25] Christian Bird, Denae Ford, Thomas Zimmermann, Nicole Forsgren, Eirini Kalliamvakou, Travis Lowdermilk, and Idan Gazit. Taking flight with Copilot. Communications of the ACM, 66(6):56–62, 2023. [26] Richard J. Torraco. Writing integrative literature reviews: Guidelines and examples. Human Resource Development Review, 4(3):356–367, 2005. [27] Hannah Snyder. Literature review as a research methodology: An overview and guidelines. Journal of Business Research, 104:333–339, 2019. [28] Jeffrey E. Froyd, Phillip C. Wankat, and Karl A. Smith. Five major shifts in 100 years of engineering education. Proceedings of the IEEE, 100:1344–1360, 2012. [29] Lisa R. Lattuca, David B. Knight, Hyun Kyoung Ro, and Brian J. Novoselich. Supporting the development of engineers’ interdisciplinary competence. Journal of Engineering Education, 106(1):71–97, 2017. 14
A PREPRINT - AUGUST 3, 2026
[30] James Prather, Brent N. Reeves, Juho Leinonen, Stephen MacNeil, Arisoa S. Randrianasolo, Brett A. Becker, Bailey Kimmel, Jared Wright, and Ben Briggs. The widening gap: The benefits and harms of generative AI for novice programmers. In Proceedings of the 2024 ACM Conference on International Computing Education Research (ICER ’24), pages 469–486. ACM, 2024. [31] James Finnie-Ansley, Paul Denny, Brett A. Becker, Andrew Luxton-Reilly, and James Prather. The robots are coming: Exploring the implications of OpenAI Codex on introductory programming. In Proceedings of the 24th Australasian Computing Education Conference (ACE ’22), pages 10–19. ACM, 2022. [32] Said Elnaffar, Farzad Rashidi, and Abedallah Zaid Abualkishik. Teaching with AI: A systematic review of chatbots, generative tools, and tutoring systems in programming education. International Journal of Learning, Teaching and Educational Research, 25(1):1–28, 2026. [33] Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N. Bennett, Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, and Eric Horvitz. Guidelines for human–AI interaction. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, pages 1–13. ACM, 2019. [34] John D. Lee and Katrina A. See. Trust in automation: Designing for appropriate reliance. Human Factors, 46(1):50–80, 2004. [35] Raja Parasuraman and Victor Riley. Humans and automation: Use, misuse, disuse, abuse. Human Factors, 39(2):230–253, 1997. [36] Raja Parasuraman, Thomas B. Sheridan, and Christopher D. Wickens. A model for types and levels of human interaction with automation. IEEE Transactions on Systems, Man, and Cybernetics—Part A: Systems and Humans, 30(3):286–297, 2000. [37] Lisanne Bainbridge. Ironies of automation. Automatica, 19(6):775–779, 1983. [38] Ben Shneiderman. Human-centered artificial intelligence: Reliable, safe & trustworthy. International Journal of Human–Computer Interaction, 36(6):495–504, 2020. [39] Ben Shneiderman. Human-Centered AI. Oxford University Press, Oxford, 2022. [40] Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S. Bernstein, and Ranjay Krishna. Explanations can reduce overreliance on AI systems during decision-making. Proceedings of the ACM on Human-Computer Interaction, 7(CSCW1):Article 129, 2023. [41] National Academies of Sciences, Engineering, and Medicine. How People Learn II: Learners, Contexts, and Cultures. The National Academies Press, Washington, DC, 2018. [42] Kristin Børte, Katrine Nesje, and Sølvi Lillejord. Barriers to student active learning in higher education. Teaching in Higher Education, 28(3):597–615, 2023. [43] John Sweller, Jeroen J. G. van Merriënboer, and Fred Paas. Cognitive architecture and instructional design: 20 years later. Educational Psychology Review, 31:261–292, 2019. [44] Barry J. Zimmerman. Becoming a self-regulated learner: An overview. Theory Into Practice, 41(2):64–70, 2002. [45] Donald A. Schön. The Reflective Practitioner: How Professionals Think in Action. Basic Books, New York, 1983. [46] Rose Luckin and Mutlu Cukurova. Designing educational technologies in the age of AI: A learning sciencesdriven approach. British Journal of Educational Technology, 50(6):2824–2838, 2019. [47] Enkelejda Kasneci, Kathrin Seßler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Jürgen Pfeffer, Oleksandra Poquet, Michael Sailer, Albrecht Schmidt, Tina Seidel, Matthias Stadler, Jochen Weller, Jochen Kuhn, and Gjergji Kasneci. ChatGPT for good? on opportunities and challenges of large language models for education. Learning and Individual Differences, 103:102274, 2023. [48] Shraddha Barke, Michael B. James, and Nadia Polikarpova. Grounded Copilot: How programmers interact with code-generating models. Proceedings of the ACM on Programming Languages, 7(OOPSLA1):85–111, 2023. [49] Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23), pages 1–22. ACM, 2023. [50] Michael C. Jensen and William H. Meckling. Theory of the firm: Managerial behavior, agency costs and ownership structure. Journal of Financial Economics, 3(4):305–360, 1976. 15
A PREPRINT - AUGUST 3, 2026
[51] Aaron Baird and Likoebe M. Maruping. The next generation of research on IS use: A theoretical framework of delegation to and from agentic IS artifacts. MIS Quarterly, 45(1):315–341, 2021. [52] Olaf Zawacki-Richter, Victoria I. Marı́n, Melissa Bond, and Franziska Gouverneur. Systematic review of research on artificial intelligence applications in higher education—where are the educators? International Journal of Educational Technology in Higher Education, 16(1):39, 2019. [53] Francesca Lucchetti, Zixuan Wu, Arjun Guha, Molly Q. Feldman, and Carolyn Jane Anderson. Substance beats style: Why beginning students fail to code with LLMs. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL-HLT 2025), Volume 1: Long Papers, pages 8541–8610. Association for Computational Linguistics, 2025. [54] ACM/IEEE-CS Joint Task Force on Computing Curricula. Software engineering 2014: Curriculum guidelines for undergraduate degree programs in software engineering. Technical report, ACM and IEEE Computer Society, 2015. [55] Bahar Memarian and Tenzin Doleck. Fairness, accountability, transparency, and ethics (FATE) in artificial intelligence (AI) and higher education: A systematic review. Computers and Education: Artificial Intelligence, 5:100152, 2023. [56] National Institute of Standards and Technology. Artificial intelligence risk management framework (AI RMF 1.0). Technical Report NIST AI 100-1, U.S. Department of Commerce, 2023. [57] European Parliament and Council of the European Union. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (artificial intelligence act). Official Journal of the European Union, L series, 2024. [58] Fengchun Miao, Wayne Holmes, Ronghuai Huang, and Hui Zhang. AI and Education: Guidance for PolicyMakers. UNESCO, Paris, 2021. [59] Brent Mittelstadt. Principles alone cannot guarantee ethical AI. Nature Machine Intelligence, 1(11):501–507, 2019. [60] Gerhard Fischer. Lifelong learning—more than training. Journal of Interactive Learning Research, 11(3):265– 294, 2000. [61] Duri Long and Brian Magerko. What is AI literacy? competencies and design considerations. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, pages 1–16. ACM, 2020. [62] Juebei Chen, Anette Kolmos, and Xiangyun Du. Forms of implementation and challenges of PBL in engineering education: A review of literature. European Journal of Engineering Education, 46(1):90–115, 2021. [63] Mohammed Akour and Mamdouh Alenezi. The enduring impact of gamification on software engineering students’ engagement. International Journal of Technology Enhanced Learning, 2024. [64] CC2020 Task Force. Computing curricula 2020: Paradigms for global computing education. Technical report, ACM and IEEE Computer Society, 2020.
16