ConceptioArchivearXiv CS
arXiv CSopen access

A Technical Typology of AI Systems in Public Administration

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

A Technical Typology of AI Systems in Public Administration Jonathan Rystrøma,1,∗, Chris Schmitzb,1 , Nathan Daviesa,d , Gerhard Hammerschmidb , Albert Meijerc , Chris Russella a University of Oxford, Oxford, UK b Hertie School, Berlin, Germany c Utrecht University, Utrecht, Netherlands d Harvard University, Cambridge, MA, USA

arXiv:2606.31755v1 [cs.CY] 30 Jun 2026

Abstract Research on artificial intelligence (AI) in the public sector often treats “AI” as a single category, neglecting technical distinctions between different AI systems. But these distinctions affect how different systems impact core public values like accountability, procedural justice, and non-discrimination. This paper argues that public administration research would benefit from more technical precision on “AI” and makes three contributions to this end. First, we introduce a typology of five categories of AI systems: handcoded, glass-box, black-box, general-purpose, and agentic systems. We calibrate the typology to public administration by grouping system types by their distinct implications for public values. Second, we evaluate technical precision in recent public administration research about AI by coding 91 highly-cited papers (2019–2025) using our typology. We find widespread imprecision: most papers (55%) leave the studied system underspecified, 31% motivate their work with a different system than they study, and 41% make more general conclusions than the studied system supports. Finally, we give practical recommendations for future research. We highlight common pitfalls to avoid, and suggest that researchers should, at a minimum, provide enough technical detail to locate the studied system in our typology. To this end, we provide a practical guide – a short set of diagnostic questions answerable from public information and without specialist technical knowledge. Keywords: artificial intelligence, public administration, typology, digital government, general-purpose AI, algorithmic governance

1. Introduction “AI” is everywhere in government – and can refer to almost everything. Consider scholarship on public-sector “chatbots” (Androutsopoulou et al., 2019): early chatbots were rulebased systems with fixed, explicitly coded “conversation trees” (Adamopoulou and Moussiades, 2020), but current systems such as ChatGPT are built on externally developed generalpurpose models, can parse ambiguity, and generate fluent, context-sensitive outputs. These systems have different affordances – such as flexible interaction – and constraints, such as externalisation of control and unverifiable training data (Bommasani et al., 2021). The chatbot is not unusual in this: across benefits eligibility, fraud detection, and case triage, the single label of “AI” is routinely used to describe systems with different governance-relevant properties. For questions of governance, what matters about an AI system is its affordances – what it makes possible or forecloses for a particular actor (Zammuto et al., 2007). Even distinguishing ‘data-driven’ AI systems from rule-driven ones, as some research does (Wang et al., 2023), can be too imprecise to clarify these affordances. For example, whether an AI system is legible even to experts depends on its technical underpinnings (Burrell, ∗ Corresponding author. Email: [email protected]; Tel: +44 01865 287210; Oxford Internet Institute, University of Oxford, Oxford OX2 6GG, United Kingdom. 1 These authors contributed equally as joint first authors.

2016): some remain inspectable while others are black boxes – a difference in what auditors can examine or citizens can contest. Similarly, general-purpose AI systems afford vastly different things than task-specific models, for example by lowering barriers to adoption and externalising control over training data. This shift in affordances was consequential enough that the European Union’s AI Act was updated in response (Gstrein et al., 2024; Wang et al., 2026) – but public administration research has not reflected it clearly. It is therefore evident that technical precision about “AI” can be helpful in public administration research – but simultaneously, not all technical distinctions between AI systems matter. It is unclear, for example, that administration scholars need to agonise over the distinction between recurrent neural networks (Sundermeyer et al., 2012) and transformers (Vaswani et al., 2017), or about whether a predictive system uses Support Vector Machines or Random Forests. Neither distinction is likely to affect governance affordances, such as how a public servant may interpret its output. Adding this technical detail may even obscure a paper’s claims, or make it less clear how widely they generalise. This paper addresses the gap these examples illustrate: the field of public administration neither has clarity on how much technical detail about AI systems is necessary, nor the tools or conventions to provide this detail. As a result, it is unclear how well past research reflects meaningful technical differences between AI systems.

We argue that PA research would benefit greatly from a modest increase in technical precision about the studied AI systems: classifying them into a typology of just five categories based on the affordances they provide government actors. To that end, we make three contributions, each answering one research question: Which technical distinctions between AI systems matter to the study of public administration (§2 – §3)? We formulate a test: a technical distinction between two AI systems is necessary whenever their implications for public values differ, and unnecessary otherwise. Arguing that no existing AI taxonomies meet this definition, we introduce our core contribution: a technical typology of five types of AI systems, constructed such that each group has distinct public-value implications. We specify how the governance implications of each group vary across five public values. How well are these technical distinctions made in recent PA research (§4 – §5)? We analyse a sample of recent highlycited papers in public administration and digital government. We find that technical imprecision is widespread, with many papers leaving the studied system unspecified, motivating their work with systems different from the ones they study, or drawing conclusions broader than their evidence supports. How can researchers and policymakers ensure technical precision in the future (§6)? We highlight common pitfalls we find in our analysis, and offer practical recommendations for future research. Specifying the system type need not be costly: we provide a handful of diagnostic questions with which studied AI systems can be placed in our typology, which can be answered using information that is usually available to researchers. The added precision, we argue, improves both the internal validity of individual studies and the cumulative development of the field.

suffers from technical imprecision – and whether such imprecision, if present, weakens its findings – is not self-evident. To find out, we first capture systematically what the field treats as “of interest,” so we can then ask whether those concerns vary across technically different systems. Public values provide such a framework. Public values are the features of government bodies that uphold good governance – equity, legitimacy, and accountability among them. Though private institutions may exhibit some of these, public-sector organisations are distinctive in their steadfast commitment to them, and upholding them is the bedrock of public administration research. Given this centrality, we take all research on AI in PA to study the interaction of AI with at least one public value. Drawing on the “good digital governance” framework of Stalenhoef et al. (2024), we map key research concerns about AI in PA onto five public values (Table 1). Collectively, these dimensions allow us to test the need for technical specificity: wherever two technically different systems bear on one of these values differently, enough specification to tell them apart becomes necessary. While these five interpretations capture a large share of the field’s work, we make no claim that they are exhaustive, given the breadth of AI’s potential effects on public administration. We revisit this limitation in §6.4. Even so, we posit that these issues capture a substantial portion of the field’s work and are therefore sufficient to demonstrate the need for greater technical precision. Adding further values or interpretations may multiply the points at which such systems diverge and may be a fruitful avenue for future research (§6.3). Below, we briefly introduce the five dimensions and their AIrelevant interpretations.

2. Requirements for a Typology We begin by establishing from prior literature why – and when – technical precision about AI systems is required, proceeding in four steps. We first formulate a test for when technical specification matters, by analysing public values: wherever two technically different AI systems bear differently on a public value, researchers should provide enough technical precision to distinguish them. Next, we motivate the need for a novel typology by reviewing existing taxonomies of AI systems and arguing that none of them meets this requirement. Third, we describe how affordance theory can be used to construct a typology that does, slicing systems precisely where their governance-relevant affordances change. Finally, using the affordance lens, we specify three ways technical imprecision may weaken PA research: underspecification of studied AI systems, mischaracterisation of prior research, and overgeneralisation of conclusions.

Dimension

Core Value

AI-relevant interpretation

Democracy

Participation

Transparency: Decision logic must be open to public scrutiny so that citizens and representatives can inspect the basis on which authority is exercised.

Procedural justice

Explainability: Consequential decisions must be reasoned and communicable to those affected, enabling meaningful contestation. Non-discrimination: Individuals must be treated equitably regardless of protected characteristics; historical data must not encode and reproduce past inequalities.

Rule of law

Human rights

Governing capability

Quality of governance

Responsibility

2.1. When Technical Precision Matters: AI and Public Values A large and fast-growing body of scholarship charts the impact of AI systems on government (Wirtz et al., 2019; ValleCruz et al., 2020; Madan and Ashok, 2023). Whether this work

Implementation capacity: The state must be able to deploy and manage AI systems in service of public purposes. Accountability: Public action must be attributable to a responsible actor who can be held answerable for it.

Table 1: Good governance dimensions, core values (Stalenhoef et al., 2024), and AI-specific interpretations common in PA research.

2

Participation. Public decisions should be open to scrutiny, so citizens and their representatives can inspect how authority is exercised (Ananny and Crawford, 2018; Kroll et al., 2017). The AI-relevant research issue is model transparency. Much work evaluates whether, and how, AI systems enable or undermine public participation – by making decision logic visible (Mökander and Schroeder, 2024; Schmitz and Bryson, 2025), opening or overwhelming new channels for public input, or embedding decisions in ways immune to public examination.

meets the requirements of our public-value framework – not because they are poorly constructed, but because each was built for a different purpose. Existing taxonomies organise AI systems using three sets of principles. First, some classify systems by their technical properties – such as learning paradigm, architecture, or the task completed (e.g. classification). The OECD Framework for the Classification of AI Systems, for example, maps systems along four contextual dimensions: people and organisations, technical characteristics, data and input, and task and output (OECD, 2022). Similarly, the INSYTE framework scores systems on eight such dimensions and renders each as a radar chart (Porter et al., 2025). A second organising principle is application domain and potentially its associated risk level, as in the EU AI Act’s tiered scheme and derived regulatory analysis (Laux et al., 2024; Buttaboni and Floridi, 2026). Application-based classification is particularly prevalent in PA taxonomies, which catalogue systems by use case and government function (Berryhill et al., 2019; Wirtz et al., 2019). The third organising principle is AI systems’ role in a decision, distinguishing AI that suggests, offloads, or supersedes a human judgement (Roehl and Hansen, 2024; König and Wenzelburger, 2020).

Procedural Justice. Administrative law in most democratic systems mandates that consequential decisions be reasoned and communicated to those affected (Wachter et al., 2017a; de Bruijn et al., 2022; Buttaboni and Floridi, 2026). A common research theme is how the explainability of AI systems bears on this duty: whether a functional explanation suffices for procedural legitimacy (Lazar, 2024), whether human-illegible decision rules leave affected citizens any meaningful way to contest a decision, and what a “right to explanation” can deliver when the available explanations are only post-hoc approximations. Human Rights. Public bodies must treat individuals affected by their actions or decisions equitably. Scholars investigate how AI may affect such non-discrimination: for example, many investigate how statistical regularities in historical training data can introduce bias by encoding and reproducing past inequalities (Barocas and Selbst, 2016; Corbett-Davies et al., 2023). Another research strand is sociotechnical design: whether AIassisted governance satisfies non-discrimination hinges substantially on how the underlying system was constructed and where bias might originate (Selbst et al., 2019; Wachter et al., 2021a; Green, 2022).

Each of these approaches serves a distinct purpose, such as regulation (Buttaboni and Floridi, 2026) or safety engineering (Porter et al., 2025). However, it is unclear whether they result in taxonomies suitable for PA analysis. Via our public value framework, we can pose a simple criterion to evaluate this: a suitable taxonomy should group technically distinct systems that share public-value implications, and separate those whose implications differ. Existing taxonomies fail this test in three ways, two of which are sketched in Figure 1. First, many underspecify: they group systems whose governance-relevant properties differ. A shared risk tier or use-case label can place a predictive policing model beside a hospital triage assistant, though the two diverge in explainability, bias mechanisms, and accountability. Second, especially technical taxonomies overspecify: they split systems too granularly, even where their public-value implications match. Crucially, having too many categories makes it unclear where meaningful distinctions lie. Specifying that a studied model is a random forest, for example, does not clarify how takeaways may transfer to neural networks. Third, many conflate functional and technical categories. A “suggesting” system, for example, could produce a single risk score or paragraphs of text, which evidently vary in governance-relevant dimensions.

Quality of Governance. We expect public organisations to be effective, efficient, and economical. Scholars frequently analyse how introducing AI requires implementation capacity: the organisational and technical competence to deploy and manage technology in service of public aims (Lawrence et al., 2023; Madan and Ashok, 2023; Neumann et al., 2024). This includes analysis of the in-house skills to procure, integrate, maintain, and oversee AI systems, of introduced dependence on external vendors and infrastructure, and of whether deployed systems increase efficiency. Responsibility. Finally, public action must be attributable to a specific public actor who can answer for it (Nissenbaum, 1996; Bovens, 2007). An active research area is accountability: because AI systems distribute decision-making across complex technical architectures and lengthy supply chains, scholars ask where accountability should land – with the official who relied on a system’s output, the agency that deployed it, or the vendor that built it – and whether existing mechanisms can still locate a responsible actor at all (Matthias, 2004; Sterz et al., 2024).

Given that none of the mentioned taxonomies were designed for PA analysis, it is understandable that none pass these tests. However, there is therefore a clear need for a typology of AI systems targeted at clarifying their public-value implications. Currently, researchers risk either underspecifying their scope by using no taxonomy at all – invoking “AI”, “algorithms”, or “automated decision-making” generically – or relying on an unsuitable existing taxonomy, which does not clearly track such properties.

2.2. Why Existing AI Taxonomies are Unsuitable Several taxonomies of AI systems already exist, within public administration and beyond. However, we argue none of them 3

2.4. Three Forms of Technical Imprecision Using the lens of affordance thresholds, we can phrase more precisely how technical imprecision on AI could weaken PA research. We theorise three potential forms of imprecision here; in our review of the field (§4), we operationalise these definitions and measure how often each occurs. (a) underspecified

(b) overspecified

(c) our approach

• Underspecification could arise when too little technical detail is provided, such that it is not clear what the affordance profile of a studied AI system is. For example, describing a system only as “AI” or an “algorithm” may not allow its specific affordances to be recovered, such that the generalisability of any claims made cannot be verified.

Figure 1: Conceptual diagram of three approaches to classifying AI systems. (a) An underspecified taxonomy fails to distinguish between systems with different governance implications. (b) An overspecified taxonomy draws too many distinctions between systems, obscuring differences in their affordances. (c) Rather than generating a novel taxonomy from scratch, our approach identifies PA-relevant affordance thresholds (solid red lines) within existing taxonomies.

• Mischaracterisation may arise when a paper motivates its approach with one type of system, then studies another. For example, a paper that opens on the dangers of opaque, black-box risk scoring and then examines a rule-based eligibility calculator ports over inaccurate assumptions about the system’s afforded transparency.

2.3. Affordance Theory as the Organising Principle Affordance theory provides a suitable organising principle for a better-calibrated taxonomy. Developed by Gibson (1979) and elaborated in organisational and information-systems research (Zammuto et al., 2007; Majchrzak and Markus, 2013; Leonardi, 2011), affordances describe what a technology makes possible or forecloses for a particular actor in a particular setting. For example, Bovens and Zouridis (2002) famously motivate the concept of screen-level bureaucracy with the affordances of a digital form over a paper one. Affordance theory reflects that a technology’s public-value impacts depend on its architecture – but are not solely defined by it. They are also shaped by the interaction between the technology and the goals, capacities, and context of those who use it. Identical algorithms can produce different organisational outcomes depending on the setting (Meijer et al., 2021), and how AI systems affect public values is a question of sociotechnical design (Schmitz and Bryson, 2025). Indeed, such organisational and contextual analysis remains vital. Throughout this work, we do not suggest that technical detail should replace such analysis, but that it must complement it: the technical architecture defines which affordances exist to begin with. For example, open-weights large language models (LLMs) can afford public organisations the processing of sensitive data where proprietary models do not, but this need not imply that they are used for that purpose (Robinson, 2026). An affordance-based typology is therefore suitable to bridge organisational and technical analysis. In §3, we construct this typology by drawing boundaries we term affordance thresholds: distinctions only between sets of AI systems with different public value-relevant affordances, as visualised in Fig. 1 (c). A technical difference between two systems that alters whether a citizen can contest a decision, or whether an auditor can inspect its logic, produces such a threshold: the typology should distinguish between them. In contrast, a shift that does not affect what actors can do – for example, one that only improves predictive accuracy – does not cross an affordance threshold, and the systems should remain in the same category.

• Overgeneralisation could occur if a paper presents its findings as more general than the evidence supports. For example, a conclusion about “AI in government” drawn from the study of a black-box model may not apply to generalpurpose models. Evidently, claims can cross affordance thresholds without being imprecise in one of these ways: a well-scoped insight from a case study of one system could readily generalise to many other types, for example. Technical precision allows us to distinguish which claims do so validly. 3. Typology of AI Systems in Public Administration This section introduces our core contribution: a technical typology of five classes of AI systems, presented in Figure 2. We define and describe each class of system, specify when a system crosses the “affordance threshold” between classes, and describe the distinct public-value implications of each class (Stalenhoef et al., 2024). These are summarized in Table 2. We deliberately do not attempt to define “AI” ourselves. All of the levels in our typology have been termed “AI” by some widely-cited papers, as we analyse in §6.1. Rather, we provide a simple but nuanced vocabulary for talking about these types of systems. 3.1. Hand-coded systems The first layer is hand-coded systems: systems whose decision rules are authored in code rather than learned from data. The canonical example is rule-based public benefit administration (Enqvist, 2024), and the category has been extensively studied in the e-governance and digital government literatures (Dunleavy, 2006; Zouridis et al., 2020). Hand-coded systems drive the shift from street-level to system-level bureaucracy (Bovens and Zouridis, 2002), in which discretion is encoded in software rather than exercised case-by-case by individual officials. 4

System Type

Technical Description

Diagnostic Question

Hand-coded

Discretion or policy encoded as explicit rules (Enqvist, 2024).

Glass-box

Rules learned from data, but legible to experts (Rudin, 2019).

Black-box

Learned logic defies expert inspection (Burrell, 2016).

General-purpose

General, externally pretrained models (Bommasani et al., 2021).

Does the system adapt a generally-trained model to a specific task?

GPT-4 via Azure for citizen queries (Bright et al., 2025).

Agentic

General-purpose model scaffolded to act over time, e.g. via tool-use (Yao et al., 2023).

Is the general-purpose model scaffolded to act autonomously over time, e.g. via tools?

Benefits agent querying registries autonomously (Ilves et al., 2025).

Illustrative example Florida ACCESS benefits platform (Citron, 2008).

Are the rules the system follows learned from data, rather than authored in code? Is it infeasible for any human expert to understand the system’s learned logic and internal functioning?

Allegheny Family Screening Tool (Vaithianathan et al., 2017). COMPAS recidivism scoring (Dressel and Farid, 2018).

Threshold: qualitative distinction ⊂

Subset: special case of the type above

Figure 2: Overview of Technical Typology of AI Systems. The five system classes (left) are separated by two kinds of relationships. Orange diamonds mark thresholds – qualitative distinctions; ⊂ markers mark subsets, where a system is a special case of the one above. Diagnostic questions distinguish each class from the class above, and can therefore be used as a specification tool, as described in §6.2.

in code; once rules begin to be derived from data, the system crosses into the glass-box layer.

Affordance: efficiency and traceability. While hand-coded systems can improve efficiency, traceability, and standardise procedures, they simultaneously diffuse accountability (Citron, 2008) and flatten local complexities, exacerbating existing legibility dynamics (Scott, 1998). Furthermore, the shift to digital systems enables large-scale data collection and analysis, with significant implications for privacy (Zuboff, 2015).

Example: Hand-coded systems U.S. state public-benefits eligibility platforms such as Florida’s ACCESS, Texas’s TIERS, and California’s CalWIN are examples of hand-coded systems. Legal scholarship describes these systems as encoding administrative policy into software rules, with caseworkers often reviewing sample outputs before finalisation. Their failures typically arose not from statistical learning but from incorrectly coded rules or policy distortions embedded in the software itself (Citron, 2008).

Challenge: complexity and bias from data-rule interaction. It is important to note that rule complexity can be immense even within hand-coded systems. While the logic is authored rather than learned, thousands of intersecting rules can still exceed human cognitive limits (Simon, 1947). Moreover, understanding a rule does not equate to understanding its interaction with realworld data. Even fully legible, hand-coded rules can introduce bias because discrimination is a function of how the system affects outcomes in practice (Wachter et al., 2021b). For instance, a seemingly neutral hand-coded rule that declines benefits if an applicant has a continuous unemployed_duration > 6 months may inadvertently discriminate against women taking maternity leave. Thus, bias and discrimination can manifest through the interaction between fixed rules and contextual realities, entirely independent of statistical learning.

3.2. Glass-box systems The second layer is glass-box systems: systems whose decisions rules are learned from data, but whose learned logic remains inspectable – at least by experts (Burrell, 2016; Rudin, 2019). Canonical methods include linear and logistic regression (Fox, 2015), decision trees (de Ville, 2013), Principal Component Analysis (Abdi and Williams, 2010), and TF-IDF (Bafna et al., 2016) – methods used for prediction, classification, and representation across applications such as childwelfare risk modelling (Hall et al., 2024), welfare-recipient predictions (Sansone and Zhu, 2023), and policy-document text analysis (Altaweel et al., 2019). This is a binary shift into machine learning; rules are no longer specified exhaustively but derived from data. The shift opens up case-level adaptation and leveraging of historical data patterns, while introducing new governance challenges around bias and accountability.

Boundary. Whether hand-coded systems truly fall under “AI” is contested. Some scholars explicitly include rule-based systems (Selten et al., 2023), and scholarship on Robotic Process Automation (RPA) often uses the language of AI. Our aim is not to settle that question, but to sharpen the vocabulary used to distinguish between different kinds of systems. Hand-coded systems remain in this layer so long as their rules are authored 5

Layer

Democracy (Participation)

Rule of Law (Justice & Rights)

Governing Capability (Quality & Responsibility)

Hand-coded

Shifts to system-level administration, flattening local complexities (Bovens and Zouridis, 2002; Scott, 1998).

Standardises procedures but expands data matching, risking privacy (Citron, 2008).

Enhances efficiency and traceability, yet diffuses accountability (Bovens and Zouridis, 2002; Citron, 2008).

Glass-box

Adapts to case variation but risks reproducing inequalities and displacing public values (D’ignazio and Klein, 2023; Wachter et al., 2021a; Green and Chen, 2021).

Legibility permits auditing of learned features, though accountability blurs (Sandvig et al., 2014).

Shared decision-making among model, data, and user creates “moral crumple zones” (Elish, 2019).

Black-box

Boosts performance in unstructured domains but restricts citizen capacity to contest decisions (Wang et al., 2023; Valle-Cruz et al., 2024).

Threatens procedural justice via justification deficits, impeding legal verification (Grimmelikhuijsen and Meijer, 2022; Rudin, 2019).

Inscrutiable internal logic obscures decision pathways and contests responsibility (Cobbe et al., 2023; Janssen et al., 2020).

General-purpose

Lowers adoption barriers via natural language, but democratisation of control remains partial (Bright et al., 2025; Hashem et al., 2025).

Persuasiveness and unfaithful explanations risk automation bias; inaccessible training data heightens privacy risks (Mayne et al., 2026; Bender et al., 2021).

Distributed responsibility and hardware dependence create vendor lock-in and complicate auditing (Brown, 2023; Mökander et al., 2024)

Agentic

Reduces citizen friction but scales demand, risking administrative overload (Ilves et al., 2025; Yun et al., 2024; Marques et al., 2025).

Diffuses discretionary power by acting proactively, challenging traditional procedural safeguards and accountability (Chan et al., 2023).

Demands runtime oversight rather than ex post evaluation; capabilities present a “jagged frontier” (Schmitz et al., 2025; Chan et al., 2024; Dell’Acqua et al., 2026).

Table 2: Summary of public-value implications of each layer of the taxonomy.

Affordance: case-level adaptation. While glass-box systems provide novel affordances – glass-box systems can adapt to historical variations and improve administrative fit where fixed rules are too coarse – learned rules are deeply entangled with their training data. They reproduce historical inequalities and encode patterns that do not reflect current public values or legal commitments (D’ignazio and Klein, 2023; Wachter et al., 2021a). This is especially problematic in public administration, since predictive accuracy is not itself the goal; decisions must also reflect normative commitments (Grimmelikhuijsen and Meijer, 2022). Decision-makers therefore risk placing undue weight on statistical outputs, even when they displace other relevant public values (Green and Chen, 2021).

ditable because scrutiny can occur at the level of learned rules and features rather than relying primarily on indirect experiments (Sandvig et al., 2014). Boundary. Glass-box systems remain in this layer so long as their learned logic is meaningfully inspectable; once that ceases to be the case, they move into the black-box layer. Example: Glass-box systems The Allegheny Family Screening Tool (AFST), used in Allegheny County, Pennsylvania, scores families’ risk of child abuse or neglect on a scale of 1–20 using a logistic regression model trained on historical child welfare records (Vaithianathan et al., 2017). Auditors and oversight bodies can inspect the model’s coefficients and understand which variables drive higher scores, making the learned logic comparatively legible.

Challenges: transparency-fairness tradeoff and responsibility. Importantly, transparency and rule complexity exist on a spectrum. While the logic of a glass-box system is theoretically inspectable, a linear regression with thousands of variables or a highly branched decision tree can easily exceed human cognitive limits (Lipton, 2018). Furthermore, transparency is often antagonistic with respect to fairness. Because glass-box systems are constrained in their complexity, the easiest way to achieve acceptable baseline performance is often to optimise for the majority while ignoring minority groups or complex edge cases (Ferry et al., 2025). This flexibility-interpretability tradeoff is a technical inevitability – and a key factor in explaining why public administrators might choose to implement more complex, black-box systems. At the same time, responsibility becomes more diffuse: when a rule is generated from data, accountability can blur between model, data, designer, and user, creating the conditions for “moral crumple zones” (Elish, 2019). Still, compared to later black-box systems, glass-box systems remain relatively au-

3.3. Black-box systems The third layer is black-box systems: systems whose learned logic resists meaningful inspection, even by experts (systems with “algorithmic opacity” following Burrell, 2016). Canonical examples include deep neural networks (Goodfellow et al., 2016), random forests (Breiman, 2001), and UMAP (McInnes et al., 2020) deployed across unstructured domains (e.g., face recognition; Rezende, 2020) and accuracy critical applications (e.g., medical imaging; Rystrøm et al., 2026a). The exact threshold between glass-box and black-box systems lies on a spectrum of algorithmic complexity: as models grow more complex, expert inspection becomes algorithmically infeasible. Consequently, approximate post-hoc explanation becomes a technical necessity (Mittelstadt et al., 2019), and audits must become indirect through behavioural testing 6

(Sandvig et al., 2014) – eroding the ability to justify, contest, or mechanistically explain individual decisions.

General-purpose systems are a subset of black-box systems, distinguished from other black-box systems by broad pre-training for general capabilities (Bommasani et al., 2021), rather than narrow, task-specific utilisation of experience (Mitchell, 2013). Consequently, the primary shift is one of externalisation. Because of the complexity of general-purpose systems, the training process, data provenance, and model design is often done by external model providers and thus no longer fully inspectable or controllable by the deploying organisation (Mökander et al., 2024).

Affordance: complex performance. What black-box systems lack in transparency, they compensate for through improved performance in high-dimensional or unstructured domains, where simpler glass-box models struggle (Krizhevsky et al., 2012). As a result, black-box systems are widely deployed in government despite their governance challenges (Valle-Cruz et al., 2024). Challenges: procedural justice and responsibility. Opacity specifically threatens procedural justice. First, it creates a justification problem: it becomes difficult to ensure legal requirements are consistently met when the decision-making logic cannot be inspected (Grimmelikhuijsen and Meijer, 2022). Second, it creates a contestation problem: citizens lack the knowledge or mechanisms to challenge decisions effectively (Wang et al., 2023). Third, it creates an explanation problem: even where a “right to explanation” is invoked, explanations are limited to local approximations rather than exact causal pathways (Wachter et al., 2017b). Opacity also makes responsibility more contested. Complex supply chains of models and data make decision pathways less reconstructible (Cobbe et al., 2023). Still, compared to later systems, black-box models are typically developed and trained within organisational boundaries, meaning that data collection and model development remain under institutional control.

Affordance: lower adoption barriers. A key affordance of externalisation is that it lowers barriers to adoption. Because much of the technical complexity sits outside the organisation, general-purpose systems can be integrated into administrative processes without specialised AI engineering expertise (Bright et al., 2025). In addition, natural-language interfaces can make existing bureaucratic systems more accessible by mediating interactions between citizens and administrative procedures (Hashem et al., 2025). However, this “democratisation” is partial: while it becomes easier to use such systems, control over their behaviour remains limited. This is reinforced by their computational requirements. Many general-purpose systems require specialised hardware and are therefore deployed via cloud infrastructure, creating further dependence on external providers and reinforcing the externalisation of control (Qiu et al., 2025). Challenges: distributed responsibility and procedural justice. This transition fundamentally alters how the system utilises experience to improve performance. While traditional machine learning uses experience to improve at a specific, narrow task, general-purpose systems leverage broad pre-training to develop versatile capabilities that are then applied to specific downstream contexts (Brown et al., 2020). In technical research, these are often referred to as ‘foundation models’ (Bommasani et al., 2021), though the term general-purpose highlights their functional role in the public sector as adaptable building blocks. By shifting the technical burden of training to external providers, these systems introduce a critical governance vulnerability: the externalisation of evaluation. Organisations may be tempted to rely on a provider’s general benchmarks rather than rigorously testing the system for the specific, local context of an administrative task (Rystrøm et al., 2026b). Externalisation primarily threatens responsibility and accountability. Responsibility is not only diffused but distributed across a complex supply chain involving model developers, platform providers, and deploying organisations, each with only partial control over outcomes (Brown, 2023). This complicates auditing; understanding a given application requires evaluating not just the downstream implementation, but also the underlying model and the opaque governance practices of its provider (Mökander et al., 2024). In practice, this makes it difficult to determine whether problematic outcomes, such as discriminatory bias, stem from the training data, the model design, or the specific prompting patterns. Furthermore, the scale and opacity of training data intensify human rights concerns regarding pri-

Boundary. Black-box systems remain in this layer so long as the underlying model is trained directly for the task at hand; once the model is instead pre-trained for a general objective and adapted to downstream tasks, the system crosses into the general-purpose layer. Example: Black-box systems Amazon’s scrapped AI resume screening system serves as an illustrative case for the technical mechanisms of bias in opaque systems. The model was trained on ten years of resumes but developed a pervasive bias against women, penalising applications that alluded to women’s colleges or sports (Dastin, 2018). Because the system relied on weak, highly contextual proxy signals within an opaque neural architecture, the bias was not located in a single, auditable rule (Raghavan et al., 2020). Instead, it emerged from the model’s interaction with historical data, demonstrating how opacity shifts the burden of oversight from inspecting legible rules to auditing behavioural outcomes.

3.4. General-purpose systems The fourth layer is general-purpose systems: systems pretrained on general tasks – such as next-token prediction – that can be adapted to diverse downstream applications through mechanisms like transfer learning or natural language instructions (Brown et al., 2020). Canonical examples include large language models (LLMs) like ChatGPT (OpenAI, 2022), various embedding models (Devlin et al., 2019), and vision models – all of which are being increasingly deployed by public sector organisations (Bright et al., 2025). 7

vacy and the automated reproduction of harmful social patterns (Bender et al., 2021). Finally, general-purpose systems pose significant challenges for procedural justice. Because many applications rely on natural-language interfaces, there is a temptation to treat the system itself as a source of explanation (Zhu et al., 2024). Yet, such linguistic outputs are not guaranteed to reflect the underlying basis of a decision (Mayne et al., 2026). The persuasive and anthropomorphic nature of these outputs can make explanations appear authoritative even when they are unfaithful (Salvi et al., 2025), potentially increasing automation bias and reinforcing the “moral crumple zones” that obscure human responsibility (Elish, 2019).

of steps over time. In the public sector, this promises to transform service delivery by bridging fragmented infrastructures; as argued by Ilves et al. (2025), agentic systems can act across heterogeneous data sources and administrative systems to streamline bureaucratic procedures. From the citizen’s perspective, agentic systems also reshape the interface with government by acting as proactive intermediaries. By helping citizens navigate complex eligibility requirements and administrative hurdles, these agents can significantly reduce the effort required to access public services (Yun et al., 2024; Jo et al., 2025).

Boundary. While general-purpose systems are versatile, they remain primarily reactive, mapping inputs to outputs. Once systems move beyond this paradigm to act proactively and interact with their environment over time, they enter the final layer of the typology.

Challenges: jagged reliability, surging demand, and runtime oversight. The reliability of these systems in public-sector environments remains a significant concern. There are currently no evaluations that accurately capture their capacity for administrative tasks (Rystrøm et al., 2026b), a problem compounded by the “jagged frontier” of agentic capabilities, which makes it difficult to predict which tasks they will perform reliably and where they will fail (Dell’Acqua et al., 2026).

Example: General-purpose systems New York City’s “MyCity” AI chatbot, launched to help business owners navigate local bureaucracy, illustrates the dangers of externalised evaluation. Powered by external foundation models, the system was deployed to provide legal and regulatory guidance without sufficient domain-specific testing. Consequently, the chatbot frequently hallucinated and advised citizens to break the law—wrongly suggesting, for instance, that employers could legally take a cut of their workers’ tips (Lecher, 2024). This case highlights the severe risks of assuming a foundation model’s general competence will safely transfer to a specific administrative context without rigorous, localised evaluation.

Furthermore, the reduction in interaction costs may substantially increase the total demand for public services, as agents can interface with government systems at scale on behalf of individuals (Marques et al., 2025). Without appropriate institutional countermeasures, this surge in automated requests may challenge the fundamental processing capacity and responsiveness of administrative systems. Finally, agentic systems introduce substantial risks for responsibility and procedural justice (Chan et al., 2023). Because these systems act autonomously across time, responsibility often shifts away from discrete moments of decision toward ongoing processes, making it harder to attribute specific outcomes to individual actors (Schmitz and Bryson, 2025). Governance therefore requires new capacities for runtime oversight – the ability to monitor, constrain, and intervene in live system behaviour as it unfolds (Chan et al., 2024). Without such mechanisms, the use of agentic systems risks diffusing discretionary power away from human decision-makers and undermining established structures of accountability. Ultimately, agentic systems mark a definitive shift from systems that produce outputs to systems that act, introducing a distinct set of governance challenges that cannot be addressed through existing static approaches alone (Chan et al., 2023).

3.5. Agentic systems The final layer of the typology is agentic systems: AI systems that can pursue complex and general goals, act with autonomy, and affect their environment (Kasirzadeh and Gabriel, 2025). Examples include large language models with access to external ‘tools’ and APIs (Yao et al., 2023) as well as autonomous vehicles (e.g., Rao and Frtunikj, 2018). Agentic systems are a subset of general-purpose systems, as they are usually created by “scaffolding” or augmenting general-purpose systems. The conceptual shift is therefore from strictly reactive input-output mapping to proactive, iterative loops of reasoning and action within open-ended environments (Yao et al., 2023). As a result, the primary governance shift is from instance-level decision-making (e.g., the correctness of a classification) to process-level action (e.g., the suitability of guardrails for an agent), requiring a move from ex-post evaluation towards runtime monitoring and intervention in ongoing system behaviour (Schmitz et al., 2025).

Example: Agentic systems While agentic systems are still in their infancy, Bürokratt from Estonia offers an early vision. Originally a hand-coded chatbot, Bürokratt is being upgraded with an LLM-based orchestration system to more intelligently handle complex citizen queries by proactively querying information from Estonia’s public sector data infrastructure. The long-term vision is for Bürokratt to be an interface to and orchestrator of different government agencies’ agent systems (Ilves et al., 2025).

Affordance: service delivery and citizen interface. This transition represents a profound shift in the evolving definition of the machine learning “task.” Rather than producing discrete, static outputs, agentic systems use experience to navigate sequences 8

4. Analysing Technical Imprecision in Public Administration Research on AI

4.2. Coding Once we have selected high-impact papers, we code how each paper specifies, motivates, and generalises about AI systems. We conduct structured qualitative coding, following established procedures for systematic content analysis (Saldaña, 2025; Krippendorff, 2019).

We next validate our typology by analysing whether it would improve the precision of existing research. To do so, we code impactful public administration and digital government papers on AI from the last seven years, evaluating whether technical specification using the typology would mitigate their imprecision (§2.4). This section details our methodology. We present the findings in §5 and discuss their significance in §6.

Coding Scheme. The primary unit of coding is not the entire paper, but a strand: a concise summary of a key claim in the paper. An example strand is “public organisations are not held to a lower responsibility standard for algorithmic versus human discrimination”. This meso-level analysis (Miles et al., 2014) has two advantages: it captures papers with more nuance, and allows systematic comparison of claims within individual papers. For each strand, we code three pieces of metadata. First, we code it as either empirical, motivation, or conclusion. Empirical strands identify which AI systems or class of systems the paper studies; motivation strands summarise how the paper positions itself in existing AI literature; and conclusion strands summarise the core claims the paper makes. Second, coders classify the AI system the strand addresses using our typology (§3). Where the paper provides insufficient technical detail to determine the system type, the strand is coded as underspecified. Third, coders assign one or more public-value dimensions from the governance framework (Stalenhoef et al., 2024). Conceptual papers and literature reviews are coded using the same procedure. Where a paper studies no concrete system, coders treat the paper’s main motivating examples or conceptualisation of ‘AI’ as the empirical strands. If this conceptualisation provides enough information to identify a system type, it is classified as such; if it deliberately identifies a broad category within an affordance threshold, it is coded as justifiably generic; otherwise, it is coded as underspecified. For analysis, the strands are aggregated on the paper level as described in §4.3. This structure enables systematic comparison between the systems used to motivate a paper, the systems actually studied, and the systems to which conclusions are applied.

4.1. Data To identify impactful papers studying AI in government and public administration, we conduct a systematic literature search using OpenAlex (Priem et al., 2022). We focus on leading journals in public administration and digital governance, selected based on their relevance and citation impact within the field (van Thiel, 2021; Heeks and Bailur, 2007). The full venue list and keyword query are reported in Appendix A; all data, code, and materials are available online.2 We include papers published between 2019 and 2025. This period ensures coverage of all categories in our typology, with a slight underrepresentation of agentic systems, which only began emerging in 2023 (Yao et al., 2023). From this pool, we select the most-cited papers per year separately for public administration and digital government venues. We use citations as a proxy for impact – a common but contested heuristic (Flyvbjerg and Turner, 2018) – and sample per year to mitigate temporal bias in citation accumulation (Bornmann and Daniel, 2008). We determine how many papers to include per venue type and year through a pre-specified stability analysis of the aggregate estimates produced by our sampling rule. We evaluate whether the three main outcomes remain stable as additional papers are added within each stream-year cell. This follows the logic of stability-based sample adequacy, in which estimates are considered sufficient once they remain within a prespecified tolerance corridor as samples are added (Schönbrodt and Perugini, 2013), while also drawing on work on incremental-sampling thresholds (Guest et al., 2020). We iteratively expand the corpus, recompute the aggregate outcomes, and estimate uncertainty via bootstrap resampling (Davison and Hinkley, 1997). Full methodological details are provided in Appendix B. We find that K = 8 papers per venue type (digital government and public administration) is sufficient to stabilise the aggregate outcomes under our sampling rule, yielding an initial set of 109 papers. We then manually screen all papers and exclude those that use AI purely as a methodological tool (e.g., using AI to predict corruption; Lima and Delen, 2020) or that do not treat AI as an empirical or theoretical subject. The final corpus consists of 91 papers. The full screening flow, including counts at each stage, is reported in Appendix A (Figure A.7).

4.3. Operationalising Imprecision Using the strand construct, we formalise the three failures introduced in §2.4. We treat each as paper-level outcomes, measured by aggregating strand-level codes. Where relevant, we apply an any-mismatch rule: a paper is flagged if at least one strand exhibits the imprecision. This reflects our interest in whether greater technical specification would improve the precision of a paper’s framing or claims. Underspecification. A paper is underspecified if any empirical strand is coded underspecified: the paper gives too little detail to place the system it studies within our typology (§3). Mischaracterisation. A paper is mischaracterised if at least one motivation strand invokes a system type that differs consequentially from the one studied – where the mismatch, not merely the wording, matters for the claim being motivated.

2 https://anonymous.4open.science/r/AITypology4PA-4FC6/

9

n=50

Overgeneralisation. A paper is overgeneralised if at least one conclusion strand reaches beyond the system type its empirical strands support, and the gap matters for the claim’s validity or policy relevance. This is the costliest failure for cumulative science: later work may build on claims that do not hold across technical contexts (Schroeder, 2020).

% of Papers

50

LLM-Assisted Extraction. We use LLM-assisted extraction to support the initial extraction of candidate strands (Dai et al., 2023; Nguyen-Trung, 2025). The LLM is used to impose a consistent preliminary structure for each paper; all coding decisions are made exclusively by human coders. Each paper is first converted into a full-text markdown representation and provided to the LLM together with the full coding prompt reproduced in Appendix A.2. We use Gemini 3.1 Flash-Lite Preview (Gemini Team et al., 2025) to produce a structured extraction for each paper. After receiving the LLM output, the assigned human coder reads the full paper and revises, adds, merges, or removes strands as required. Coders independently make all judgements regarding system classification, publicvalue dimensions, mischaracterisation, and overgeneralisation, and do not receive LLM-generated suggestions, such that reported rates depend on human judgement alone. The appendix codebook (Appendix C) is the prompt used for LLM-assisted extraction, reproduced verbatim. It defines both the preliminary extraction task given to the LLM and the annotation guidance used by human coders. The codebook specifies the typology labels, the public-value dimensions (following Stalenhoef et al., 2024), the definition of each strand type, and the decision rules for identifying consequential mischaracterisation and overgeneralisation.

40 30 20 10 0

n=19 n=12

n=11 n=2

n=4

n=0

x x eric ose gentic ded ified -bo -bo A pec ly Gen and-co Glass Black al Purp s r e b H er n Und ustifia e G J Figure 3: Underspecification. 55% of papers provide insufficient information to determine which system is empirically studied. The most commonly studied system is black-box systems, with agentic systems completely unstudied.

novel ones. Overturned flags lower the reported rate, while uncounted misses can only raise the true rate (Begg and Greenes, 1983). The reported figures are therefore a conservative estimate of the prevalence of imprecision in the corpus. However, our design trades off against reviewer blinding: because adjudication is triggered by a flag, the second coder knows an error was proposed. Of the 148 strands flagged by the primary coder, 128 were retained on adjudication, and 20 (14%) were overturned. This non-trivial but modest rate is consistent with adjudication working as a genuine refinement. 5. Findings

Scheme Validation and Refinement. The coding scheme was piloted on a subset of 10 papers coded by all authors, after which the codebook was refined to improve conceptual clarity and consistency (Mayring, 2015; Schreier, 2012). The remaining papers were randomly assigned to authors for independent coding. Ambiguous cases were recorded during coding and, after reliability assessment, discussed among the authors and resolved by consensus. These consensus decisions form the final dataset used for the analysis below.

We find significant imprecision across all three analysed categories. The results below present the overall rates and their relation to public values and typology dimensions. We find no changes in rate over time (Fig. 5). Summary statistics and figure-generation scripts are available in the project repository.4 5.1. Underspecification Of 91 coded papers, 50 (55%) are underspecified: across all empirical references to the studied system, there is insufficient information to classify it with certainty. Fig. 3 shows the number of analysed papers empirically studying each type of system in our typology. Among fully specified papers, black-box systems are the most commonly studied category (N=19). In contrast, generalpurpose systems are relatively understudied. Only 11 papers explicitly analyse general-purpose systems empirically, despite their growing prominence (Straub et al., 2023). No papers are classified as studying agentic systems in our corpus. Papers mentioning ‘agents’ primarily engage with these systems at a conceptual level (e.g., “cognitive robots” in Wirtz et al., 2019) or in relation to physical automation (e.g., drones in Straub et al., 2024), rather than contemporary LLM-based agents

Adjudication. The judgements driving our paper-level outcomes are validated through codebook-grounded adjudication (Krippendorff, 2019).3 For each paper, a second author reassesses every strand whose value sets a paper-level flag – motivation strands flagged as mischaracterised, conclusion strands as overgeneralised, and the empirical strands of any underspecified paper – against the codebook and the paper text, retaining a flag only where its documented decision rule is met and removing it otherwise. Residual disagreements are settled by a third author (O’Connor and Joffe, 2020). We design adjudication conservatively: second coders can only remove imprecision flags set by the first coder, not add 3 A double-coded subset large enough to estimate inter-coder agreement with usable precision was not feasible given corpus size and per-paper coding cost; a power analysis is provided in the repository.

4 https://anonymous.4open.science/r/AITypology4PA-4FC6/

10

0.50 0.25

6% (n=16)

0.00

16% (n=58)

14% (n=21)

27% (n=71)

12% (n=50)

ts ce lity ice ion righ pat sibi nan just i r n n l c e i a o a t v p ur Par f go Hum Res ced ty o i l Pro a Qu

31% (n=91)

50 0

2019

2020

2021

Underspecified

l Tota

2022

2023

2024

Publication Year

Mischaracterised

2025

Overgeneralised

Figure 5: Trends in rates. We find no significant changes in any specification category over time.

Public Sector Value

Figure 4: Mischaracterisation. Proportion of papers that have mismatches between systems mentioned in the motivation and the systems empirically studied. In total, 31% of papers have mischaracterised strands. Error-bars are 95% Wilson (1927) scores.

(Ilves et al., 2025). However, as our citation-weighted sampling structurally disadvantages recent work (§6.4), some of this absence could reflect citation lag, as discussed in §6.1.3. 5.2. Mischaracterisation 31% of coded papers mischaracterise AI systems: they exhibit at least one consequential mismatch between motivating and empirically analysed systems. Fig. 4 shows the proportion of papers that make at least one mischaracterised claim within each governance dimension. We see statistically similar rates across value dimensions.

44% (n=2) (n=4) 26% 18% 58% 41% Total (n=50) (n=19) (n=11) (n=12) (n=91) 12% (n=2) (n=3) 25% (n=3) (n=4) 16% Participation (n=16) (n=32) (n=8) 25% (n=2) (n=2) 0% 17% 20% 18% Procedural justice (n=28) (n=10) (n=6) (n=5) (n=49) 20% 17% Human rights (n=5) (n=1) (n=6) 37% (n=1) (n=4) 33% 18% 50% 38% Quality of governance (n=43) (n=15) (n=11) (n=12) (n=80) 22% (n=1) (n=3) 14% 20% 50% 27% Responsibility (n=36) (n=14) (n=10) (n=12) (n=71) x d d o ox e e i ose eneric Total d ecif ss-b ck-b -co urp Und

p ers

d

Han

Gla

Bla

lP

era

Gen

1.0

Fraction overgeneralised

0.75

% of Papers

Fraction mischaracterised

100

1.00

0.8 0.6 0.4 0.2 0.0

G

Empirical AI System Classification

Figure 6: Overgeneralisation. Heatmap between overgeneralisation for system type (X-axis) and public value (Y-axis). Outer cells indicate marginals. In total, 41% of papers overgeneralise.

5.3. Overgeneralisation 41% of coded papers make at least one claim which is more general than their empirics justify. Fig. 6 maps instances of overgeneralisation across typology dimensions and public values. We find significant rates of overgeneralisation in every cell with enough data to make statistical claims. Generally, papers with underspecified systems (column 1), or that address “AI” generically (column 6), are more likely to make overgeneralised claims. The only exception is black-box systems (middle column), which also has a high prevalence. We discuss this further in §6.1.1. Claims about the quality of governance are most likely to be overgeneralised. This category covers practical claims about implementation, such as organisational factors in AI use, or the tasks for which AI systems are used. These vary more frequently across technically different systems than the more fundamental and conceptual claims in other public-value categories.

We therefore make two contributions with the aim of improving the technical precision of future work on AI in public administration. First, in §6.1, we highlight three common types of pitfall we find in our analysis – both to illustrate practically how these harm knowledge development, and to help researchers avoid them in the future. Second, in §6.2 we give practical recommendations for future research on AI in the public sector, which we believe greatly help technical precision – without requiring researchers to have either deep technical knowledge or closer access to studied systems. 6.1. Patterns of Imprecisions Across the analysed papers, we find three prominent patterns of imprecision. These include confusion introduced by the use of generic terms (§6.1.1), overreliance on research about blackbox systems (§6.1.2), and a failure to “future-proof” claims, evidenced by their inapplicability to agentic systems (§6.1.3).

6. Discussion

6.1.1. Generic Terms The single biggest driver of technical imprecision we find is the indiscriminate use of broad, generic, or ambiguous terms, such as “AI", “machine learning", “algorithmic decisionmaking" (ADM), or “chatbot". Three types of issues result. First, most broad terms can refer to systems across the typology, such that they invite overgeneralisation – in other words, authors use generic language but refer to specific systems. For

Our analysis indicates that public administration and digital government research about “AI” often overlooks technical distinctions that matter for governance. Sorting studied systems into a technical typology of just five categories suggests remarkable potential for more precision. As developed in our theory (§2), such imprecision should be avoided because it harms the field’s development of cumulative knowledge. 11

example, David et al. (2025) attribute to AI a set of “distinguishing features” – adaptive capacity, management of complex tasks, automation of decisions – without specifying which systems have them; the claim cannot be assessed because the referent is left open. Andrews (2019) similarly conflates “algorithms”, which conventionally span hand-coded and learned systems, with “machine learning”, collapsing a threshold across which transparency and accountability differ sharply (see 3). Second, many of these terms have imprecise or contested definitions in themselves. Most notably, as discussed above, “AI” is taken by some authors to include complex, but hand-coded rule-based systems, such as robotic process automation (RPA), while others take it as synonymous with “machine learning” – covering only the second tier in our typology onwards. Combined with underspecification, such ambiguity can even cast doubt on whether studied systems are “AI” at all, and therefore on the AI-specificity of derived claims. Surveys of publicsector documents highlight this same issue in registers of federal AI applications (Khan et al., 2024) and AI policy initiatives van Noordt et al. (2025).Where authors do not resolve these ambiguities, it is unclear what system types they draw from and map to. Finally, the conception of some terms has advanced as technology has progressed. Take the term “chatbot”: although the conversational user interface has remained similar, in the past decade chatbots have evolved from hand-coded “conversation tree” systems to generally capable, general-purposepowered agents (Adamopoulou and Moussiades, 2020; Ilves et al., 2025). Reducing these vastly different systems to their interface is imprecise. For example, Aoki (2020) studies chatbots they label “narrow AI” without establishing whether the chatbots follow hand-coded conversation trees or use black-box NLP intent recognition. Ju et al. (2023) note the higher fluency of GPT-like systems but design guidelines on assumptions that predate the externalisation these systems presuppose. It appears plausible that this imprecision is driven by the term “chatbot” being established even as the affordances of the underlying technology have changed drastically. Where these terms are defined and scoped clearly, their use can, of course, be appropriate: for example, discussion of the “intransparency of AI” may hold across all systems learned from data (Bullock et al., 2020; Lazar, 2024). It may even be required to use such terms, to reflect analysis of their use or perception: vignette experiments, for example, may reasonably describe a system as “AI-based software” to test what participants infer. But derived claims can still overgeneralise: Gesk and Leyer (2022) state that technical classifications “are therefore not elaborated here”, despite motivating the study with the opacity and undocumented rules of black-box systems – affordances its generic stimulus never instantiates.

other layers (Figs. 3, 6); where papers underspecify the studied system, we most frequently speculate that it is black-box. However, the affordance profile of black-box systems is narrow: they are usually trained separately by each organisation on their own data, purpose-bound, and they produce numeric or binary outputs, such as risk scores, likelihood estimates, or yes/no decisions. Imprecisions frequently result from overextending claims made about black-box systems. For example, Wang et al. (2025) draw general conclusions about “algorithmic” decisionmaking from a study whose effects on participation plausibly depend on the level of transparency, abandoning the ruledriven/data-driven distinction the same authors drew in Wang et al. (2023). Wirtz et al. (2019) present opacity, training-data bias, and autonomous learning as universal challenges of AI, even though their own application table includes rule-based systems to which these do not apply; their claims about implementation capacity and accountability hold cleanly only for the black-box layer. Chen et al. (2024) extend a functional typology (Makasi et al., 2022) developed before the proliferation of general-purpose models, but do not register the change in the skills required to audit and govern such systems (Mökander et al., 2024). These overextensions span most public value dimensions, but often share three patterns. First, claims on participation and procedural justice are often only valid for systems with the explainability affordance of black-box systems. These produce a singular, quantitative output, and “explainability” is taken to mean an understanding of model internals that produce it, e.g. generated via explainable AI (XAI) techniques (Mowbray et al., 2023). In contrast, LLMs may produce long text outputs – which can contain testable explanations in themselves, and therefore be institutionally valid without any understanding of model internals (Schmitz and Bryson, 2025). Second, claims on quality of governance are often overindexed on the technical or organisational specifics of blackbox models. For example, large volumes of high-quality data are often named as a requirement to “train AI”, but externally procured GPAI systems do not require any internal training data. Alon-Barkat et al. (2025) find that in-house development raises perceived responsibility relative to outsourcing, but treat internalisation as a free choice – whereas general-purpose systems carry inherent externalisation pressures relevant for implementation capacity, so the finding may not hold where the model is developed elsewhere (§3.4). Last, claims on responsibility from black-box models can underestimate the complexity of accountability allocation in modern AI supply chains (Brown, 2023). Black-box systems invite the assumption that data and model training are both internal to the organisation. Further, there is a difference in the type of AI system outputs citizens and officials interact with: an LLMgenerated text explanation, for example, may be more persuasive to a decision-maker than a single numeric score (Salvi et al., 2025), calling into question conclusions about, e.g., automation bias (Alon-Barkat and Busuioc, 2023). For example, Keppeler et al. (2025) study human–AI ensembles using a black-box tool but generalise their conclusions to “AI advice”

6.1.2. Overextrapolation from Black-Box Systems A second common pitfall is overextrapolation of conclusions that were drawn based on the study of black-box systems. Black-box systems are prominent: they are the most studied category and the empirical basis for many conclusions about 12

in general. The overreliance on black-box systems likely has historical drivers. Much of the fundamental literature on “AI in government” was published between 2019 and 2022 (Aarab et al., 2025), when such systems formed the frontier of AI capabilities (Brown et al., 2020). Indeed, a black-box quantitative risk scoring model likely caused the canonically referenced Dutch benefit scandal, which spurred an explosion of work in the field (Peeters and Widlak, 2023). The 2022 “general-purpose shift” driven by the introduction of ChatGPT then introduced a new class of system with drastically different affordances (§3) and regulatory and societal implications (Wang et al., 2026) – shortly after the canon developed.

these remain valuable contributions, and it is clear how their insights map to agents. 6.2. Recommendations for Public Administration Research Our work demonstrates that PA researchers should strive to improve the durability and generalisability of their findings by being more technically precise about AI. However, in so doing, they may encounter practical challenges: access to detailed information can be difficult, they may rely on surveys or interviews with non-experts, or they may lack the necessary technical background. We provide three sets of practical recommendations. Recommendations for Specifying AI Systems 1. Explicitly Specify AI System Types

6.1.3. Inapplicability to Agentic Systems (“Future-proofing”) A third form of imprecision we find is failure to address agentic systems, the newest layer of the typology. This takes two forms: some claims made in work published before agentic systems proliferated do not translate to them, and the field empirically so far does not study their deployment. The shift from general-purpose to agentic systems affects affordances across all public values (§3.5), but most consequentially responsibility, because of the implications for human oversight. Moving from reactive input-output mapping to proactive action over time (§3.5) moves oversight from the expost evaluation of discrete outputs to the runtime monitoring of ongoing processes (Schmitz et al., 2025; Chan et al., 2024). Accountability must be allocated for extended courses of action, rather than in discrete moments of decision, diffusing discretionary power away from identifiable actors (Chan et al., 2023; Schmitz and Bryson, 2025). Findings about the accountability of general-purpose chatbots – where a human can review each output – do not transfer to agentic systems that act across system boundaries without per-step review, since the oversight point has moved. We flag imprecisions in many papers because they make general claims about “AI” that are invalidated by this affordance boundary. For example, as Busuioc (2021) highlights, whether technical transparency solves accountability questions is a question of bureaucratic and process design. That interventions such as XAI improve perceived accountability for singlepoint decisions, therefore, does not express anything about their impact on the accountability of multi-turn agent actions. Further, across the reviewed papers, we find no study of agentic systems themselves (see Fig. 3). Given the recency of these systems, this is unsurprising – technical research on agents is accumulating, but little of it speaks to public administration (Rystrøm et al., 2026b). Beyond the specifics of agentic systems, this failure mode highlights how technical precision also contributes to making claims “future-proof”: as AI systems change and improve, claims about “AI” are more likely to age poorly than those with clear system types. For example, Wang et al. (2023) experimentally compare rule-driven (hand-coded) and data-driven (blackbox) decision-making, and Keppeler (2024) likewise grounds its study of disclosure effects in black-box systems. Both of

• Situate the system under study within a structured typology, such as the one presented here. Its design serves as a specification checklist: answering the four diagnostic questions in Fig. 2 places a system in exactly one class. • Add as much technical detail as necessary to clarify the affordances of the system, e.g. the specific name of studied LLMs – but no more. • Consider including a concrete diagram, system visualisation, or practical example of the system in use, helping readers quickly assess the system’s affordances and scope. 2. Use Proxy Indicators and Flag Uncertainty • Where technical detail is unavailable, approximate the affordances of the system with proxy indicators, such as the data used to train the AI model or its precise type of inputs and outputs. • Explicitly highlight any remaining uncertainty about technical specifics, rather than generalising to “AI”. 3. Scope Relevance of Past Work and Conclusions • Before drawing on past work, attempt to determine the AI system studied in it, and judge whether its affordances allow meaningful translation. • When drawing conclusions, be explicit about what types of AI systems you expect your claims to generalise to.

1. Explicitly specify AI system types. Scholars should specify the type of AI system they study, such as by placing it in the typology we propose. This does not require exhaustive technical detail, just enough specificity for readers to understand the affordance profile of the system. In Fig. 2, we provide four diagnostic questions. Answering these top-to-bottom maps an AI system to exactly one class. For typical PA cases, each of these is answerable from publicly available information about the system as deployed – without access to source code or model architecture. Our typology as presented is a minimum bound on technical specificity (§2), but for some topics, more technical detail may be warranted. Many systems also combine layers – a blackbox system embedded in a hand-coded decision system, say. 13

While we discuss how the typology could be expanded below (§6.3), individual authors may use a simple affordance-based litmus test to decide how much detail to include: would adding this detail distinguish between two systems with meaningfully different affordances? For example, different LLMs perform differently on publicsector tasks (Rystrøm et al., 2026b). Authors studying an LLMbased chatbot should therefore err towards naming the model used to clarify its affordances (specific to a version, e.g. “Gemini 3.1 Flash-Lite Preview”, which we use above), rather than referring to “an LLM”.

or system studied. Public administration is deeply familiar with phrasing such scope conditions: scholars are careful about whether findings depend on a particular institutional setting, administrative tradition, policy sector, or level of government. The same practice should be commonplace for technical reach. To “future-proof” claims, a practical solution may be to scope them to “currently available” AI systems. 6.3. Further Research Our typology serves two purposes: it exemplifies in general that technical precision about AI beyond the current standard is necessary, and it enables such precision for current systems. This focus suggests two promising strands for future research. First, it may be fruitful to detail out the typology we introduce – both “horizontally” by adding more nuanced publicvalue dimensions, and “vertically” by distinguishing more granularly between system types. For example, agentic systems have “degrees of agenticness” (Kasirzadeh and Gabriel, 2025) and vary in their autonomy, goal-directedness, and impact. These degrees may affect the public-sector affordances that different agentic systems have. Second, as AI systems evolve, research on their public-value implications should keep pace. Newer systems may have novel affordance profiles compared to current ones. For example, three potentially consequential developments in AI research are the increasing agenticness of AI systems discussed above, continual learning methods – which produce AI systems whose internal structure constantly updates, rather than being static after training (Yu et al., 2026), and embodiment, the integration of general-purpose systems with physical hardware (Firoozi et al., 2025). Each of these advances, and others that may emerge, could produce systems with novel affordance profiles, and PA research should analyse how these match or differ from past ones.

2. Use proxy indicators and flag uncertainty. Where practical challenges prevent the above specification, authors should a) use proxy indicators to approximate affordance profiles, and b) highlight any uncertainty that remains. Proxy indicators about AI systems may be available even if the above technical detail is not. These may include: • The type of data used to train the AI model, and who trained it. • The way the AI model is hosted and accessed by the organisation (e.g. on-premise vs. remotely). • The model’s or system’s input and output types – such as a single risk score or a free-text explanation. • Information about the system’s performance, such as its classification accuracy or benchmark results. • If the AI model or system is a third-party product, its name or vendor. As we theorise (§3) and demonstrate (§6.1), each of these indicators readily provides affordance-relevant information, and should therefore not be written off as irrelevant or overly technical. Finally, should uncertainty remain, describing that uncertainty is more informative than an undifferentiated generalisation to “AI”. Doing so conveys the maximal intended scope of claims, eases (or allows) retroactive specification, and “futureproofs” statements.

6.4. Limitations Beyond possible extensions in future research, we highlight three possible limitations of our work. Case Selection. Our analysis draws on a specific sample – the most highly cited papers shaping public administration and digital government scholarship on AI (2019–2025) – which may not represent the field as a whole. Our sample may exhibit different patterns than one composed of less-cited or more applied work. Specifically, citation-weighted sampling may over-represent conceptual and review work relative to applied case studies (Table A.4). It also structurally disadvantages recent work, which may partly explain the scarcity of papers on general-purpose and agentic systems.

3. Scope relevance of past work and conclusions. Technical precision should not only be applied to the AI system at hand: researchers should apply similar precision both when drawing on past work on AI in PA, and when concluding beyond the systems studied. To avoid mischaracterised motivation, researchers should attempt to typologise the AI systems which past work studies, and judge whether core claims translate. For example, a paper on algorithmic transparency studying black-box systems may provide valuable framing for a paper studying a general-purpose system, but the exact transparency techniques employed may not translate. As our methodology shows (§4), such analysis is possible retroactively in many cases. Similarly, scholars should specify for which types of AI systems they expect their conclusions to hold. If scoped well, conclusions can evidently be more general than the single case

Coding. Because our flagging is conservative – every positive is adjudicated by a second coder – the reported mischaracterisation and overgeneralisation prevalences are also conservative. Further, coding errors remain possible despite our measures to prevent them: we report adjudication rates and a power analysis, and reach no unresolved disagreement about codes in adjudication. 14

References

Detail and Currency. As discussed above (§6.3), there are still unexplored implications of our typology, and it will require updating as novel AI systems are introduced. We are explicit about these bounds and suggest both directions for future work.

Aarab, A., El Marzouki, A., Boubker, O., El Moutaqi, B., 2025. Integrating AI in public governance: A systematic review. Digital 5, 59. URL: https://www.mdpi.com/2673-647 0/5/4/59, doi:10.3390/digital5040059.

7. Conclusion

Abdi, H., Williams, L.J., 2010. Principal component analysis. WIREs Computational Statistics 2, 433–459. URL: https: //onlinelibrary.wiley.com/doi/abs/10.1002/wics .101, doi:10.1002/wics.101.

The expansion of AI in public administration has spurred a robust and valuable body of research. As our structural review demonstrates, this existing literature provides an essential foundation for understanding how algorithmic systems interact with core public values such as democratic participation, procedural justice, and governing capability. However, the conceptual tools used to classify these systems must keep pace with their technological architectures without getting swept away by a torrent of technical distinctions. But stronger technical specification of AI system types is a worthwhile investment. Retaining the umbrella term “AI” without technical clarification produces underspecification, internal inconsistency, and overgeneralisation that weaken otherwise sound findings. Avoiding these methodological pitfalls does not require public administration scholars to adopt highly granular engineering taxonomies. It only requires anchoring our definitions to affordance thresholds – the points at which a technical shift fundamentally alters what governance actors can or cannot do. By applying just a slight increase in specificity, researchers can significantly extend the transferability and applicability of their claims, ensuring that insights drawn from one context are reliably mapped to the right systems in the future.

Adamopoulou, E., Moussiades, L., 2020. Chatbots: History, technology, and applications. Machine Learning with Applications 2, 100006. URL: https://www.sciencedirec t.com/science/article/pii/S2666827020300062, doi:10.1016/j.mlwa.2020.100006. Alon-Barkat, S., Busuioc, M., 2023. Human–AI interactions in public sector decision making: “automation bias” and “selective adherence” to algorithmic advice. Journal of Public Administration Research and Theory 33, 153–169. URL: https://doi.org/10.1093/jopart/muac007, doi:10.1093/jopart/muac007. Alon-Barkat, S., Busuioc, M., Schwoerer, K., Weißmüller, K.S., 2025. Algorithmic discrimination in public service provision: Understanding citizens’ attribution of responsibility for human versus algorithmic discriminatory outcomes. Journal of Public Administration Research and Theory 35, 469– 488. URL: https://academic.oup.com/jpart/artic le/35/4/469/8249873, doi:10.1093/jopart/muaf024.

Declaration of generative AI and AI-assisted technologies in the manuscript preparation process

Altaweel, M., Bone, C., Abrams, J., 2019. Documents as data: A content analysis and topic modeling approach for analyzing responses to ecological disturbances. Ecological Informatics 51, 82–95. URL: https://www.sciencedirec t.com/science/article/pii/S1574954118303364, doi:10.1016/j.ecoinf.2019.02.014.

Large language models are a central part of the methodology as described in §4. Specifically, we use Gemini 3.1 Flash-Lite to extract structured information as part of our qualitative coding pipeline. All judgments and assessments were made solely by the authors, with no LLM-generated suggestions. All author judgements and LLM-extracted strands are available in the project repository. Furthermore, Claude Code was used to assist in creating the plots and figures. All code was reviewed and validated by the authors. ChatGPT and Claude were used for light copy-editing. The authors take full responsibility for all content and materials.

Ananny, M., Crawford, K., 2018. Seeing without knowing: Limitations of the transparency ideal and its application to algorithmic accountability. New Media & Society 20, 973– 989. URL: https://doi.org/10.1177/146144481667 6645, doi:10.1177/1461444816676645. Andrews, L., 2019. Public administration, public leadership and the construction of public value in the age of the algorithm and ‘big data’. Public Administration 97, 296–310. URL: https://onlinelibrary.wiley.com/doi/abs/ 10.1111/padm.12534, doi:10.1111/padm.12534. Androutsopoulou, A., Karacapilidis, N., Loukis, E., Charalabidis, Y., 2019. Transforming the communication between citizens and government through AI-guided chatbots. Government Information Quarterly 36, 358–367. URL: https: //linkinghub.elsevier.com/retrieve/pii/S074062 4X17304008, doi:10.1016/j.giq.2018.10.001. 15

Aoki, N., 2020. An experimental study of public trust in AI chatbots in the public sector. Government Information Quarterly 37, 101490. URL: https://linkinghub.e lsevier.com/retrieve/pii/S0740624X1930406X, doi:10.1016/j.giq.2020.101490.

Bright, J., Enock, F., Esnaashari, S., Francis, J., Hashem, Y., Morgan, D., 2025. Generative AI is already widespread in the public sector: Evidence from a survey of UK public sector professionals. Digital Government: Research and Practice 6, 1–13. URL: https://dl.acm.org/doi/10.1145 /3700140, doi:10.1145/3700140.

Bafna, P., Pramod, D., Vaidya, A., 2016. Document clustering: TF-IDF approach, in: 2016 International Conference on Electrical, Electronics, and Optimization Techniques (ICEEOT), pp. 61–66. doi:10.1109/ICEEOT.2016.7754 750.

Brown, I., 2023. Allocating Accountability in AI Supply Chains. Technical Report. Ada Lovelace Institute. URL: https://www.adalovelaceinstitute.org/resource/ ai-supply-chains/.

Barocas, S., Selbst, A.D., 2016. Big Data’s Disparate Impact. URL: https://papers.ssrn.com/abstract=2477899 ., doi:10.2139/ssrn.2477899, arXiv:2477899.

Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., . . . , Amodei, D., 2020. Language models are few-shot learners, in: Proceedings of the 34th International Conference on Neural Information Processing Systems, Curran Associates Inc., Red Hook, NY, USA. pp. 1877–1901. URL: https: //dl.acm.org/doi/10.5555/3495724.3495883.

Begg, C.B., Greenes, R.A., 1983. Assessment of diagnostic tests when disease verification is subject to selection bias. Biometrics 39, 207. URL: https://www.jstor.org/stab le/2530820?origin=crossref, doi:10.2307/2530820, arXiv:2530820.

Bullock, J., Young, M.M., Wang, Y.F., 2020. Artificial intelligence, bureaucratic form, and discretion in public service. Information Polity 25, 491–506. URL: https://www.medr a.org/servlet/aliasResolver?alias=iospress&doi =10.3233/IP-200223, doi:10.3233/IP-200223.

Bender, E.M., Gebru, T., McMillan-Major, A., Shmitchell, S., 2021. On the dangers of stochastic parrots: Can language models Be too big?, in: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pp. 610–623.

Burrell, J., 2016. How the machine ‘thinks’: Understanding opacity in machine learning algorithms. Big Data & Society 3, 2053951715622512. URL: https://journals.sagep ub.com/doi/10.1177/2053951715622512, doi:10.117 7/2053951715622512.

Berryhill, J., Heang, K.K., Clogher, R., McBride, K., 2019. Hello, World: Artificial Intelligence and Its Use in the Public Sector. OECD Working Papers on Public Governance 36. OECD Publishing. URL: https://ideas.repec.org/p/ oec/govaaa/36-en.html, doi:10.1787/726fd39d-en.

Busuioc, M., 2021. Accountable artificial intelligence: Holding algorithms to account. Public Administration Review 81, 825–836. URL: https://onlinelibrary.wiley.com/ doi/abs/10.1111/puar.13293, doi:10.1111/puar.132 93.

Bommasani, R., Hudson, D.A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M.S., Bohg, J., Bosselut, A., Brunskill, E., 2021. On the opportunities and risks of foundation models.

Buttaboni, C., Floridi, L., 2026. A regulatory taxonomy of AI opacity in the EU: Rethinking transparency, traceability, interpretability, and explainability. AI and Ethics 6, 100. URL: https://link.springer.com/10.1007/s43681-025 -00940-0, doi:10.1007/s43681-025-00940-0.

Bornmann, L., Daniel, H.D., 2008. What do citation counts measure? A review of studies on citing behavior. Journal of Documentation 64, 45–80. URL: https://doi.org/10.1 108/00220410810844150, doi:10.1108/002204108108 44150.

Chan, A., Ezell, C., Kaufmann, M., Wei, K., Hammond, L., Bradley, H., Bluemke, E., Rajkumar, N., Krueger, D., . . . , Anderljung, M., 2024. Visibility into AI agents, in: The 2024 ACM Conference on Fairness, Accountability, and Transparency, ACM, Rio de Janeiro Brazil. pp. 958–973. URL: https://dl.acm.org/doi/10.1145/3630106.36589 48, doi:10.1145/3630106.3658948.

Bovens, M., 2007. Analysing and Assessing Accountability: A Conceptual Framework. European Law Journal 13, 447–468. URL: https://onlinelibrary.wiley.com/doi/abs/ 10.1111/j.1468-0386.2007.00378.x, doi:10.1111/j. 1468-0386.2007.00378.x. Bovens, M., Zouridis, S., 2002. From Street-Level to SystemLevel Bureaucracies: How Information and Communication Technology is Transforming Administrative Discretion and Constitutional Control. Public Administration Review 62, 174–184. URL: https://onlinelibrary.wiley.com/ doi/abs/10.1111/0033-3352.00168, doi:10.1111/00 33-3352.00168.

Chan, A., Salganik, R., Markelius, A., Pang, C., Rajkumar, N., Krasheninnikov, D., Langosco, L., He, Z., Duan, Y., . . . , Maharaj, T., 2023. Harms from increasingly agentic algorithmic systems, in: Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, Association for Computing Machinery, New York, NY, USA. pp. 651–666. URL: https://dl.acm.org/doi/10.1145/3593013.3 594033, doi:10.1145/3593013.3594033.

Breiman, L., 2001. Random forests. Machine learning 45, 5– 32. doi:10.1023/A:1010933404324. 16

Chen, T., Gasco Hernández, M., Esteve Laporta, M., 2024. The adoption and implementation of artificial intelligence chatbots in public organizations: Evidence from U.S. state governments. American Review of Public Administration 54, 255–270. URL: https://www.scopus.com/pages/pub lications/85170831025, doi:10.1177/027507402312 00522.

de Ville, B., 2013. Decision trees. WIREs Computational Statistics 5, 448–455. URL: https://onlinelibr ary.wiley.com/doi/abs/10.1002/wics.1278, doi:10.1002/wics.1278. Dell’Acqua, F., McFowland, E., Mollick, E., Lifshitz, H., Kellogg, K.C., Rajendran, S., Krayer, L., Candelon, F., Lakhani, K.R., 2026. Navigating the jagged technological frontier: Field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality. Organization Science URL: https://pubsonline.inf orms.org/doi/full/10.1287/orsc.2025.21838, doi:10.1287/orsc.2025.21838.

Citron, D.K., 2008. Technological due process. Washington University Law Review 85, 1249–1313. Cobbe, J., Veale, M., Singh, J., 2023. Understanding accountability in algorithmic supply chains, in: 2023 ACM Conference on Fairness Accountability and Transparency, ACM, Chicago IL USA. pp. 1186–1197. URL: https: //dl.acm.org/doi/10.1145/3593013.3594073, doi:10.1145/3593013.3594073.

Devlin, J., Chang, M.W., Lee, K., Toutanova, K., 2019. BERT: Pre-training of deep bidirectional transformers for language understanding, in: Burstein, J., Doran, C., Solorio, T. (Eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Association for Computational Linguistics, Minneapolis, Minnesota. pp. 4171–4186. URL: https: //aclanthology.org/N19-1423/, doi:10.18653/v1/N1 9-1423.

Corbett-Davies, S., Gaebler, J.D., Nilforoshan, H., Shroff, R., Goel, S., 2023. The measure and mismeasure of fairness. Journal of Machine Learning Research 24. Dai, S.C., Xiong, A., Ku, L.W., 2023. LLM-in-the-loop: Leveraging large language model for thematic analysis, in: Bouamor, H., Pino, J., Bali, K. (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2023, Association for Computational Linguistics, Singapore. pp. 9993– 10001. URL: https://aclanthology.org/2023.find ings-emnlp.669/, doi:10.18653/v1/2023.findings-e mnlp.669.

D’ignazio, C., Klein, L.F., 2023. Data Feminism. MIT press. URL: https://books.google.com/books?hl=en&lr= &id=rHOdEAAAQBAJ&oi=fnd&pg=PR9&dq=data+feminis m+klein&ots=mXwrpceedi&sig=MI7_gJ8l1kdTlo-GEU Xh2kRrskk.

Dastin, J., 2018. Insight - amazon scraps secret AI recruiting tool that showed bias against women. Reuters URL: https: //www.reuters.com/article/world/insight-amazo n-scraps-secret-ai-recruiting-tool-that-showe d-bias-against-women-idUSKCN1MK0AG/.

Dressel, J., Farid, H., 2018. The accuracy, fairness, and limits of predicting recidivism. Science Advances 4, eaao5580. URL: https://www.science.org/doi/full/10.1126/sciad v.aao5580, doi:10.1126/sciadv.aao5580.

David, A., Yigitcanlar, T., Desouza, K., Mossberger, K., Cheong, P.H., Corchado, J., Beeramoole, P.B., Paz, A., 2025. Public perceptions of responsible AI in local government: A multi-country study using the theory of planned behaviour. Government Information Quarterly 42, 102054. URL: https://www.sciencedirect.com/science/ar ticle/pii/S0740624X25000486, doi:10.1016/j.giq. 2025.102054.

Dunleavy, P., 2006. Digital Era Governance: IT Corporations, the State, and e-Government. Oxford University Press, Oxford. doi:10.1093/acprof:oso/9780199296194.001. 0001. Elish, M.C., 2019. Moral Crumple Zones: Cautionary Tales in Human-Robot Interaction. Engaging Science, Technology, and Society 5, 40–60. URL: https://estsjournal.org/ index.php/ests/article/view/260, doi:10.17351/e sts2019.260.

Davison, A.C., Hinkley, D.V., 1997. Bootstrap Methods and Their Application. Cambridge Series in Statistical and Probabilistic Mathematics, Cambridge University Press, Cambridge. URL: https://www.cambridge.org/core /books/bootstrap- methods- and- their- appli cation/ED2FD043579F27952363566DC09CBD6A, doi:10.1017/CBO9780511802843.

Enqvist, L., 2024. Rule-based versus AI-driven benefits allocation: GDPR and AIA legal implications and challenges for automation in public social security administration. Information & Communications Technology Law 33, 222–246. URL: https://www.tandfonline.com/doi/full/10. 1080/13600834.2024.2349835, doi:10.1080/13600834 .2024.2349835.

de Bruijn, H., Warnier, M., Janssen, M., 2022. The perils and pitfalls of explainable AI: Strategies for explaining algorithmic decision-making. Government Information Quarterly 39, 101666. URL: https://www.sciencedirec t.com/science/article/pii/S0740624X21001027, doi:10.1016/j.giq.2021.101666.

Ferry, J., Aïvodji, U., Gambs, S., Huguet, M.J., Siala, M., 2025. Taming the triangle: On the interplays between fairness, interpretability, and privacy in machine learning. Computational Intelligence 41, e70113. URL: https://onlineli 17

brary.wiley.com/doi/abs/10.1111/coin.70113, doi:10.1111/coin.70113.

//doi.org/10.1093/ppmgov/gvac008, doi:10.1093/pp mgov/gvac008.

Firoozi, R., Tucker, J., Tian, S., Majumdar, A., Sun, J., Liu, W., Zhu, Y., Song, S., Kapoor, A., . . . , Schwager, M., 2025. Foundation models in robotics: Applications, challenges, and the future. The International Journal of Robotics Research 44, 701–739. URL: https://doi.org/10.1177/ 02783649241281508, doi:10.1177/0278364924128150 8.

Gstrein, O.J., Haleem, N., Zwitter, A., 2024. General-purpose AI regulation and the European union AI act. Internet Policy Review 13. URL: https://policyreview.info/articl es/analysis/general-purpose-ai-regulation-and -ai-act, doi:10.14763/2024.3.1790. Guest, G., Namey, E., Chen, M., 2020. A simple method to assess and report thematic saturation in qualitative research. PLOS ONE 15, e0232076. URL: https://journals.plo s.org/plosone/article?id=10.1371/journal.pone. 0232076, doi:10.1371/journal.pone.0232076.

Flyvbjerg, B., Turner, J.R., 2018. Do classics exist in megaproject management? International Journal of Project Management 36, 334–341. URL: http://arxiv.org/abs/17 10.09678, doi:10.1016/j.ijproman.2017.07.006, arXiv:1710.09678.

Hall, S.F., Sage, M., Scott, C.F., Joseph, K., 2024. A systematic review of sophisticated predictive and prescriptive analytics in child welfare: Accuracy, equity, and bias. Child and Adolescent Social Work Journal 41, 831–847. URL: https://doi.org/10.1007/s10560-023-00931-2, doi:10.1007/s10560-023-00931-2.

Fox, J., 2015. Applied Regression Analysis and Generalized Linear Models. Sage Publications. Francis, J.J., Johnston, M., Robertson, C., Glidewell, L., Entwistle, V., Eccles, M.P., Grimshaw, J.M., 2010. What is an adequate sample size? Operationalising data saturation for theory-based interview studies. Psychology & Health 25, 1229–1245. doi:10.1080/08870440903194015.

Hashem, Y., Bright, J., Chakraborty, S., 2025. Mapping the potential: Generative AI and public sector work URL: http s://apo.org.au/node/330966.

Gemini Team, Anil, R., Borgeaud, S., Alayrac, J.B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A.M., Hauth, A., . . . , Vinyals, O., 2025. Gemini: A family of highly capable multimodal models. URL: http://arxiv.org/ab s/2312.11805, doi:10.48550/arXiv.2312.11805, arXiv:2312.11805.

Heeks, R., Bailur, S., 2007. Analyzing e-government research: Perspectives, philosophies, theories, methods, and practice. Government Information Quarterly 24, 243–265. URL: ht tps://www.sciencedirect.com/science/article/pi i/S0740624X06000943, doi:10.1016/j.giq.2006.06. 005.

Gesk, T.S., Leyer, M., 2022. Artificial intelligence in public services: When and why citizens accept its usage. Government Information Quarterly 39, 101704. URL: https: //www.sciencedirect.com/science/article/pii/S0 740624X22000375, doi:10.1016/j.giq.2022.101704.

Ilves, L., Kilian, M., Parazzoli, S.M., Peixoto, T.C., Velsberg, O., 2025. The Agentic State: Rethinking Government for the Era of Agentic AI. Technical Report. Global Government Technology Centre Berlin and The World Bank. Janssen, M., Brous, P., Estevez, E., Barbosa, L.S., Janowski, T., 2020. Data governance: Organizing data for trustworthy artificial intelligence. Government Information Quarterly 37, 101493. URL: https://www.sciencedirect.com/scie nce/article/pii/S0740624X20302719, doi:10.1016/ j.giq.2020.101493.

Gibson, J.J., 1979. The Ecological Approach to Visual Perception. Houghton Mifflin Comp, Boston, Mass. Goodfellow, I., Bengio, Y., Courville, A., Bengio, Y., 2016. Deep Learning. volume 1. MIT press Cambridge. Green, B., 2022. The flaws of policies requiring human oversight of government algorithms. Computer Law & Security Review 45, 105681. URL: https://linkinghub.e lsevier.com/retrieve/pii/S0267364922000292, doi:10.1016/j.clsr.2022.105681.

Jo, J., Zhang, H., Cai, J., Goyal, N., 2025. AI trust reshaping administrative burdens: Understanding trust-burden dynamics in LLM-assisted benefits systems, in: Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, Association for Computing Machinery, New York, NY, USA. pp. 1172–1183. URL: https: //dl.acm.org/doi/10.1145/3715275.3732077, doi:10.1145/3715275.3732077.

Green, B., Chen, Y., 2021. Algorithmic Risk Assessments Can Alter Human Decision-Making Processes in High-Stakes Government Contexts. Proceedings of the ACM on HumanComputer Interaction 5, 1–33. URL: https://dl.acm.o rg/doi/10.1145/3479562, doi:10.1145/3479562.

Ju, J., Meng, Q., Sun, F., Liu, L., Singh, S., 2023. Citizen preferences and government chatbot social characteristics: Evidence from a discrete choice experiment. Government Information Quarterly 40, 101785. URL: https: //www.sciencedirect.com/science/article/pii/S0 740624X22001216, doi:10.1016/j.giq.2022.101785.

Grimmelikhuijsen, S., Meijer, A., 2022. Legitimacy of algorithmic decision-making: Six threats and the need for a calibrated institutional response. Perspectives on Public Management and Governance 5, 232–242. URL: https: 18

Kasirzadeh, A., Gabriel, I., 2025. Characterizing AI Agents for Alignment and Governance. URL: http://arxiv.org/ abs/2504.21848, doi:10.48550/arXiv.2504.21848, arXiv:2504.21848.

AAAI/ACM Conference on AI, Ethics, and Society, Association for Computing Machinery, New York, NY, USA. pp. 606–652. URL: https://dl.acm.org/doi/10.1145/3 600211.3604701, doi:10.1145/3600211.3604701.

Keppeler, F., 2024. No thanks, dear AI! Understanding the effects of disclosure and deployment of artificial intelligence in public sector recruitment. Journal of Public Administration Research and Theory 34, 39–52. URL: https://ac ademic.oup.com/jpart/article/34/1/39/7174960, doi:10.1093/jopart/muad009.

Lazar, S., 2024. Legitimacy, Authority, and Democratic Duties of Explanation, in: Sobel, D., Wall, S. (Eds.), Oxford Studies in Political Philosophy Volume 10. 1 ed.. Oxford University Press, Oxford, pp. 28–56. URL: https://academic.oup .com/book/56337/chapter/445461225, doi:10.1093/ oso/9780198909460.003.0002.

Keppeler, F., Borchert, J., Pedersen, M.J., Lehmann Nielsen, V., 2025. How ensembling AI and public managers improves decision-making. Journal of Public Administration Research and Theory 35, 261–276. URL: https://academic.oup .com/jpart/article/35/3/261/8116003, doi:10.109 3/jopart/muaf009.

Lecher, C., 2024. NYC’s AI Chatbot Tells Businesses to Break the Law. The Markup URL: https://themarkup.org/ar tificial-intelligence/2024/03/29/nycs-ai-cha tbot-tells-businesses-to-break-the-law. Leonardi, P.M., 2011. When Flexible Routines Meet Flexible Technologies: Affordance, Constraint, and the Imbrication of Human and Material Agencies1. MIS Quarterly 35, 147– 167. URL: https://doi.org/10.2307/23043493, doi:10.2307/23043493.

Khan, M.S., Shoaib, A., Arledge, E., 2024. How to promote AI in the US federal government: Insights from policy process frameworks. Government Information Quarterly 41, 101908. URL: https://www.sciencedirect.com/science/ar ticle/pii/S0740624X23001089, doi:10.1016/j.giq. 2023.101908.

Lima, M.S.M., Delen, D., 2020. Predicting and explaining corruption across countries: A machine learning approach. Government Information Quarterly 37, 101407. URL: https: //www.sciencedirect.com/science/article/pii/S0 740624X19302473, doi:10.1016/j.giq.2019.101407.

König, P.D., Wenzelburger, G., 2020. Opportunity for renewal or disruptive force? How artificial intelligence alters democratic politics. Government Information Quarterly 37, 101489. URL: https://www.sciencedirect. com/science/article/pii/S0740624X1930245X, doi:10.1016/j.giq.2020.101489.

Lipton, Z.C., 2018. The mythos of model interpretability. Communications of The Acm 61, 36–43. URL: https: //doi.org/10.1145/3233231, doi:10.1145/3233231. Madan, R., Ashok, M., 2023. AI adoption and diffusion in public administration: A systematic literature review and future research agenda. Government Information Quarterly 40, 101774. URL: https://www.sciencedirect. com/science/article/pii/S0740624X22001101, doi:10.1016/j.giq.2022.101774.

Krippendorff, K., 2019. Content Analysis: An Introduction to Its Methodology. SAGE Publications, Inc. URL: https: //methods.sagepub.com/book/mono/content-analy sis-4e/toc, doi:10.4135/9781071878781. Krizhevsky, A., Sutskever, I., Hinton, G.E., 2012. ImageNet classification with deep convolutional neural networks, in: Advances in Neural Information Processing Systems, Curran Associates, Inc. URL: https://papers.nips.cc/paper _files/paper/2012/hash/c399862d3b9d6b76c8436e9 24a68c45b-Abstract.html.

Majchrzak, A., Markus, M.L., 2013. Technology Affordances and Constraints Theory (of MIS) URL: https://doi.org/ 10.4135/9781452276090.n282, doi:10.4135/97814522 76090.n282. Makasi, T., Nili, A., Desouza, K.C., Tate, M., 2022. A typology of chatbots in public service delivery. IEEE Software 39, 58– 66. URL: https://ieeexplore.ieee.org/document/9 405373/, doi:10.1109/MS.2021.3073674.

Kroll, J.A., Huey, J., Barocas, S., Felten, E.W., Reidenberg, J.R., Robinson, D.G., Yu, H., 2017. Accountable algorithms. University of Pennsylvania Law Review 165, 633.

Marques, J.D., Duarte, A.V., de Carvalho, A.M.M., Rocha, G., Martins, B., Oliveira, A.L., 2025. Leveraging LLMs to streamline the review of public funding applications, in: Potdar, S., Rojas-Barahona, L., Montella, S. (Eds.), Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track, Association for Computational Linguistics, Suzhou (China). pp. 2041–2060. URL: https://aclanthology.org/2025.emnlp-indus try.143/, doi:10.18653/v1/2025.emnlp-industry.14 3.

Laux, J., Wachter, S., Mittelstadt, B., 2024. Trustworthy artificial intelligence and the european union AI act: On the conflation of trustworthiness and acceptability of risk. Regulation & Governance 18, 3–32. URL: https://online library.wiley.com/doi/abs/10.1111/rego.12512, doi:10.1111/rego.12512. Lawrence, C., Cui, I., Ho, D., 2023. The bureaucratic challenge to AI governance: An empirical assessment of implementation at U.S. federal agencies, in: Proceedings of the 2023 19

Matthias, A., 2004. The responsibility gap: Ascribing responsibility for the actions of learning automata. Ethics and Information Technology 6, 175–183. URL: https: //doi.org/10.1007/s10676-004-3422-1, doi:10.1 007/s10676-004-3422-1.

Mowbray, A., Chung, P., Greenleaf, G., 2023. Explainable AI (XAI) in Rules as Code (RaC): The DataLex approach. Computer Law & Security Review 48, 105771. URL: https: //linkinghub.elsevier.com/retrieve/pii/S026736 4922001145, doi:10.1016/j.clsr.2022.105771.

Mayne, H., Kang, J.S., Gould, D., Ramchandran, K., Mahdi, A., Siegel, N.Y., 2026. A positive case for faithfulness: LLM self-explanations help predict model behavior. URL: http: //arxiv.org/abs/2602.02639, doi:10.48550/arXiv.2 602.02639, arXiv:2602.02639.

Neumann, O., Guirguis, K., Steiner, R., 2024. Exploring artificial intelligence adoption in public organizations: A comparative case study. Public Management Review 26, 114–141. URL: https://www.tandfonline.com/doi/full/10. 1080/14719037.2022.2048685, doi:10.1080/14719037 .2022.2048685.

Mayring, P., 2015. Qualitative content analysis: Theoretical background and procedures, in: Bikner-Ahsbahs, A., Knipping, C., Presmeg, N. (Eds.), Approaches to Qualitative Research in Mathematics Education: Examples of Methodology and Methods. Springer Netherlands, Dordrecht, pp. 365– 380. URL: https://doi.org/10.1007/978-94-017-9 181-6_13, doi:10.1007/978-94-017-9181-6_13.

Nguyen-Trung, K., 2025. ChatGPT in thematic analysis: Can AI become a research assistant in qualitative research? Quality & Quantity 59, 4945–4978. URL: https://doi.org/ 10.1007/s11135-025-02165-z, doi:10.1007/s11135 -025-02165-z. Nissenbaum, H., 1996. Accountability in a computerized society. Science and Engineering Ethics 2, 25–42. URL: https://doi.org/10.1007/BF02639315, doi:10.100 7/BF02639315.

McInnes, L., Healy, J., Melville, J., 2020. UMAP: Uniform manifold approximation and projection for dimension reduction. URL: http://arxiv.org/abs/1802.03426, arXiv:1802.03426.

van Noordt, C., Medaglia, R., Tangi, L., 2025. Policy initiatives for artificial intelligence-enabled government: An analysis of national strategies in Europe. Public Policy and Administration 40, 215–253. URL: https://research.cbs.d k/en/publications/policy-initiatives-for-a rtificial-intelligence-enabled-government/, doi:10.1177/09520767231198411.

Meijer, A., Lorenz, L., Wessels, M., 2021. Algorithmization of bureaucratic organizations: Using a practice lens to study how context shapes predictive policing systems. Public Administration Review 81, 837–846. URL: https://online library.wiley.com/doi/abs/10.1111/puar.13391, doi:10.1111/puar.13391. Miles, M.B., Huberman, A.M., Saldana, J., 2014. Qualitative Data Analysis: A Methods Sourcebook. SAGE Publications, Inc, Los Angeles London New Delhi Singapore Washington DC.

O’Connor, C., Joffe, H., 2020. Intercoder reliability in qualitative research: Debates and practical guidelines. International Journal of Qualitative Methods 19, 1609406919899220. URL: https://journals.sagepub.com/doi/10.1177 /1609406919899220, doi:10.1177/1609406919899220.

Mitchell, T.M., 2013. Machine Learning. McGraw-Hill Series in Computer Science. nachdr. ed., McGraw-Hill, New York.

OECD, 2022. OECD Framework for the Classification of AI Systems. OECD Digital Economy Papers 323. OECD. URL: https://www.oecd.org/en/publications/oecd-fra mework-for-the-classification-of-ai-systems_c b6d9eca-en.html, doi:10.1787/cb6d9eca-en.

Mittelstadt, B., Russell, C., Wachter, S., 2019. Explaining explanations in AI, in: Proceedings of the Conference on Fairness, Accountability, and Transparency, Association for Computing Machinery, New York, NY, USA. pp. 279–288. URL: https://doi.org/10.1145/3287560.3287574, doi:10.1145/3287560.3287574.

OpenAI, 2022. ChatGPT: Optimizing language models for dialogue. URL: https://openai.com/blog/chatgpt/.

Mökander, J., Schroeder, R., 2024. Artificial intelligence, rationalization, and the limits of control in the public sector: The case of tax policy optimization. Social Science Computer Review 42, 1359–1378. URL: https://doi.org/10.117 7/08944393241235175, doi:10.1177/08944393241235 175.

Page, M.J., McKenzie, J.E., Bossuyt, P.M., Boutron, I., Hoffmann, T.C., Mulrow, C.D., Shamseer, L., Tetzlaff, J.M., Akl, E.A., . . . , Moher, D., 2021. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ , n71URL: https://www.bmj.com/lookup/doi/10.11 36/bmj.n71, doi:10.1136/bmj.n71.

Mökander, J., Schuett, J., Kirk, H.R., Floridi, L., 2024. Auditing large language models: A three-layered approach. AI and Ethics 4, 1085–1115. URL: https://doi.org/10.1 007/s43681-023-00289-2, doi:10.1007/s43681-023 -00289-2.

Peeters, R., Widlak, A.C., 2023. Administrative exclusion in the infrastructure-level bureaucracy: The case of the dutch daycare benefit scandal. Public Administration Review 83, 863–877. URL: https://onlinelibrary.wiley.com/ doi/10.1111/puar.13615, doi:10.1111/puar.13615. 20

Porter, Z., Calinescu, R., Lim, E., Hodge, V., Ryan, P., Burton, S., Habli, I., Lawton, T., McDermid, J., . . . , Zou, J., 2025. INSYTE: A Classification Framework for Traditional to Agentic AI Systems. ACM Transactions on Autonomous and Adaptive Systems 20, 1–39. URL: https://dl.acm.o rg/doi/10.1145/3760424, doi:10.1145/3760424.

with Deep Learning, PMLR. URL: https://openreview .net/forum?id=DuRUqZgwk8. Rystrøm, J., Schmitz, C., Korgul, K., Batzner, J., Russell, C., 2026b. Agent benchmarks fail public sector requirements, in: IASEAI 2026, arXiv. URL: http://arxiv.org/ abs/2601.20617, doi:10.48550/arXiv.2601.20617, arXiv:2601.20617.

Priem, J., Piwowar, H., Orr, R., 2022. OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts. URL: http://arxiv.org/abs/2205.01833, doi:10.48550/arXiv.2205.01833, arXiv:2205.01833.

Saldaña, J., 2025. The Coding Manual for Qualitative Researchers. 5e ed., Sage, London Thousand Oaks, California.

Qiu, T., He, Z., Chugh, T., Kleiman-Weiner, M., 2025. The lock-in hypothesis: Stagnation by algorithm, in: FortySecond International Conference on Machine Learning. URL: https://openreview.net/forum?id=mE1M62 6qOo.

Salvi, F., Horta Ribeiro, M., Gallotti, R., West, R., 2025. On the conversational persuasiveness of GPT-4. Nature Human Behaviour 9, 1645–1653. URL: https://www.nature.c om/articles/s41562-025-02194-6, doi:10.1038/s4 1562-025-02194-6.

Raghavan, M., Barocas, S., Kleinberg, J., Levy, K., 2020. Mitigating bias in algorithmic hiring: Evaluating claims and practices, in: Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, Association for Computing Machinery, New York, NY, USA. pp. 469–481. URL: ht tps://dl.acm.org/doi/10.1145/3351095.3372828, doi:10.1145/3351095.3372828.

Sandvig, C., Hamilton, K., Karahalios, K., Langbort, C., 2014. Auditing algorithms: Research methods for detecting discrimination on internet platforms. Data and Discrimination: Converting Critical Concerns into Productive Inquiry 22, 4349–4357. Sansone, D., Zhu, A., 2023. Using machine learning to create an early warning system for welfare recipients*. Oxford Bulletin of Economics and Statistics 85, 959–992. URL: https://onlinelibrary.wiley.com/doi/10.111 1/obes.12550, doi:10.1111/obes.12550.

Rao, Q., Frtunikj, J., 2018. Deep learning for self-driving cars: Chances and challenges, in: Proceedings of the 1st International Workshop on Software Engineering for AI in Autonomous Systems, Association for Computing Machinery, New York, NY, USA. pp. 35–38. URL: https: //dl.acm.org/doi/10.1145/3194085.3194087, doi:10.1145/3194085.3194087.

Schmitz, C., Bryson, J., 2025. A moral agency framework for legitimate integration of AI in bureaucracies (extended abstract). Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society 8, 2292–2293. URL: https://ojs. aaai.org/index.php/AIES/article/view/36714, doi:10.1609/aies.v8i3.36714.

Rezende, I.N., 2020. Facial recognition in police hands: Assessing the ‘clearview case’ from a European perspective. New Journal of European Criminal Law 11, 375–389. URL: https://doi.org/10.1177/2032284420948161, doi:10.1177/2032284420948161. Robinson, N., 2026. Open to open-source AI? Navigating AI model choice in public sector agencies. Government Information Quarterly 43, 102133. URL: https://www.scienc edirect.com/science/article/pii/S0740624X26000 304, doi:10.1016/j.giq.2026.102133.

Schmitz, C., Rystrøm, J., Batzner, J., 2025. Oversight structures for agentic AI in public-sector organizations, in: Kamalloo, E., Gontier, N., Lu, X.H., Dziri, N., Murty, S., Lacoste, A. (Eds.), Proceedings of the 1st Workshop for Research on Agent Language Models (REALM 2025), Association for Computational Linguistics, Vienna, Austria. pp. 298–308. URL: https://aclanthology.org/2025.re alm-1.21/.

Roehl, U.B.U., Hansen, M.B., 2024. Automated, administrative decision-making and good governance: Synergies, tradeoffs, and limits. Public Administration Review 84, 1184– 1199. URL: https://onlinelibrary.wiley.com/doi/ abs/10.1111/puar.13799, doi:10.1111/puar.13799.

Schönbrodt, F.D., Perugini, M., 2013. At what sample size do correlations stabilize? Journal of Research in Personality 47, 609–612. URL: https://www.sciencedirect.com/sc ience/article/pii/S0092656613000858, doi:10.101 6/j.jrp.2013.05.009.

Rudin, C., 2019. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence 1, 206–215. URL: https://www.nature.com/articles/s42256-019-0 048-x, doi:10.1038/s42256-019-0048-x.

Schreier, M., 2012. Qualitative Content Analysis in Practice. SAGE Publications Ltd. URL: https://methods.sage pub.com/book/mono/qualitative-content-analysi s-in-practice/toc, doi:10.4135/9781529682571. Schroeder, R., 2020. Big data and cumulation in the social sciences. Information, Communication & Society 23, 1593–1607. URL: https://www.tandfonline.co

Rystrøm, J., Fu, Z., Russell, C., 2026a. OxEnsemble: Fair ensembles for low-data classification, in: Medical Imaging 21

van Thiel, S., 2021. Research Methods in Public Administration and Public Management: An Introduction. 2 ed., Routledge, London. doi:10.4324/9781003196907.

m/doi/full/10.1080/1369118X.2019.1594334, doi:10.1080/1369118X.2019.1594334. Scott, J.C., 1998. Seeing Like a State: How Certain Schemes to Improve the Human Condition Have Failed. Yale University Press. URL: https://www.jstor.org/stable/j.ctt1n q3vk, arXiv:j.ctt1nq3vk.

Vaithianathan, R., Putnam-Hornstein, E., Jiang, N., Nand, P., Maloney, T., 2017. Developing Predictive Models to Support Child Maltreatment Hotline Screening Decisions: Allegheny County Methodology and Implementation. Technical Report. Centre for Social Data Analytics.

Selbst, A.D., Boyd, D., Friedler, S.A., Venkatasubramanian, S., Vertesi, J., 2019. Fairness and abstraction in sociotechnical systems, in: Proceedings of the Conference on Fairness, Accountability, and Transparency, ACM, Atlanta GA USA. pp. 59–68. URL: https://dl.acm.org/doi/10.1145/328 7560.3287598, doi:10.1145/3287560.3287598.

Valle-Cruz, D., Criado, J.I., Sandoval-Almazán, R., RuvalcabaGomez, E.A., 2020. Assessing the public policy-cycle framework in the age of artificial intelligence: From agenda-setting to policy evaluation. Government Information Quarterly 37, 101509. URL: https://www.sciencedirect. com/science/article/pii/S0740624X20302884, doi:10.1016/j.giq.2020.101509.

Selten, F., Robeer, M., Grimmelikhuijsen, S., 2023. ‘just like I thought’: Street-level bureaucrats trust AI recommendations if they confirm their professional judgment. Public Administration Review 83, 263–278. URL: https://onli nelibrary.wiley.com/doi/10.1111/puar.13602, doi:10.1111/puar.13602.

Valle-Cruz, D., García-Contreras, R., Gil-Garcia, J.R., 2024. Exploring the negative impacts of artificial intelligence in government: The dark side of intelligent algorithms and cognitive machines. International Review of Administrative Sciences 90, 353–368. URL: https://doi.org/10.1177/00 208523231187051, doi:10.1177/00208523231187051.

Simon, H.A., 1947. Administrative Behavior. Macmillan Company.

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \., Polosukhin, I., 2017. Attention is all you need, in: Advances in Neural Information Processing Systems, pp. 5998–6008.

Stalenhoef, F., Oostvogel, J., Ruijer, E., Meijer, A., 2024. Een dialoog voor de borging van goed digitaal bestuur: Ontwikkeling van het instrument ‘van principes naar acties’ met scenario-based design thinking. Bestuurswetenschappen 78, 21–39. URL: https://www.boomportaal. nl/doi/10.5553/Bw/016571942024078002004, doi:10.5553/Bw/016571942024078002004.

Wachter, S., Mittelstadt, B., Floridi, L., 2017a. Why a right to explanation of automated decision-making does not exist in the general data protection regulation. International Data Privacy Law 7, 76–99. URL: https://doi.org/10.109 3/idpl/ipx005, doi:10.1093/idpl/ipx005.

Sterz, S., Baum, K., Biewer, S., Hermanns, H., LauberRönsberg, A., Meinel, P., Langer, M., 2024. On the Quest for Effectiveness in Human Oversight: Interdisciplinary Perspectives, in: The 2024 ACM Conference on Fairness Accountability and Transparency, ACM, Rio de Janeiro Brazil. pp. 2495–2507. URL: https://dl.acm.org/doi/10.11 45/3630106.3659051, doi:10.1145/3630106.3659051.

Wachter, S., Mittelstadt, B., Russell, C., 2017b. Counterfactual explanations without opening the black box: Automated decisions and the GDPR. Harv. JL & Tech. 31, 841. URL: https://heinonline.org/hol-cgi-bin/get_pdf.c gi?handle=hein.journals/hjlt31&section=29.

Straub, V.J., Hashem, Y., Bright, J., Bhagwanani, S., Morgan, D., Francis, J., Esnaashari, S., Margetts, H., 2024. AI for bureaucratic productivity: Measuring the potential of AI to help automate 143 million UK government transactions. URL: http://arxiv.org/abs/2403.14712, doi:10.48550/arXiv.2403.14712, arXiv:2403.14712.

Wachter, S., Mittelstadt, B., Russell, C., 2021a. Bias preservation in machine learning: The legality of fairness metrics under EU non-discrimination law. West Virginia Law Review URL: https://researchrepository.wvu.edu/w vlr/vol123/iss3/4/, doi:10.2139/ssrn.3792772. Wachter, S., Mittelstadt, B., Russell, C., 2021b. Why fairness cannot be automated: Bridging the gap between EU nondiscrimination law and AI. Computer Law & Security Review 41, 105567. URL: https://www.sciencedirec t.com/science/article/pii/S0267364921000406, doi:10.1016/j.clsr.2021.105567.

Straub, V.J., Morgan, D., Bright, J., Margetts, H., 2023. Artificial intelligence in government: Concepts, standards, and a unified framework. Government Information Quarterly 40, 101881. URL: https://www.sciencedirect. com/science/article/pii/S0740624X23000813, doi:10.1016/j.giq.2023.101881.

Wang, G., Guo, Y., Zhang, W., Xie, S., Chen, Q., 2023. What type of algorithm is perceived as fairer and more acceptable? A comparative analysis of rule-driven versus datadriven algorithmic decision-making in public affairs. Government Information Quarterly 40, 101803. URL: https:

Sundermeyer, M., Schlüter, R., Ney, H., 2012. LSTM neural networks for language modeling, in: Thirteenth Annual Conference of the International Speech Communication Association. doi:10.21437/Interspeech.2012-65. 22

S. (Eds.), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 5: Tutorial Abstracts), Association for Computational Linguistics, Mexico City, Mexico. pp. 19–25. URL: https: //aclanthology.org/2024.naacl-tutorials.3/, doi:10.18653/v1/2024.naacl-tutorials.3.

//www.sciencedirect.com/science/article/pii/S0 740624X23000035, doi:10.1016/j.giq.2023.101803. Wang, G., Zhang, Z., Xie, S., Guo, Y., 2025. Province of origin, decision-making bias, and responses to bureaucratic versus algorithmic decision-making. Public Administration Review 85, 1738–1756. URL: https://onlinelibrary.wiley. com/doi/10.1111/puar.13928, doi:10.1111/puar.139 28.

Zouridis, S., van Eck, M., Bovens, M., 2020. Automated discretion, in: Evans, T., Hupe, P. (Eds.), Discretion and the Quest for Controlled Freedom. Springer International Publishing, Cham, pp. 313–329. URL: https://doi.org/10 .1007/978-3-030-19566-3_20, doi:10.1007/978-3-0 30-19566-3_20.

Wang, J., Selbst, A.D., Barocas, S., Venkatasubramanian, S., 2026. Distinguishing task-specific and general-purpose AI in regulation, in: Proceedings of the Symposium on Computer Science and Law, Association for Computing Machinery, New York, NY, USA. pp. 185–197. URL: https: //dl.acm.org/doi/10.1145/3788646.3789523, doi:10.1145/3788646.3789523.

Zuboff, S., 2015. Big other: Surveillance capitalism and the prospects of an information civilization. Journal of Information Technology 30, 75–89. URL: http://journals.sag epub.com/doi/10.1057/jit.2015.5, doi:10.1057/ji t.2015.5.

Wilson, E.B., 1927. Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association 22, 209–212. URL: http://www.tandfonlin e.com/doi/abs/10.1080/01621459.1927.10502953, doi:10.1080/01621459.1927.10502953. Wirtz, B.W., Weyerer, J.C., Geyer, C., 2019. Artificial intelligence and the public sector—applications and challenges. International Journal of Public Administration 42, 596–615. URL: https://www.tandfonline.com/doi/full/10. 1080/01900692.2018.1498103, doi:10.1080/01900692 .2018.1498103.

Appendix A. Data, coding, and corpus summary This appendix details corpus construction and coding. Figure A.7 summarises the screening process and Table A.5 summarises the coding dimensions. The corpus was assembled from all venues listed in the study configuration (Table A.3), using OpenAlex as the retrieval source, a title- and abstract-based keyword filter, and a citation-based annual sampling rule. The OpenAlex query string, venue list, and corpus metadata are provided as machine-readable files at https://anonymous.4ope n.science/r/AITypology4PA-4FC6/.

Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K.R., Cao, Y., 2023. ReAct: Synergizing reasoning and acting in language models, in: The Eleventh International Conference on Learning Representations. URL: https://openreview .net/forum?id=WE_vluYUL-X. Yu, D., Zhang, X., Chen, Y., Liu, A., Zhang, Y., Yu, P.S., King, I., 2026. Recent advances of multimodal continual learning: A comprehensive survey. IEEE Transactions on Neural Networks and Learning Systems , 1–21URL: https://ieee xplore.ieee.org/abstract/document/11456498, doi:10.1109/TNNLS.2026.3658485.

Appendix A.1. Corpus screening The corpus was assembled in four stages. First, we retrieved all articles published between 2018 and 2025 from the configured public administration and digital government venues listed in venues.yaml. Second, retrieved records were deduplicated by DOI where available. Third, we applied a keyword filter to titles and abstracts reconstructed from the OpenAlex abstract inverted index. Fourth, from the matched set we selected the eight most-cited papers per year separately for public administration venues and digital government venues, to balance the corpus across the two journal streams. The resulting set was then manually screened to exclude papers that used AI only as a methodological tool or did not treat AI as an empirical or theoretical subject. Figure A.7 presents this process in PRISMA-like form (Page et al., 2021). A total of 7,922 records were retrieved from OpenAlex and deduplicated by DOI. Of these, 325 matched the keyword filter. Applying the citation-based annual sampling rule yielded 122 papers, from which 21 were excluded during manual screening, leaving a final analysed sample of 101 papers.

Yun, L., Yun, S., Xue, H., 2024. Improving citizen-government interactions with generative artificial intelligence: Novel human-computer interaction strategies for policy understanding through large language models. PLOS One 19, e0311410. URL: https://journals.plos.org/plos one/article?id=10.1371/journal.pone.0311410, doi:10.1371/journal.pone.0311410. Zammuto, R.F., Griffith, T.L., Majchrzak, A., Dougherty, D.J., Faraj, S., 2007. Information Technology and the Changing Fabric of Organization. Organization Science 18, 749–762. URL: https://doi.org/10.1287/orsc.1070.0307, doi:10.1287/orsc.1070.0307. Zhu, Z., Chen, H., Ye, X., Lyu, Q., Tan, C., Marasovic, A., Wiegreffe, S., 2024. Explanation in the era of large language models, in: Zhang, R., Schneider, N., Chaturvedi, 23

Table A.3: Journals included in the literature review

Stream

Journal

Public Administration

Public Administration Review Journal of Public Administration Research and Theory Public Administration Governance Public Management Review Perspectives on Public Management and Governance American Review of Public Administration Administration & Society International Journal of Public Administration Public Performance & Management Review Journal of Policy Analysis and Management International Review of Administrative Sciences

Digital Government

Government Information Quarterly Information Polity International Journal of Electronic Government Research Digital Government: Research and Practice

pling rule. This is not a test of thematic saturation in the sense of determining whether additional papers would yield new concepts or codes. Instead, the analysis evaluates whether the paper’s main aggregate findings are stable to the inclusion of additional papers within the same sampling frame. The procedure is therefore closer to stability-based sample-size assessment, where estimates are judged adequate once they remain within a specified tolerance corridor (Schönbrodt and Perugini, 2013), while also drawing on methodological work that operationalises saturation through explicit stopping rules and incremental sampling criteria (Francis et al., 2010; Guest et al., 2020). The bootstrap stability analysis is implemented in saturation_analysis.py, available at https://anonymou s.4open.science/r/AITypology4PA-4FC6/.

Appendix A.2. Coding procedure and variables We conducted structured manual coding of all papers in the final sample. The coding scheme was jointly piloted by all three authors on 10 papers and refined iteratively before full coding began. Candidate quotations were surfaced using an LLM (see repository), after which one author read each paper in full and coded all relevant references using the final scheme. Multiple rows were created when a quotation referenced multiple systems, and multiple PA-relevance labels were allowed where applicable. Each extracted reference was coded along three dimensions: AI system classification, role in paper, and PA relevance. AI system classification used the typology described in the main text; role in paper distinguished Motivation, Empirical, and Conclusion; and PA relevance used the second-level dimensions of the good digital governance framework. Coders also recorded a brief justification for each code. Table A.5 summarises these dimensions. The three paper-level outcomes reported in the main text were derived from these strand-level codings. A paper was classified as underspecified if all of its empirical references were coded as Underspecified. A paper was classified as mischaracterised if, within a given PA-relevance dimension, at least one motivation strand referred to a different system type from the one studied empirically, and as overgeneralised if, within a dimension, at least one conclusion strand did so. Papers studying multiple empirical system types could contribute to multiple empirical categories. The paper-level results table and the analysis scripts implementing these rules are available in the repository linked above.

For each K ∈ {1, . . . , 8}, we form a corpus by taking the top-K papers from each stream-year combination, compute the three aggregate outcomes introduced in §4.3—underspecification, mischaracterisation, and overgeneralisation—and apply two pre-specified criteria. First, local stability requires the point estimates at K = 6, 7, 8 to lie within 3 percentage points of each other. Second, flat trajectory slope requires a linear fit over K ∈ {5, . . . , 8} to have a slope of at most 0.5 percentage points. These criteria operationalise the requirement that adding further papers within the sampling rule should not materially change the aggregate estimates. However, the exact values are somewhat arbitrary; the substantive evidence is the visual convergence as shown in Fig. B.8. Uncertainty at each K is calculated by block-bootstrapping stream-year combinations with replacement over 1,000 iterations to produce 95% bands, following the general use of bootstrap resampling to quantify sampling variability around estimated quantities (Davison and Hinkley, 1997). Both criteria are met for all three outcomes at K = 8—see Fig. B.8. This provides an empirical bound on how much adding further papers within the same sampling rule would shift our aggregate find-

Appendix B. Stability analysis of aggregate estimates We assess sampling adequacy through a pre-specified stability analysis of the aggregate outcomes produced by our sam24

Table A.4: Descriptive summary of the analysed corpus

Panel A. Final sample composition

N

Final analysed papers Public administration venues Digital government venues

91 42 49

Published in 2018 Published in 2019 Published in 2020 Published in 2021 Published in 2022 Published in 2023 Published in 2024 Published in 2025

0 9 10 10 14 16 17 15

Panel B. Empirical paper-level system classifications

N

Hand-coded Glass-box Black-box General-purpose Agentic Underspecified

2 4 19 11 0 50

Panel C. Paper-level outcome flags

N

Underspecified Mischaracterised Overgeneralised

50 28 37

Panel D. Empirics type

N

Case study Survey Vignette experiment Experiment Systematic literature review Conceptual framework Other

22 10 21 4 23 10 1

ings, and supports treating K = 8 as adequate for the substantive claims we make. The stability result is conditional on the sampling scope—top-cited PA and digital government venues, 2019–2025—and does not extend to claims about scholarship outside this scope.

Produce a structured analysis covering three ,→ sections: **Empirics**, **Motivation**, and ,→ **Conclusions/Claims**. Do Empirics first, as ,→ your classification there anchors the judgements you make in the other two sections. ,→ ---

Appendix C. Codebook

### **TAXONOMY OF AI SYSTEM TYPES**

You are analyzing a public administration research ,→ paper for **technical precision in its treatment ,→ of AI**. The core question throughout is: does ,→ the paper's use of "AI" as an undifferentiated ,→ category cause analytical problems — in its ,→ motivation, its empirical claims, or its ,→ conclusions?

Use this typology consistently throughout. When ,→ classifying systems, always ask: what is the most ,→ precise classification the paper's evidence ,→ actually supports? Default to a more conservative ,→ classification when in doubt. ### 1. Hand-coded

--### **YOUR TASK**

25

Identification Records retrieved from OpenAlex across all configured venues and years, deduplicated by DOI (n = 7922)

Table A.5: Summary of coding dimensions

Records matching title/abstract keyword filter (n = 284)

Dimension

Values

Unit

AI system classification

Hand-coded; Glass-box; Black-box; General-purpose; Agentic; Justifiably Generic; Underspecified

Paper

Role in paper

Motivation; Empirical; Conclusion

Strand

PA relevance

Participation; Procedural justice; Human rights; Quality of governance; Responsibility; None

Strand

Records selected by citation-based annual sampling rule (n = 109) Top 8 per year from public administration venues Top 8 per year from digital government venues Underspecification

Rate (%)

100

Records excluded in manual screening (n = 16) AI used only as a methodological tool AI not treated as an empirical or theoretical subject

n=16

50

n=31

n=45 n=57

Mischaracterisation

n=67 n=77 n=85 n=89

n=16

n=16

0

1 2 3 4 5 6 7 8

Overgeneralisation

n=31

n=45 n=57 n=67 n=77 n=85 n=89

1 2 3 4 5 6 7 8

K (papers per cell)

n=31 n=45 n=57

n=67 n=77 n=85 n=89

1 2 3 4 5 6 7 8

Figure B.8: Stability analysis. Aggregate rates for our three main analytical constructs as we increase our sampling criteria. All constructs meet our stability criteria at K = 8.

Included Final analysed sample (n = 91)

Machine learning systems that learn from data but ,→ produce interpretable, human-readable models. ,→ Experts can inspect and understand the model's ,→ decision logic.

Figure A.7: Corpus construction and screening procedure. Articles were retrieved from all configured venues using OpenAlex, deduplicated, filtered using title- and abstract-based keyword matching, and then sampled using a citationbased rule selecting the eight most-cited papers per year separately for public administration and digital government venues. The resulting set was manually screened to exclude papers that used AI only as a methodological tool or did not treat AI as an empirical or theoretical subject.

- **Key characteristics:** Data-driven, interpretable ,→ output, simple learned rules - **Examples:** Logistic regression, linear ,→ regression, decision trees, PCA, TF-IDF scoring - **Tip:** If the paper describes a system as ,→ statistically trained but whose outputs can be ,→ understood or audited by experts, classify as ,→ glass-box.

Traditional software systems where all rules and ,→ logic are explicitly written by humans. The ,→ system's behaviour is fully determined by its ,→ code and does not change based on data.

--### 3\. Black box systems

- **Key characteristics:** Explicit rules, stable ,→ behaviour, fully interpretable, no learned ,→ components - **Examples:** Rule-based benefit eligibility ,→ systems, tax calculation software, structured ,→ decision trees implemented as code - **Tip:** If the paper describes a system that ,→ applies fixed rules to determine outcomes (e.g. ,→ "the system decides who gets benefits based on ,→ income thresholds"), classify as Hand-coded.

Machine learning systems that trade interpretability ,→ for performance. The internal logic of the model ,→ cannot be directly inspected, requiring post-hoc ,→ explanation methods. - **Key characteristics:** High performance, opaque ,→ internal logic, operates on structured or ,→ unstructured data, post-hoc explainability ,→ required - **Examples:** Random forests, neural networks, NLP ,→ classifiers, predictive analytics tools, judicial ,→ outcome prediction systems, algorithmic workforce ,→ management tools

--### 2\. Glass-box systems

26

- **Tip:** If the paper describes a system using ,→ terms like "machine learning", "NLP", "pattern ,→ detection", or "predictive analytics" without ,→ claiming interpretability, classify as black-box.

**Classification rules:** * **Be aggressive about "Underspecified"**. If the ,→ paper names a concrete system or application (a ,→ chatbot, a "decision support tool", a "risk ,→ scoring system") but does not provide enough ,→ technical detail to determine the system type, ,→ classify it as Underspecified even if inference ,→ seems plausible. "Sounds like supervised ,→ learning" is not sufficient — you need the paper ,→ to provide a basis. The bar for moving out of ,→ Underspecified is: the paper explicitly names a ,→ technique (e.g. "neural network", "logistic ,→ regression", "rule-based"), describes a decision ,→ logic that makes the type inferrable with high ,→ confidence (e.g. "programmed rules to exclude ,→ claimants" → Hand-coded), it names a specific ,→ tool which is clearly identifiable, or the system ,→ is so well-known externally that classification ,→ is unambiguous (e.g. facial recognition → ,→ Black-box). * **"Justifiably generic"** applies in two narrow cases only: (a) the paper is explicitly ,→ conceptual about a well-defined class of systems ,→ (e.g. "automated decision-making" as a class), or ,→ (b) the empirical design itself precludes ,→ system-level specificity (e.g. a survey vignette ,→ that presents "an AI system" to respondents — you ,→ cannot add technical detail to what respondents ,→ saw). Even in case (b), the researcher should ,→ still demonstrate clarity about what they mean. ,→ * **For motivation**, use the same typology. "Generic ,→ across a class of systems" is appropriate when ,→ the paper invokes "AI", "ADM", or "machine ,→ learning" without specifying a system — but note ,→ this as an imprecision where it matters.

--### 4\. General-purpose systems Large models pre-trained on massive datasets ,→ (typically internet-scale) and adapted to ,→ specific tasks via fine-tuning or prompting. ,→ These models are not built or managed by the ,→ deploying organisation. - **Key characteristics:** Pre-trained on unknown or ,→ large-scale data, adapted via prompting or ,→ fine-tuning, general-purpose, opaque training ,→ data, generates text/images/other outputs - **Examples:** GPT-3, GPT-4, ChatGPT, BERT-based ,→ systems, large vision-language models - **Tip:** If the paper describes a system as a ,→ "large language model", "generative AI", or ,→ "GPT", or notes that it generates human-like text ,→ from prompts, classify as General-purpose. --### 5\. Agentic System AI systems characterised by autonomy, the ability to interact with their environment, and the capacity ,→ to pursue complex goals across multiple steps, ,→ often using external tools or APIs. ,→ - **Key characteristics:** Autonomy, tool use, ,→ multi-step goal pursuit, environmental ,→ interaction, broad generality - **Examples:** LLM-based agents with web search or ,→ database access, automated workflow systems, ,→ multi-agent pipelines - **Tip:** If the paper describes a system that acts ,→ independently in an environment, uses tools, or ,→ completes multi-step tasks without human ,→ intervention at each step, classify as Agentic.

--### **PUBLIC VALUE DIMENSIONS** Map all motivations, empirics, and conclusions to one ,→ or more of the following dimensions: ### Participation *(Democracy)*

---

The reference engages with how the AI system affects ,→ citizens' or stakeholders' ability to engage ,→ with, influence, or be included in public ,→ decision-making processes.

### 6\. Underspecified Use this classification when the paper references an ,→ AI system without providing sufficient detail to ,→ determine its technical nature, or uses AI as a ,→ generic concept.

- **Sub-values:** Responsiveness, Inclusion, ,→ Transparency, Collaboration - **Indicators:** Citizens' ability to contest or ,→ appeal automated decisions; inclusive design of ,→ AI systems; transparency of decision processes to affected parties; collaborative governance of AI ,→ ,→ deployment

- **Examples:** "AI tools", "automated ,→ decision-making systems", "new technologies" ,→ without further specification - **Tip:** If in doubt between two categories, note ,→ both and explain your reasoning. Only use ,→ Underspecified when no reasonable classification ,→ can be inferred.

27

- **Sub-values:** Agility, Expertise, Carefulness, ,→ Security, Effectiveness, Efficiency, ,→ Independence, Risk awareness - **Indicators:** Whether AI enhances or substitutes ,→ administrative expertise and competence; agility ,→ or adaptability of AI-supported governance; ,→ operational security of AI systems; effectiveness ,→ and efficiency of AI-supported public services; ,→ awareness and management of AI-related risks; independence from vendor lock-in ,→ - **Example:** A paper examines how adopting a ,→ commercial AI platform created problematic ,→ dependency on a private vendor, undermining ,→ administrative independence.

- **Example:** A paper argues that fully automated ,→ benefit decisions eliminate the case-worker ,→ interaction through which applicants could raise ,→ concerns, reducing meaningful participation. --### Procedural justice *(Rule of law)* The reference engages with the legal and legitimate ,→ functioning of governance — whether the AI system ,→ operates through fair legal procedures that are ,→ suitable, explainable, non-discriminatory, and ,→ user-friendly. - **Sub-values:** Suitability, Explainability, ,→ Proportionality, User-friendliness, ,→ Disputability, Solution-oriented approach - **Indicators:** Whether AI use is legally and ,→ procedurally suitable for the decision at stake; ,→ whether outputs can be explained to affected ,→ individuals; non-discriminatory operation of the ,→ system; user-friendliness of AI-mediated ,→ services; availability of redress mechanisms - **Example:** A paper argues that a black-box ,→ risk-scoring tool used in parole decisions cannot ,→ provide legally adequate explanations, violating ,→ procedural justice requirements.

--### Responsibility *(Governing capability)* The reference engages with how the AI system is ,→ embedded in the broader system of checks and ,→ balances — whether accountability is clear, ,→ decisions are verifiable, and ultimate human ,→ responsibility is preserved. - **Sub-values:** Accountability, Verifiability, ,→ Human final responsibility, Integrity, Continuity - **Indicators:** Clear assignment of responsibility ,→ for AI decisions; auditability and verifiability ,→ of outputs; ensuring a human remains ultimately ,→ responsible for decisions taken with AI support; ,→ integrity of AI-supported processes; operational ,→ continuity - **Example:** A paper argues that when an AI system ,→ makes a harmful welfare decision, it is unclear ,→ whether responsibility lies with the procuring ,→ agency, the vendor, or the individual official, ,→ creating a responsibility gap.

--### Human rights *(Rule of law)* The reference engages with whether the AI system ,→ affects individual rights and freedoms, with ,→ particular emphasis on data protection and human ,→ autonomy. - **Sub-values:** Non-discrimination, Freedom of ,→ expression, Privacy, Human autonomy, Human ,→ dignity - **Indicators:** Data protection concerns; erosion ,→ of individual autonomy through automated nudging ,→ or profiling; threats to human dignity in ,→ automated interactions; surveillance; chilling ,→ effects on expression - **Example:** A study finds that an automated ,→ content moderation system disproportionately ,→ suppresses political speech by minority groups, ,→ threatening freedom of expression.

--### None The reference has no substantive engagement with any ,→ of the six PA relevance dimensions. --### **SECTION 1: EMPIRICS** Classify the system(s) the paper actually studies ,→ empirically. This section anchors everything ,→ else.

--### Quality of governance *(Governing capability)*

For each system or system group:

The reference engages with whether the AI system ,→ affects the ability of the governance system to ,→ adapt to change and function effectively and ,→ competently.

* Give it a short name * Classify it using the taxonomy above, defaulting to ,→ Underspecified * Write 2–3 sentences explaining your classification, ,→ citing what the paper does and does not say

28

* Provide 1–3 direct quotes from the paper supporting ,→ the classification

### **SECTION 3: CONCLUSIONS / CLAIMS** Identify the paper's **core conclusions or claims** ,→ (typically 3–5). For each:

If the paper has no empirical component (pure ,→ conceptual/literature review), state this ,→ explicitly and explain what the closest analogue ,→ to an "empirical object" is. A conceptual paper's ,→ 'empirical' object is its main motivating example ,→ or conceptualisation of AI. A claim about 'AI ,→ systems with property P' is classified by which ,→ typology layer P uniquely identifies, or as ,→ Justifiably Generic if P crosses no affordance ,→ thresholds, otherwise Underspecified.

* **One-liner**: A concise summary of the claim (e.g. ,→ "AI tools reduce frontline discretion", "citizen ,→ co-design prevents algorithmic harm"). * **Public value dimension(s)**: Map to the taxonomy ,→ above. * **Supporting quotes**: 2–4 direct quotes from the ,→ paper stating or supporting this conclusion. * **(C) Overgeneralisation flag \[0/1\]**: Does the ,→ conclusion generalise beyond what the empirics ,→ can support, due to technical imprecision? Flag 1 ,→ if yes, with an explanation. * The standard here is: does the conclusion apply ,→ the finding uniformly to "AI" when the empirics ,→ only covered a specific technical subset — and ,→ would that matter for the conclusion's validity ,→ or policy applicability? * **Give leeway** for the very common case where conclusions mention "AI" generically but the ,→ underlying finding is not meaningfully ,→ ,→ distorted by this. Only flag 1 where the overgeneralisation is consequential: where ,→ applying the finding to a different system type ,→ (e.g. extending a finding about Black-box ML to ,→ hand-coded rule systems, or vice versa) would ,→ be a gross mistake, or where the conclusion is ,→ the kind that sounds much stronger than the ,→ technical evidence warrants. ,→ * A conclusion about process or governance (e.g. ,→ "citizen involvement improves ethical ,→ outcomes") gets somewhat more leeway than a ,→ conclusion about technical properties (e.g. ,→ "these systems are typically opaque"), since ,→ the former may apply across system types more ,→ plausibly.

--### **SECTION 2: MOTIVATION** Identify **2–5 key motivating strands** — the main ,→ reasons the paper gives for why its research ,→ question matters. For each strand: * **One-liner**: A concise description of what the ,→ strand is (e.g. "AI tools curbing individual ,→ discretion", "opacity of black-box models ,→ undermining accountability"). Can be longer if ,→ needed. * **Public value dimension(s)**: Map to one of the ,→ five in the taxonomy above. * **Supporting quotes**: 2–5 direct quotes from the ,→ paper that anchor this strand. * **AI system classification**: Classify the AI system(s) invoked in this motivating strand, ,→ using the taxonomy. If the paper is generic here, ,→ note "Generic across a class of systems" and flag ,→ it. ,→ * **(A) Mismotivation flag \[0/1\]**: Does this ,→ strand motivate with a type of AI system that is ,→ meaningfully different from what the paper ,→ actually studies (as classified in Section 1)? ,→ Flag 1 if yes, with an explanation. The two main ,→ failure modes are: * **(a) Imprecision**: the strand invokes AI generically in a way that obscures relevant ,→ differences (e.g. treating black-box ML and ,→ ,→ hand-coded rule systems as the same governance ,→ problem). * **(b) Mismotivation proper**: a motivating claim rests on one system type (e.g. Black-box ML) ,→ ,→ but the paper then studies a different type ,→ (e.g. hand-coded rules), making the motivation ,→ and the empirics technically mismatched. * **Give leeway** for the common case where a paper ,→ invokes "AI" broadly in its framing but the ,→ actual gap between what it motivates with and ,→ what it studies is minor or conventional. Only ,→ flag 1 when the mismatch is consequential — ,→ i.e. it would materially change the ,→ interpretation of the paper's contribution or ,→ its policy implications.

--### **ADDITIONAL GUIDANCE** **On technical classification generally:** The ,→ typology distinguishes systems by their internal ,→ logic and the nature of the inference they perform, not by their application domain. A ,→ ,→ welfare benefit calculator can be Hand-coded; a ,→ welfare risk scorer can be Black-box. A "chatbot" ,→ could be Hand-coded (rule-tree), Black-box ,→ (LLM/transformer), or Underspecified (no ,→ information given). Do not allow domain labels ("AI in healthcare", "policing AI") to substitute ,→ ,→ for technical classification.

---

29

**On the relationship between sections:** The ,→ Empirics classification is your anchor. ,→ Mismotivation (A) asks whether the motivation is ,→ technically consistent with what was studied. ,→ Overgeneralisation (C) asks whether the ,→ conclusions go beyond what the empirics showed. ,→ Both require you to compare against your Empirics ,→ classification. **On quoting:** All quotes must be verbatim from the ,→ paper. Do not paraphrase into quote marks. **On the empirics type field:** Provide a short label ,→ only (e.g. "case study", "survey", "vignette ,→ experiment", "systematic literature review", ,→ "conceptual/framework"). A fuller taxonomy will ,→ be applied later.

30

Record · ID 324945 · SHA-256 34f0901b7cdfc481
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.