arXiv:2604.22679v1 [cs.CY] 24 Apr 2026
H OW S UPPLY C HAIN D EPENDENCIES C OMPLICATE B IAS M EASUREMENT AND ACCOUNTABILITY ATTRIBUTION IN AI H IRING A PPLICATIONS
Gauri Sharma Mila Quebec AI Institute Montreal, Quebec, Canada [email protected]
Maryam Molamohammadi Mila Quebec AI Institute Montreal, Quebec, Canada [email protected]
A BSTRACT The increasing adoption of AI systems in hiring has raised concerns about algorithmic bias and accountability, prompting regulatory responses including the EU AI Act, NYC Local Law 144, and Colorado’s AI Act. While existing research examines bias through technical or regulatory lenses, both perspectives overlook a fundamental challenge: modern AI hiring systems operate within complex supply chains where responsibility fragments across data vendors, model developers, platform providers, and deploying organizations. This paper investigates how these dependency chains complicate bias evaluation and accountability attribution. Drawing on literature review and analysis of regulatory frameworks, we demonstrate that fragmented responsibilities create two critical problems. First, bias emerges from component interactions rather than isolated elements, yet proprietary configurations prevent integrated evaluation. When discrimination occurs, causal responsibility spreads across actors, creating attribution challenges. A resume parser may function without bias independently but contribute to discrimination when integrated with specific ranking algorithms and filtering thresholds. Second, information asymmetries mean employers (deploying organizations) bear legal responsibility without technical visibility into vendor-supplied algorithms, while vendors control implementations without meaningful disclosure requirements or industry specificity. Each stakeholder may believe they are compliant; nevertheless, the integrated system may produce biased outcomes. Analysis of implementation ambiguities reveals these challenges in practice. We propose multi-layered interventions including system-level audits evaluating integrated deployments, vendor guidelines requiring fairness approach disclosure, continuous monitoring mechanisms, and detailed documentation across dependency chains. Our findings reveal that effective governance requires coordinated action across technical, organizational, and regulatory domains to establish meaningful accountability in distributed development environments. Keywords AI hiring systems, algorithmic bias, accountability, supply chain dependencies, fairness evaluation, regulatory compliance, vendor guidelines, bias auditing
1
Introduction
Modern AI hiring systems are assembled from components built by different vendors, resume parsers, ranking algorithms, scoring models, and deployment infrastructure, each controlled by different actors with different obligations. The increasing adoption of these systems has sparked debate about fairness, accountability, and algorithmic bias, yet a fundamental challenge remains underexamined: when responsibility is distributed across data vendors, model developers, platform providers, and deploying organizations, bias evaluation and accountability attribution become structurally difficult. Work on algorithmic accountability in AI supply chains [1, 2] and on bias in “plug-and-play AI services” [3] has grown substantially. Yet these insights have not been connected to the specific legal obligations and regulatory frameworks
A PREPRINT
that govern high-stakes automated decision-making. AI hiring systems offer a useful starting point: they involve welldocumented bias harms [4, 5], multi-vendor pipelines [1], and active regulatory attention across multiple jurisdictions [6, 7]. Moreover, hiring decisions are documented as a key determinant of economic opportunity and social mobility [8]. We therefore investigate (RQ) How do supply chain dependencies complicate bias evaluation and accountability attribution? We focus on AI hiring systems because they offer a well-researched instance of a general problem. Hiring involves established bias harms, active multi-jurisdictional regulation, and documented multi-vendor pipelines, making it a useful starting point for understanding challenges that appear across high-stakes AI deployments more broadly. Modular development, information asymmetries, and misalignment between legal responsibility and technical control are not unique to hiring. The same structural conditions appear in healthcare, finance, education, and criminal justice [9, 6]; hiring is where these dynamics are currently best documented and most actively regulated. Over 250 commercial HR AI tools (Figure 1) now span the hiring pipeline [10]. Research has developed along two parallel perspectives: technical approaches including bias detection and fairness-aware algorithms [11], and regulatory frameworks covering compliance requirements and governance mechanisms [12]. Both treat systems as discrete and auditable, controlled by a single responsible party. Evidence shows these systems harm job seekers, particularly marginalized groups [13], and whether they reduce bias remains unclear [4, 14]. Recruiters indicate reservations about data accuracy and limited control over the systems [15]. Vendor claims about bias reduction are rarely supported by independent evidence [5], and knowledge disparities mean some candidates can game systems while others remain unaware automated screening exists [16, 17]. Regulatory frameworks increasingly treat automated hiring as high-risk, with the EU AI Act requiring strict transparency and oversight [18, 19], yet neither technical nor regulatory approaches account for how responsibility fragments across supply chains in practice. Modern AI hiring systems exist within complex supply chains (Figure 1) where different stakeholders contribute training data, foundational models, fine-tuning services, and deployment infrastructure. This makes responsibility difficult to attribute [1], creating the “problem of many hands” [20, 21]. Employers depend on vendors while lacking tools to assess model training or bias mitigation [22]. Vendors claim neutrality while employers rely on vendor-assured compliance [5]. Table 1 summarizes these power asymmetries. New York City’s Local Law 144 [23], the first U.S. law requiring bias audits for automated hiring tools, resulted in what researchers termed “null compliance,” with only 5% of employers posting required reports [24]. Traditional approaches examine systems in isolation, assuming a single accountable party [25, 26]. They fail to address how vendor-platform-user dependencies hide bias origins and spread responsibility [1]. We argue that siloed technical and regulatory solutions are insufficient to address these supply chain interdependencies [27]. Prior work on algorithmic supply chain accountability [1, 2] and bias in plug-and-play AI services [3] establishes that responsibility fragments across distributed development environments. We extend this work in three directions. First, we show that existing legal frameworks do not merely fail to address multi-vendor pipelines but actively generate accountability gaps by assigning liability based on assumptions of employer control that these pipelines do not support. Second, we treat bias attribution in multi-vendor hiring systems as structurally difficult rather than merely technically complex, in that even with full cooperation among parties, no single actor possesses sufficient visibility to audit the integrated system. Third, where prior supply chain accountability work treats legal frameworks as context, we show they are part of the problem. Drawing on literature review and regulatory analysis, we demonstrate that these dynamics create blind spots in bias detection, prevent meaningful accountability even when all parties attempt to comply, and undermine both technical and regulatory fairness measures. Effective solutions must therefore address sociotechnical complexity and distributed development rather than targeting any single actor or component.
2
Methodology
This paper combines a structured literature review with regulatory analysis to investigate how supply chain dependencies complicate bias evaluation and accountability attribution in AI hiring systems. We conducted a structured literature review across three databases: Web of Science Core Collection, Scopus, and SSRN. Search terms combined concepts across three domains: AI hiring systems (e.g., “algorithmic hiring,” “automated recruitment,” “resume screening”), bias and fairness (e.g., “algorithmic bias,” “fairness metrics,” “disparate impact”), and accountability (e.g., “supply chain accountability,” “algorithmic auditing,” “vendor responsibility”). We limited results to English language publications and applied no date restriction given the recency of the field. Records were screened by title and abstract for relevance to AI hiring, bias evaluation, or accountability in algorithmic systems, and excluded if they did not directly address these topics. We also identified additional literature through references cited in highly relevant works. We additionally reviewed practitioner and policy literature on AI governance in hiring, including regulatory impact assessments, audit reports, and industry analyses, to complement the academic literature. 2
A PREPRINT
We analyzed three primary regulatory frameworks governing AI hiring systems: the EU AI Act, New York City Local Law 144, and Colorado’s Artificial Intelligence Act. Analysis focused on how each framework assigns responsibility across supply chain actors, what compliance mechanisms it establishes, and where gaps or ambiguities arise in multi-vendor deployment contexts. Regulatory texts were read alongside secondary literature on AI governance and employment discrimination law to identify tensions between legal assumptions and the structural realities of multi-vendor pipelines.
3
AI Use Cases in the Hiring Pipeline
This section examines AI hiring systems to understand both the bias landscape these systems create and the technical dependencies through which bias propagates. We document systematic discrimination patterns across demographics and organizational contexts (see Figure 1 for bias sources) (Section 3.1), then examine how software dependencies distribute responsibility in ways that enable these patterns to persist (Section 3.2). Commercial HR AI tools (Figure 1) span the complete hiring pipeline [4, 10]. Applications include job description generation and candidate sourcing at early stages, resume parsing and matching during screening, personality prediction and game-based assessments during evaluation [28, 29, 30, 31], video interview analysis at later stages [32, 33, 34, 35], and retention prediction systems post-hire. The screening stage receives the most research attention, as it is the first algorithmic encounter candidates face in the pipeline. Modern resume screening systems employ BERT variants for parsing and matching tasks [36, 37, 38, 39], graph neural networks [40] or use hierarchical attention mechanisms [41]. Studies around deployment show processing speeds of approximately one resume per second [42], with hybrid systems combining multiple NLP approaches [43]. Emerging practices include using large language models to generate synthetic training data [44], raising concerns about laundered bias where discrimination becomes even more difficult to detect and attribute [45]. Beyond resume analysis, screening uses personality prediction [28, 46], game-based assessments [29, 30], and multimodal video analysis examining verbal content, facial expressions, and vocal characteristics [34, 35]. These diverse technical approaches, often integrated from multiple vendors, create complex dependencies where bias sources become difficult to isolate.
Figure 1: AI-enabled use cases in hiring workflow and examples of bias sources, derived from reviewed literature and market research. 3.1 3.1.1
Bias and Discrimination Systematic Discrimination Across Demographics
Studies document systematic racial bias [47] and gender-domain [48, 49, 50] associations across multiple LLMs used for resume screening. These biased recommendations alter human decision-making and limit recruiter autonomy [51], and field evidence confirms that candidates with Black-associated names receive fewer callbacks [11, 52]. Disability bias remains underexamined [53, 54]. Standard “fairness through unawareness” approaches fail candidates with disabilities or multiple marginalized identities, as these characteristics remain inferable through proxy variables [55]. Multiple marginalized identities lead to greater disadvantages than single identities [56, 4, 48], challenging fairness frameworks that treat protected categories independently. Multimodal video analysis introduces additional bias, where superficial features such as headscarves or visible background items alter personality scores [34, 35] in ways that disproportionately affect marginalized candidates, raising concerns about emotional labor burdens and privacy violations [57]. 3
A PREPRINT
3.1.2
Knowledge Disparities and Organizational Gaps
Recruiters sought AI to reduce cognitive biases but expressed concerns about data accuracy, system opacity, and loss of contextual judgment [17, 15, 58]. HR professionals need both technical skills to evaluate AI and people skills to maintain human elements [59]. Bias persists because systems learn from historically biased decisions, creating feedback loops where technical fixes alone cannot address organizational discrimination sources [60]. Candidate research reveals disparate awareness, some understand algorithmic screening while others don’t know automation is involved [16]. Evidence suggests AI does not democratize opportunity; hiring success for young seekers remains predicted by family income and referrals rather than algorithmic resume optimization [61]. Candidates perceived AI as less fair than humans, citing lack of human connection and explanation [13, 62, 63]. These bias patterns and knowledge gaps are amplified by the structure of AI supply chains. 3.2
Software Dependencies and Accountability Challenges
AI hiring systems are produced within supply chains of interconnected actors sharing data [1]. Challenges include distributed responsibility, limited visibility, and regulatory gaps [1], creating the paradox where automation intended to reduce bias may instead amplify it [14]. Infrastructure providers depend minimally on individual customers while customers depend heavily on providers, creating power asymmetries [1, 9]. When harms occur, causal factors and decision-making authority misalign [1], fragmenting accountability across stakeholders. Table 1 summarizes how control and power distribute unevenly across these actors, with each stakeholder possessing authority in some dimensions while lacking it in others. Creating AI systems involves collaborative effort across actors with unequal resources [9]. Current interventions assume end-to-end visibility, but distributed supply chains mean work outside one’s module feels beyond developer control [2]. Table 1: Stakeholder Roles and Power Asymmetries in AI Hiring Systems, based on analysis of regulatory frameworks, literature, and industry practice. Stakeholder Applicants/Unions Vendors Deployers Regulatory Bodies Auditors
Control/Power File complaints, contest decisions, opt-out (with unknown consequences) Design algorithms, determine technical feasibility, customize implementations Vendor selection, build vs. buy decisions, configuration choices Enforcement authority, regulation design
Limitations Lack access to algorithmic details, scoring methods, or filtering criteria Work with limited data, lack visibility into deployment contexts Cannot inspect vendor algorithms, limited technical expertise Struggle with technical complexity and cross-jurisdictional gaps Independent evaluation, third-party as- Limited access to proprietary systems and sessment training data
Organizations often use and reuse datasets with limited insight into the development choices behind them [64]. Without accountability mechanisms, bias can propagate as datasets are incorporated into hiring systems. Vendors may emphasize procedural fairness metrics, including the four-fifths rule, a legally established standard under US employment law [65], though researchers have noted it may not fully address outcomes for rejected candidates [66]. Independent evidence supporting vendor bias-reduction claims remains limited [5], and independent audits are rare [67]. Employers often receive high-level summaries rather than detailed system documentation, making it difficult to meaningfully evaluate tools they deploy despite bearing legal liability [22, 68]. Vendors, meanwhile, control implementations while having limited visibility into how their tools perform across different deployment contexts [69]. These challenges are further complicated by legal frameworks designed for traditional employment practices, where employers control hiring end-to-end.
4
Where does the legal accountability lie?
Integrating AI hiring systems into anti-discrimination frameworks reveals tensions between how laws assign responsibility and how systems are built. This section examines legal accountability from stakeholder perspectives, using ‘employer,’ ‘deployer,’ and ‘organization’ interchangeably to refer to entities that use AI hiring systems. As Table 1 illustrates, each stakeholder possesses control in some dimensions while facing limitations in others. Despite clear regulatory frameworks, accountability remains ambiguous due to dependency chains and limited deployer control. 4
A PREPRINT
4.1
Traditional Employment Law and Its Assumptions
Traditional anti-discrimination frameworks including Title VII [70], the Americans with Disabilities Act [71], and the Age Discrimination in Employment Act [72] prohibit discrimination based on race, gender, disability, and age. The Uniform Guidelines on Employee Selection Procedures [73] operationalize these requirements through the four-fifths rule, under which selection rates falling below 80% of the highest group’s rate constitute adverse impact requiring employers to validate procedures as job-related business necessities [65]. These frameworks assume employers control hiring end-to-end, making discrimination traceable to a single actor with both the responsibility and authority to correct it [14, 74]. This assumption breaks with AI systems, employers deploy tools they did not develop, built on models they did not train, using algorithms they cannot inspect, creating a paradox of legal responsibility without technical control or visibility. 4.2
AI-Specific Regulations: Divergent Approaches to Screening Accountability
Jurisdictions have introduced AI-specific regulations to distribute accountability. These reveal distinct strategies for assigning responsibility when screening systems discriminate. Regulations use varying terminology. The European Union’s AI Act refers to ‘providers,’ U.S. regulations use ‘developers’ or ‘vendors,’ though all denote entities that create and supply AI hiring tools. The EU AI Act classifies AI systems used for recruitment and candidate screening as “high-risk” [18, 19], requiring providers to use representative datasets, document sources, implement risk management, and undergo conformity assessments [75]. Deployers must ensure human oversight, monitor outcomes, conduct impact assessments, and notify applicants of AI screening, though implementation challenges emerge when systems trained on one company’s data exhibit bias when deployed by another with different applicant pools. Colorado’s Artificial Intelligence Act (effective June 2026) creates parallel obligations for developers and deployers, requiring impact assessment documentation and Attorney General notification within 90 days of discovering discriminatory outcomes, though necessary information for assessments remains undefined. New York City’s Local Law 144 (effective July 2023) diverges by placing all obligations on employers, requiring annual bias audits, published summaries, and candidate notifications, while creating no vendor obligations and allowing vendors to decline audit participation by citing trade secrets [76]. These contrasting approaches have produced similar failures: regulations targeting deployers without system access, or vendors without deployment visibility, both struggle to establish meaningful accountability, and analysis of 116 filed audits found widespread under-reporting through missing demographic data, opaque methods, and misaligned metrics [77]. 4.3
Implementation Ambiguity
Even when AI-specific regulations attempt to distribute accountability, their implementation reveals fundamental ambiguities [78]. The four-fifths rule [65] measures final hiring outcomes yet regulations remain unclear whether compliance is measured at screening, interviews, or final hiring stages, and following screening audits does not guarantee fair outcomes. When NYC LL 144 requires auditing “the automated tool”, the definition becomes unclear in systems integrating parsers, matching algorithms, ranking systems, and filtering thresholds [79, 67]. Analysis of filed audits found this scope ambiguity enables under-reporting through missing demographic data, opaque methods, and measuring demographic composition of advanced candidates rather than differential filtering of qualified candidates [77]. Attribution ambiguity emerges from multi-vendor integration where determining the failed component is difficult as parsers, matching algorithms, scoring thresholds, and employer configurations all contribute, while platforms like LinkedIn or Indeed control applicant data needed for audits without any obligation to share it. Temporal ambiguity further undermines accountability as parsers update, algorithms retrain, and thresholds adjust continuously, yet regulations assume static systems auditable annually and provide no guidance on whether incremental changes require reassessment. Accountability disappears across these scope, attribution, and temporal ambiguities, underscoring the need for rigorous auditing frameworks [74].
5
Discussion
Sections 3 and 4 reveal that addressing bias and accountability requires coordinated interventions across multiple domains. Technical solutions alone cannot overcome information and access asymmetries. Regulatory mandates alone cannot account for dynamic systems. Organizational policies alone cannot bridge the gap between legal obligations and technical capabilities. Our analysis identified five recurring bias and accountability gaps from supply chain dependencies: conflicting fairness definitions (Section 5.1.1), dynamic deployment contexts (Section 5.1.2), fragmented jurisdictional requirements (Section 5.2), legal-technical translation ambiguities (Section 5.3), and distributed responsibility (Sections 5.4–5.5). We propose multi-layered interventions spanning technical, regulatory, and organizational domains, organized around these challenge areas. 5
A PREPRINT
5.1
Fairness Challenges in AI Hiring Systems
Before addressing regulatory compliance or vendor responsibilities, it is necessary to first establish what ‘fairness’ means in multi-vendor deployments. Modern AI hiring systems face two interconnected fairness challenges of incompatible fairness definitions and dynamic deployment contexts that shift fairness properties over time. 5.1.1
Establishing rationale behind fairness metrics
The fairness impossibility theorem shows that common fairness metrics cannot be simultaneously optimized [80]: demographic parity requires equal selection rates, equalized odds demands equal error rates, and predictive parity ensures equal accuracy of predictions [81]. When vendors developing different components each optimize for different metrics independently, integrated systems satisfy no coherent fairness standard [82, 83, 84]. In practice, a resume parser optimizing for fairness through unawareness may pair with a ranking algorithm optimizing for demographic parity and an employer threshold satisfying the four-fifths rule [65], creating systems where each component appears compliant yet the integration discriminates. Regulations worsen this by mandating fairness without specifying which definition. The EU AI Act requires bias detection without metric guidance, Colorado requires reasonable care without defining standards, and NYC requires demographic parity measurement that may satisfy audits while other discrimination forms persist. Employers bear legal responsibility yet cannot inspect vendor fairness approaches, allowing each party to claim compliance while integrated systems discriminate. Table 2 outlines recommendations for aligning these fairness metrics. Table 2: Recommendations for Fairness Metric Alignment Recommendation
Component
Stakeholder(s)
System-level audit (evaluating fairness metric conflicts; predeployment; after system updates)
Document which fairness metric each component uses
Auditors, Vendors
Assess whether metric conflicts produce discriminatory outcomes Verify the combined system satisfies hiring context requirements
Auditors, Employers
Disclose which fairness definitions their tools optimize for
Vendors
Disclose technical methods used Disclose incompatibilities with other fairness approaches
Vendors Vendors
Vendor disclosure guidelines (fairness definitions and incompatibilities; at contracting; update as needed)
5.1.2
Employers,
Auditors, Employers
Managing dynamic fairness
Fairness in AI hiring is not static but shifts dynamically with applicant pools, skill requirements, and organizational contexts. For example, a resume screening system achieving acceptable fairness when processing 1,000 applications monthly may produce disparate impact when processing 10,000 during hiring surges. Technology sector hiring sees demographic composition shift following layoffs, while university recruiting cycles introduce temporal patterns where screening encounters different populations [85]. Skill requirements evolve over time, and systems trained on outdated taxonomies may disadvantage candidates with newer skills, producing discrimination independent of algorithmic bias. Temporal changes can also invalidate recourse mechanisms, where advice that once improved outcomes becomes ineffective as systems evolve [86]. The system’s technical operation remains unchanged, but fairness emerges from how algorithms interact with input data distributions. Current regulations assume static fairness certifiable annually, yet fairness properties change continuously [86]. The EU AI Act requires updating risk management when systems are “substantially modified” without defining this threshold. Colorado mandates annual assessments without guidance on when fairness drift requires interim reassessment. Vendoremployer dependencies amplify this since vendors control updates and retraining yet contracts rarely require fairness stability across changes. When vendors push updates without assessing fairness implications, employers must accept altered fairness properties. Employers cannot diagnose whether fairness degradation stems from applicant pool changes or algorithmic updates. We propose continuous monitoring strategies to address this in Table 3. 6
A PREPRINT
Table 3: Recommendations for Managing the Dynamicity of Fairness Recommendation
Component
Stakeholder(s)
Periodic auditing (as fairness properties change; ongoing; after major changes)
Audit when applicant pool characteristics change substantially
Auditors, Employers
Audit after vendor system updates Audit following changes to job requirements or screening thresholds Conduct audits at least quarterly Compare current performance against baseline fairness metrics to detect drift
Auditors, Employers Auditors, Employers
Review job descriptions to ensure inclusive language
Employers
Ensure recognized qualifications reflect current job needs Incorporate new credentials and emerging skills Identify and correct proxy discrimination through obsolete skill weights
Employers
Track post-hire performance to validate screening predictions across demographic groups
Employers, Auditors
Collect recruiter feedback on alignment between AI recommendations and human assessments Collect candidate surveys about perceived fairness Use continuous monitoring to detect emerging bias between audit cycles
Employers
Skill taxonomy and job description reviews (preventing biased proxies; quarterly)
Continuous feedback mechanisms (post-hire tracking and surveys; continuous)
5.2
Auditors, Employers Auditors, Employers
Employers Employers, Auditors
Employers Employers, Auditors
Fragmented Compliance Across Jurisdictions
Different states impose varying AI audit requirements, disclosure mandates, and bias testing standards, creating challenges for multi-jurisdictional employers and vendors. These divergent frameworks create operational complications [6, 7]. When screening systems integrate components from multiple vendors operating under different state regulations, compliance fragments. The resume parser vendor may conduct audits meeting NYC requirements while the ranking algorithm vendor provides Colorado documentation. Employers deploying both components should combine separate compliance efforts, yet no regulation addresses aggregating compliance across multi-vendor systems. Table 4 provides recommendations for multi-jurisdictional disclosure. Audit methodologies and fairness metrics vary across jurisdictions [6]. NYC requires measuring selection rates across demographic groups, Colorado requires impact assessments without specifying metrics, and the EU requires datasets “free of errors” with “specific attention to bias detection” without fairness definition guidance. These approaches may optimize for incompatible fairness definitions. A screening system satisfying one jurisdiction’s requirements may inadequately address another’s concerns, while notification requirements also differ in timing, specificity, and required information. 5.3
Translating Compliance Requirements Into Technical Implementations
Legal language around bias and discrimination often does not map cleanly to technical metrics. Employment discrimination law uses terms like adverse impact, disparate treatment, and business necessity. These concepts emerged from decades of case law addressing human decision-making. Translating them into technical specifications for algorithmic systems requires interpretive choices that regulations leave unresolved [87], highlighting the need for formal frameworks that bridge legal and computational definitions [88]. This creates a gap where regulations mandate compliance without specifying how to achieve it technically. When NYC LL 144 requires a bias audit, it does not specify whether to use actual applicant data or synthetic resumes, which demographic categories beyond race/ethnicity/sex to examine, or what selection rate differences constitute unacceptable bias. This ambiguity permeates technical implementation. The EU AI Act’s requirement for “sufficiently representative” datasets could mean representative of the general population, the labor force in relevant occupations, actual applicant pools, or qualified candidate pools where each interpretation produces different technical requirements and potential 7
A PREPRINT
Table 4: Recommendations for Multi-Jurisdictional Compliance Recommendation
Component
Stakeholder(s)
Vendor compliance documentation (multi-jurisdictional obligations; at contracting; update as regulations change)
Document which state and international regulations the tool addresses
Vendors
Document audit methodologies and their alignment with jurisdiction-specific requirements Document which fairness metrics are measured and whether they satisfy varying state standards Identify gaps requiring deployer-side configuration for full compliance
Vendors
Track applicant jurisdictions based on job location and residence
Employers, Vendors
Provide jurisdiction-appropriate AI-use notifications with required timing and content Document disclosure delivery for compliance verification Automate disclosure management across jurisdictions
Employers
Applicant disclosure systems (meeting state-level requirements; at application; ongoing)
Vendors Vendors
Employers Employers, Vendors
discrimination patterns. A general population dataset may underrepresent qualified candidates in specialized roles, while an actual applicant pool dataset may encode historical discrimination. Similarly, ‘bias detection’ encompasses diverse methods ranging from demographic parity measurements to complex causal inference or training data reweighting. Different methods optimize for incompatible fairness metrics, yet regulations provide no guidance on which technical approaches satisfy legal requirements. Table 5 details audit recommendations to verify this legal-technical alignment. These translation gaps create multiple problems. Regulations requiring “automated employment decision tool” audits do not clarify whether this means auditing components separately or only integrated systems, allowing employers to define audit scope narrowly. Multi-stakeholder studies of algorithmic management reveal persistent challenges in aligning software implementations with legal requirements [78]. Regulations mandate annual audits while systems change continuously, yet whether incremental changes require reassessment remains undefined. Information asymmetries prevent verification as vendors provide high-level fairness summaries while employers lack access to training data, model architectures, or algorithmic logic needed to assess whether implementations satisfy regulatory requirements. When discriminatory outcomes emerge, these gaps prevent clear attribution. When responsibility fragments across multiple parties, the integrated system can produce discriminatory outcomes even when each component meets its individual compliance obligations [1, 21]. 5.4
Accountability Complexities Through Dependency Chains
Supply chains in AI hiring involve technical dependency chains that integrate components from multiple sources, including base models, fine-tuning datasets, ranking algorithms, and deployment infrastructure. This makes bias attribution fundamentally difficult [1]. Information asymmetries mean those with the greatest need to understand systems (employers, applicants) have the least access, while vendors have the weakest disclosure incentives. Platforms take this further by controlling the applicant data employers need for mandated audits without any obligation to share it. This dependency structure creates two categories of problems, which we term evaluation convolution and attribution convolution. Evaluation convolution describes the condition where interconnected components cannot be assessed for bias in isolation because discriminatory outcomes only emerge from their interaction. For example, a resume parser, a ranking algorithm, and an employer-set threshold may each appear unbiased separately yet produce discrimination together. Yet proprietary configurations prevent full deployment analysis since vendors protect algorithmic logic as trade secrets while employers configure systems to their specific needs. Attribution convolution describes the condition where, even once discriminatory outcomes are detected, no party can identify which component or decision caused them, because each actor sees only the segment of the pipeline they control and no one holds an end-to-end view. When screening rejects qualified candidates from protected groups at higher rates, identifying the source requires determining whether bias originated in training data, feature extraction, matching algorithms, ranking calibration, filtering thresholds, or their interactions. Regulations assign obligations by actor rather than by interaction, while causal 8
A PREPRINT
Table 5: Recommendations for Legal-Technical Translation Recommendation
Component
Stakeholder(s)
System-level audit (verifying technical–legal alignment; predeployment; after major updates)
Verify whether technical implementations address the legal requirements they claim to satisfy
Auditors
Document which interpretations were chosen when translating ambiguous legal language into technical specifications Assess whether chosen technical approaches suit the deployment context Verify how requirements such as “free of errors” were operationalized (e.g., validation, outlier removal, bias correction)
Auditors
Document how each applicable legal requirement was technically operationalized
Vendors
Document which fairness metrics and biascorrection methods were implemented and why Document testing procedures used to validate legal compliance Document known limitations where the system may not fully satisfy legal standards
Vendors
Verify whether new features or data sources introduced compliance risks
Auditors
Verify whether technical implementations still satisfy evolving legal requirements Align audit frequency with system change velocity (e.g., quarterly for frequent updates)
Auditors
Vendor technical documentation (pre-deployment compliance verification; before deployment)
Periodic re-auditing (ongoing compliance verification; quarterly or annually depending on update frequency)
Audit trail documentation (enabling attribution; continuous; retained for legal accountability)
Auditors Auditors
Vendors Vendors
Auditors, Employers
Log which system version was active for each screening decision
Employers, Vendors
Log configuration settings and filtering thresholds applied Log vendor updates and their timing Log any manual overrides or customizations
Employers Vendors Employers
responsibility spreads across so many components that attribution becomes impractical [2]. Table 6 maps specific audit and documentation responsibilities across the supply chain. These two categories produce a third problem, remediation convolution, where identified discrimination cannot be corrected because the party bearing legal responsibility to act and the party holding technical control to act are different actors with no obligation to coordinate. Vendors cannot remediate deployment configurations, while employers cannot modify proprietary model architectures, locking bias into the system. Temporal dynamics deepen these convolutions where systems evolve continuously through updates, retraining, and configuration changes, yet vendors may not retain historical model versions and employers may not log configuration settings. Technical, organizational, commercial, and regulatory factors combine to produce a fourth category, accountability convolution, where responsibility is distributed across enough parties that no single actor possesses both the visibility and the control required to ensure the integrated system is fair (Figure 2). 5.5
Accountability Fragmentation Across Stakeholders
Accountability fragments differently across employers, vendors, applicants, and regulatory bodies based on their distinct positions in the AI hiring ecosystem, as outlined in Table 1. Each stakeholder faces unique constraints and information asymmetries that limit effective oversight [89, 90]. Employers bear legal liability (e.g., Title VII) for discriminatory filtering yet often lack access to the training data, weights, or architectures needed to evaluate vendor algorithms. At the same time, vendors may lack reliable demographic data or have no legal right to use applicant data in model development. These asymmetries run in both directions, making meaningful auditing difficult for all parties involved. 9
A PREPRINT
Figure 2: Beyond data and model bias, information asymmetries, vendor-dependent supply chains, legal-technical misalignment, and fragmented visibility converge to produce evaluation, attribution, remediation, and accountability convolutions in AI hiring systems. Table 6: Recommendations for Mapping Accountability in Dependency Chains Recommendation
Component
Stakeholder(s)
System-level audit (tracing complete dependency chains; end-toend; during system evaluation)
Map all components from training data through deployment infrastructure
Auditors, Vendors
Identify which vendor or party controls each component Test components both in isolation and in integration to detect where bias emerges Document interactions that produce discriminatory outcomes
Auditors
Document each party’s contribution (training data, model architecture, fine-tuning, configuration)
Vendors, Employers
Document dependencies on upstream inputs and downstream specifications Maintain version history and update timing
Vendors, Employers
Disclose whether applicant data is shared with thirdparty services
Vendors
Disclose data retention periods and purposes Disclose whether data is used to retrain models affecting other clients Disclose data handling when vendor relationships terminate
Vendors Vendors
Specify which party remediates bias in training data versus deployment configuration
Employers, Vendors
Specify remediation timelines when discrimination is identified Specify cost allocation for corrective actions Specify coordination processes for multi-party fixes
Employers, Vendors
Standardized documentation (enabling attribution across touchpoints; standardized across the supply chain)
Data flow disclosure (vendor usage throughout supply chain; at contracting; update as practices change)
Contractual responsibility allocation (addressing remediation convolution; at contracting; enforced during incidents)
Auditors Auditors
Vendors, Employers
Vendors
Employers, Vendors Employers, Vendors
Table 7 outlines operational protocols and training to bridge this gap. When discriminatory patterns emerge, substantial switching costs combine with inability to attribute bias to vendor components versus employer configurations. Vendors possess detailed knowledge about algorithm training and which applicant groups their systems filter out, yet consistently frame their role as providing neutral technology. When discriminatory outcomes emerge, vendors point to employer deployment choices, data limitations, or technical constraints while protecting algorithms as trade secrets. Platforms 10
A PREPRINT
use similar strategies but can deny employers the applicant data needed for mandated audits, preventing determination of whether discrimination originates from platform filtering or employer tools. Vendors routinely make bias-reduction claims without evidence [5], and independent audits remain extremely rare [67]. Job applicants hold legal rights to challenge discriminatory screening, but exercising these rights requires knowing how they were evaluated. Only a small percentage of employers post required notices, and even when applicants know AI screened them, they lack information about extracted factors, scoring methods, or thresholds [16]. Regulatory bodies verify compliance through documentation and audit reports rather than direct system access, making it difficult to catch inaccurate reporting, and regulators depend on complaints and audits from the very parties whose accountability failures enabled the harm. Cross-border deployments, jurisdictional gaps, and regulatory design failures where large employers fall outside a law’s scope by definition further limit enforcement effectiveness. These fragmented structures enable each party to claim compliance while integrated systems discriminate, creating distributed responsibility without accountability [21, 20]. Addressing these barriers requires coordinated action spanning technical monitoring, organizational guidelines, and regulatory controls rather than relying on any individual mechanism or voluntary commitment.
Table 7: Recommendations for Mapping Accountability Across Stakeholders Recommendation
Component
Stakeholder(s)
Employer preparedness (training and response protocols; predeployment; ongoing)
Train recruiters on how screening algorithms filter candidates and weight factors
Employers
Train recruiters on system limitations and known bias patterns Train recruiters on when and how to override AI recommendations Train recruiters on legal responsibility for discriminatory outcomes Establish thresholds that trigger system suspension Establish vendor escalation procedures Establish interim manual screening processes Define decision authority for system modifications Re-evaluate affected past candidates when bias is discovered
Employers, Vendors
Specify vendor responsibility for training data quality and model bias
Employers, Vendors
Specify deployer responsibility for configuration and monitoring Allocate liability when discrimination results from component interactions
Employers, Vendors
Disclose which hiring stages use automated screening
Employers
Disclose general categories of factors considered by algorithms Provide processes for contesting decisions and requesting human review Provide contact information for questions Log which AI tools evaluated each application and at what stage
Employers
Enable internal auditor escalation for bias patterns
Employers, Auditors
Notify vendors when deployment reveals discrimination Maintain audit trails combining feedback and system records
Employers
Vendor contractual requirements (responsibility allocation and disclosure obligations; at contracting)
Applicant transparency and recourse (disclosure requirements and accessible documentation; at application; retained for accountability)
Continuous monitoring systems (multi-stakeholder feedback and audit trails; continuous)
11
Employers Employers Employers Employers Employers Employers Employers
Employers, Vendors
Employers Employers Employers
Employers, Auditors
A PREPRINT
5.6
An Illustrative Example
The following example is fictional but draws on structural conditions documented throughout this paper. Employer E licenses a resume screening platform from Vendor V. Vendor V’s platform integrates a matching model originally developed by Company M, a specialist AI firm that Vendor V acquires eighteen months into the contract. Before the acquisition, Company M has no direct relationship with Employer E and no obligations under the contract. After the acquisition, Company M’s model is integrated into Vendor V’s platform, but the version active at deployment differs from the one described in the original fairness documentation. Employer E configures filtering thresholds based on Vendor V’s onboarding recommendations and treats Vendor V’s component-level fairness summary as sufficient due diligence. No integrated audit occurs at any stage. Recruiter R eventually notices that Candidate C and others from certain demographic groups are being filtered out at disproportionate rates. Whether this pattern predates the acquisition, results from it, or emerges from the interaction of Company M’s model with Vendor V’s platform and Employer E’s thresholds cannot be determined. No party holds an end-to-end view of the pipeline. This is evaluation convolution: the integrated system cannot be assessed because each actor sees only their own segment. When Employer E escalates, attribution convolution follows: Vendor V points to Employer E’s threshold configurations; Employer E points to the model’s feature weightings provided by Vendor V; the design choices embedded in Company M’s model before the acquisition are not visible to either party. Even if the source were identified, remediation convolution would apply: fixing the system requires coordinated changes across three parties with no established process for doing so. Employer E’s annual audit examined components in isolation, satisfying regulatory requirements while leaving the integrated system unevaluated. Vendor V declines broader audit participation, citing trade secrets. Company M, now absorbed into Vendor V, points to the transfer of responsibility at acquisition. The platform holding Candidate C’s application data has no obligation to share it. Each party has met its individual obligations. The system continues to screen out Candidate C. This is accountability convolution. What could have prevented this outcome is what Section 5 proposes: system-level audits of integrated deployments, remediation protocols covering multi-party coordination including provisions for mid-contract acquisitions, vendor disclosure of fairness approaches at contracting, and ongoing monitoring rather than periodic snapshots. These mechanisms did not exist here, and they are not currently required by law. This example captures a broader pattern where bias sources multiply across supply chain actors while regulatory and accountability frameworks fail to keep pace, leaving no single intervention able to ensure the integrated system is fair.
6
Conclusion
AI hiring systems exemplify a broader challenge in algorithmic bias and accountability where bias sources, fairness definitions, regulatory requirements, and dependency chains fragment both bias detection and accountability across the integrated system in ways that existing governance frameworks cannot address. When multiple vendors contribute components to integrated systems, bias evaluation and accountability attribution become structurally difficult rather than merely technically complex. Current approaches, whether technical audits of isolated components or regulatory mandates targeting single actors, cannot overcome the evaluation, attribution, remediation, and accountability convolutions created by supply chain dependencies. Our analysis reveals that effective governance requires fundamental shifts. Effective governance requires auditing integrated deployments rather than isolated components, establishing shared standards in contracts that give all parties a common basis for accountability, and updating regulations to account for multi-party technical misalignment. Without coordinated interventions spanning these domains, stakeholders can satisfy individual regulatory requirements while the complete system exhibits bias. Our analysis has limitations. We focus on AI hiring systems primarily in the United States and European Union, and accountability dynamics may differ in regulatory contexts with varying legal requirements and enforcement mechanisms. Our emphasis on resume screening reflects where algorithmic hiring has received the most scrutiny, though other stages of the hiring pipeline pose their own challenges and warrant further investigation. The interventions we propose are overarching normative mechanisms, intended to apply regardless of deployment context, rather than empirically validated recommendations. More specific guidance would require detailed knowledge of the organizational context, regional regulatory environment, and use case in question, as well as access to proprietary contracts, system architectures, and deployment data that vendors protect as trade secrets. To validate and operationalize these mechanisms in practice, they would need to be tailored accordingly. That inaccessibility is itself a product of the fragmented pipelines, information asymmetries, and misaligned legal responsibilities this paper documents. The interventions we propose are therefore starting points rather than complete solutions, intended as a foundation that practitioners and policymakers can expand and adapt to their specific settings. The deeper problem this paper 12
A PREPRINT
documents is not specific to hiring: wherever AI systems are assembled from components built by different actors under different obligations, the bias and accountability problems we document will follow. Hiring makes this visible because the regulation is active, the harms are documented, and the pipeline is well mapped. The same convolutions will surface in healthcare, criminal justice, education, and finance as modular development and global supply chains become the norm rather than the exception. Future research should examine whether the mechanisms proposed here can function given commercial incentives discouraging transparency, technical complexities resisting attribution, and power asymmetries concentrating control with vendors. Developing governance approaches that establish meaningful accountability in distributed development environments is a prerequisite for AI systems that deliver on promises of fairness and equitable opportunity.
Generative AI Usage Statement The authors used ChatGPT free version for grammar and style editing, and LaTeX formatting assistance. No AIgenerated content was used for substantive writing, analysis, or ideas. The authors are fully responsible for all intellectual contributions and the accuracy of this work.
References [1] Jennifer Cobbe, Michael Veale, and Jatinder Singh. Understanding accountability in algorithmic supply chains. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’23, page 1186–1197, New York, NY, USA, 2023. Association for Computing Machinery. [2] David Gray Widder and Dawn Nafus. Dislocated accountabilities in the “ai supply chain”: Modularity and developers’ notions of responsibility. Big Data & Society, 10(1):20539517231177620, 2023. [3] Kornel Lewicki, Michelle Seng Ah Lee, Jennifer Cobbe, and Jatinder Singh. Out of context: Investigating the bias and fairness concerns of “artificial intelligence as a service”. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, New York, NY, USA, 2023. Association for Computing Machinery. [4] Alessandro Fabris, Nina Baranowska, Matthew J. Dennis, David Graus, Philipp Hacker, Jorge Saldivar, Frederik Zuiderveen Borgesius, and Asia J. Biega. Fairness and bias in algorithmic hiring: A multidisciplinary survey. ACM Trans. Intell. Syst. Technol., 16(1), January 2025. [5] Manish Raghavan, Solon Barocas, Jon Kleinberg, and Karen Levy. Mitigating bias in algorithmic hiring: evaluating claims and practices. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, FAT* ’20, page 469–481, New York, NY, USA, 2020. Association for Computing Machinery. [6] Sacha Alanoca, Shira Gur-Arieh, Tom Zick, and Kevin Klyman. Comparing apples to oranges: A taxonomy for navigating the global landscape of ai regulation. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’25, page 914–937, New York, NY, USA, 2025. Association for Computing Machinery. [7] Lewin Schmitt. Mapping global ai governance: a nascent regime in a fragmented landscape. AI and Ethics, 2, 05 2022. [8] NATALIE SHEARD. Algorithm-facilitated discrimination: a socio-legal study of the use by employers of artificial intelligence hiring systems. Journal of Law and Society, 52(2):269–291, 2025. [9] Ian Brown. Allocating accountability in AI supply chains. Technical report, Ada Lovelace Institute, June 2023. [10] Ben Eubanks. Artificial Intelligence for HR: Use AI to Support and Develop a Successful Workforce. Kogan Page, London, UK, 2nd edition, 2022. [11] Zhisheng Chen. Ethics and discrimination in artificial intelligence-enabled recruitment practices. Humanities and Social Sciences Communications, 10:Article 567, 2023. [12] Anna Lena Hunkenschroer and Christoph Luetge. Ethics of ai-enabled recruiting and selection: A review and research agenda. Journal of Business Ethics, 178(4):977–1007, 2022. [13] Cassidy Pyle, Kat Roemmich, and Nazanin Andalibi. U.s. job-seekers’ organizational justice perceptions of emotion ai-enabled interviews. Proc. ACM Hum.-Comput. Interact., 8(CSCW2), November 2024. [14] Ifeoma Ajunwa. The paradox of automation as anti-bias intervention. Cardozo L. Rev., 41:1671, 2019. [15] Lan Li, Tina Lassiter, Joohee Oh, and Min Kyung Lee. Algorithmic hiring in practice: Recruiter and hr professional’s perspectives on ai use in hiring. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, AIES ’21, page 166–176, New York, NY, USA, 2021. Association for Computing Machinery. 13
A PREPRINT
[16] Lena Armstrong, Jayne Everson, and Amy J. Ko. Navigating a black box: Students’ experiences and perceptions of automated hiring. In Proceedings of the 2023 ACM Conference on International Computing Education Research - Volume 1, ICER ’23, page 148–158, New York, NY, USA, 2023. Association for Computing Machinery. [17] Mitra Lashkari and Jinghui Cheng. “finding the magic sauce”: Exploring perspectives of recruiters and job seekers on recruitment bias and automated tools. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, New York, NY, USA, 2023. Association for Computing Machinery. [18] European Union. Annex iii: High-risk ai systems referred to in article 6(2). EU Artificial Intelligence Act (Regulation (EU) 2024/1689), 2024. Accessed: 2026-01-05. [19] European Parliament and Council of the European Union. Annex III: High-risk AI systems referred to in article 6(2). regulation (EU) 2024/1689 of the european parliament and of the council on artificial intelligence (artificial intelligence act). Official Journal of the European Union. https://eur-lex.europa.eu/legal-content/ EN/TXT/HTML/?uri=OJ:L_202401689, 2024. Accessed: 2026-01-05. [20] Dennis F. Thompson. Moral responsibility of public officials: The problem of many hands. American Political Science Review, 74(4):905–916, 1980. [21] Helen Nissenbaum. Accountability in a computerized society. Science and Engineering Ethics, 2(1):25–42, March 1996. [22] PwC. Responsible ai and third-party risk management: what you need to know. Industry report, PricewaterhouseCoopers, 2025. [23] New York City Council. Local law 144: Automated employment decision tools. New York City Administrative Code, 2021. Effective July 5, 2023. [24] Lucas Wright, Roxana Mika Muenster, Briana Vecchione, Tianyao Qu, Pika (Senhuang) Cai, Alan Smith, Comm 2450 Student Investigators, Jacob Metcalf, and J. Nathan Matias. Null compliance: Nyc local law 144 and the challenges of algorithm accountability. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’24, page 1701–1713, New York, NY, USA, 2024. Association for Computing Machinery. [25] Inioluwa Deborah Raji, Andrew Smart, Rebecca N. White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes. Closing the ai accountability gap: defining an end-to-end framework for internal algorithmic auditing. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, FAT* ’20, page 33–44, New York, NY, USA, 2020. Association for Computing Machinery. [26] Abeba Birhane, Ryan Steed, Victor Ojewale, Briana Vecchione, and Inioluwa Deborah Raji. Ai auditing: The broken bus on the road to ai accountability, 2024. [27] Claudio Novelli, Mariarosaria Taddeo, and Luciano Floridi. Accountability in artificial intelligence: what it is and how it works. AI & Society, 39(4):1871–1882, August 2024. [28] Eric Grunenberg, Heinrich Peters, Matt J. Francis, Mitja D. Back, and Sandra C. Matz. Machine learning in recruiting: predicting personality from cvs and short text responses. Frontiers in Social Psychology, Volume 1 2023, 2024. [29] Franziska Leutner, Sonia-Cristina Codreanu, Suzanne Brink, and Theodoros Bitsakis. Game based assessments of cognitive ability in recruitment: Validity, fairness and test-taking experience. Frontiers in Psychology, 13:942662, 2023. [30] Richard N Landers, Michael B Armstrong, Andrew B Collmus, Salih Mujcic, and Jason Blaik. Theory-driven game-based assessment of general cognitive ability: Design theory, measurement, prediction of performance, and test fairness. Journal of Applied Psychology, 107(10):1655, 2022. [31] Pedro J Ramos-Villagrasa, Elena Fernandez-del Rio, and Angel Castro. Game-related assessments for personnel selection: A systematic review. Frontiers in Psychology, 13:952002, 2022. [32] Monideepa Tarafdar, Irina Rets, Lindsey Zuloaga, and Nathan Mondragon. How hirevue created “glass box” transparency for its ai application. MIS Quarterly Executive, 24(1):47–65, 2025. [33] Josh Liff, Nathan Mondragon, Cari Gardner, Christopher J Hartwell, and Adam Bradshaw. Psychometric properties of automated video interview competency assessments. Journal of Applied Psychology, 109(6):921, 2024. [34] Brandon M. Booth, Louis Hickman, Shree Krishna Subburaj, Louis Tay, Sang Eun Woo, and Sidney K. D’Mello. Bias and fairness in multimodal machine learning: A case study of automated video interviews. In Proceedings of the 2021 International Conference on Multimodal Interaction, ICMI ’21, page 268–277, New York, NY, USA, 2021. Association for Computing Machinery. 14
A PREPRINT
[35] Eleanor Drage and Kerry Mackereth. Does ai debias recruitment? race, gender, and ai’s “eradication of difference”. Philosophy & technology, 35(4):89, 2022. [36] Changmao Li, Elaine Fisher, Rebecca Thomas, Steve Pittard, Vicki Hertzberg, and Jinho D. Choi. Competencelevel prediction and resume & job description matching using context-aware transformer models. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8456–8466, Online, November 2020. Association for Computational Linguistics. [37] XiaoWei Li, Hui Shu, Yi Zhai, and ZhiQiang Lin. A method for resume information extraction using bert-bilstmcrf. In 2021 IEEE 21st International Conference on Communication Technology (ICCT), pages 1437–1442, 2021. [38] Vedant Bhatia, Prateek Rawat, Ajit Kumar, and Rajiv Ratn Shah. End-to-end resume parsing and finding candidates for a job description using bert, 2019. [39] Dor Lavi, Volodymyr Medentsiy, and David Graus. consultantbert: Fine-tuned siamese sentence-bert for matching jobs and job seekers, 2021. [40] Shuqing Bian, Xu Chen, Wayne Xin Zhao, Kun Zhou, Yupeng Hou, Yang Song, Tao Zhang, and Ji-Rong Wen. Learning to match jobs with resumes from sparse interaction data using multi-view co-teaching network. 2020. [41] Chuan Qin, Hengshu Zhu, Tong Xu, Chen Zhu, Chao Ma, Enhong Chen, and Hui Xiong. An enhanced neural network approach to person-job fit in talent recruitment. ACM Trans. Inf. Syst., 38(2), February 2020. [42] Asmita Deshmukh and Anjali Raut. Applying bert-based nlp for automated resume screening and candidate ranking. volume 12, pages 591–603. Springer, 2025. [43] Gurushantha Murthy G R, Shinu Abhi, and Rashmi Agarwal. A hybrid resume parser and matcher using regex and ner. In 2023 International Conference on Advances in Computation, Communication and Information Technology (ICAICCIT), pages 24–29, 2023. [44] Panagiotis Skondras, Panagiotis Zervas, and Giannis Tzimas. Generating synthetic resume data with large language models for enhanced job description classification. Future Internet, 15(11):363, 2023. [45] Sierra Wyllie, Ilia Shumailov, and Nicolas Papernot. Fairness feedback loops: Training on synthetic data amplifies bias. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’24, page 2113–2147, New York, NY, USA, 2024. Association for Computing Machinery. [46] Louis Hickman, Nigel Bosch, Vincent Ng, Rachel Saef, Louis Tay, and Sang Eun Woo. Automated video interview personality assessments: Reliability, validity, and generalizability investigations. Journal of Applied Psychology, 107(8):1323–1351, 2022. [47] Johann D. Gaebler, Sharad Goel, Aziz Huq, and Prasanna Tambe. Auditing the use of language models to guide hiring decisions, 2024. [48] Lena Armstrong, Abbey Liu, Stephen MacNeil, and Danaë Metaxa. The silicon ceiling: Auditing gpt’s race and gender biases in hiring. In Proceedings of the 4th ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, EAAMO ’24, New York, NY, USA, 2024. Association for Computing Machinery. [49] Kyra Wilson and Aylin Caliskan. Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval, page 1578–1590. AAAI Press, 2024. [50] Jiafu An, Difang Huang, Chen Lin, and Mingzhu Tai. Measuring gender and racial biases in large language models: Intersectional evidence from automated resume evaluation. PNAS Nexus, 4(3):pgaf089, March 2025. [51] Kyra Wilson, Mattea Sim, Anna-Maria Gueorguieva, and Aylin Caliskan. No thoughts just ai: Biased llm hiring recommendations alter human decision making and limit human autonomy, 2025. [52] Patrick Kline, Evan K Rose, and Christopher R Walters. Systemic discrimination among large u.s. employers*. The Quarterly Journal of Economics, 137(4):1963–2036, 06 2022. [53] Kate Glazko, Yusuf Mohammed, Ben Kosa, Venkatesh Potluri, and Jennifer Mankoff. Identifying and improving disability bias in gpt-based resume screening. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’24, page 687–700, New York, NY, USA, 2024. Association for Computing Machinery. [54] Maarten Buyl, Christina Cociancig, Cristina Frattone, and Nele Roekens. Tackling algorithmic disability discrimination in the hiring process: An ethical, legal and technical analysis. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’22, page 1071–1082, New York, NY, USA, 2022. Association for Computing Machinery. 15
A PREPRINT
[55] Nicholas Tilmes. Disability, fairness, and algorithmic bias in ai recruitment. Ethics and Inf. Technol., 24(2), June 2022. [56] Elisabeth K. Kelan. Algorithmic inclusion: Shaping the predictive algorithms of artificial intelligence in hiring. Human Resource Management Journal, 34(3):694–707, 2024. [57] Alexis Shore Ingber and Nazanin Andalibi. Emotion ai in job interviews: Injustice, emotional labor, identity, and privacy. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’25, page 1–17, New York, NY, USA, 2025. Association for Computing Machinery. [58] Zhisheng Chen. Collaboration among recruiters and artificial intelligence: removing human prejudices in employment. Cognition, Technology & Work, 25(1):135–149, 2023. [59] Melika Soleimani, Ali Intezari, James Arrowsmith, David J Pauleen, and Nazim Taskin. Reducing ai bias in recruitment and selection: an integrative grounded approach. The International Journal of Human Resource Management, pages 1–36, 2025. [60] Wilberforce Murikah, Jeff Kimanga Nthenge, and Faith Mueni Musyoka. Bias and ethics of ai systems applied in auditing - a systematic review. Scientific African, 25:e02281, 2024. [61] Lena Armstrong and Danaé Metaxa. Navigating automated hiring: Perceptions, strategy use, and outcomes among young job seekers. Proc. ACM Hum.-Comput. Interact., 9(2), May 2025. [62] Agata Mirowska and Laura Mesnet. Preferring the devil you know: Potential applicant reactions to artificial intelligence evaluation of interviews. Human Resource Management Journal, 32(2):364–383, 2022. [63] Md Sajjad Hosain, Mohammad Bin Amin, Gouranga Chandra Debnath, and Md Atikur Rahaman. The use of artificial intelligence (ai) in the hiring process: Job applicants’ perceptions of procedural justice. Computers in Human Behavior Reports, 19:100713, 2025. [64] Ben Hutchinson, Andrew Smart, Alex Hanna, Remi Denton, Christina Greer, Oddur Kjartansson, Parker Barnes, and Margaret Mitchell. Towards accountability for machine learning datasets: Practices from software engineering and infrastructure. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, page 560–575, New York, NY, USA, 2021. Association for Computing Machinery. [65] Equal Employment Opportunity Commission and U.S. Department of Justice and Office of Personnel Management and U.S. Department of Labor and U.S. Department of the Treasury. Questions and answers to clarify and provide a common interpretation of the Uniform Guidelines on Employee Selection Procedures, March 1979. Federal Register, Vol. 44, No. 43, Friday, March 2, 1979. Citation: Title VII, 29 CFR Part 1607. [66] Päivi Seppälä and Magdalena Małecka. Ai and discriminative decisions in recruitment: Challenging the core assumptions. Big Data & Society, 11(1):20539517241235872, 2024. [67] Christo Wilson, Avijit Ghosh, Shan Jiang, Alan Mislove, Lewis Baker, Janelle Szary, Kelly Trindel, and Frida Polli. Building and auditing fair algorithms: A case study in candidate screening. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, page 666–677, New York, NY, USA, 2021. Association for Computing Machinery. [68] Marco Barone. Third-party ai risk: A holistic approach to vendor assessment, February 2024. [69] Avi Gesser, Matt Kelly, Johanna Skrzypczyk, Tigist Kassahun, Jarrett Lewis, and Joshua A. Goland. Good ai vendor risk management is hard, but doable. Debevoise Data Blog, sep 2024. Accessed: 2026-01-12. [70] United States Congress. Title VII of the Civil Rights Act of 1964. Public Law 88-352, 42 U.S.C. § 2000e et seq., 1964. As amended by the Civil Rights Act of 1991 and the Lilly Ledbetter Fair Pay Act of 2009. [71] United States Congress. Americans with disabilities act of 1990. S. 933, 101st Cong., Public Law 101-336, 42 U.S.C. § 12101 et seq., July 1990. Introduced by Sen. Tom Harkin (D-IA) on May 9, 1989; passed Senate Sept. 7, 1989; signed into law July 26, 1990. [72] United States Congress. Age discrimination in employment act of 1967. Public Law 90-202, 29 U.S.C. § 621 et seq., 1967. As amended by the Older Workers Benefit Protection Act of 1990 and the Civil Rights Act of 1991. [73] Equal Employment Opportunity Commission. PART 60-3—Uniform Guidelines on Employee Selection Procedures (1978). 41 CFR Part 60-3, August 1978. Authority: Secs. 201, 202, 203, 203(a), 205, 206(a), 301, 303(b), and 403(b) of E.O. 11246; as amended by sec. 715 of Civil Rights Act of 1964, as amended (42 U.S.C. 2000(e)-14). Source: 43 FR 38295, 38314. [74] Ifeoma Ajunwa. An auditing imperative for automated hiring. Harvard Journal of Law & Technology, 34(1), 2021. 16
A PREPRINT
[75] European Parliament and Council of the European Union. Article 10: Data and data governance. Regulation (EU) 2024/1689 of the European Parliament and of the Council on Artificial Intelligence (Artificial Intelligence Act), 2024. Official Journal of the European Union. [76] Brenda Leong Jey Kumarasamy. Practical considerations for bias audits under NYC Local Law 144, 2023. Accessed: 2025. [77] Marissa Kumar Gerchick, Ro Encarnación, Cole Tanigawa-Lau, Lena Armstrong, Ana Gutiérrez, and Danaé Metaxa. Auditing the audits: Lessons for algorithmic accountability from local law 144’s bias audits. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’25, page 29–44, New York, NY, USA, 2025. Association for Computing Machinery. [78] Jonathan Lynn, Rachel Y. Kim, Sicun Gao, Daniel Schneider, Sachin S. Pandya, and Min Kyung Lee. Regulating algorithmic management: A multi-stakeholder study of challenges in aligning software and the law for workplace scheduling. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’25, page 547–572, New York, NY, USA, 2025. Association for Computing Machinery. [79] Lara Groves, Jacob Metcalf, Alayna Kennedy, Briana Vecchione, and Andrew Strait. Auditing work: Exploring the new york city algorithmic bias audit regime. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’24, page 1107–1120, New York, NY, USA, 2024. Association for Computing Machinery. [80] Manish Raghavan. Inherent tradeoffs in the fair determination of risk scores, 2023. [81] Andrew Bell, Lucius Bynum, Nazarii Drushchak, Tetiana Zakharchenko, Lucas Rosenblatt, and Julia Stoyanovich. The possibility of fairness: Revisiting the impossibility theorem in practice. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’23, page 400–422, New York, NY, USA, 2023. Association for Computing Machinery. [82] Shira Mitchell, Eric Potash, Solon Barocas, Alexander D’Amour, and Kristian Lum. Algorithmic fairness: Choices, assumptions, and definitions. Annual Review of Statistics and Its Application, 8:141–163, 2021. [83] Reuben Binns. Fairness in machine learning: Lessons from political philosophy. In Sorelle A. Friedler and Christo Wilson, editors, Proceedings of the 1st Conference on Fairness, Accountability and Transparency, volume 81 of Proceedings of Machine Learning Research, pages 149–159. PMLR, 23–24 Feb 2018. [84] Derek Leben. Normative principles for evaluating fairness in machine learning. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, AIES ’20, page 86–92, New York, NY, USA, 2020. Association for Computing Machinery. [85] U.S. Equal Employment Opportunity Commission. High tech, low inclusion: Diversity in the high tech workforce and sector, 2014–2022. Technical report, U.S. Equal Employment Opportunity Commission, September 2024. [86] Giovanni De Toni, Stefano Teso, Bruno Lepri, and Andrea Passerini. Time can invalidate algorithmic recourse. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’25, page 89–107, New York, NY, USA, 2025. Association for Computing Machinery. [87] Mathias Hanson, Gregory Lewkowicz, and Sam Verboven. Engineering the law-machine learning translation problem: developing legally aligned models, 2026. [88] Holli Sargeant and Måns Magnusson. Formalising anti-discrimination law in automated decision systems. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’25, page 181–194, New York, NY, USA, 2025. Association for Computing Machinery. [89] Lars Lindkvist and Sue Llewellyn. Accountability, responsibility and organization. Scandinavian Journal of Management, 19(2):251–273, 2003. [90] G. Fahey and F. Köster. Means, ends and meaning in accountability for strategic education governance. OECD Education Working Papers 204, OECD Publishing, Paris, 2019.
17