FAIR-Compute A Roadmap for Fair and Efficient Allocation of Federated Digital Research Infrastructure Final Report to the National Federated Compute Services (NFCS) Flexible Fund
arXiv:2607.28290v1 [cs.DC] 30 Jul 2026
Project short code: FAIRC
Principal Investigator Professor Konstantinos (Kostas) E. Zachariadis School of Economics and Finance, Queen Mary University of London
Project team Dr Ahmed Sayed Professor Dimitris Fotakis Angeliki Mathioudaki Wan Shuen Siaw Final project report
•
Co-lead — Electronic Engineering & Computer Science, QMUL Advisor — Electrical & Computer Engineering, NTUA Researcher (simulation & model) — NTUA & ICCS Researcher (roadmap & field data) — Mathematical Sciences, QMUL Submitted July 2026
FAIR-Compute
NFCS Flexible Fund — Final Report
Contents Executive Summary
3
1 Introduction and Approach
3
2 Methodology and Evidence Base
4
2.1
Stakeholder interviews and follow-up . . . . . . . . . . . . . . . . . . . . . . . . .
4
2.2
Demand-side stakeholder survey
. . . . . . . . . . . . . . . . . . . . . . . . . . .
5
2.3
Literature review . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
6
2.4
Simulation study: rationale . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
6
3 The UK and International Research-Computing Landscape
6
3.1
Infrastructure tiers and federation models . . . . . . . . . . . . . . . . . . . . . .
6
3.2
Occupancy, utilisation, and the economic case . . . . . . . . . . . . . . . . . . . .
7
3.3
Where allocation breaks down . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
7
3.4
Demand-side findings . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
9
3.5
International comparison and lessons . . . . . . . . . . . . . . . . . . . . . . . . .
9
3.6
Sustainability and carbon . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
4 A Model of HPC Allocation
10
4.1
System, jobs, and the descriptive scheduler . . . . . . . . . . . . . . . . . . . . . 10
4.2
Strategic reports and user utility . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
4.3
Offline welfare benchmark . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
5 Simulation Study
11
5.1
Data, synthetic generation, and framework . . . . . . . . . . . . . . . . . . . . . . 11
5.2
Experiment 1 — Priority weights as policy levers . . . . . . . . . . . . . . . . . . 14
5.3
Experiment 2 — Practical rules versus the offline benchmark . . . . . . . . . . . 15
5.4
Experiment 3 — Strategic reporting and scheduling externalities . . . . . . . . . 16
5.5
Experiment 4 — Federation as routing under congestion . . . . . . . . . . . . . . 17
5.6
Summary of findings . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
6 Recommendations for the NFCS Roadmap
20
7 Future Work and Research Directions
24
8 Conclusion
25
A Strategic-reporting mechanisms
25
B Evaluation metrics
26 1
FAIR-Compute
NFCS Flexible Fund — Final Report
C Inferred QoS-tier proxy
26
D Federation policies and complete results
26
E Stakeholder engagement programme
27
F Software and reproducibility
27
Team and Acknowledgements
27
References
30
2
FAIR-Compute
NFCS Flexible Fund — Final Report
Executive Summary As demand for high-performance computing (HPC), high-throughput computing, and data storage grows, the way scarce compute is allocated — not just how much exists — has become a decisive factor in the productivity of UK research. FAIR-Compute studies allocation in a federated Digital Research Infrastructure (DRI) as a problem at the intersection of algorithmic game theory, transport economics, and HPC scheduling. Our central observation is simple: once a shared system must decide who runs, when, and under what evidence, its scheduler settings cease to be a purely technical matter and become policy. The project combined three strands of evidence: a landscape review of UK and international allocation practice (including EuroHPC, WLCG, ACCESS, JASMIN, and DiRAC) supported by a stakeholder survey; a mechanism-design model of allocation under strategic and uncertain user reports; and a simulation study built on the public Fresco/Anvil workload trace and controlled synthetic stress tests. Three results recur across all three strands. First, allocation records measure occupancy (resources reserved) rather than utilisation (useful work done), so the system cannot currently answer the question UKRI and DSIT most want answered — whether resources deliver productive value. Second, simple, transparent scheduling heuristics perform close to an offline full-information benchmark once tuned, so the near-term opportunity is to tune and instrument existing schedulers rather than replace them. Third, federation is genuinely valuable but behaves like uncoordinated (“selfish”) routing when left unmanaged: it can minimise average delay while quietly concentrating load — and potentially harm — on the receiving system. This mirrors a classic effect from road-traffic economics (Braess’s paradox), where letting everyone independently pick the fastest route can leave the whole network worse off — so coordinated placement rules are needed to contain it. From this evidence we derive eight recommendations for the NFCS roadmap, categorised Must / Should / Could Have and spanning a five-year horizon. The foundation is measurement: a unified data-collection standard (FAIRC-1) and an occupancy-versus-utilisation efficiency KPI (FAIRC-2), both Must-haves on which the others depend. These are followed by practical, low-risk governance improvements — publishing Slurm tuning guidance (FAIRC-4), a tiered rapid-access route paired with Community Centres of Excellence (FAIRC-7), cross-system benchmarking and exchange rates (FAIRC-6), and ex-post verification of requested-versusconsumed resources (FAIRC-8) — and two forward-looking pilots: value-aware prioritisation (FAIRC-3) and coordinated governance of federated job mobility (FAIRC-5). If the roadmap adopts one thing first, it should be measurement: almost everything else depends on data that does not yet exist in a consistent form.
1. Introduction and Approach The UK is moving from a set of independently governed HPC systems towards a federated National Compute ecosystem in which access, platform services, and infrastructure are increasingly disaggregated and coordinated centrally. This transition raises an allocation question that existing, single-institution scheduling heuristics were not designed to answer: how should a heterogeneous, multi-owner federation match a diverse supply of compute to a heterogeneous, strategically-behaving demand, fairly and efficiently, when the information needed to decide is imperfect and privately held? FAIR-Compute frames this as a mechanism-design problem rather than a purely engineering one. Allocation has two layers that must be considered together: the scheduler, which fixes ordering and placement (FIFO, backfilling, aging); and the policy mechanism, which shapes incentives and priorities (fairshare, quality-of-service tiers, walltime penalties, credits). Users
3
FAIR-Compute
NFCS Flexible Fund — Final Report
may misreport their needs, either strategically (to gain priority) or through genuine uncertainty about runtimes, and the resulting behaviour — most visibly walltime “padding” — is best treated as a permanent feature of the system to be modelled, not a bug to be wished away. The project was organised around three complementary activities, and the report follows their logic. Section 2 sets out our methodology and evidence base in detail — the stakeholder interviews and their follow-up, the demand-side survey, the literature review, and the design of the simulation study — so that the basis for every later claim is transparent. Section 3 presents the resulting landscape review: the UK and international allocation ecosystem, the economic evidence, and the operational frictions surfaced by our case studies. Section 4 develops the formal model of allocation under strategic and uncertain reporting that underpins the simulations. Section 5 reports the simulation study and its four experiments in full, with methods, results, and figures. Section 6 states the eight roadmap recommendations in the NFCS output format, each tied back to this evidence. Section 7 sets out the future-work agenda the project opens up, and Section 8 concludes. The appendices collect the technical detail — evaluation metrics, the workload proxy, the full federation policies and results, the engagement programme, and reproducibility notes — so that the main narrative stays readable while the underlying work remains fully documented.
2. Methodology and Evidence Base FAIR-Compute combined three strands of evidence, each chosen to compensate for the blind spots of the others. A landscape and literature review established what is known — from both academic mechanism-design theory and operational practice in the UK and abroad. Stakeholder engagement (semi-structured interviews and a demand-side survey) grounded that theory in how allocation actually works, and fails, on live national systems. And a simulation study let us run controlled, counterfactual experiments that no single operator could run on a production system. Triangulating across the three is what gives the recommendations in Section 6 their weight: each is supported by theory, by operator testimony, and, where possible, by simulation evidence.
2.1 Stakeholder interviews and follow-up The primary qualitative evidence comes from a series of semi-structured interviews conducted between February and June 2026 with operators, directors, and policymakers spanning the breadth of the UK ecosystem. In total we held in-depth sessions with seven organisations: QMUL ITS Research (operators of the Apocrita Tier-3 cluster); the STFC IRIS / IRIS-SCRC allocation team; the JASMIN data-intensive facility (NERC); the DiRAC distributed HPC service for theoretical physics and cosmology; the EPCC at Edinburgh; the UKRI Digital Research Infrastructure (DRI) programme; and the Department for Science, Innovation and Technology (DSIT) public-compute division, including a dedicated session on the AI Research Resource (AIRR). Several stakeholders were consulted more than once as our questions sharpened; Appendix E summarises the full engagement programme. Each interview followed a prepared agenda covering user behaviour and demand (estimation accuracy, walltime padding, strategic job-splitting), mechanism and tier design (preemption logic, pricing and cost-recovery, Slurm priority weights), and the transition to federation (identity mapping, cross-charging, data gravity). Crucially, interviews were treated as an iterative process rather than one-off data collection. After each session the findings were written up and mapped onto the formal model’s variables — for example, IRIS-SCRC’s penalty loop and DiRAC’s under-utilisation review thresholds were mapped to the historical usage state that
4
FAIR-Compute
NFCS Flexible Fund — Final Report
reappears in the model (Section 4) and in recommendation FAIRC-8; JASMIN’s storage-binding constraint motivated the dual-resource extension noted under FAIRC-2; and the recurring “occupancy versus productive use” theme became a through-line of the whole project. Where operators raised open questions — most importantly how to define “productive use” and how to share pseudonymised job logs — these were carried back into subsequent sessions and into formal data requests, several of which remain in progress and directly motivate FAIRC-1. Analysis was thematic rather than statistical, which suits a small number of deep expert interviews. Findings were synthesised into recurring themes (occupancy versus utilisation; the administrative bottleneck; batch versus interactive tension; contested fairness; the difficulty of measuring outcomes) and, importantly, cross-checked across independent sources: a claim was only treated as robust when it recurred across organisations with different incentives — for instance, the observation that heuristics “work fairly well but we don’t know how far from optimum” surfaced independently at an institutional cluster, a national credit-based service, and within DSIT. This triangulation is what licenses the report to generalise from a modest number of interviews. The data-request process. A specific, and instructive, strand of the engagement was the attempt to obtain real job-scheduling data to calibrate the simulations. We requested pseudonymised historical job logs (submission and actual runtimes, resource requests, queue and scheduling decisions) and, where possible, anonymised “reason-for-rejection” records, and we discussed acceptable sharing terms (pseudonymisation at source, removal of system-identifying labels). Despite constructive engagement, no UK dataset could be released on workable terms within the project window — a combination of privacy, administrative, and effort constraints that operators themselves flagged as a barrier. This experience is not a footnote but a finding: the absence of a routine, agreed data-sharing pathway is precisely the gap that recommendation FAIRC-1 is designed to close, and it is why the simulation study (Section 5) had to fall back on public US workload data.
2.2 Demand-side stakeholder survey To complement the operator-facing interviews with a demand-side view, we ran an exploratory stakeholder survey with a user-facing focus. The instrument comprised 16 questions covering resource-estimation behaviour, allocation and fairness perceptions, scheduling strategies, and federation challenges, and was administered through Microsoft Forms. Data collection ran over a two-month period from April to June 2026. The survey was distributed through a combination of conference QR codes (including at the UKRI NFCS Spring Conference in Bristol), LinkedIn posts, and direct professional networks, allowing respondents to participate electronically. A total of ten responses (N = 10) were collected. Respondents comprised system administrators, HPC facility managers, and research software engineers (n = 6), together with funder, policy, and institutional stakeholders (n = 4); several reported interacting with more than one class of infrastructure (national supercomputers, university clusters, and cloud HPC). Respondents are therefore best understood as people who support, manage, or represent users — and who observe user behaviour daily across many accounts — rather than as a representative sample of end users, and we interpret the responses accordingly. We note two further limitations transparently. First, the survey did not capture respondents’ Research Council affiliation (e.g. EPSRC, STFC, BBSRC), which limits any analysis by funding community. Second, the sample is small, so the findings should be read as exploratory and qualitative rather than statistically representative of the wider UK community. Responses were analysed thematically; the resulting themes — persistent walltime overestimation, strategic buffering, the absence of any 5
FAIR-Compute
NFCS Flexible Fund — Final Report
single agreed definition of fairness, and the predominantly organisational (rather than technical) nature of federation barriers — are reported in Section 3.3 and corroborate the interview evidence.
2.3 Literature review In parallel we conducted a structured review of the state of the art, organised into four areas that together frame the allocation problem: the engineering status quo (Slurm-style priority scheduling, fairshare, and backfilling); markets with money for heterogeneous data-centre resources; coordination mechanisms and the price of anarchy in congested, self-interested systems; and mechanism design without money, including artificial-currency and priority-based approaches. This review both positioned FAIR-Compute against existing work and supplied the theoretical scaffolding — congestion games, welfare benchmarks, and incentive compatibility — used in the model and simulations that follow.
2.4 Simulation study: rationale Finally, because no operator can ethically run counterfactual scheduling experiments on a live national system, and because pseudonymised UK job logs could not be obtained on workable terms within the project window, we built a simulation framework on public workload data. Its design, data handling, and four experiments are documented in full in Section 5; here we simply flag the guiding principle. We deliberately separate real-trace replay (which validates realism but is low-contention) from controlled synthetic stress tests (which create the contention under which allocation policy actually bites), and we label every result accordingly, so that the reader always knows whether a number reflects observed behaviour or a sandbox counterfactual.
3. The UK and International Research-Computing Landscape 3.1 Infrastructure tiers and federation models UK research computing is organised in tiers: international-scale facilities (Tier 0, e.g. IsambardAI, EuroHPC, PRACE); the national flagship service (Tier 1, ARCHER2, which ends in 2026 and is to be replaced by the new Edinburgh-based Next National Supercomputing Service, NNSS, from 2027); regional and specialist hubs (Tier 2, e.g. Bede, Baskerville, Cirrus); and institutional systems (Tier 3, e.g. QMUL’s Apocrita) [6]. Researchers typically progress up this ladder as their computational needs grow, but each tier is separately governed, which fragments access. Four organisational models coordinate such systems: independent (local control, but fragmented); grid (shared resources, but complex middleware, exemplified by the Worldwide LHC Computing Grid, WLCG [25]); cloud (elastic, but with cost and interconnect limits for tightlycoupled HPC); and federated (coordination with retained institutional autonomy). Federation is the UK’s chosen direction. International practice offers concrete mechanisms: the EuroHPC Federation Platform provides federated single sign-on (via the MyAccessID identity service) together with cross-system allocation and resource management through the Waldur platform [8, 24]; Open OnDemand lowers onboarding friction through a common web portal [13]; and the US ACCESS programme operates a credit-and-exchange model in which portable allocation credits are converted between heterogeneous resources using benchmark-derived exchange rates [1, 2]. The four models trade integration against autonomy in different ways (Table 1), and each leaves 6
FAIR-Compute
NFCS Flexible Fund — Final Report
a characteristic residue — fragmentation, middleware overhead, cost, or governance complexity. Federation is attractive precisely because it aims to keep local control while sharing the services (identity, allocation, software) that most benefit from coordination. Table 1: Organisational models for coordinating HPC systems. Model
Strength
Weakness
Independent Grid Cloud
Full local control Resource sharing across sites Elastic, on-demand provisioning
Federated
Coordination with retained autonomy
Fragmentation; multiple accounts Middleware complexity; steep workflows Cost and interconnect limits for tightlycoupled HPC Governance and cross-charging complexity
Recent UKRI proposals point to a disaggregated stack — access, platform services, and compute as independent layers behind one “single pane of glass” portal — coordinated by Community Centres of Excellence (CCEs) that broker allocations across federated systems and provide training and research-software-engineering (RSE) support, alongside tiered access with rapid “Seedcorn / Bright Idea” routes [7, 23].
3.2 Occupancy, utilisation, and the economic case A recurring theme across every operational interview is the gap between occupancy — resources reserved or blocked for a job — and utilisation — resources actually performing useful computation. A system can show high occupancy and low utilisation (reserved-but-idle capacity), which is pure opportunity cost. The natural efficiency measure is therefore utilisation divided by occupancy, yet current accounting almost universally charges for occupancy, which rewards over-requesting and walltime padding. Published economic evaluations show that HPC investment delivers strong returns, but they are scarce and methodologically inconsistent (Table 2). EPSRC’s evaluation of its federated Tier-1/Tier-2 programme reports a benefit–cost ratio (BCR) of 6.5–19.5 [6]. The Met Office programme reports a 9:1 cost–benefit ratio against the do-nothing option, based on a net present social value of £13.74bn (with a 25% optimism bias) on a £1.2bn investment; because that ratio is not a simple benefit-to-investment quotient, it is not directly comparable with the EPSRC figure [22]. Industry estimates warrant particular caution: for NVIDIA’s privately-funded Cambridge-1, Frontier Economics estimated an economic value of roughly £600M (∼$831M) over ten years on a $100M investment [10], and other cited figures vary. Even at face value, these numbers do not establish that federation yields higher returns than independent provision — only that HPC broadly pays off and that comparable, standardised ROI evidence is missing. That measurement gap is itself an argument for recommendations FAIRC-1 and FAIRC-2.
3.3 Where allocation breaks down Our interviews and case studies surface a consistent set of structural frictions that any federated allocation mechanism must confront. The administrative bottleneck. Across facilities, manual, peer-review-based allocation — not hardware — is the binding scaling constraint. IRIS-SCRC and DiRAC both operate multimonth allocation cycles: a call opens, a technical team reviews requested hardware, a committee decides, and letters issue months later. Reviews are labour-intensive — operators described allocation reviews consuming on the order of 60–100 person-hours, and DiRAC assessing roughly 7
FAIR-Compute
NFCS Flexible Fund — Final Report
Table 2: Published benefit–cost ratios for public HPC infrastructure. The ratios derive from different evaluation methodologies and should not be read as a direct comparison between infrastructure models (see text); the purpose of the table is to show that published studies consistently find substantial economic value from HPC investment across different contexts. We report the headline ratios and their basis rather than derived investment/benefit figures, which vary by source. NVIDIA Cambridge-1 is discussed in the text as an industry estimate. Case study
BCR
EPSRC Tier-1 & Tier-2 HPC (federated)
6.5–19.5 : 1
Met Office supercomputing (independent)
9:1
Basis and source Ten-year return on EPSRC’s HPC investment; independent evaluation by London Economics for EPSRC [6]. Ratio against the do-nothing option; net present social value £13.74bn (25% optimism bias) on a £1.2bn investment [22].
67 applications per round under a “double jeopardy” process that judges scientific merit and computational justification independently. As AI and interdisciplinary demand grows, this human review becomes the scaling limit, producing slow cycles, high overhead, and limited responsiveness. This is the direct motivation for a tiered rapid-access route (FAIRC-7). Occupancy versus productive use. Every operator distinguished, in their own terms, between blocking a resource and using it. Standard accounting charges for occupancy (core-hours reserved) rather than utilisation (work delivered), which rewards conservative over-requesting and walltime padding. Worse, allocation units correlate poorly with scientific output: as one director put it, the amount of science obtained from a given allocation is “poorly correlated with the allocation units being used,” and knowing the relationship in advance requires prior domain knowledge that most users lack. Interactive and software-defined environments sharpen the problem, because they must stay “on” to preserve state and therefore block resources even when idle. Batch versus interactive, and the technology shift. Legacy batch processing guarantees fair access and system stability but is poorly suited to code development, cluster-optimisation research, and cybersecurity work. Software-defined environments (containers, VMs) let users treat a cluster as a “larger laptop” through a single pane of glass, lowering the barrier to entry — but hardware is currently partitioned statically between batch and interactive modes, and making that boundary dynamic is itself an unsolved allocation problem. Contested fairness. There is no single agreed definition of “fair.” Stakeholders variously invoke scientific merit, proportionality to funding or contribution, equal access, and alignment with national strategic priorities. Fairness is therefore a multi-dimensional, contested objective, not a single quantity to optimise — a point that recurs in the survey findings below and shapes the model in Section 4. Contrasting facility models. The frictions manifest differently depending on the binding constraint, as our two most detailed case studies show (Table 3). JASMIN (NERC) is a dataintensive facility of roughly 55,000 CPU cores and 60–80 PB of tiered storage, where storage occupancy, not compute, is the scarce resource; compute on its Lotus cluster runs under a lighttouch FairShare model per consortium, and spare capacity is exposed to the wider federation as a residual claimant. DiRAC, serving theoretical physics, astronomy, and cosmology across four 8
FAIR-Compute
NFCS Flexible Fund — Final Report
hardware-software co-designed sites, is compute-bound: it allocates CPU/GPU-hour credits through peer review, expires them quarterly on a “use-it-or-lose-it” basis, links priority to historical use, and triggers follow-up review below 50% (small) or 80% (large) utilisation. These are not competing designs but rational responses to different scarcities — which is precisely why a federation needs mechanisms flexible enough to span both. Table 3: Two contrasting national facility models, from our operator interviews. The binding constraint drives the allocation mechanism. JASMIN (data-intensive)
DiRAC (compute-intensive)
Primary constraint Compute allocation Temporal policy Behavioural lever
Storage occupancy (60–80 PB) Light-touch FairShare per consortium Historical-need consortium quotas Minimal; “no monopoly” principle
Federation posture
Residual claimant (shares spare compute)
Compute (CPU/GPU-hours) Credit-based, peer-reviewed Quarterly “use-it-or-lose-it” credits Under-use review (50%/80% thresholds) Specialised, co-designed per site
3.4 Demand-side findings The demand-side stakeholder survey (Section 2; N = 10) — answered largely by those who support and represent users rather than by end users themselves — corroborates the operator picture from the other side of the interface. Three themes stand out. On resource estimation, respondents reported systematic walltime overestimation with inconsistent accuracy, confirming a persistent information asymmetry: users know more about their workloads than the scheduler does, and they hedge that uncertainty through conservative buffering and backfill-oriented submission. On fairness, responses revealed no consensus definition — merit, funding-proportionality, equal access, and national-priority alignment were all invoked, reinforcing that fairness is contested rather than given. On federation, the barriers respondents raised were as much organisational (governance, cross-charging, data movement) as technical. These demand-side frictions fall hardest on particular groups. Early-career researchers without an established PI or local RSE support, and non-traditional HPC disciplines (biology, medicine, the social sciences) that increasingly need large-scale compute but lack established workflows, face the steepest structural barriers: complex onboarding, multiple accounts and proposals, command-line and batch-scheduler expertise, and fragmented identity and software stacks. The consequence is both underused national infrastructure and higher barriers to entry for exactly the interdisciplinary communities the federation is meant to broaden — the demand-side case for CCEs and a rapid-access tier (FAIRC-7).
3.5 International comparison and lessons Other national and international systems offer concrete design lessons. The Worldwide LHC Computing Grid demonstrates cost-sharing by pledge and, critically, a maturing benchmarking discipline: its move from the HS06 benchmark to HEPScore — containerised real experiment workloads, reproducible to better than 1% across x86 and ARM — is the template for the exchange-rate framework we recommend (FAIRC-6) [11]. The US ACCESS programme shows that portable, benchmark-derived credits can let researchers move computational value across heterogeneous machines without re-applying, decoupling allocation approval from machine selection. Further afield, China and Japan illustrate the two poles of the curiosity-versus-mission tension: China’s centralised, state-led model prioritises strategic mission workloads, whereas 9
FAIR-Compute
NFCS Flexible Fund — Final Report
Japan’s federated HPCI balances mission and curiosity-driven access through a peer-reviewed, quota-plus-priority system [4, 12, 19]. The UK’s disaggregated, CCE-coordinated direction is a deliberate attempt to occupy the middle ground — integration with retained autonomy — and the mechanisms in this report are aimed squarely at making that middle ground work.
3.6 Sustainability and carbon Allocation efficiency is also an energy and carbon question. Idle-but-reserved capacity (high occupancy, low utilisation) draws power without producing science, and operating costs are dominated by energy: DiRAC, for instance, reports roughly £2M of a £5.5M annual operating budget as electricity. Raising realised utilisation per unit of reserved capacity — the KPI proposed in FAIRC-2 — therefore lowers energy and carbon per unit of research output, and curbing walltime over-reservation (FAIRC-8) has the same effect. We do not quantify carbon in tonnes of CO2 equivalent (tCO2 e) in this study, which is a limitation; we recommend that the efficiency KPI be reported alongside energy consumption so that the carbon impact of allocation policy can be tracked consistently across the federation.
4. A Model of HPC Allocation This section sets out the formal model that gives the project its analytical spine. It has three parts: a description of the system and how a production scheduler actually allocates resources; a decision-theoretic account of how strategic, uncertain users choose what to report; and an offline welfare benchmark against which practical rules can be measured. The simulations in Section 5 implement the descriptive scheduler and the offline benchmark directly; the strategic-utility analysis (Section 4.2) frames the mechanism-design questions that motivate recommendations FAIRC-3 and FAIRC-8 and is, in part, a direction for future work. Throughout, the guiding stance — shared with our operator interviewees — is that strategic behaviour such as walltime padding is a permanent feature of the system to be modelled, not a defect to be assumed away.
4.1 System, jobs, and the descriptive scheduler A cluster provides a per-time-unit capacity L = (L1 , . . . , Ld ) across d resource types (CPU cores, GPUs, memory, . . . ). Each job j is a tuple βj = (aj , rj , lj′ ) of arrival time aj , resource requirement rj ∈ Nd , and reported walltime lj′ ; its actual runtime lj is realised only during execution, and a job is terminated if lj > lj′ . An allocation assigns a start time sj ≥ aj ; P completion is cj = sj + lj ; feasibility requires j xj (t) ≤ L for all t. Production schedulers such as Slurm [20] do not solve an explicit welfare objective; they rank eligible jobs by a weighted priority score Priorityj (t) = wage Agej (t) + wfs Fairi(j) (t) + wqos QoSj + wsize Sizej + wpart Partj + · · · , (1) where Agej (t) = t − aj and the fairshare term decreases with a user’s recent usage through a public history state Hi (t) that decays exponentially [26]. The non-negative weights w· are set by administrators, not users. Among eligible jobs E(t) = {j waiting : rj ≤ Lleft (t)} the highest-priority job is scheduled first, with smaller jobs backfilled when they do not delay it.
4.2 Strategic reports and user utility Users are typically uncertain about runtimes at submission, knowing only a distribution lj ∼ Fj over possible execution times. The reported walltime lj′ therefore functions as self-selected 10
FAIR-Compute
NFCS Flexible Fund — Final Report
insurance: requesting a larger walltime reduces the probability that the job is killed before completion, but typically increases queue delay (a larger reservation is harder to place and to backfill) and may expose the user to ex-post penalties. The value a user derives from a completed job degrades with delay. Writing the waiting time Wj = sj −aj , we model completion value as vj (Wj ) = Vj fj (Wj ) for a non-increasing urgency function fj ; the linear special case vj (Wj ) = Vj − κj Wj captures a constant marginal disutility of waiting κj (we reserve cj for completion times). Crucially, Wj itself depends on lj′ through the scheduler, so a larger safety buffer carries an implicit upfront cost. Adding a mechanism-imposed ex-post penalty hj (lj′ , lj , Hi(j) ) that depends on the reporting gap and on the user’s public history, expected utility is h
i
E[Uj (lj′ )] = E Vj − κj Wj − hj (lj′ , lj , Hi(j) ) ,
(2)
so the reported walltime is a genuine strategic choice, trading termination risk against queue delay and penalty exposure. The specific termination rules (kill-at-walltime versus preemptand-requeue) and penalty forms — an overestimation charge, or a fairshare adjustment that routes the cost through the history state Hi(j) , as IRIS-SCRC and DiRAC do in practice — are set out in Appendix A. The strategic-reporting mechanism itself is not exercised in the present simulations; Experiment 3 empirically illustrates the related system-wide externalities that arise from runtime information reporting, and the potential loss of scientific value under value-blind scheduling. This part of the model is the mechanism-design foundation for recommendation FAIRC-8 and for the research programme in Section 7.
4.3 Offline welfare benchmark To measure how far practical rules sit from an ideal, we use an offline benchmark that schedules with full information. With binary variables yj,t (job j runs at t) and zj,t (job j completes at P P t), subject to arrival (yj,t = 0 for t < aj ), capacity ( j rj,k yj,t ≤ Lk ), completion ( tτ =aj yj,τ ≥ P lj zj,t ), and single-completion ( t zj,t = 1) constraints, the benchmark maximises total value net of waiting cost, max
X
Vj − κj Wj ,
Wj =
X
(t − aj − lj ) zj,t .
(3)
t
j∈J
This weighted-flow-time / value objective is the reference against which FIFO, Slurm-like, and tuned heuristics are compared in Section 5. One practical caveat applies to how the benchmark should be read. Although it is offline and full-information by construction, solving it exactly on the workload windows used for comparison is computationally demanding, and the solver did not always certify global optimality, instead terminating with a small but non-zero optimality gap. The offline results reported in Section 5 should therefore be understood as the best feasible solution found within the available computational limits — with an objective value close to the solver’s best bound — rather than as a proven global optimum. This distinction also helps explain why a heuristic can occasionally appear to match or outperform the offline benchmark on individual reported metrics.
5. Simulation Study 5.1 Data, synthetic generation, and framework Data. No internal UK HPC trace was available on workable terms within the project window, so the study uses the public Fresco dataset [9, 16] — roughly 20.9 million jobs across 75 months 11
FAIR-Compute
NFCS Flexible Fund — Final Report
from three US academic systems — focusing on Purdue’s Anvil cluster [21], the newest and, crucially, the only Slurm-based system, over June 2022–May 2023. Anvil records exactly what a scheduling simulation needs: submit, start and end times, requested resources and walltime, queue/partition, account identifiers, and completion status. Because Fresco is distributed as monthly files, account-level history can be discontinuous at month boundaries; we therefore treat each month as a self-contained trace for descriptive and low-pressure analysis, and carry state explicitly where history matters. Since the Fresco trace does not expose the true ACCESS project type, a per-account QoS proxy is inferred from annualised usage and mapped to the four ACCESS-style tiers (Explore, Discover, Accelerate, Maximize), giving the simulator realistic account-scale heterogeneity (Appendix C). Descriptively, the Anvil workload is bursty in time and highly skewed across accounts (Figure 1): demand concentrates in working hours, and a small number of accounts submit a disproportionate share of jobs — both features that any realistic synthetic workload must preserve.
(a) Arrivals per hour of day
(b) Jobs per account (top 30)
Figure 1: Descriptive statistics of the Anvil workload: bursty temporal demand (left) and a heavy-tailed distribution of jobs across accounts (right).
Synthetic workload construction. The selected Anvil months are low-contention relative to Anvil-scale capacity, so replaying them directly is useful for realism but cannot stress a scheduler. We therefore learn the workload’s structure and regenerate controlled synthetic traces. Jobs are first grouped by clustering on scheduling-relevant features (requested cores, hosts, runtime, partition, QoS), with cluster quality checked by the elbow/inertia and silhouette diagnostics (Figure 2); a surrogate classifier is then trained to reproduce the cluster labels, and its SHAP attributions confirm the groups are driven by genuine scheduling features — the shared and whole-node partitions, requested cores, runtime, and host count — rather than artefacts (Figure 3) [15]. Synthetic jobs are sampled from this learned structure and given arrival times under three regimes: an even-fill regime (denser, regular demand), a bursty-real regime (preserving the real burst structure), and a bursty-fill regime that adds jobs into quiet periods for stress tests, with volume multipliers to tune pressure. The framework supports all three, but no results using the even-fill regime are reported here: the contention experiments below use the bursty-real and bursty-fill traces. The synthetic traces thus preserve realistic job structure while letting us control the one thing the real months lack — contention. The framework replays the same workload under four rule families — FIFO (transparent baseline), Slurm-like weighted priority (operational heuristic), multi-objective-tuned weights, and an offline ILP benchmark — and, in Experiment 4, under federation policies. Outcomes are scored on system efficiency (utilisation, throughput, makespan), user-facing delay (waiting time, turnaround, slowdown), and fairness (account/job-size slowdown gap and Jain indices). Requested walltime is treated as both a scheduling input and a behavioural signal, isolated by replaying under actual-runtime versus requested-walltime occupation modes. We address four research questions: how weights map to objectives (RQ1); how far heuristics sit from the offline benchmark (RQ2); how strategic reporting creates externalities (RQ3); and how federation balances global gains against local-user protection (RQ4).
12
FAIR-Compute
NFCS Flexible Fund — Final Report
(a) Elbow/inertia
(b) Cluster statistics
Figure 2: Clustering diagnostics used to choose and validate the number of workload groups before synthetic generation.
Figure 3: SHAP feature importance for the surrogate workload classifier. The learned clusters are driven by scheduling-relevant features (partition, cores, runtime, hosts), confirming the synthetic traces preserve realistic structure.
13
FAIR-Compute
NFCS Flexible Fund — Final Report
Scheduling rules and occupation modes. The rule families compared span the full spectrum from transparent to optimal (Table 4). Because requested walltime is simultaneously a scheduling input and a behavioural signal, we replay each workload under three occupation modes to separate the two effects: (i) actual runtime used for both priority and resource holding (the best-information case); (ii) requested walltime for priority, actual runtime for occupation (isolating the ordering distortion); and (iii) requested walltime for both (the pessimistic blocking case that also captures over-reservation). Comparing modes lets us attribute outcomes to queue re-ordering versus resource over-holding. Table 4: Scheduling rules compared across the simulation study. Rule
Role
FIFO Slurm-like priority Tuned priority (pymoo/NSGA-II) Offline ILP (PyJobShop) Federation broker
Transparent baseline; easy to explain, optimises nothing. Operationally realistic weighted score (age, fairshare, size, QoS). Weights searched on the efficiency–fairness trade-off surface. Full-information welfare benchmark (weighted flow time / value). Cross-system placement policies (Experiment 4).
Evaluation metrics. Outcomes are scored on three axes. System efficiency: platform utilisation, throughput, and makespan. User-facing delay: waiting time (submission to start), turnaround (submission to completion), and slowdown (turnaround divided by runtime, which normalises delay by job length). Fairness: the account-level and job-size slowdown gap (max−min), and Jain’s index applied both to allocated CPU service (a desirable quantity) and to delay (a cost). A good schedule combines low average slowdown with a high slowdown-Jain index; a high Jain index alone can simply mean everyone is served equally badly, which is why we always report fairness alongside the average level of delay. Formal definitions of every metric are collected in Appendix B.
5.2 Experiment 1 — Priority weights as policy levers Treating the Slurm weights over age, fairshare, job size, and QoS as decision variables, we search the trade-off surface with the NSGA-II genetic algorithm via pymoo [3, 5, 18], scoring each candidate on utilisation, waiting time, slowdown, and account-level slowdown gap. On the low-contention real trace, weights mainly redistribute delay; under synthetic contention they become genuine policy levers. The search returns a Pareto set of non-dominated policies (Table 5): no single configuration is best on every objective — one maximises utilisation and fairness, another minimises waiting time, another slowdown, with a balanced compromise in between. The operational implication is that choosing weights is setting policy, and there is no universally correct setting. Table 5: Representative Pareto-optimal weight policies on the synthetic filled workload (Experiment 1). No policy dominates on all objectives. Policy P1 (utilisation / fairness) P5 (waiting time) P6 (balanced) P8 (slowdown)
Utilisation
Avg wait (min)
Avg slowdown
Account gap
0.874 0.871 0.869 0.865
379.5 342.2 358.0 354.8
13.32 8.57 8.22 8.19
552.1 565.0 561.5 561.6
To expose why multi-objective tuning matters, we also ran four single-objective searches, each 14
FAIR-Compute
NFCS Flexible Fund — Final Report
targeting only one metric (Table 6). A reading note: because the evolutionary search evaluates a finite set of candidate weight configurations, each row reports the best configuration found for its target, not a guaranteed optimum of that metric — which is why the slowdownfocused run happened to land on a configuration with slightly higher utilisation (0.919) than the utilisation-focused run (0.914). The qualitative pattern is nonetheless clear: the utilisationfocused configuration incurs a visible cost in waiting time, slowdown, and account fairness, while the delay- and fairness-oriented configurations sacrifice very little utilisation while improving user-facing outcomes. In the multi-objective run, extreme policies of this kind do not survive as Pareto solutions because waiting time and slowdown are protected simultaneously. The practical lesson is direct: optimising one scheduling goal in isolation quietly damages the others, so operators need explicit multi-objective guidance rather than a single default weight vector. Table 6: Single-objective priority-weight tuning on the synthetic filled workload (Experiment 1). Each row reports the best configuration the search found for one target metric; note the collateral damage to the others. Objective optimised
Utilisation
Avg wait (min)
Avg slowdown
Account gap
Maximise utilisation Minimise waiting time Minimise slowdown Minimise account gap
0.914 0.909 0.919 0.911
1058.5 965.6 961.0 960.4
70.00 62.17 62.13 61.83
833.0 768.2 765.8 764.2
5.3 Experiment 2 — Practical rules versus the offline benchmark We replay a high-pressure workload under FIFO, Slurm-like size-priority, a tuned policy, and an offline ILP (minimising weighted flow time) implemented with PyJobShop [14, 17]. FIFO is clearly worst. The tuned and size-priority heuristics actually match or slightly beat the offline ILP on average waiting time — unsurprising, since the ILP optimises weighted flow time rather than raw wait, and is itself solved only to a small optimality gap (Section 4.3) — while the ILP retains the lowest average slowdown; the heuristics improve sharply on FIFO in delay and fairness — at a slight cost in raw utilisation — without dominating the full-information benchmark overall (Table 7). The message for operators is that transparent, Slurm-compatible heuristics are “good enough” once tuned: the near-term priority is to tune existing schedulers, not to replace them. Table 7: Practical rules versus the offline ILP benchmark on the synthetic filled workload (Experiment 2). Rule Offline ILP (best found) FIFO Size-priority Tuned (P5)
Utilisation
Avg wait (min)
Avg slowdown
Account gap
0.766 0.755 0.737 0.737
83.5 112.7 81.1 82.8
1.76 4.11 1.86 2.05
17.4 102.1 17.4 24.1
It is equally instructive to run the same comparison on the real Anvil trace, which is lowcontention: there, platform utilisation is pinned at roughly 0.25 for every rule, and ILP, FIFO and the priority variants become nearly indistinguishable on delay and fairness (Table 8). When there are rarely enough queued jobs to compete, the scheduler has little to decide, so priority weights mostly redistribute small amounts of delay rather than changing aggregate outcomes. This is not a null result but a scoping result: it demonstrates that scheduling policy only bites 15
FAIR-Compute
NFCS Flexible Fund — Final Report
under contention, which is exactly why the controlled synthetic stress tests — not the sparse real months — are the right instrument for evaluating allocation mechanisms, and why a highcontention real UK trace (FAIRC-1) would be so valuable. The fairness picture (Figure 4) tells the same story: under load the distributional differences between rules become visible, with FIFO and a QoS-heavy rule producing the largest account-level imbalance and size-priority the most even service. Table 8: Practical rules on the low-contention real Anvil trace (Experiment 2). Utilisation is capped by available demand (≈0.25), so the rules barely differ — scheduling policy only bites under contention. Rule Offline ILP (best found) FIFO Size-priority Fairshare-priority QoS-priority
Utilisation
Avg wait (min)
Avg slowdown
Account gap
0.252 0.252 0.252 0.252 0.252
2.26 4.41 2.27 3.14 5.32
1.11 1.53 1.11 1.36 1.78
6.8 30.3 6.8 25.6 50.1
(a) Real trace (low contention)
(b) Synthetic filled (high contention)
Figure 4: Account-level fairness (Jain index) across scheduling rules. Under low contention (left) the rules are nearly indistinguishable; under contention (right) the distributional consequences of policy choice emerge.
5.4 Experiment 3 — Strategic reporting and scheduling externalities Two sub-experiments probe private information. In the runtime-perturbation study, reducing the runtimes of 20% of jobs improves aggregate delay (weighted flow time falls from 3.85M to 3.42M; average wait from 222 to 195 minutes), but Figure 5 shows the improvement is not uniform: many unchanged jobs (blue) also shift because shorter occupation releases capacity and the scheduler re-ranks the queue. Runtime information thus creates a queue externality — one user’s report changes other users’ waiting times. The aggregate effect is summarised in Table 9: the fixed-batch schedule, which keeps the original job order but uses the shorter runtimes, improves more than the dynamic scheduler that re-ranks the queue, showing that some of the capacity released by shorter jobs is dissipated by the reordering itself, and the offline benchmark reveals the remaining gap to full-information scheduling. In the deadline-and-value study, each job carries a soft deadline and a scientific value; the offline benchmark preserves all 2388 value units and completes all 500 jobs on time, whereas the 16
FAIR-Compute
NFCS Flexible Fund — Final Report
Table 9: Aggregate effect of reducing the runtimes of 20% of jobs (Experiment 3). Lower is better; “global” rows are the offline full-information benchmark. Schedule Baseline, dynamic layered Baseline, offline global Perturbed, dynamic layered Perturbed, fixed-batch Perturbed, offline global
Weighted flow time
Avg wait (min)
3,853,622 2,926,142 3,416,793 3,264,625 2,513,687
222.0 89.5 194.6 172.8 65.6
Figure 5: Job-level waiting times under runtime perturbation. Red points are jobs whose runtimes were reduced; blue points are unchanged. Blue points moving off the diagonal show that unchanged jobs are delayed or advanced by queue re-ranking — a system-wide externality. online rules recover only 1552–1612 units and finish around 340 jobs on time — despite having lower average waiting time and higher utilisation (Table 10, Figure 6). Operational metrics can therefore be misleading: a schedule may look efficient while destroying a large fraction of scientific value. This motivates value-aware prioritisation — with the caveat that if value or urgency is self-reported, the mechanism must be incentive-compatible. Table 10: Deadline-value scheduling under the full-deadline value metric (Experiment 3). Online rules look efficient on wait/utilisation yet lose substantial value. Scheduler Offline PyJobShop (best found) FIFO FIFO + deadline priority FIFO + dynamic value priority
On-time jobs
Avg wait (min)
Avg slowdown
Utilisation
500 341 340 339
1055.3 557.4 563.3 562.4
1.61 47.04 46.72 46.33
0.614 0.824 0.819 0.820
5.5 Experiment 4 — Federation as routing under congestion Two autonomous systems A and B each run a local FIFO scheduler; a broker decides at submission whether a job stays home or moves, under policies of increasing coordination: no federation, naive (go to the lowest estimated completion time), switching-cost (a delay “toll” on remote execution), minimum-gain (move only if the saving exceeds a threshold), and utilisation-protected (move only if the receiver has spare capacity). The natural lens is a congestion-game analogy: naive placement resembles unpriced selfish routing, in which adding capacity to an uncoordinated network can worsen outcomes (Braess’s paradox). By analogy, uncoordinated placement may concentrate load on receiving systems and create risks for their local users. Appendix D gives the full policy definitions and the complete results across all four scenarios; the main text reports selected rows. 17
FAIR-Compute
NFCS Flexible Fund — Final Report
Figure 6: Achieved scientific value by scheduler. The offline benchmark preserves the full available value; value-blind online rules recover only around two-thirds. Federation helps most under asymmetric load (Table 11): naive routing cuts average wait from 127.8 to 4.0 minutes. But naive is best on aggregate delay precisely because it is selfish — it consumes shared spare capacity aggressively and concentrates load on the receiving system. The switching-cost, minimum-gain, and utilisation-protected rules trade a little global delay to filter weak transfers and steer movement into the genuine relief direction (B → A), protecting the receiving system’s local users. For example, utilisation protection raises average wait from 4.0 to 9.0 minutes but cuts A → B transfers from 5,974 to 1,748 while improving slowdown from 1.39 to 1.18. The finding is deliberately honest: on aggregate metrics naive currently looks best. The federation results are consistent with the selfish-routing analogy — uncoordinated placement reduces aggregate delay while generating large, sometimes bidirectional, transfer volumes that may concentrate load on receiving systems — but the present simulations do not directly quantify harm to local users; they provide directional evidence for the need to examine congestion and local-service risks. The case for coordination therefore rests on protecting local service and stability as the federation scales, and quantifying local harm precisely is part of the remaining work. Table 11: Asymmetric-load federation (system A idle, B busy). Naive minimises delay but moves work indiscriminately; coordinated rules protect the receiver. Policy No federation Naive Switching cost Utilisation-protected
Avg wait (min)
Avg slowdown
A→B
B →A
127.8 4.0 18.9 9.0
13.33 1.39 2.17 1.18
0 5,974 739 1,748
0 13,402 22,466 8,751
Federation still helps under balanced high load, because two systems are rarely congested in lockstep (Table 12); one temporarily has usable capacity while the other spikes. Here naive again gives the lowest average wait, but its unpriced character is clearer: it moves large volumes in both directions (A → B and B → A), whereas switching cost filters weak transfers and lets the direction of net relief emerge. In the heterogeneous case, where system B carries the GPUs, naive over-uses the specialised resource; the toll relieves that pressure and lets CPU-suitable work flow back to A. Finally, the minimum-gain and utilisation-protected families (Table 13) show that protection is a dial, not a binary choice between isolation and a free-for-all. Minimum-gain removes lowbenefit transfers, keeping most of the delay benefit while sharply cutting cross-system traffic; adding a switching cost makes movement more directional, concentrating it in the genuine relief direction rather than simply reducing it. The utilisation-protected policy is designed to shield a receiving system’s own users, limiting transfers when the receiver is under pressure at the cost 18
FAIR-Compute
NFCS Flexible Fund — Final Report
Table 12: Balanced high-load federation: both systems under comparable pressure. Naive minimises wait but moves work in both directions; switching cost concentrates movement. Avg wait
Avg slow.
A→B
B →A
No federation Naive Switching cost
133.8 51.3 71.6
16.50 8.88 11.81
0 11,770 7,179
0 13,497 14,692
No federation Naive Switching cost
20.0 1.3 9.0
3.83 1.21 1.28
0 14,502 7,633
0 9,062 12,073
Scenario
Policy
Homogeneous
Heterogeneous
of a little more delay; the relative effect of the two families varies across scenarios (minimumgain yields fewer transfers in some cases, as Table 13 shows), so differences in transfer volumes should be read as directional rather than as a fixed ordering. Figure 7 visualises the effect — switching cost suppresses weak A → B transfers while allowing strong B → A relief around B’s demand spike. Table 13: Conservative federation rules (heterogeneous scenarios). Coordination trades a little delay for markedly less — and more directional — cross-system movement, protecting receiving systems. Avg wait
Avg slow.
A→B
B →A
Naive Minimum-gain Utilisation-protected
0.2 0.3 0.5
1.00 1.00 1.00
6,822 2,801 3,051
9,108 7,080 7,682
Naive Minimum-gain Utilisation-protected
1.3 1.0 3.8
1.21 1.16 1.50
14,502 7,514 7,454
9,062 5,125 5,075
Scenario
Policy
Low A / high B
Balanced
5.6 Summary of findings The simulations support four roadmap-relevant conclusions: (i) scheduler weights are policy levers with genuine trade-offs, so operators need guidance, not a single default (FAIRC-4); (ii) tuned transparent heuristics approach the offline benchmark, so tune before replacing (FAIRC4); (iii) private runtime and value information create system-wide externalities that only proper measurement can expose (FAIRC-8, FAIRC-1/2), and, under a synthetic value model, valueblind scheduling can leave a large fraction of scientific value unrealised (FAIRC-3); and (iv) federation is valuable but must be coordinated to avoid congestion-style harm (FAIRC-5), which in turn needs fair cross-system exchange rates (FAIRC-6). All quantitative results on contention derive from synthetic stress tests (a clearly-labelled sandbox), because a highcontention UK trace was not available — the strongest single argument for FAIRC-1. Configuration and gap analysis. The federation study abstracts to two autonomous systems (each 300 nodes × 128 CPU cores, with one GPU per node on the heterogeneous system B; 30-day synthetic traces), using a switching-cost penalty as a stand-in for network transfer cost and not modelling inter-site data movement or authentication. Experiment 1 and the runtimeperturbation study use a representative 7,000-job high-pressure window (simulated platform: 300 nodes × 128 cores, i.e. 38,400 CPU cores), while the deadline-value study uses a 500-job window to keep the offline optimisation tractable. All reported simulation figures are point
19
FAIR-Compute
NFCS Flexible Fund — Final Report
Figure 7: Cross-system transfers over time under naive versus switching-cost routing (asymmetric load). The switching cost suppresses weak A → B movement and concentrates B → A relief around system B’s demand spike — behavioural evidence of route regulation rather than mere volume reduction. estimates from these representative windows rather than averages over multiple random seeds, and should be read as indicative magnitudes rather than precise expected values. The principal gap is the absence of a high-contention real UK trace: the contention results are therefore counterfactual sandbox experiments, and validating them on pseudonymised UK job logs — enabled by FAIRC-1 — is the necessary next step before any operational adoption. Implementation and reproducibility notes for the whole framework are collected in Appendix F.
6. Recommendations for the NFCS Roadmap The eight recommendations below follow the NFCS output guidance: each carries the FAIRC short code, a Must/Should/Could-Have category, recommendation text with rationale, an indicative timescale within the five-year roadmap, and a resource estimate. FTE and cost figures are first-pass planning estimates for discussion, drawing on operational benchmarks gathered in interviews, and should be confirmed before adoption. They are ordered by suggested implementation sequence. FAIRC-1 — Unified data-collection standard
Must Have
Recommendation. Establish a common, mandatory data-collection and reporting standard for all federated systems, capturing per job a minimal agreed schema: requested versus consumed resources (cores, memory, GPU, walltime), queue/partition and QoS, account and Research-Council attribution, completion status, and — where a request is refused — an anonymised reason-for-rejection. Reporting should be automated, pseudonymised at source, and aggregated through the planned single-pane-of-glass access layer. Rationale. Every other recommendation, and any future allocation mechanism, depends on data that does not currently exist in a consistent form. Across our interviews the same gap recurred: operators “don’t measure enough or anything now,” outputs are weakly tracked, and monitoring quality varies widely. Our own simulation work was forced onto public US data precisely because UK job logs could not be obtained on workable terms. Without a shared schema the federation cannot answer whether a resource is used for the
20
FAIR-Compute
NFCS Flexible Fund — Final Report
purpose it was granted, and productively. This is the foundational, lowest-regret investment, and underpins interoperability across heterogeneous National Compute Resources. Timescale. Year 1, 12 months. Foundational — a dependency for FAIRC-2, 3, 5 and 8. Pilot the schema on 2–3 willing systems in the first 6 months, then extend. Resources. ∼1.5 FTE (1.0 Research Software Engineer; 0.5 data/policy analyst); ∼£140,000. Aligns with the platform-services / accounting layer of the disaggregated stack. FAIRC-2 — Occupancy-vs-utilisation KPI
Must Have
Recommendation. Adopt a standard efficiency metric across federated systems, defined as realised utilisation divided by reserved occupancy (e.g. delivered FLOP or core-hours against reserved core-hours), reported alongside existing occupancy-based accounting and paired with at least one outcome indicator so that “productive use” is measured rather than assumed. Rationale. Current accounting charges for occupancy, not utilisation; a system can show high occupancy and low utilisation (reserved-but-idle), which is pure opportunity cost. DSIT and UKRI repeatedly framed their open question as defining and measuring “productive use” and ROI; DiRAC already treats 85–90% utilisation as an economic necessity. A standard occupancy–utilisation signal gives the federation a comparable basis for valuefor-money judgements and exposes the over-reservation that walltime padding produces. JASMIN shows storage occupancy can itself be the binding constraint, so metrics should extend to storage efficiency where appropriate. Timescale. Year 1–2, 12 months. Depends on FAIRC-1 telemetry. Storage-bound systems should report a parallel storage-occupancy measure. Resources. ∼0.5 FTE (data analyst) plus operator effort to expose telemetry; ∼£50,000. Builds on existing monitoring rather than new infrastructure. FAIRC-4 — Tune Slurm weights as policy levers
Should Have
Recommendation. Treat scheduler priority weights (age, fairshare, job size, QoS) as policy instruments, and publish practical guidance mapping common policy objectives — utilisation, responsiveness, slowdown, account fairness — to weight configurations. Retain existing schedulers; do not mandate replacement. Rationale. Our simulations show that, under contention, priority-weight choices act as genuine allocation policy: no single configuration is best on all objectives — the tuning yields a set of trade-off options, each best on a different goal (utilisation, waiting time, fairness), with no single overall winner. Tuned, transparent Slurm-compatible heuristics land close to the offline full-information benchmark, so they are “good enough” for most current loads and should be tuned, not replaced — a low-cost, low-risk improvement operators can adopt immediately. Timescale. Year 1, 9 months. Short-term; benefits from FAIRC-1 for validation but does not require it. Resources. ∼0.5 FTE (postdoctoral researcher) with operator review; ∼£45,000. Directly usable by Tier-2/Tier-3 operators and CCEs.
21
FAIR-Compute
NFCS Flexible Fund — Final Report
FAIRC-7 — Tiered rapid access and CCEs
Should Have
Recommendation. Reduce the administrative barrier on two complementary fronts: introduce a lightweight rapid-access route (“Seedcorn / Bright Idea”) for small, exploratory, or early-career requests alongside standard peer review; and resource Community Centres of Excellence (CCEs) as the community-facing coordination layer, brokering allocations across systems and providing training and RSE support. Rationale. Manual peer review is the binding administrative constraint, not hardware — operators described allocation reviews consuming 60–100 person-hours and DiRAC assessing ∼67 applications per round under a “double jeopardy” process. Uniform heavyweight review does not scale and disproportionately excludes early-career researchers and non-traditional disciplines. A rapid tier lowers the transactional barrier (monitored via FAIRC-1 rather than gated by review); CCEs address the structural barrier by pairing brokerage with training and RSE support, complementing the automated mechanisms in FAIRC-1, 2 and 8 with the human coordination layer federation needs. Timescale. Phased Years 1–3. Rapid route: Year 1, ∼6 months. CCE coordination: Years 2–3, alongside the NCR rollout. Resources. Rapid route ∼0.3 FTE, ∼£30,000/year; CCE coordination leverages UKRI’s planned CCE investment (define the allocation-brokerage interface, ∼0.3 FTE, ∼£30,000). FAIRC-6 — Cross-system benchmarking & exchange rates
Should Have
Recommendation. Adopt a federation-wide benchmarking and exchange-rate framework so that allocations are portable across heterogeneous systems, using domainrepresentative workloads rather than a single synthetic benchmark, with published conversion factors. Rationale. Portable allocation — central to federation and to the ACCESS-style credit model — requires a defensible way to compare a core-hour on one system with a corehour on another. The WLCG’s move from the HS06 benchmark to HEPScore — which uses containerised real experiment workloads and is reproducible to better than 1% across x86 and ARM processors — is the proven template [11]; ACCESS already operationalises performance-equivalence multipliers. Without a transparent exchange basis, cross-system mobility stalls or embeds hidden inequities between communities and architectures (notably CPU versus GPU). Timescale. Year 1–2, 12 months. Medium-term; a precondition for portable credits and for FAIRC-5. Benefits from FAIRC-1. Resources. ∼1.0 FTE (benchmarking engineer) with domain-community input; ∼£95,000. Aligns with WLCG/HEPScore, ACCESS, and EuroHPC/Waldur accounting. FAIRC-8 — Ex-post verification & usage history
Should Have
Recommendation. Introduce ex-post verification comparing requested against consumed resources, feeding a transparent account-level usage-history state into future allocation and priority, so that persistent over-requesting or sustained under-use is gently penalised and accurate forecasting rewarded. Keep penalties graduated and appealable. Rationale. This codifies practice operators already use informally: IRIS-SCRC accounts for prior under/over-use in future allocations, and DiRAC triggers follow-up review below 50% (small) or 80% (large) utilisation and decays credits quarterly. Our model rep-
22
FAIR-Compute
NFCS Flexible Fund — Final Report
resents this as a usage-history state that links a user’s past behaviour to their future priority. Formalising it federation-wide discourages walltime padding and hoarding — the over-reservation FAIRC-2 measures — while a sandbox or experience-based grace period protects newcomers who genuinely cannot estimate their needs. Verification depends on FAIRC-1 telemetry. Timescale. Year 2, 9 months. Medium-term. Depends on FAIRC-1; complements FAIRC-2. Resources. ∼1.0 FTE (0.5 mechanism-design researcher; 0.5 RSE); ∼£90,000. Extends IRIS-SCRC and DiRAC practice. FAIRC-3 — Pilot value-aware prioritisation (PVR)
Could Have
Recommendation. Pilot an allocation mechanism that accounts for the scientific value lost when time-sensitive jobs wait, using a Projected Value Remaining (PVR) measure that models the non-increasing value of a job over time, evaluated against current schedulers before any production use. Design the interface so users can express time-sensitivity simply, without disclosing complex private information. Rationale. Current schedulers observe resource requests but not whether a computation is still useful if it completes late; our Experiment 3 shows a highly-utilised schedule can still lose a large fraction of scientific value. PVR lets the federation compare mechanisms by preserved value, connecting allocation to the “productive use” and mission-priority questions DSIT raised. The principal risk is incentive compatibility — self-declared urgency invites exaggeration — so it must be tested under truthful and strategic reporting, hence a pilot rather than a deployment. Timescale. Year 2–3, 18 months. Depends on FAIRC-1 and FAIRC-2; should precede any federation-wide value-based pricing. Resources. ∼2.0 FTE (1.0 postdoctoral researcher; 1.0 senior RSE); ∼£230,000. Aligns with the portal and mission-led weighting. FAIRC-5 — Govern federated job mobility
Could Have
Recommendation. Before opening unmetered shared channels between clusters, put in place coordinated placement rules — minimum-gain thresholds, receiving-system utilisation protection, and switching-cost “tolls” — that allow a job to move only when the benefit is clear and the receiving system’s local users are protected. Treat naive “send each job wherever it looks fastest” routing as a baseline to improve on, not a target. Rationale. Federation lets an overloaded system borrow spare or specialised capacity, but our Experiment 4 shows uncoordinated, individually-greedy placement behaves like selfish routing: it minimises global delay yet can harm the receiving system’s local users and risk congestion under load spikes — the computing analogue of Braess’s paradox. Coordinated rules recover most of the benefit while limiting local harm and concentrating transfers in the genuine relief direction. We note honestly that naive currently looks best on aggregate metrics; the case for coordination rests on protecting local service and stability, and strengthening the local-harm evidence is part of the remaining work. Timescale. Year 3–4, 18 months. Long-term; depends on FAIRC-6 and FAIRC-1; runs with the phased expansion of the federated portfolio. Resources. ∼1.5 FTE (1.0 systems/network engineer; 0.5 economist/game theorist);
23
FAIR-Compute
NFCS Flexible Fund — Final Report
∼£150,000. Aligns with EuroHPC/Waldur meta-scheduling.
7. Future Work and Research Directions The recommendations in Section 6 are deliberately grounded in what can be done now with existing infrastructure. They also open onto a richer research agenda that FAIR-Compute is well placed to pursue, and which we believe the NFCS should be ambitious about, because the questions our interviews and simulations raised are exactly the ones a maturing federation will have to answer. From heuristics to provable guarantees. Experiment 2 shows tuned Slurm-like rules landing close to the offline benchmark, but only empirically, on particular windows. A natural next step is theoretical: establishing approximation guarantees for transparent priority rules under stated assumptions on workload structure and contention. If simple, operator-maintainable heuristics can be proven to approximate welfare, fairness, or slowdown objectives, the federation gains a principled — not merely observed — licence to keep the schedulers it already runs. Value-aware allocation and incentive compatibility. The value experiment (Experiment 3) and recommendation FAIRC-3 point to a Projected Value Remaining mechanism, but the deep problem is incentive compatibility: if priority depends on self-declared urgency or value, users will exaggerate. The research programme here is to design mechanisms — potentially with artificial currencies, credits, or carefully bounded priority tokens — under which truthful reporting of time-sensitivity is (approximately) optimal, and to characterise the efficiency–fairness frontier such mechanisms can reach. This connects directly to the classical literature on mechanism design without money and to dynamic assignment with limited manipulability. Federation as a priced congestion game. Experiment 4 treats federation through a congestion-game lens but stops at simple tolls. The fuller agenda is to quantify the price of anarchy of uncoordinated cross-system routing on realistic federated topologies, and to design market-clearing tolls or shadow prices that provably realign self-interested placement with global welfare while protecting receiving systems. Doing this well requires the local-harm metrics we flag as outstanding, and it is where transport economics and algorithmic game theory have the most to offer the DRI. Multi-resource clearing and dynamic boundaries. JASMIN’s storage-binding constraint (Section 3.3) motivates extending the single-resource model to a dual-resource (compute-andstorage) clearing problem, and the batch-versus-interactive tension motivates treating the boundary between batch and software-defined environments as itself a dynamic allocation decision rather than a static partition. Both are concrete, high-value modelling problems surfaced directly by operators. Learning, newcomers, and carbon. Finally, several softer directions recur in the interviews: learning user “types” and workload profiles over time so that history-based mechanisms (FAIRC-8) treat newcomers fairly through sandbox or experience-based grace periods; rewarding accurate estimation with bonus credits; and making allocation carbon-aware, reporting and
24
FAIR-Compute
NFCS Flexible Fund — Final Report
eventually optimising tonnes of CO2 equivalent alongside utilisation. Each is a modest extension individually, and collectively they describe a federation that not only allocates efficiently but learns, includes, and decarbonises as it scales. Underpinning all of this is a single dependency we cannot over-state: data. Every direction above becomes sharper, and testable, the moment the federation can supply real, pseudonymised job logs at scale. That is why FAIRC-1 is both the most mundane and the most enabling recommendation in this report, and why, given access to such data, we are eager to move the simulation evidence from sandbox to operational validation and to pursue the mechanism-design questions above in earnest.
8. Conclusion FAIR-Compute set out to treat federated compute allocation as an economic and mechanismdesign problem, not only an engineering one, and to test that lens against real practice and simulation. The consistent message is that the UK federation should measure first: without a shared data standard and an efficiency KPI, neither operators nor funders can tell whether resources are used productively, and none of the more ambitious mechanisms can be evaluated. Beyond measurement, the near-term wins are unglamorous but real — tune the transparent schedulers already in place, add a rapid-access tier, and standardise cross-system benchmarking — while the genuinely novel research frontier is value-aware allocation and the careful governance of federated mobility, where uncoordinated “selfish” routing can quietly undermine the very efficiency federation is meant to deliver. We would welcome the opportunity to pressuretest these recommendations against DSIT and UKRI policy constraints and, given access to real UK job logs, to move the simulation evidence from sandbox to operational validation.
A. Strategic-reporting mechanisms This appendix develops the strategic-reporting part of the model (Section 4.2). It is the mechanism-design foundation for recommendation FAIRC-8 and the future work of Section 7, and is not exercised by the simulations, which implement only the descriptive scheduler and the offline benchmark. Termination and requeue. How an over- or under-estimate is resolved depends on the mechanism. Under a kill-at-walltime rule the job is terminated the moment it exceeds its reservation, so its completion time is Cj = sj + lj if lj ≤ lj′ and Cj = ∞ otherwise (an over-run yields no value). Under a gentler preempt-and-requeue rule the job is not lost but delayed, incurring a requeue penalty Rj : Cj =
sj + lj ,
lj ≤ lj′ ,
s + l′ + R + (l − l′ ), j
j
j
25
j
j
lj > lj′ .
(4)
FAIR-Compute
NFCS Flexible Fund — Final Report
Penalty mechanisms. The ex-post penalty hj (lj′ , lj , Hi(j) ) can take three canonical forms, each mapping onto observed practice: (
Ujover = (
Ujfs = (
Ujprm =
Vj − κj Wj − αover (lj′ − lj )+ , 0, Vj − κj Wj − gj (Hi(j) ), −gj (Hi(j) ),
lj ≤ lj′ , lj > lj′ ,
(overestimation charge)
lj ≤ lj′ , lj > lj′ ,
Vj − κj Wj , Vj − κj (Wj + Rj ) − αprm (lj − lj′ )+ ,
(fairshare adjustment) lj ≤ lj′ , lj > lj′ .
(preempt-and-requeue)
The overestimation charge prices the excess reservation directly; the fairshare adjustment routes the penalty through the evolving history state Hi(j) = G(Hi(j) , lj′ , lj ), so that accounting by reported rather than actual usage lowers the user’s future priority — exactly the loop IRISSCRC and DiRAC operate; and preempt-and-requeue replaces the harsh kill with a bounded delay, partially insuring users against genuine uncertainty while preserving a deterrent against gross over-reporting. Characterising truthful, welfare-improving mechanisms of this kind, with and without artificial currency, is the mechanism-design programme flagged in Section 7.
B. Evaluation metrics For completeness, we define the metrics used throughout Section 5. For a job j with submission time aj , start sj , completion cj , and runtime lj : the waiting time is Wj = sj −aj ; the turnaround is cj − aj ; and the slowdown is (cj − aj )/lj , which normalises delay by job length so that a short job made to wait is penalised more than a long one. Platform utilisation is the fraction of available core-time actually used over the evaluation window; throughput is jobs completed per unit time; makespan is the time to clear the workload. For fairness we report, per account and per job-size class, the slowdown gap (max−min average slowdown) and Jain’s index J(x) = P P ( i xi )2 / n i x2i , applied both to allocated CPU service (a desirable quantity, where a higher index means more even sharing) and to delay (a cost, where a high index can also mean uniformly poor service). Because a high Jain index on delay is ambiguous, we always read it together with the average level of delay: the target is low average slowdown and a high slowdown-Jain index.
C. Inferred QoS-tier proxy The public Fresco trace records account identifiers and per-job resources but not the true ACCESS project type. To give the simulator realistic account-scale heterogeneity we infer a proxy tier from observed demand: each account’s monthly usage is annualised (CPU jobs measured in core-hours; whole-node jobs in node-equivalent core-hours at 128 cores/node; high-memory and GPU jobs charged with multipliers) and mapped to the four ACCESS-style tiers (Table 14). The proxy should be read as an inferred demand tier, not the true administrative QoS; conclusions are not sensitive to the exact cut-points.
D. Federation policies and complete results The federation broker (Experiment 4) is invoked at each job’s arrival; it builds the feasible execution options, estimates each option’s completion time from the unfinished work already on the target system, applies the active policy, and submits the job. The policies compared are defined in Table 15; the switching cost is a fixed component plus 10% of the selected runtime, 26
FAIR-Compute
NFCS Flexible Fund — Final Report
Table 14: ACCESS-style QoS-tier proxy inferred from annualised account demand. Proxy tier
Annualised usage proxy
Interpretation
Explore (1) Discover (2) Accelerate (3) Maximize (4)
≤ 400,000 (400,000, 1,500,000] (1,500,000, 3,000,000] > 3,000,000
Small-scale or exploratory demand Moderate demand Larger mid-scale demand Highest inferred demand tier
applied only to remote execution. Table 16 gives the complete results across all four workload scenarios, extending the selected rows shown in Section 5. Table 15: Federation placement policies (Experiment 4), in increasing order of coordination. Policy
Rule
No federation Naive Switching-cost
Jobs run only on their home system. Place on the eligible system with the lowest estimated completion time. Choose the lowest estimated completion time plus a remote switching toll (fixed + 10% of runtime). Move remotely only if the estimated saving exceeds a threshold θ. Move only if minimum-gain holds and the receiving system stays below a utilisation ceiling Umax .
Minimum-gain Utilisation-protected
E. Stakeholder engagement programme Table 17 summarises the organisations engaged through interviews and follow-up during February– June 2026. Engagement was iterative: findings from earlier sessions shaped the agendas of later ones, and operational insights were mapped to the model variables and recommendations as described in Section 2.
F. Software and reproducibility The simulation framework is implemented in Python. Multi-objective priority-weight search uses pymoo with the NSGA-II algorithm [3, 5, 18]; the offline welfare benchmark is a constraintprogramming model solved with PyJobShop [14, 17]; workload-cluster interpretability uses SHAP [15]. Single-system experiments use a representative 7,000-job high-pressure window on a simulated platform of 300 nodes × 128 cores (38,400 CPU cores); the deadline-value study uses a 500-job window to keep the offline optimisation tractable; the federation study uses two such systems with 30-day synthetic traces. All contention results are point estimates from these windows and are labelled as synthetic sandbox experiments throughout. The eventdriven federation simulator underlying Experiment 4 — covering homogeneous and heterogeneous (CPU/GPU) two-system federation, the workload scenarios, switching costs, and the local-user-protection rules — is openly available at https://github.com/angelamath/HPC_ federation_simulator.
Team and Acknowledgements FAIR-Compute is led by Professor Konstantinos (Kostas) E. Zachariadis (Principal Investigator, School of Economics and Finance, Queen Mary University of London), with Dr Ahmed Sayed (co-lead, Electronic Engineering and Computer Science, QMUL) and Professor Dimitris Fotakis 27
FAIR-Compute
NFCS Flexible Fund — Final Report
Table 16: Complete federation results across all workload scenarios (Experiment 4). “Hom.”/“Het.” denote homogeneous (CPU-only) and heterogeneous (CPU+GPU) hardware. Avg wait
Avg slow.
A→B
B →A
Asym., Hom.
No federation Naive Switching cost Minimum-gain Utilisation-protected
127.8 4.0 18.9 4.7 9.0
13.33 1.39 2.17 1.61 1.18
0 5,974 739 1,748 1,748
0 13,402 22,466 8,316 8,751
Asym., Het.
No federation Naive Switching cost Minimum-gain Utilisation-protected
9.5 0.2 8.2 0.3 0.5
1.60 1.00 1.06 1.00 1.00
0 6,822 1,371 2,801 3,051
0 9,108 20,408 7,080 7,682
Bal., Hom.
No federation Naive Switching cost Minimum-gain Utilisation-protected
133.8 51.3 71.6 58.0 69.6
16.50 8.88 11.81 9.78 10.20
0 11,770 7,179 5,274 3,117
0 13,497 14,692 6,116 4,435
Bal., Het.
No federation Naive Switching cost Minimum-gain Utilisation-protected
20.0 1.3 9.0 1.0 3.8
3.83 1.21 1.28 1.16 1.50
0 14,502 7,633 7,514 7,454
0 9,062 12,073 5,125 5,075
Scenario
Policy
Table 17: Interview and engagement programme (roles anonymised to organisation level). Organisation
Focus of engagement
QMUL ITS Research (Apocrita) STFC IRIS / IRIS-SCRC
Tier-3 operations, cost-recovery, fairshare, walltime behaviour RSAP allocation cycle, ex-post penalty loops, computing-model bidding Data-intensive allocation, storage governance, residual-claimant federation Credit-based allocation, “use-it-or-lose-it”, under-utilisation review National service transition, federation coordination, reporting Roadmap, ROI/KPIs, disaggregated stack, CCEs AIRR allocation, “productive use”, policy objectives
JASMIN (NERC) DiRAC EPCC / NFCS coordination UKRI Digital Research Infrastructure DSIT public-compute division
28
FAIR-Compute
NFCS Flexible Fund — Final Report
(advisor, Electrical and Computer Engineering, National Technical University of Athens). Angeliki Mathioudaki (NTUA and ICCS) led the simulation study and model; Wan Shuen Siaw (School of Mathematical Sciences, QMUL) led the landscape review and field data. We thank the many operators and policymakers who gave their time in interviews and the stakeholder survey, including colleagues at QMUL ITS Research, IRIS, JASMIN, DiRAC, EPCC, DSIT, and the UKRI Digital Research Infrastructure programme. We acknowledge the support of the National Federated Compute Services NetworkPlus, who have funded this work through the Engineering and Physical Sciences Research Council (EPSRC) under Grant No. EP/Z534493/1 (NFCS Network+ Flexible Fund).
References [1] ACCESS. ACCESS Project Types. project-types, 2026. Accessed: 2026-07-01.
https://allocations.access-ci.org/
[2] ACCESS. ACCESS Resources. https://allocations.access-ci.org/resources, 2026. Accessed: 2026-07-01. [3] Julian Blank and Kalyanmoy Deb. Pymoo: Multi-objective optimization in python. Ieee access, 8:89497–89509, 2020. [4] China Academy of Information and Communications Technology (CAICT). White paper on China’s computing power development index. English translation (2024), Center for Security and Emerging Technology, 2022. https://cset.georgetown.edu/wp-content/ uploads/t0581_china_compute_index_2022_EN.pdf. [5] Kalyanmoy Deb, Amrit Pratap, Sameer Agarwal, and TAMT Meyarivan. A fast and elitist multiobjective genetic algorithm: Nsga-ii. IEEE transactions on evolutionary computation, 6(2):182–197, 2002. [6] EPSRC. The impact of epsrc’s investments in high performance computing infrastructure: Final report. Technical report, UK Research and Innovation, 2019. https://www.ukri.org/wp-content/uploads/2022/07/ EPSRC-050722-ImpactEPSRCInvestmentsHighPerformanceComputingInfrastructure. pdf. [7] ETP4HPC. Federation of computing infrastructures. Technical report, European Technology Platform for High-Performance Computing, 2025. [8] EuroHPC Joint Undertaking. Eurohpc federation platform. https://www.eurohpc-ju. europa.eu/supercomputers/eurohpc-federation-platform_en, 2026. Accessed: 202606-30. [9] FRESCO Project. FRESCO: Open repository and analysis of system usage data. https: //www.frescodata.xyz/, 2026. Accessed: 2026-05-04. [10] Frontier Economics. The economic impact of NVIDIA cambridge-1, 2021. Estimated economic value approx. £600M (approx. $831M) over ten years. [11] Domenico Giordano et al. HEPScore: A new CPU benchmark for the WLCG. In EPJ Web of Conferences (CHEP), 2023. Adopted by WLCG April 2023, replacing HEP-SPEC06; reproducible to better than 1% across x86 and ARM. https://arxiv.org/abs/2306. 08118.
29
FAIR-Compute
NFCS Flexible Fund — Final Report
[12] High Performance Computing Infrastructure (HPCI). Overview of HPCI. https://www. hpci-office.jp/en/about_hpci/what_is_hpci, 2024. Accessed: 2026-07-17. [13] David Hudak, Doug Johnson, Alan Chalker, Jeremy Nicklas, Eric Franz, Trey Dockendorf, and Basil L McMichael. Open ondemand: A web-based client portal for hpc centers. Journal of Open Source Software, 3(25):622, 2018. [14] Leon Lan and Joost Berkhout. Pyjobshop: Solving scheduling problems with constraint programming in python. arXiv preprint arXiv:2502.13483, 2025. [15] Scott Lundberg and SHAP contributors. SHAP: Shapley additive explanations documentation. https://shap.readthedocs.io/en/latest/, 2026. Accessed: 2026-05-12. [16] Joshua McKerracher, Preeti Mukherjee, Rajesh Kalyanam, and Saurabh Bagchi. Fresco: A public multi-institutional dataset for understanding hpc system behavior and dependability. In Practice and Experience in Advanced Research Computing 2025: The Power of Collaboration, pages 1–6. 2025. [17] PyJobShop Developers. PyJobShop Documentation. https://pyjobshop.org/stable/ index.html, 2026. Accessed: 2026-07-01. [18] pymoo Developers. pymoo: Multi-objective Optimization in Python. https://pymoo. org/, 2026. Accessed: 2026-07-01. [19] Depei Qian and Zhongzhi Luan. High performance computing development in China: A brief review and perspectives. Computing in Science & Engineering, 21(1):6–16, 2018. doi:10.1109/MCSE.2018.2875367. [20] SchedMD. Slurm workload manager documentation, 2026. URL https://slurm.schedmd. com/documentation.html. Accessed: 2026-03-14. [21] X Carol Song, Preston Smith, Rajesh Kalyanam, Xiao Zhu, Eric Adams, Kevin Colby, Patrick Finnegan, Erik Gough, Elizabett Hillery, Rick Irvine, et al. Anvil-system architecture and experiences from deployment and early user operations. In Practice and experience in advanced research computing 2022: Revolutionary: Computing, connections, you, pages 1–9. 2022. [22] UK Government. Met office supercomputing 2020+ programme: Accounting officer assessment. Technical report, Government Major Projects Portfolio, GOV.UK, 2022. Net present social value £13.74bn (25% optimism bias); 9:1 cost–benefit ratio against the do-nothing option. [23] UK Research and Innovation. National compute ecosystem town hall. UKRI Digital Research Infrastructure programme, 2026. June 2026. [24] Waldur. Waldur: Cloud and hpc management platform. https://waldur.com/, 2026. Accessed: 2026-06-30. [25] WLCG Collaboration. Memorandum of understanding for the worldwide lhc computing grid. https://wlcg.web.cern.ch/organisation-mou/sample-mou, 2020. Accessed: 2026-06-30. [26] Jason Yalim. Toward dynamically controlling slurm’s classic fairshare algorithm. In Practice and Experience in Advanced Research Computing 2020: Catch the Wave, pages 538– 542. 2020.
30