ConceptioArchivearXiv CS
arXiv CSopen access

Affinity Tailor: Dynamic Locality-Aware Scheduling at Scale

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
operating-systemsvirtualization
operating systems, kernel, virtualization

arXiv:2604.27915v1 [cs.OS] 30 Apr 2026

Affinity Tailor: Dynamic Locality-Aware Scheduling at Scale Jin Xin Ng

Ori Livneh

Richard O’Grady

Josh Don

[email protected] Google, USA

[email protected] Google, USA

[email protected] Google, USA

[email protected] Google, USA

Peng Ding

Samuel Grossman

Luis Otero

Chris Kennelly

[email protected] Google, USA

[email protected] Google, USA

[email protected] Google, USA

[email protected] Google, USA

David Lo

Carlos Villavieja

[email protected] Google, USA

[email protected] Google, USA

Abstract

schedulers continuously interleaving threads from distinct applications on the same physical cores. This interleaving is inefficient due to the continuous loss of microarchitectural state. Every time the kernel schedules a new thread onto a core, it effectively begins diluting the previous thread’s instruction execution and data access history, thus degrading the hit rates across the memory hierarchy (L1, L2, and last-level caches) and reducing the accuracy of core-level predictive hardware structures, such as the branch predictor and hardware prefetchers. Concurrently, memory bandwidth per core has stagnated in recent server architectures, while workloads have become memory-intensive [13]. This disparity leaves shared memory interconnects highly vulnerable to saturation caused by traffic from antagonistic workloads. Thread scheduling is a primary determinant of application performance [21]. Modern OS schedulers—such as Linux’s Completely Fair Scheduler (CFS) [25] and Earliest Eligible Virtual Deadline First (EEVDF) scheduler [7, 29]—prioritize work-conservation. A strictly work-conserving scheduler operates on the axiom that no CPU core should sit idle if a runnable thread exists. Consequently, the OS constantly multiplexes threads across cores to minimize immediate queueing. While recent advances have made scheduling policies more extensible [10, 11], work-conservation remains a dominant paradigm. Work-conserving approaches maximize theoretical CPU cycle utilization but are increasingly at odds with modern hardware. Over the last several generations of server CPUs, the cadence of Moore’s law has driven core counts exponentially higher while memory bandwidth per core has stagnated. On modern processor architectures composed of numerous chiplet-based core complexes, non-uniform cache hierarchies pose significant inefficiencies for application performance as memory traffic must traverse shared memory interconnects [33], often with fixed bandwidth ceilings. Even when the bandwidth ceilings are not reached, elevated memory bandwidth results in higher memory latency [13]. Simultaneously, software applications spend a significant portion

Modern large multicore systems often run multiple workloads that share CPUs under schedulers such as Linux CFS. To keep CPUs busy, these schedulers load-balance runnable work, causing each workload to execute on many cores. This weakens locality at the microarchitectural level: workloads lose reuse in caches, branch predictors, and prefetchers, and interfere more with one another—especially on chiplet-based systems, where spreading execution across cores also spreads it across LLC boundaries. A natural alternative is strict CPU partitioning, but hard partitions leave capacity idle when workloads do not fully use their reserved CPUs. We present Affinity Tailor, a userspace-guided kernel scheduling system built on a key insight: the kernel can preserve locality for workloads that share CPUs by treating demandsized, topologically compact CPU sets as affinity hints rather than hard partitions. A userspace controller estimates each workload’s CPU demand online and assigns a preferred CPU set sized to that demand, chosen to be as disjoint as possible from other workloads while spanning as few LLC domains as possible. The kernel then uses this set as an affinity hint, steering threads toward those CPUs while still allowing execution elsewhere when needed to preserve utilization. Deployed at Google, Affinity Tailor delivers geometric-mean per-CPU throughput gains of 12% on chiplet-based systems and 3% on non-chiplet systems over Linux CFS. Furthermore, faster execution reduces memory residency, yielding per-GB throughput gains of 3-7%. Our findings suggest that future schedulers should treat spatial locality as a first-class objective, even at the expense of work-conservation.

1

Introduction

As modern datacenter processors aggressively scale core counts, individual applications increasingly struggle to saturate massive hardware topologies. To maximize hardware utilization, hyperscalers aggressively co-locate hundreds of workloads per machine [2, 3], resulting in operating system 1

Affinity Tailor: Dynamic Locality-Aware Scheduling at Scale

of cycles backend-bound [16, 28], i.e., waiting for data to load. When a work-conserving scheduler migrates a thread between processor cores, it forces cache lines to move across the processor, further congesting shared interconnects and degrading the performance of all co-located workloads. Operating systems provide mechanisms like cpusets to strictly limit thread execution to specific cores. However, static partitioning is ill-suited for modern datacenters for two reasons. First, workloads are notoriously bursty; strict limits cause severe latency penalties during spikes in parallelism. Second, datacenter operators frequently overcommit [2, 3] resources—the sum of CPU resources sold on a machine exceeds the physical machine capacity—making it impossible to assign each application a disjoint set of CPUs. Thus, while cpusets can preserve microarchitectural state, they are fundamentally incompatible with the bursty, overcommitted environment of modern datacenters. In the literature, hardware-assisted resource partitioning mechanisms [4, 5, 20, 27, 32] have been discussed as a means of managing the use of shared system resources between co-located applications, but these systems do not govern thread-to-core placement. Separately, the Nest scheduler [17] concentrates threads onto a dense set of warm cores to exploit higher turbo frequencies, but lacks the notion of per-application isolation. Cache Aware Scheduling patches [6] in the Linux kernel propose LLC-affinity heuristics in the load balancer, but operate with minimal insight into applications’ execution histories. None of these systems provide the dynamic, locality-aware, application-isolating scheduling needed in modern warehouse-scale datacenters. We introduce Affinity Tailor, an OS scheduling architecture that opportunistically maximizes spatial locality without strictly sacrificing work-conservation. Affinity Tailor introduces a novel Linux kernel mechanism, Preferred Cores, to provide soft affinity. Threads are drawn to dynamically sized, "hot" execution domains during nominal load, thus improving spatial locality and better retaining microarchitectural state. Specifically, Affinity Tailor promotes thread execution on "hot" preferred cores—featuring primed caches, prefetchers, and an accurate branch predictor. The Preferred Cores mechanism acts as a permeable boundary, permitting threads to burst onto external cores during momentary parallelism spikes to prevent excessive thread queuing. Affinity Tailor utilizes two distinct core allocation strategies. We initially developed a chiplet-granularity algorithm for split-LLC architectures. By applying soft affinity to pack applications into the absolute minimum number of required chiplets, this algorithm sought to reduce cross-boundary migrations to alleviate shared interconnect saturation. Our initial fleet deployments revealed unexpected efficiency gains in per-core predictive structures—specifically L1/L2 caches and branch predictors—independent of the LLC. Driven by these insights, we developed a secondary, fine-grained core

allocation algorithm tailored specifically for monolithic-LLC architectures. Unlike previously proposed microsecond-scale scheduling systems [9, 12, 14, 15, 18, 19, 26] that replace the OS scheduler with custom scheduling stacks, Affinity Tailor functions transparently within the kernel to maximize spatial locality, and is compatible with arbitrary server hardware and software applications. We deployed and evaluated both algorithms of Affinity Tailor across thousands of machines in Google’s global fleet over one week across a highly diverse mix of latency-sensitive user-facing services and throughput-oriented batch workloads. Our evaluation demonstrates that Affinity Tailor improves aggregate application throughput by up to 12% perCPU and up to 7% per-GB memory, enabling systems to harvest most of the performance benefits of isolated scheduling domains without sacrificing work-conservation. In summary, this paper makes the following key contributions: • We present Preferred Cores, a novel Linux kernel mechanism providing cgroup-based soft affinity that preserves work-conservation. • We introduce a userspace system that dynamically sizes soft affinity regions using short-horizon demand predictions. • We deployed Affinity Tailor in Google’s production fleet, evaluating it across four distinct server platforms and observing aggregate application throughput improvements of up to 12% per-CPU and up to 7% per-GB memory on highly utilized machines. • We empirically demonstrate that aggressively loadbalancing threads to prevent immediate queueing—the foundational principle of modern OS schedulers—is actively detrimental on modern hardware and software architectures.

2

Background and Motivation

Affinity Tailor is motivated by three converging hardware and operational trends: the economic realities of warehousescale computing, the constraints of modern multi-core topologies, and the limitations of existing operating system isolation mechanisms. 2.1

Economics of Overcommitment

Hyperscalers operate under tight economic constraints: energy, power, and physical hardware are scarce resources [1, 20, 23]. Furthermore, organic growth in computational demand routinely outpaces the physical supply of newly procured hardware. While modern datacenter processors now feature massive topologies with 256 logical cores per socket, individual application sizes have largely not followed this trend. Some 2

Affinity Tailor: Dynamic Locality-Aware Scheduling at Scale

the same physical server to maximize hardware utilization. Cluster managers employ a technique known as overcommitment [2, 3] to ensure even higher load factors—allowing the sum of requested CPUs to frequently exceed the sum of physical CPUs on the system. Statistical models are used to ensure that individual applications are likely to have access to their requested CPUs. Crucially, high machine utilization exacerbates underlying physical bottlenecks, such as in memory bandwidth saturation and overheads through the loss of microarchitectural state.

workloads simply cannot scale proportionally due to Amdahl’s Law, as vertical scaling is ultimately bottlenecked by serial computation. However, many others are restricted by operational constraints, such as strict reliability and failover requirements that favor distributed instances, or by the necessity of provisioning fragmented capacity to handle diurnal traffic patterns [30, 31]. Figure 2 shows that the vast majority of applications in Google’s fleet request fewer than 10 CPUs. 300

Logical Cores

256

2.2 200

The aggressive multi-tenancy required by modern economics is fundamentally at odds with the trajectory of modern hardware architectures. The most critical structural bottleneck exacerbated by overcommitment is memory bandwidth. Over the last several generations of server CPUs, the ratio of memory bandwidth available per CPU core has stagnated.

192

128

128

128 112

100 72 56

Microarchitectural Interference

56

0 2017 2018 2019 2020 2021 2022 2023 2024 2025 Year (Server Generations)

MemBW per CPU (GB/s)

2

Percentage of Instances (limit < 𝑥)

Figure 1. Logical CPU cores per socket across recent server generations, showing a 4.6x increase from 56 to 256 CPUs over the last eight years.

1

1.77 1.56

1.52

1.5 1.44

1.46 1.36

1.26

1

1.3

0.96

0.96

0.5 2017 2018 2019 2020 2021 2022 2023 2024 2025 Year (Server Generations)

0.8 0.6

Figure 3. Memory bandwidth per logical core across recent server generations. Relative to the rise in core counts, memory bandwidth per core has remained comparatively flat while per-core compute efficiency has generally improved.

0.4 0.2 0 0.1

1

10

This degradation is compounded by operating systems’ ineffective preservation of spatial locality. When an OS scheduler such as CFS strictly adheres to work-conservation, it aggressively load-balances threads across the entire processor socket to prevent immediate thread queueing. While theoretically optimal for cycle utilization, this constant thread migration induces severe microarchitectural interference: Core-Level Predictive Structures The most immediate consequence of aggressive thread migration is the continuous pollution of core-local predictive state. Every time a thread is migrated, it leaves behind its "warm" execution history. As the thread begins executing on a new, "cold" remote core, it suffers a steep increase in branch mispredictions and L1/L2 cache misses while simultaneously evicting the state of the previous tenant. Furthermore, modern datacenter processors depend on highly tuned cache replacement and prefetching algorithms

100

Normalized CPU Limit (Log Scale)

Figure 2. The proportion of application instances executing at or below a given normalized CPU limit in Google’s fleet. CPU limits are normalized to represent equivalent computational power across server generations. To bridge this massive core-count vs. application size gap and minimize Total Cost of Ownership (TCO), cluster managers like Google’s Borg [31] must rely on aggressive multitenancy. Dedicated servers are only economically viable in limited circumstances such as truly latency-critical workloads, as they come with a price premium. For general services, hundreds of disparate workloads are co-located onto 3

Affinity Tailor: Dynamic Locality-Aware Scheduling at Scale

2.3

Weighted Median Weighted IQR Box

To combat microarchitectural interference, one seemingly obvious solution is to enforce strict execution isolation via the Linux kernel’s native cpuset subsystem. This imposes hard affinity, strictly binding threads to rigid CPU boundaries. However, these static solutions categorically fail at the scale of modern datacenters for two primary reasons. First, hard affinity is mathematically incompatible with aggressive resource overcommitment. In overcommitted environments where the aggregate requested CPU limits exceed physical machine capacity, it is impossible to assign disjoint CPU sets to all tenants without violating the pigeonhole principle. Attempting to overlap hard affinity boundaries forces arbitrary contention on shared CPU. Because workloads peak at unpredictable intervals, co-locating tightly-bound workloads risks causing throughput violations due to severe thread queueing. Workloads could be starved of requested CPU resources while cores outside the overlapping boundary remain idle. Furthermore, relying on userspace cluster agents to resolve these violations by altering the assigned CPU sets is ineffective; their second-scale reaction times are incapable of mitigating microsecond-scale workload bursts. Second, even in the absence of overcommitment, strict CPU limits are hostile to microsecond-scale traffic bursts. Fleet telemetry from Google’s production servers in Figure 5 indicates that individual applications routinely rely on bursting well beyond their requested CPU limits. Unpredictable spikes in parallelism quickly saturate strictly bounded CPU sets. Under hard affinity, this saturation results in immediate, localized thread queueing, generating unacceptable tail-latency spikes for user-facing services. To accommodate these bursty workloads without stranding capacity, operating systems seeking to improve spatial locality require dynamic, permeable boundaries rather than rigid jails.

LLC MPKI

10

5

0 10 −4

10 −3

10 −2

10 −1

100

Hard Affinity & Bursting

101

Context Switches per Million Instructions

Figure 4. Impact of context switching rates on last-level cache misses per-kilo-instruction (LLC MPKI) in Google’s fleet. The trend demonstrates that as context switching rates increase, LLC MPKI increases, highlighting the penalties of frequent thread re-scheduling.

to mask latency [13]. Aggressive thread migration effectively blinds these predictive mechanisms; without a stable execution history, their accuracy drops, leading to degradations in overall processor efficiency. Additionally, individual cores are forced to handle diverse working sets from distinct workloads, further degrading the efficacy of these predictive structures. Last-Level Cache (LLC) Beyond private caches, thread migration degrades the shared LLC. As demonstrated in Figure 4, we observe a clear correlation between increased context switching rates and higher LLC misses per-kiloinstruction. This is particularly severe in chiplet-based designs where the LLC is split into discrete slices. When a thread is load-balanced across a chiplet boundary, its previously established LLC cache lines are effectively orphaned, forcing the hardware to fetch data from remote cache slices. Memory Bandwidth As migrating threads suffer localLLC misses due to degraded LLC performance, they trigger a flood of cross-chiplet, and main-memory reads. In modern architectures, constant data movement saturates the shared memory interconnects, starving all co-located tenants of the already scarce per-core memory bandwidth. Memory bandwidth congestion is severely compounded by hardware prefetchers, which aggressively issue memory fetches in an attempt to hide data access latency. As recent work such as Limoncello [13] has demonstrated, in bandwidth-constrained environments, aggressive hardware prefetching becomes actively detrimental; modern systems must dynamically disable these prefetchers under load to improve overall throughput. Conversely, minimizing crosschiplet data movement inherently reduces memory interconnect saturation, thus allowing hardware prefetchers to remain active and effective for longer durations.

3

System Architecture

Affinity Tailor comprises three components: a kernel mechanism that implements preferential thread steering, a demand predictor that estimates short-term CPU requirements, and a topologically-aware core allocation algorithm that assigns disjoint affinity regions to the containers on a given machine. 3.1

The Preferred Cores Kernel Mechanism

Preferred Cores is a novel Linux kernel mechanism which operates on a per-cgroup basis, providing a cgroupfs interface to specify a set of preferred CPUs. Unlike cpusets, which impose hard affinity by strictly restricting thread placement, Preferred Cores introduces a soft affinity heuristic. The scheduler favors these cores to maximize locality but retains the flexibility to use non-Preferred Cores. The custom scheduling mechanism is built into a component that we named Core-Aware Scheduling (CAS). CAS was inserted into the thread wakeup and load-balancing paths in 4

Affinity Tailor: Dynamic Locality-Aware Scheduling at Scale Thread Wakes Up

Fast Path Scan intersection of cpuset & preferred_cores

Idle core found?

Yes

Enqueue on Preferred Core

No

Slow Path Scan remaining cpuset

Idle core found?

Yes

Enqueue on Non-preferred Core

No

Figure 5. Density map of maximum per-second CPU usage within each 5-minute interval versus requested CPU limit, showing that workloads frequently burst well beyond their nominal CPU allocations.

Queue Thread

Figure 6. Core-Aware Scheduling (CAS) wakeup decision tree. CAS first scans a container’s preferred cores for an idle CPU and falls back to the broader runnable cpuset only when no preferred core is available, preserving work-conservation while biasing placement toward locality.

the Linux scheduler. When a thread becomes runnable, CAS evaluates placement using a tiered approach: • Fast path (preferred cores scan): CAS intersects the container’s runnable cpuset with its dynamically assigned Preferred Cores mask. It scans this subset for an idle1 core. If found, the thread is enqueued there immediately. • Slow path (runnable cores scan): If an idle core is not found in the fast path, i.e., if capacity within the Preferred Cores mask is fully saturated, CAS falls back to scanning the broader runnable cpuset for an idle core.

agents,such as those used by Borg, utilize complex machinelearning models to inform machine-level overcommitment capabilities [3], we find that such models produce forecasts at time horizons too long for the sizing of Preferred Cores masks. The peak CPU utilization of individual containers over a 24-hour horizon tends to be near-or-equal to their requested limits, resulting in the same pigeonhole problem described in Section 2.1. We observed that while the aggregate requested CPU limits on Google machines routinely exceed physical capacity, the aggregate actual CPU utilization consistently remains well below the hardware’s physical capacity. Therefore, we reasoned that we could reasonably size soft affinity regions based on the recent high-percentile CPU utilization of containers. Specifically, the cluster agent samples a container’s CPU utilization every second. Let 𝑢𝑖 represent the CPU utilization measured during the 1-second interval 𝑖. For a trailing 5minute window ending at time 𝑡 (consisting of 300 discrete 1-second measurements), we formally define 𝐷𝑒𝑚𝑎𝑛𝑑 (𝑡) as the 𝑝-th percentile of this set of observations:   𝐷𝑒𝑚𝑎𝑛𝑑 (𝑡) = Percentile𝑝 {𝑢𝑡 −𝜏 | 𝜏 ∈ [1, 300]}

During load-balancing, CAS prevents threads queued within Preferred Cores from being migrated to non-Preferred Cores, and aggressively migrates both running and runnable threads towards idle Preferred Cores at regular intervals. The design of CAS strictly preserves work-conservation, as runnable threads are still permitted to use any available CPU, irrespective of the configured Preferred Cores. Notably, if all containers can be given disjoint sets of Preferred Cores, threads of distinct containers would be strongly "attracted" to those disjoint regions, thereby minimizing spillover into the shared CPU regions and maximizing spatial locality. 3.2

Demand-Based Dynamic Region Sizing

The efficacy of soft affinity depends on the sizing and assignment of the Preferred Cores masks. While modern cluster

By utilizing a high percentile (e.g., 𝑝 = 99), this scheme allows us to ensure the soft affinity regions are sized to comfortably match or exceed the actual CPU demand for the vast majority of time, without encountering the pigeonhole

1We additionally use a set of heuristics to identify a core which is likely to

become idle soon, in similar fashion to CFS. 5

Affinity Tailor: Dynamic Locality-Aware Scheduling at Scale

3.3.1 Chiplet-Based / Split-LLC Systems. We initially designed Algorithm 1 of Affinity Tailor for processor architectures that feature numerous core complexes (chiplets) with their own individual LLC. Data crossing chiplet boundaries must travel through shared memory interconnects, resulting in higher access latencies. The chiplet boundaries impose a steep performance cliff, making it suited as a natural partitioning mechanism to which we aligned our algorithm’s allocation units. Algorithm 1 first schedules workloads with CPU demand under 1 chiplet in ascending order of demand, as singlechiplet assignment obviates the need for inter-chiplet assignment. It then proceeds with scheduling remaining workloads in descending order of demand to minimize fragmentation of large containers. The algorithm performs an ascending combinatorial search to find the absolute minimum number of chiplets required to satisfy each container’s demand. When multiple valid chiplet sets of the same size exist, the algorithm breaks ties by selecting the set with the largest residual capacity, effectively inflating the Preferred Cores regions to provide the most additional capacity for bursting. We find that the combinatorial search is viable, since the number of chiplets per processor remains small, ranging only up to 16 chiplets. A combinatorial search is used due to additional core allocation policy constraints not discussed in this paper.

problem. We integrated this predictor into our userspace cluster agent daemon, Borglet [31]. Figure 7 shows the correlation between the trailing 5minute p99 CPU usage and the subsequent 5-minute p99 CPU usage of containers, indicating that our method projects near-term demand with high accuracy.

Figure 7. Density map of p99 workload CPU utilization across consecutive 5-minute intervals. The concentration near the 𝑌 = 𝑋 line indicates that recent p99 utilization is strongly predictive of near-term demand (𝑅 2 of 0.934).

3.3.2 Monolithic-LLC Systems. We later developed Algorithm 2 of Affinity Tailor for processor architectures that feature a monolithic LLC. In this architecture, the performance cliff imposed by LLC boundaries does not exist, since the LLC is shared equally among all cores. We reasoned that the dominant source of interference on these systems are core-local structures such as the L1/L2 caches, branch predictors, and prefetchers. Therefore, we designed the algorithm to have core-granularity allocation units. The design of this algorithm was also driven by the significant core-level effects we observed during the evaluation of the chiplet-based Algorithm 1. Algorithm 2 calculates a per-socket scaling factor, equal to the ratio of available physical cores to the aggregate predicted demand of containers assigned to the socket. This scaling factor is used to enlarge the Preferred Cores regions, providing for a buffer region to absorb bursts of CPU utilization without spilling. We plan to unify these two algorithms in future work, which will be discussed in Section 7.

3.2.1 Peak Parallelism Handling. As seen in Figure 5, datacenter workloads are typically bursty, with peak parallelism that often exceeds their requested limits and average CPU usage. For that reason, while the scheme described above would ensure soft affinity regions are sized to accommodate average CPU usage, it would fail to handle microsecond-scale bursts of CPU usage. To minimize thread spillover beyond the soft affinity region when applications’ parallelism exceeds average usage, we enlarge the assigned soft affinity regions to better loadbalance usage across the remainder of the processor. Doing so adds a buffer for applications to spill threads, instead of immediately spilling to shared CPU regions. This scheme is feasible since the sum of actual usage rarely exceeds physical capacities; in most of the fleet, we typically see machine utilization below 50%. The specific algorithms are discussed in the following section. 3.3

4

Topology-Aware Core Allocation

Using the per-container demand predictions, the Borglet core allocator assigns containers disjoint Preferred Cores while respecting the processor’s physical topology. The allocation granularity varies by the dominant source of interference on platform-specific architectures.

Evaluation Methodology

We deployed Affinity Tailor across Google’s production fleet. Our evaluation baseline is a heavily modified, latency-optimized version of the Linux Completely Fair Scheduler (CFS). While recent work on scheduling architectures frequently leverages userspace dataplanes or kernel-bypass frameworks [9, 12, 14, 6

Affinity Tailor: Dynamic Locality-Aware Scheduling at Scale

Algorithm 1 Affinity Tailor for Split-LLC Systems

and realistic comparison with the deployable state-of-the-art in Google’s production environment. Each evaluated machine runs hundreds of live production services, ranging from highly latency-sensitive user-facing production applications (e.g., Search, Spanner) to throughputoriented background workloads. As Affinity Tailor is specifically targeted at latency-sensitive applications, we report results exclusively for this class of workloads. Our evaluation spans four distinct hardware architectures:

1: Input: Set of containers 𝑇 , Set of chiplets 𝐶 2: Output: Chiplet soft affinity masks for containers 3: 𝐶𝑐𝑎𝑝 [𝑐] ← Initial CPU capacity for each chiplet 𝑐 ∈ 𝐶 4: 𝐶𝑎𝑝𝑐ℎ𝑖𝑝𝑙𝑒𝑡 ← CPU capacity of a single chiplet 5: 𝑇𝑠𝑖𝑛𝑔𝑙𝑒 ← {𝑡 ∈ 𝑇 | Demand(𝑡) ≤ 𝐶𝑎𝑝𝑐ℎ𝑖𝑝𝑙𝑒𝑡 } 6: 𝑇𝑚𝑢𝑙𝑡𝑖 ← 𝑇 \ 𝑇𝑠𝑖𝑛𝑔𝑙𝑒 7: SortedT ← SortAscending(𝑇𝑠𝑖𝑛𝑔𝑙𝑒 ) 8: SortedT ← SortedT ∪ SortDescending(𝑇𝑚𝑢𝑙𝑡𝑖 ) 9: for each container 𝑡 ∈ SortedT do 10: 𝑑𝑡 ← Demand(𝑡) 11: 12: 13: 14: 15: 16: 17: 18: 19: 20: 21: 22:

• Platforms 1, 2, 3 are successive generations of outof-order multicores featuring split-LLC chiplet topologies. • Platform 4 is a recent out-of-order multicore featuring a monolithic LLC.

BestSet ← ∅ // Search for smallest chiplet combos fitting demand for 𝑖 = 1 to |𝐶 | do Sets𝑖 ← All combinationsÍof 𝐶 of size 𝑖 ValidSets ← {𝑠 ∈ Sets𝑖 | 𝑥 ∈𝑠 𝐶𝑐𝑎𝑝 [𝑥] ≥ 𝑑𝑡 } if ValidSets ≠ ∅ then // Tie-break: Maximize capacity Í for bursts BestSet ← arg max𝑠 ∈ValidSets ( 𝑥 ∈𝑠 𝐶𝑐𝑎𝑝 [𝑥]) break end if end for

To quantify fleet-wide efficiency, we measure application throughput, which represents the number of requests served by the application per unit of time. We utilize hardware performance counters to capture detailed microarchitectural metrics, specifically LLC references per kilo-instruction (RPKI), LLC misses per kilo-instruction (MPKI), and branch predictor MPKI, in addition to overall system memory bandwidth utilization. We use LLC RPKI as a proxy metric for the combined efficacy of the L1 and L2 data caches.

23: if BestSet ≠ ∅ then 24: ApplySoftAffinity(𝑡, BestSet) 25: UpdateCapacities(𝐶𝑐𝑎𝑝 , BestSet, 𝑑𝑡 ) 26: end if 27: end for

5

Evaluation

5.1 Preferred Core Residency We first validate that Affinity Tailor successfully confines threads to their Preferred Cores. We define Preferred Core Residency (PCR) as the fraction of an application’s total CPU time spent executing within its assigned Preferred Cores. Figure 8 plots the observed PCR across the evaluated platforms. 𝑇𝑝𝑟𝑒 𝑓 𝑒𝑟𝑟𝑒𝑑 𝑃𝐶𝑅 = 𝑇𝑡𝑜𝑡𝑎𝑙

Algorithm 2 Affinity Tailor for Monolithic-LLC Systems 1: Input: Set of containers 𝑇 2: Output: CPU affinity masks for eligible containers 3: 𝐶𝑝𝑢𝑏𝑙𝑖𝑐 ← Identify all non-reserved CPU cores 4: 𝐴𝑐𝑜𝑟𝑒𝑠 ← Count of non-reserved cores 5: // Step 1: Compute scaling factor Í 6: 𝐷𝑡𝑜𝑡𝑎𝑙 ← 𝑡 ∈𝑇 Demand(𝑡) 7: 𝐹𝑠𝑐𝑎𝑙𝑒 ← 𝐴𝑐𝑜𝑟𝑒𝑠 /𝐷𝑡𝑜𝑡𝑎𝑙 8: // Step 2: Compute and assign cores to each container

Across the evaluated chiplet-based platforms, we see that 72% of total CPU time was executed with a PCR exceeding 60%, while approximately 60% of CPU time achieved a PCR greater than 80%. In contrast, the monolithic-LLC Platform 4 exhibited a comparatively lower aggregate PCR. These findings suggest that Algorithm 1 provisions a larger Preferred Cores region for individual applications, thereby better accommodating thread spilling during microsecond-scale bursts than the proportional scaling approach of Algorithm 2.

9: for each container 𝑡 ∈ 𝑇 do 10: 𝐷𝑠𝑐𝑎𝑙𝑒𝑑 ← ⌈Demand(𝑡) × 𝐹𝑠𝑐𝑎𝑙𝑒 ⌉ 11: 𝐶𝑎𝑙𝑙𝑜𝑐 ← Pick 𝐷𝑠𝑐𝑎𝑙𝑒𝑑 cores from 𝐶𝑝𝑢𝑏𝑙𝑖𝑐 12: 𝐶𝑝𝑢𝑏𝑙𝑖𝑐 ← 𝐶𝑝𝑢𝑏𝑙𝑖𝑐 \ 𝐶𝑎𝑙𝑙𝑜𝑐 13: ApplySoftAffinity(𝑡, 𝐶𝑎𝑙𝑙𝑜𝑐 ) 14: end for

5.2

15, 18, 19, 26], deploying such systems in a warehouse-scale setting involves significant friction. Specifically, they require the use of custom runtimes [8, 9, 15, 26], compiler-level instrumentation [12], or new hardware features [14, 18, 19]. In contrast, Affinity Tailor is explicitly designed to retain broad compatibility across processor generations and applications. Evaluation against CFS provides the most accurate

Application Throughput Impact

Figure 9 shows that the deployment of Affinity Tailor yields substantial application throughput improvements across all evaluated platforms. We observed aggregate per-CPU throughput improvements ranging from 3.6% on Platform 4 up to 11.7% on Platform 3. When evaluating the aggregate throughput impact per-GB of memory, we see parallel gains 7

Affinity Tailor: Dynamic Locality-Aware Scheduling at Scale

5.3.2 Branch Prediction Accuracy. Across all evaluated platforms, we observed a reduction in Branch Predictor Misses Per Kilo-Instruction (Branch MPKI), ranging from a 0.5% reduction on Platforms 2, up to a 2.2% reduction on Platform 1. We believe these reductions arise from improved retention of execution history by the branch predictors, since Affinity Tailor reduces context switching of disparate workloads onto the same physical core. Thus, Affinity Tailor allows execution units to maintain deep, highly accurate and relevant historical state.

Proportion of CPU ≥ PCR

1 0.8 0.6 0.4

Platform 1 Platform 2 Platform 3 Platform 4

0.2 0

0

0.2

0.4

0.6

0.8

5.3.3 Last-Level Cache Effects. The impact of Affinity Tailor is most pronounced at the LLC level, particularly on split-LLC systems where chiplet boundaries act as severe performance cliffs. This is evidenced by steep reductions in both LLC MPKI and overall LLC Miss Rates. On Platform 2, LLC miss rates dropped by 7%, and on Platform 3, they dropped by 26%. These results are expected, since Algorithm 1 was designed to minimize cross-chiplet sharing of cache lines. On Platform 4, which has a monolithic LLC, we reason that the 1.1% and 0.3% improvement in LLC MPKI and LLC miss rates, respectively, are downstream effects of improved L1/L2 cache efficacy and improved branch predictor accuracy.

1

Preferred Core Residency (PCR)

Figure 8. The proportion of total CPU cycles that achieve at least a given Preferred Core Residency (PCR) threshold. Platform 4, which uses Algorithm 2 has lower PCR, indicating that Algorithm 1 is more effective at accommodating bursts.

ranging from 2.9% to 7.5%. Faster execution shortens memory residency and reduces memory footprints. Our telemetry reveals that these throughput improvements are correlated with improved microarchitectural efficiency, discussed in the following subsection. 5.3

5.3.4 Memory Bandwidth and Prefetcher Activation. The reductions in LLC misses naturally lead to significantly less main memory access, thereby reducing memory bandwidth utilization. As shown in Figure 10, Affinity Tailor reduces high-percentile (P90 and P99) memory bandwidth utilization across all evaluated platforms. We note that Platform 2 shows an increase in average memory bandwidth utilization, and observe that this is an experimental artifact by CPU-load-based dynamic load balancers, which direct additional work to the experiment group due to its improved efficiency. At Google, we employ Limoncello [13] to dynamically disable hardware prefetchers when memory bandwidth approaches the saturation threshold, to avoid severe memory access latency increases. Therefore, the aforementioned reduction in memory bandwidth saturation triggers a compounding microarchitectural benefit: sustained hardware prefetcher activation. Because Affinity Tailor lowers main memory-bound traffic, it keeps the system below the hardware prefetcher disablement threshold longer, allowing the prefetchers to remain active longer. As illustrated in Figure 11, the percentage of time hardware prefetchers remained enabled increased when memory bandwidth utilization on the platform decreased. Ultimately, this allows the system to reclaim the latency-hiding benefits of hardware prefetchers that are typically lost under high memory bandwidth utilization, thereby improving application performance.

Microarchitectural Improvements

5.3.1 Core-Level Cache Efficacy. We evaluate Last-Level Cache References Per Kilo-Instruction (LLC RPKI) as an inverse proxy for L1 and L2 cache hit rates; because a reference is only issued to the LLC when a memory access misses in both core-private caches, a lower LLC RPKI indicates improved L1/L2 efficacy. Across the evaluated platforms, LLC RPKI exhibited varying behavior dictated by the specific Affinity Tailor algorithm in use. On Platform 4, which utilizes the fine-grained Algorithm 2, we observe a 0.9% decrease in LLC RPKI. Algorithm 2 generally provisions smaller, strictly disjoint Preferred Core regions, ensuring that distinct applications rarely interleave on the same physical cores. We reason that the per-application working sets remains tightly bound to disjoint sets of coreprivate caches, allowing better retention of relevant cache lines without requiring trips to the LLC. Conversely, on Platforms 2 and 3—which utilize the chipletgranularity Algorithm 1—LLC RPKI stayed flat or increased by up to 0.8%. Because Algorithm 1 allocates in multiples of whole chiplets, the Preferred Cores regions are larger and shared between more applications. Consequently, it is more likely for threads of distinct applications to be interleaved on the same core. Furthermore, CAS lacks some of the CPUgranularity wakeup affinity mechanisms that are present in CFS. 8

Affinity Tailor: Dynamic Locality-Aware Scheduling at Scale

Throughput Uplift (%)

% Change vs. Baseline

15 0.8

0.1

N/A*N/A*N/A*

0 −0.5

−2.2

-0.9 -1.1 -0.3 −0.6

−1.8

-7.3 -7.3

−10 −20

-24.4 -25.6

−30

11.7

10 7.6

7.5

7.3

4.6

5

4.1 3.6 2.9

0 Platform 1

Platform 2

Platform 3

LLC RPKI

LLC MPKI

LLC Miss Rate

Platform 4

1 3 2 4 orm tform tform tform f t Pla Pla Pla Pla

Branch MPKI per-CPU

*We were unable to collect these metrics on Platform 1.

per-GB

Figure 9. Performance impact of Affinity Tailor. The left plot details the effect on cache and branch predictor misses, while the right plot highlights the aggregate application throughput uplift (per-CPU and per-GB). We see significant throughput gains across all evaluated platforms, correlated with improved microarchitectural efficacy. 70 60 0

Prefetcher Enabled %

Change in MemBW Util (%)

1

−1 −2 −3 −4

40 30 20 10

−5

0 Avg

Platform 1

P90

Platform 2

P99

Platform 3

Platform 1 Platform 4

Platform 2

Affinity Tailor Disabled

Figure 10. Change in average, P90, and P99 memory bandwidth utilization percentage under Affinity Tailor. Tail memory bandwidth utilization declines across all evaluated platforms.

5.4

50

Platform 3

Affinity Tailor Enabled

Figure 11. Impact of Affinity Tailor on hardware prefetcher enablement with Limoncello in-use. Prefetchers remain enabled for a larger fraction of time when Affinity Tailor reduces memory bandwidth utilization.

Thread Scheduling Latency

A consequence of Affinity Tailor’s current implementation is an increase in thread scheduling latency. By prioritizing the CAS fast-path—which restricts initial thread placement to a container’s dynamically sized Preferred Cores—the scheduler bypasses the aggressive load-balancing heuristics present in the baseline CFS that immediately seek out idle cores across the broader socket. Because CAS currently lacks these mature latency optimizations, threads frequently experience localized queueing while waiting for a preferred core to become available, rather than being instantly migrated to a

remote idle core. Threads that are scheduled outside their Preferred Cores are also aggressively migrated back to their Preferred Cores. Figure 12 illustrates the percentage increase in P99 scheduling latency (the time a thread spends in the runnable state waiting for CPU dispatch). As expected from the absence of aggressive idle-search optimizations, the deployment of Affinity Tailor causes substantial latency regressions in the tail. Across all evaluated platforms, we observe 99th percentile scheduling latencies increasing by as much as 17%. 9

Latency Increase (% Diff)

Affinity Tailor: Dynamic Locality-Aware Scheduling at Scale

development, and Plugsched [22] decouples the Linux scheduler into a loadable module. Resource Partitioning and Throttling: Prior systems have extensively studied shared resource contention through the lens of partitioning and throttling. 𝐶𝑃𝐼 2 [32] utilizes hardware performance counters to continuously monitor, detect, and throttle antagonistic workloads. Heracles [20] employs feedback controllers to safely isolate latency-critical workloads from best-effort background workloads by dynamically managing CPU, memory, and network resources. PARTIES [5] introduced QoS-aware partitioning of LLC ways, memory bandwidth, and cores for multiple services. CLITE [27] and OLPart [4] improved upon PARTIES with Bayesian optimization and online learning, respectively. These systems govern shared resource allocation (cache ways, memory bandwidth, time-on-core); Affinity Tailor governs thread-to-core placement.

20

10

0 P99

Platform 1

Platform 2

Platform 3

Platform 4

Figure 12. A comparison of the percentage increase in thread scheduling latency at the 99th percentile when Affinity Tailor is enabled. We observed significantly increased tail scheduling latencies across all evaluated platforms.

Discussion and Future Directions

7.1

Unifying Affinity Tailor Algorithms

Our evaluation of Affinity Tailor revealed a notable divergence in microarchitectural behavior between our two Preferred Cores allocation algorithms. Algorithm 1 (chipletgranularity) successfully mitigated cross-chiplet LLC thrashing on split-LLC architectures (Platforms 2 and 3). However, it yielded minimal improvements in core-private L1/L2 cache hit rates, as evidenced by stagnant to increased LLC RPKI. Conversely, Algorithm 2 (core-granularity) deployed on the monolithic-LLC Platform 4 demonstrated strong L1/L2 cache efficacy improvements by tightly packing threads onto disjoint sets of CPUs. This dichotomy highlights a clear path for future optimization: unifying the two algorithms into a tiered, hierarchical allocation algorithm. For massively parallel, chipletbased processors, the cluster agent should first bin applications into the optimal minimum set of chiplets to minimize cross-chiplet and main-memory-bound traffic. Subsequently, within those selected chiplets, the agent should apply the fine-grained, demand-scaled core allocation logic to better govern core-level predictive state sharing. By combining these approaches, future systems may simultaneously harvest the LLC improvements of coarse-grained chiplet-based partitioning and the deep, core-level state preservation of fine-grained partitioning.

In traditional operating-system design, as exemplified by CFS and EEVDF, increases in thread-queueing delay are typically avoided, as it is assumed to directly degrade application performance. However, our fleet-wide evaluation reveals that despite an "undesirable" increase in scheduling latency, overall application throughput significantly improves (as detailed in Section 5.2).

6

7

Related Work

Locality-Aware System Design: Affinity Tailor builds upon a history of chiplet-aware design practices. The Nest scheduler [17] concentrates threads onto a dense set of warm cores to exploit higher turbo frequencies, thus improving performance and energy efficiency in lightly-loaded environments. It does not incorporate per-application demand forecasting or target overcommitted multi-tenant environments. Recent proposals like Cache Aware Scheduling [6] seek to implement LLC affinity directly within the mainline Linux kernel’s load balancer, using purely kernel load metrics without additional context from userspace. Zhou et al. [33] showed that sharding TCMalloc’s transfer caches into chiplet-local structures improved application throughput. Extensible Scheduling Frameworks: There have been many recent efforts to modularize the Linux kernel’s scheduling policies, which would allow for faster implementation of Affinity Tailor’s Preferred Cores policy. The sched_ext framework [10] was merged into Linux 6.12, enabling the use of dynamically-loaded BPF scheduling policies. ghOSt [11] delegates scheduling to userspace agents with sharedmemory queues. Enoki [24] supports rapid kernel scheduler

7.2 Towards Locality-Aware Scheduling Perhaps the most profound architectural insight from the deployment of Affinity Tailor is the observation that aggressively migrating threads to minimize immediate queueing is a flawed optimization target in hyperscale datacenters built on hardware with deeply stateful microarchitectures. Our findings provide compelling evidence that a strict adherence to work-conservation—a foundational principle in modern 10

Affinity Tailor: Dynamic Locality-Aware Scheduling at Scale

schedulers like CFS and EEVDF—can be counterproductive on modern hardware and software architectures. Currently, Affinity Tailor implements improved spatial locality via a cooperative hardware-software architecture: a userspace daemon calculates the boundaries, and the kernel enforces the Preferred Cores soft affinity fast-path through CAS. However, these results suggest that operating system schedulers could benefit from a more fundamental evolution. Future OS schedulers may need to natively treat spatial locality and microarchitectural state “warmth” as first-class scheduling parameters. Such schedulers could dynamically weigh the microsecond penalty of localized queueing against the latency costs of remote memory fetches, cold branch predictors, and degraded hardware prefetchers. Shifting the core scheduling paradigm from aggressive load-balancing to a dynamic, locality-aware approach represents a promising frontier in datacenter efficiency. Conversely, these results suggest a parallel evolution path in hardware design. Because deeply stateful predictors are vulnerable to frequent context switching, future architectures might benefit from explicitly accommodating these OS scheduling realities. Hardware designers could explore mechanisms for rapid state recovery rather than presuming prolonged, consistent execution of applications.

locality would allow overcommitted systems to more effectively apportion scarce resources.

8

Conclusion

As hyperscalers aggressively overcommit hardware, conventional operating system schedulers inadvertently exacerbate microarchitectural interference by prioritizing workconservation, constantly dispersing workloads across massive processor topologies and destroying core-local predictive state. Affinity Tailor is a dynamic scheduling architecture utilizing a novel Preferred Cores kernel mechanism to establish permeable, per-application, soft affinity boundaries sized to near-term demand forecasts. Deployed across Google’s global production fleet, Affinity Tailor confines execution to warm microarchitectural domains while still permitting bursts onto external cores, improving aggregate application throughput by up to 12% per-CPU and 7% per-GB memory. Ultimately, our deployment demonstrates that the microsecond-scale queueing penalty incurred by waiting for a preferred core is eclipsed by the execution speedups gained from inheriting a hot microarchitectural environment, suggesting that future operating system designs should consider prioritizing spatial locality over strict work-conservation.

Acknowledgments 7.3

Affinity Tailor was made possible by numerous Google engineering teams across many years. We would like to especially thank Yiyan Lin and Sundar Dev for their early work in productionizing Preferred Cores. We also thank Xiangling Kong, Nan Deng, Zhiyuan Liu, Li Li, Dagang Wei, Trang Tran, Sahil Shekhawat, Shiyu Hu, Peilin Ye, Corentin Pescheloche, Steve Zekany, Colin Rioux, Darryl Gove, Aleksei Shchekotikhin, Nilay Vaish, Patrick Xia, Wanying Lu, Tae Jun Ham, Haiming Liu, and Akanksha Jain for their guidance and support. Furthermore, we are grateful to Sotiris Apostolakis, Tipp Moseley, and Sree Kodakara for their invaluable feedback. Gemini was utilized to generate sections of this Work, including text, figures, and citations.

Cross-Stack Layering for Scheduling

While Affinity Tailor demonstrates the value of coordinating userspace demand prediction with kernel-enforced soft affinity, its current implementation applies spatial locality heuristics uniformly across all configured workloads. Operating systems, functioning purely at the hardware abstraction layer, lack semantic understanding of the applications they schedule. Consequently, the kernel alone is ill-equipped to determine which workloads actually benefit from microarchitectural state preservation. Future scheduling architectures can embrace deeper crossstack layering, leveraging the rich contextual metadata accessible to userspace cluster agents like Borglet. The agent possesses comprehensive knowledge of a job’s planned resource constraints, historical execution phases, and overall workload archetype. These signals can directly inform thread-scheduling decisions. For instance, a lightweight RPC forwarding service is primarily network I/O bound; tightly packing its threads into a constrained set of Preferred Cores may artificially bottleneck its processing throughput without yielding particularly meaningful cache-hit improvements. The agent could proactively classify workloads by their affinity-sensitivity, supplying these semantic hints to Affinity Tailor. In turn, Affinity Tailor can restrict spatial locality heuristics to stateheavy, compute-bound applications, while allowing stateless or network-bound microservices to aggressively loadbalance across the socket. This targeted application of spatial

References [1] Luiz André Barroso, Urs Hölzle, and Parthasarathy Ranganathan. 2018. The datacenter as a computer: Designing warehouse-scale machines. Morgan & Claypool Publishers. doi:10.1007/978-3-031-01761-2 [2] Salman A Baset, Long Wang, and Chunqiang Tang. 2012. Towards an understanding of oversubscription in cloud. In 2nd USENIX Workshop on Hot Topics in Management of Internet, Cloud, and Enterprise Networks and Services (Hot-ICE 12). [3] Noman Bashir, Nan Deng, Krzysztof Rzadca, David Irwin, Sree Kodak, and Rohit Jnagal. 2021. Take it to the limit: peak predictiondriven resource overcommitment in datacenters. In Proceedings of the Sixteenth European Conference on Computer Systems. 556–573. doi:10.1145/3447786.3456259 [4] Ruobing Chen, Haosen Shi, Yusen Li, Xiaoguang Liu, and Gang Wang. 2023. OLPart: Online learning based resource partitioning for colocating multiple latency-critical jobs on commodity computers. In Proceedings of the Eighteenth European Conference on Computer Systems. 11

Affinity Tailor: Dynamic Locality-Aware Scheduling at Scale

347–364. doi:10.1145/3552326.3567490 [5] Shuang Chen, Christina Delimitrou, and José F Martínez. 2019. PARTIES: QoS-Aware Resource Partitioning for Multiple Interactive Services. In Proceedings of the twenty-fourth international conference on architectural support for programming languages and operating systems. 107–120. doi:10.1145/3297858.3304005 [6] Tim Chen, Peter Zijlstra, and Yu Chen. 2026. Cache Aware Scheduling. Linux Kernel Mailing List (LKML). https://lkml.org/lkml/2026/2/10/ 1369 [7] Jonathan Corbet. 2023. An EEVDF CPU scheduler for Linux. https: //lwn.net/Articles/925371/. [8] Joshua Fried, Gohar Irfan Chaudhry, Enrique Saurez, Esha Choukse, Íñigo Goiri, Sameh Elnikety, Rodrigo Fonseca, and Adam Belay. 2024. Making kernel bypass practical for the cloud with junction. In 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24). 55–73. [9] Joshua Fried, Zhenyuan Ruan, Amy Ousterhout, and Adam Belay. 2020. Caladan: Mitigating interference at microsecond timescales. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). 281–297. [10] Tejun Heo, David Vernet, Josh Don, and Barret Rhoden. 2024. sched_ext: The BPF Extensible Scheduler Class. Linux Plumbers Conference 2024, Vienna, Austria. https://lpc.events/event/18/sessions/ 192/ [11] Jack Tigar Humphries, Neel Natu, Ashwin Chaugule, Ofir Weisse, Barret Rhoden, Josh Don, Luigi Rizzo, Oleg Rombakh, Paul Turner, and Christos Kozyrakis. 2021. ghOSt: Fast & flexible user-space delegation of linux scheduling. In Proceedings of the ACM SIGOPS 28th Symposium on Operating Systems Principles. 588–604. doi:10.1145/3477132.3483542 [12] Rishabh Iyer, Musa Unal, Marios Kogias, and George Candea. 2023. Achieving Microsecond-Scale Tail Latency Efficiently with Approximate Optimal Scheduling. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles (SOSP ’23). 466–481. doi:10.1145/3600006.3613136 [13] Akanksha Jain, Hannah Lin, Carlos Villavieja, Baris Kasikci, Chris Kennelly, Milad Hashemi, and Parthasarathy Ranganathan. 2024. Limoncello: Prefetchers for scale. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3. 577–590. doi:10.1145/3620666.3651373 [14] Yuekai Jia, Kaifu Tian, Yuyang You, Yu Chen, and Kang Chen. 2024. Skyloft: A General High-Efficient Scheduling Framework in User Space. In Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles (SOSP ’24). 265–279. doi:10.1145/3694715.3695973 [15] Kostis Kaffes, Timothy Chong, Jack Tigar Humphries, Adam Belay, David Mazières, and Christos Kozyrakis. 2019. Shinjuku: Preemptive Scheduling for {𝜇second-scale} Tail Latency. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19). 345–360. [16] Svilen Kanev, Juan Pablo Darago, Kim Hazelwood, Parthasarathy Ranganathan, Tipp Moseley, Gu-Yeon Wei, and David Brooks. 2015. Profiling a warehouse-scale computer. In Proceedings of the 42nd annual international symposium on computer architecture. 158–169. doi:10.1145/2749469.2750392 [17] Julia Lawall, Himadri Chhaya-Shailesh, Jean-Pierre Lozi, Baptiste Lepers, Willy Zwaenepoel, and Gilles Muller. 2022. Os scheduling with nest: Keeping tasks close together on warm cores. In Proceedings of the Seventeenth European Conference on Computer Systems. 368–383. doi:10.1145/3492321.3519585 [18] Yueying Li, Nikita Lazarev, David Koufaty, Tenny Yin, Andy Anderson, Zhiru Zhang, G Edward Suh, Kostis Kaffes, and Christina Delimitrou. 2024. LibPreemptible: Enabling Fast, Adaptive, and Hardware-Assisted User-Space Scheduling. In 2024 IEEE International Symposium on HighPerformance Computer Architecture (HPCA). IEEE, 922–936. doi:10. 1109/HPCA57654.2024.00075

[19] Jiazhen Lin, Youmin Chen, Shiwei Gao, and Youyou Lu. 2024. Fast core scheduling with userspace process abstraction. In Proceedings of the ACM SIGOPS 30th symposium on operating systems principles. 280–295. doi:10.1145/3694715.3695976 [20] David Lo, Liqun Cheng, Rama Govindaraju, Parthasarathy Ranganathan, and Christos Kozyrakis. 2015. Heracles: Improving resource efficiency at scale. In Proceedings of the 42nd annual international symposium on computer architecture. 450–462. doi:10.1145/2749469. 2749475 [21] Jean-Pierre Lozi, Baptiste Lepers, Justin Funston, Fabien Gaud, Vivien Quéma, and Alexandra Fedorova. 2016. The Linux Scheduler: a Decade of Wasted Cores. In Proceedings of the Eleventh European Conference on Computer Systems (EuroSys ’16). 1–16. doi:10.1145/2901318.2901326 [22] Teng Ma, Shanpei Chen, Yihao Wu, Erwei Deng, Zhuo Song, Quan Chen, and Minyi Guo. 2023. Efficient scheduler live update for linux kernel with modularization. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3. 194–207. doi:10.1145/3582016. 3582054 [23] Cy McGeady, Joseph Majkut, Barath Harithas, and Karl Smith. 2025. The Electricity Supply Bottleneck on U.S. AI Dominance. https://www. csis.org/analysis/electricity-supply-bottleneck-us-ai-dominance. [24] Samantha Miller, Anirudh Kumar, Tanay Vakharia, Ang Chen, Danyang Zhuo, and Thomas Anderson. 2024. Enoki: High velocity linux kernel scheduler development. In Proceedings of the Nineteenth European Conference on Computer Systems. 962–980. doi:10. 1145/3627703.3629569 [25] Ingo Molnár. 2007. Modular scheduler core and completely fair scheduler. https://lwn.net/Articles/230501/. [26] Amy Ousterhout, Joshua Fried, Jonathan Behrens, Adam Belay, and Hari Balakrishnan. 2019. Shenango: Achieving high CPU efficiency for latency-sensitive datacenter workloads. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19). 361–378. [27] Tirthak Patel and Devesh Tiwari. 2020. CLITE: Efficient and QoSAware Co-Location of Multiple Latency-Critical Jobs for Warehouse Scale Computers. In 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 193–206. doi:10.1109/ HPCA47549.2020.00025 [28] Akshitha Sriraman and Abhishek Dhanotia. 2020. Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at Hyperscale. In Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems (Lausanne, Switzerland) (ASPLOS ’20). Association for Computing Machinery, New York, NY, USA, 733–750. doi:10.1145/3373376.3378450 [29] Ion Stoica, Hussein Abdel-Wahab, Kevin Jeffay, Sanjoy K. Baruah, Johannes E. Gehrke, and C. Greg Plaxton. 1996. A proportional share resource allocation algorithm for real-time, time-shared systems. In Proceedings of the 17th IEEE Real-Time Systems Symposium. 288–299. doi:10.1109/REAL.1996.563725 [30] Muhammad Tirmazi, Adam Barker, Nan Deng, Md E Haque, Zhijing Gene Qin, Steven Hand, Mor Harchol-Balter, and John Wilkes. 2020. Borg: the next generation. In Proceedings of the fifteenth European conference on computer systems. 1–14. doi:10.1145/3342195.3387517 [31] Abhishek Verma, Luis Pedrosa, Madhukar Korupolu, David Oppenheimer, Eric Tune, and John Wilkes. 2015. Large-scale cluster management at Google with Borg. In Proceedings of the Tenth European Conference on Computer Systems. 1–17. doi:10.1145/2741948.2741964 [32] Xiao Zhang, Eric Tune, Robert Hagmann, Rohit Jnagal, Vrigo Gokhale, and John Wilkes. 2013. CPI2: CPU performance isolation for shared compute clusters. In Proceedings of the 8th ACM European Conference on Computer Systems. 379–391. doi:10.1145/2465351.2465388 [33] Zhuangzhuang Zhou, Vaibhav Gogte, Nilay Vaish, Chris Kennelly, Patrick Xia, Svilen Kanev, Tipp Moseley, Christina Delimitrou, and 12

Affinity Tailor: Dynamic Locality-Aware Scheduling at Scale

Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3. 192–206. doi:10.1145/3620666.3651350

Parthasarathy Ranganathan. 2024. Characterizing a memory allocator at warehouse scale. In Proceedings of the 29th ACM International

13

Record · ID 155402 · SHA-256 2fb15d91fe5914a3
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.