SlideDP: Scaling Host-Resident LLM Fine-Tuning Across Multiple GPUs Ruijia Yang1 Shiyuan Lin1 Yulong Ao2 Zhiyu Li2 Yingli Zhao2 Xianduo Li2 Yonghua Lin2 Zeyi Wen1,*
1 The Hong Kong University of Science and Technology (Guangzhou) 2 Beijing Academy of Artificial Intelligence
Abstract
(a) PCIe-only
Host-resident layer streaming enables full-parameter LLM fine-tuning beyond GPU memory, but data-parallel ranks compete for shared host resources. Replicated transfers amplify traffic, while strong scaling can expose host work as computation windows shrink. We present SlideDP, a synchronous data-parallel runtime for shared-host multi-GPU systems. It maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks. An analytical step-time model characterizes resource bottlenecks and pipeline exposure; runtime measurements guide communication, chunking, and activation policies under a GPU memory budget. In matchedbatch sweeps, SlideDP achieves geometric-mean throughput ratios of 1.46–2.64× over SlideFormer, MegaTrain, and ZeROOffload. On four H100s, SlideDP approaches GPU-resident FSDP2 throughput for Qwen3-14B at a smaller batch size. With a larger batch, it processes over 1M tokens per step and exceeds FSDP2’s measured peak throughput by 11.2%. Separately, it supports 256K-token sequences for the same model and fine-tunes Qwen2.5-72B on four RTX 4090 GPUs. Project page: https://github.com/RegiaYoung/SlideDP.
PCIe fabric
1
(c) NVSwitch fabric
(b) Partial NVLink bridge PCIe fabric
Host
·····
·····
PCIe fabric
Host
Host
·····
NVSwitch
Fast intra-pair links.
GPUs share PCIe fabric. SlideFormer
High-bandwidth P2P fabric.
MegaTrain
ZeRO-Offload
(d) Per-GPU batch 16 32
SlideDP
(e) Global batch 256 30
20 20 10
0
0 1
2
4
8
Number of GPUs (ranks)
OOM
10 OOM OOM
Excess step time (s)
arXiv:2609.34162v1 [cs.DC] 28 Sep 2026
{ryang379,slin954}@connect.hkust-gz.edu.cn [email protected] [email protected] {zyli,ylzhao,xdli,ylin}@baai.ac.cn
1
2
4
8
Number of GPUs (ranks)
Figure 1. (a–c) GPU interconnects differ, but ranks share host resources. (d, e) Excess step time for Qwen3-14B on an eight-A800 node under weak scaling and strong scaling. SlideDP targets multi-GPU training with shared host resources, coordinating parameter delivery, gradient aggregation, and layer-wise updates across GPU ranks. Data parallelism (DP) preserves complete layer computation within each rank, matching the execution model of layer streaming. Tensor parallelism partitions operators within layers and introduces inter-GPU communication [29], which can complicate host–device overlap (§2). A direct DP extension therefore appears straightforward, but GPUs in one node share a fixed host resource pool. Increasing active GPUs adds consumers of CPU compute, DRAM bandwidth, and host–device I/O without proportionally increasing host service capacity. Shared-PCIe contention has also been studied in offloaded training [6]. Two challenges follow. The first is avoidable amplification: a direct DP extension can replicate each layer transfer across 𝑅 ranks and return 𝑅 gradients for host-side reduction, multiplying demand on shared host resources. SlideFormer’s public multi-GPU implementation and MegaTrain’s multi-GPU
Introduction
Full-parameter fine-tuning updates every weight in a pretrained LLM, but mixed-precision Adam maintains substantial model states, including BF16 tensors and FP32 optimizer states. For a model with 𝑁 parameters, these states require approximately 16𝑁 bytes [17, 26], which can exceed GPU memory capacity or leave insufficient space for activations even when distributed across GPUs within one node. Hostresident layer streaming addresses this constraint by keeping persistent update states in CPU memory and streaming layers through a small reusable GPU window [13, 35, 37]. This shifts pressure from GPU memory to host memory and the CPU–GPU path, making efficiency depend on a hiding condition: host-side parameter delivery, gradient return, and CPU optimizer updates must overlap with independent GPU computation and complete before their results are needed. ∗ Corresponding author: [email protected].
1
mode follow this data-movement pattern [35, 37]. The second is residual pipeline exposure: even after this duplication is removed, CPU updates, necessary transfers, and cross-rank and buffer dependencies remain; under strong scaling, the computation window available to hide this remaining work can shrink. Figure 1(d,e) illustrates the resulting execution overhead under weak and strong scaling. We report excess step time over a per-GPU-batch-matched compute reference, providing an aggregate indicator of exposed communication, host-side work, and runtime coordination overhead. The design problem is therefore to reduce redundant traffic and schedule the remaining shared-host work so that it overlaps with GPU computation. Remaining exposed work depends on both hardware topology (Figure 1(a–c)) and workload, so there is no single best communication path. Different routes, such as sharded parameter delivery or GPU-side gradient reduction, introduce different trade-offs between communication cost and overlap opportunity. On A800, removing NVLink bridges reverses the preferred delivery route in host-bound workloads, while route differences remain small in GPU-bound workloads (§4). This sensitivity motivates communication-policy selection adapted to both hardware and workload. Partitioned offload already reduces redundant state movement: ZeRO-Offload and ZeRO-Infinity partition training states across ranks and reduce unnecessary host–device transfers [27, 28]. Our focus is different: coordinating parameter delivery, gradient return, and synchronous layer-wise CPU updates within a bounded GPU window when all ranks share a fixed host resource budget. Reducing logical payload alone is therefore insufficient; communication routes and execution schedules must jointly satisfy shared-resource constraints and parameter-version dependencies. Two design principles follow. First, persistent host-state layout should be decoupled from parameter delivery and gradient reduction, allowing one authoritative host state while selecting communication routes according to topology and workload. Second, available GPU memory should be allocated according to its effect on exposed pipeline work, since fixed memory policies can be suboptimal across execution regimes. Retaining activations can reduce host–device traffic, while preserving intermediate tensors can reduce recomputation; their benefits depend on the active bottleneck and pipeline interaction. Communication and activation policies should therefore be guided by shared-resource demands and pipeline dependencies, and validated through measurements. We present SlideDP, a synchronous DP runtime for fullparameter LLM fine-tuning on one multi-GPU node. It makes the following three contributions: 1. A shared-host runtime for layer-streaming DP. We design a runtime that coordinates parameter delivery, gradient reduction and return, and layer-wise CPU updates within a bounded GPU working window. It maintains one authoritative host copy of master weights and optimizer states
while supporting replicated or sharded parameter delivery over the same persistent state layout. GPU-side reduction returns one aggregated gradient per layer, while chunking overlaps parameter conversion with transfers and gradient return with CPU updates. The runtime coordinates these operations under per-layer parameter-version and update dependencies to preserve synchronous DP semantics. 2. An analytical step-time model of shared-host scaling. We model a training step using shared-resource service demands and cross-layer, cross-rank execution dependencies, including readiness, buffer reuse, and parameter versions. The model characterizes host-bound and GPU-bound execution and explains when reducing traffic shortens a step, when shrinking computation windows expose host work, and why route choices depend on both topology and workload. We examine these effects through comparisons of communication routes, GPU counts, and chunk sizes. 3. Pipeline-aware policy selection with measured cost. AutoPolicy configures communication routes, chunk sizes, and activation layouts for the workload and hardware. Elastic Checkpointing supplies per-layer choices for retention, offloading, and recomputation. Staged profiling selects routes and chunks, screens activation candidates, and calibrates their costs in a mini-pipeline before assigning policies to layers and validating the configuration. We report policy gains across execution regimes together with selection overhead and the steps needed to amortize it. We implement SlideDP in PyTorch Distributed and evaluate it on RTX 4090, A800, and H100 platforms. Matchedbatch sweeps show geometric-mean throughput ratios of 1.46–2.64× over SlideFormer [35], MegaTrain [37], and ZeROOffload [28]. Across 4B–72B models on four H100s, it reaches approximately 85–92% of compute-only reference throughput projected from single-GPU measurements. Qwen3-14B throughput approaches GPU-resident FSDP2 [21, 39] at a smaller batch and exceeds its measured peak at a larger one. For Qwen3-32B at 64 sequences per GPU, it maintains 95– 98% weak-scaling efficiency across two to eight A800s. On four H100s, SlideDP fine-tunes Qwen3-14B with over 1M tokens per step without gradient accumulation and, separately, with 256K-token sequences. It also fine-tunes Qwen2.5-72B on a PCIe-only workstation with four RTX 4090 GPUs.
2
Background and Motivation
2.1
Training States in Full-Parameter Fine-Tuning
Full-parameter fine-tuning stresses the memory hierarchy because each training step maintains multiple categories of states with different lifetimes and access patterns. Table 1 summarizes the major states in mixed-precision Adam training [11, 17]. For a model with 𝑁 parameters, BF16/FP16 compute weights and gradients each require 2𝑁 bytes. FP32 master weights and the two FP32 Adam moment buffers contribute another 4𝑁 + 8𝑁 bytes. With memory-efficient 2
attention [3], activation memory scales as 𝑂 (𝑛ℎ𝑠𝑏) for fixed architectural ratios, where 𝑛, ℎ, 𝑠, and 𝑏 denote layer count, hidden dimension, sequence length, and batch size. Thus, the overall training footprint grows along two major scaling dimensions: model size increases persistent parameter and optimizer states, while longer sequences and larger batches increase activation memory.
without reducing per-GPU memory [12]. ZeRO [26] and FSDP [39] instead partition training states across ranks. Depending on the sharding level, this reduces optimizer, gradient, and parameter memory but adds reconstruction and synchronization: parameters are typically materialized through all-gather before computation, while gradients are synchronized and resharded through reduce-scatter after backward. Sharded ownership with CPU/NVMe residency. Offloading extends this sharded design by moving selected states out of GPU memory. ZeRO-Offload moves optimizer states and computation to the CPU [28], while ZeRO-Infinity extends the residency hierarchy to CPU and NVMe [27]. FSDP also supports CPU offload for sharded states. Other systems further explore profiling-based state placement and execution policy selection [5, 34]. When sharding and offloading are combined, relieving the capacity problem introduces an execution-coordination problem. Sharding introduces GPU-side collectives for parameter materialization and gradient synchronization or redistribution, while offloading adds CPU–GPU transfers and host-side optimizer work. On a multi-GPU node, these operations can contend for shared PCIe, host-memory bandwidth, and CPU resources while overlapping with GPU computation. Reducing state redundancy or moving state to a larger memory tier therefore does not by itself determine step time; execution also depends on how this work maps onto shared resources and how much of it remains exposed in the pipeline. This execution question becomes central once persistent states are anchored in host memory and GPUs materialize only the states needed for the current computation, which is the setting we introduce next.
Table 1. Training-state memory footprint in mixed-precision full-parameter fine-tuning. State
Precision
Compute weights BF16/FP16 Gradients BF16/FP16 Master weights FP32 Adam moments FP32 Activations BF16/FP16 Total estimate
–
Size
Main use
2𝑁 2𝑁 4𝑁 8𝑁 𝑂 (𝑛ℎ𝑠𝑏)
FWD/BWD Sync/update Update Update BWD
16𝑁 + 𝑂 (𝑛ℎ𝑠𝑏) Training step
Existing techniques reduce individual components of the training footprint. Activation checkpointing trades recomputation for lower activation memory [1], while kernel optimizations reduce temporary memory overhead [3, 8, 31]. Parameter-efficient and optimizer-state methods further reduce trainable or optimizer memory [4, 9, 14, 16, 38]. However, these techniques do not determine how persistent states and transient execution data should be placed and coordinated across devices, leaving state movement and runtime scheduling as separate challenges. Parallelism distributes training computation and states across GPUs, with different paradigms targeting different execution regimes. Tensor parallelism partitions operators across GPUs but introduces communication on computation critical paths [29]. Pipeline parallelism partitions layers into stages but requires additional scheduling and may suffer from pipeline bubbles [10, 19]. Data parallelism preserves identical layer computation across ranks while synchronizing training states. Because DP does not repartition operators within a layer, it preserves the layer boundaries and local computation schedule on which host-memory-centric streaming relies. Therefore, the central question for an efficient DP runtime is how to own, move, and update the states in Table 1 under limited GPU memory. 2.2
2.3
Host-Memory-Centric Layer Streaming
Host-memory-centric systems shift the runtime problem from memory placement to heterogeneous scheduling. Once GPUs are used as transient BF16 compute workers over bounded active layers, efficiency depends on whether hostside optimizer updates, host–device transfers, gradient movement, and GPU forward/backward computation can be coordinated at a granularity that hides CPU and PCIe latency. Layer-wise update overlap. StrongHold [30] shows that a layer’s CPU optimizer update can begin once its gradients are ready and overlap with backward computation of earlier layers. Its Megatron-based model-parallel offloading stack, however, does not provide a DP-first shared-host runtime. Host-memory-centric streaming. LoHan, SlideFormer, and MegaTrain further explore host-memory-centric finetuning with different emphases [13, 35, 37]. They share the idea that persistent parameters and optimizer states can reside in host memory while GPUs materialize only tensors required by active layers. LoHan focuses on activation traffic and recomputation trade-offs, while SlideFormer and MegaTrain extend the paradigm toward multi-GPU execution. These advances establish important execution mechanisms;
GPU-Centric Data Parallelism and Offloading
We use state ownership to describe the residency and management of authoritative training states. Existing DP abstractions remain rank-centric: states are replicated or partitioned across data-parallel ranks, with CPU and NVMe serving as additional residency tiers. Replicated and sharded state ownership. Distributed Data Parallel (DDP) replicates parameters, gradients, and optimizer on each GPU rank and synchronizes gradients, 3
H2D
D2H
Host work
replicated replicated sharded sharded
CPU GPU CPU GPU
𝑅𝑞 ℓ 𝑃ℓ 𝑅𝑞 ℓ 𝑃ℓ 𝑞ℓ 𝑃ℓ 𝑞ℓ 𝑃ℓ
𝑅𝐺 ℓ 𝐺ℓ 𝑅𝐺 ℓ 𝐺ℓ
reduce + 1 update 1 update reduce + 1 update 1 update
RTX 4090 A800 H100 8
200 100 0
9
1616
9 14
19
28
1
5
2
4
Concurrent GPUs
200
Parameters per layer 101M 275M 218M 488M
4
100
0
8
0
1
2
4
8
16 32
Physical CPU cores
Figure 2. Shared-host microbenchmarks. (a) Per-GPU cast+H2D for an 800M-parameter layer under concurrent GPUs. (b) CPU Adam core number scaling on Xeon 8468.
coordinating them under synchronous DP on a shared host remains challenging. Pipeline-oriented multi-GPU streaming. RoundPipe demonstrates multi-GPU streaming through pipeline parallelism and CPU offloading [15]. Its overlap of CPU updates with subsequent computation relies on asynchronous optimizer updates and delayed parameter versions, allowing the next iteration to proceed before CPU updates finish. Standard synchronous fine-tuning restores the optimizer-update dependency and can reintroduce stalls. Extending host-memory-centric layer streaming to multiGPU DP requires more than running multiple single-GPU instances. The challenge is coordinating multiple GPU ranks that share one authoritative host state under synchronous data-parallel semantics. A DP runtime must jointly manage parameter delivery, gradient aggregation, CPU updates, and execution dependencies across ranks. The following observations characterize how shared resources and overlap opportunities shape their costs. 2.4
300
(b) CPU Adam scaling Effective GB/s (est.)
Reduction
Cast+H2D / GPU (ms)
Delivery
(a) Transfer contention
Gparam/s
Table 2. Host traffic per iteration for a streamed layer under four datapaths over one shared host state. 𝑃ℓ /𝐺 ℓ : parameter and gradient bytes; 𝑞 ℓ : parameter loads per iteration; 𝑅: rank count. Sharded delivery adds an all-gather per load. Padding and replicated tied-weight tails are excluded.
than provisioned independently for each rank. On an A800 node, the cast-and-copy of an 800M-parameter layer sustains 15.6 GB/s per GPU when one GPU copies but 8.7 GB/s per GPU when two copy concurrently (Figure 2a). On the H100 host, CPU Adam reaches an estimated 182–184 GB/s at 16 physical cores across 101–488M-parameter layers. Increasing to 32 cores improves update speed by only 8.5–10.0% (Figure 2b). These diminishing returns complement the transfer results: increasing concurrent users or CPU workers does not provide proportional growth in shared-host service capacity, making rank-dependent host amplification a scaling concern. Removing avoidable amplification still leaves host work that the pipeline must hide. Observation 2—Residual pipeline exposure. Host-side work becomes exposed when its completion exceeds the available hiding window before downstream operations, such as buffer reuse or the next parameter dispatch, require the result. The window may include remaining backward computation and early computation in the next iteration. For example, under fixed-global-batch strong scaling, increasing 𝑅 reduces each rank’s batch and can shrink this window, although compute time need not scale exactly as 1/𝑅. Exposure depends on shared-host service capacity, queueing, and cross-rank readiness, as captured by the step-time model (Section 3.4). Figure 1(e) illustrates this scaling challenge. For Qwen3-14B at a fixed global batch of 256, increasing the A800 GPU count from four to eight reduces the matched compute reference from 27.2 to 14.4 s, while SlideFormer’s step time remains near 43 s. SlideDP reduces step time from 29.6 to 15.3 s, demonstrating effective overlap as the per-rank computation window shrinks. Under weak scaling, the local computation window is largely preserved, but more ranks can still increase shared-host contention. Configurations dominated by exposed host/transfer work are host-bound; those dominated by accelerator computation are GPU-bound. Exposed cost therefore depends not only on remaining work, but also on how execution routes map it onto shared resources.
Data Parallelism Is Not Free on a Shared Host
Single-GPU host-memory-centric systems make one GPU efficient by hiding host work behind that GPU’s computation. Moving from one GPU to 𝑅 GPUs creates two primary pressures on this hiding condition: avoidable host demand can grow with 𝑅, while strong scaling can shrink the computation window available to hide the remaining work. Neither effect is captured by logical byte counts alone: traffic maps onto physical paths with finite service rates, and pipeline dependencies determine which service becomes exposed in step time. We make four observations along this chain; §3 turns them into a design and §4 measures them. Observation 1—Avoidable amplification. A direct DP extension causes host demand to grow with GPU rank count, while host service capacity remains fixed. Table 2 accounts for the per-layer traffic of the four datapaths a DP runtime can choose from. A direct extension of a single-GPU engine— replicated delivery and CPU-side reduction—multiplies both directions by 𝑅 and adds a host-side reduction that competes with the optimizer update for DRAM bandwidth. These resources are shared within the host/NUMA domain rather 4
scaling (Section 3.4). Elastic Checkpointing exposes activation retention, offloading, and recomputation choices (Section 3.5); AutoPolicy measures their pipeline effects and selects a memory-feasible configuration (Section 3.6).
Observation 3: the right delivery route depends on topology and regime. Sharded delivery cuts each rank’s parameter payload to 𝑃ℓ /𝑅 but adds an all-gather on the GPU interconnect. On a 4×A800 node with NVLink bridges and Qwen3-14B at 16 sequences per GPU, sharded delivery is 17% faster than replicated; after the bridges are removed, replicated is 26% faster. At 8 sequences per GPU the gaps are 33% and 19%. In GPU-bound eight-GPU workloads (64–256 sequences per GPU), the two routes differ by under 2.5% (§4). A wrong route matters most when transfers are exposed, which is why the choice cannot be made from topology alone. GPU memory allocation likewise shapes the remaining pipeline exposure. Observation 4—The value of spare HBM depends on exposed pipeline work. Bounded layer streaming can leave substantial HBM headroom. At 32 sequences per GPU, the fixed Qwen3-14B configuration in Table 3 uses 15.0 GiB of peak CUDA reserved memory on an 80-GB H100. This headroom can reduce different sources of exposed work: retaining saved activations on the GPU removes offload and reload traffic, whereas retaining additional intermediate results reduces recomputation. Because these choices reduce different components of the execution pipeline, their benefit depends on which component is exposed. Because spare HBM can reduce different exposed costs depending on the active bottleneck, memory allocation itself becomes a runtime policy decision rather than a fixed checkpointing choice. Fixed memory policies can therefore be suboptimal across regimes. Together, these observations show that shared-host DP requires coordinated decisions across state movement, execution scheduling, and memory allocation rather than independent optimization of individual components.
3
Design of SlideDP
3.1
Overview
3.2
Shared Host State and Communication Routes
SlideDP separates persistent state ownership from communication routes. For each of 𝐿 layers, the host holds one shared copy of FP32 parameters Θ𝑙32 and Adam moments (𝑚𝑙 , 𝑣𝑙 ). The rank-0 optimizer worker updates this state once per iteration. Three invariants preserve synchronous DP: all ranks compute a layer with the same parameter version; its update consumes the gradient aggregated across all ranks; and its next-iteration dispatch waits for that update. GPU windows hold temporary copies, so changing the delivery route does not change state ownership or update semantics. Parameter dispatch. For a streamed layer with 𝑃𝑙 BF16 parameter bytes, full dispatch sends 𝑃𝑙 to each rank; sharded dispatch sends 𝑃𝑙 /𝑅 to each rank and reconstructs the layer with all-gather. If the layer is loaded 𝑞𝑙 times per iteration, aggregate H2D traffic is 𝑅𝑞𝑙 𝑃𝑙 or 𝑞𝑙 𝑃𝑙 , respectively, ignoring padding. Most layers reload parameters for backward after forward copies leave the window; 𝑞𝑙 counts actual loads. Replicated tied-weight tails are accounted for separately. Both routes leave the host master state unsharded. Gradient return. For a 𝐺𝑙 -byte communicated gradient, GPU reduction sums rank-local gradients to rank 0, which averages and returns one gradient to the host. CPU reduction returns all local gradients, using 𝑅𝐺𝑙 D2H bytes before hostside aggregation. Both paths perform one optimizer update. For sharded layers with fixed 𝑞𝑙 , sharded dispatch with GPU reduction therefore keeps aggregate parameter and gradient payload independent of 𝑅. Activations, collectives, and replicated exceptions contribute additional resource demands. Routes and physical resources. The best route depends on the resources carrying its traffic. Parameter shards traverse shared host paths, while reconstruction uses GPU interconnects. NVLink or NVSwitch can make all-gather cheaper than repeated H2D; a slower collective path can favor full dispatch. CPU reduction also competes with Adam for host resources. The runtime selects delivery and reduction paths for these shared resources; Sections 4.4 and 4.5 evaluate demand reduction and delivery-route sensitivity. Control threads, the update worker, and pinned buffers retain the same CPU/NUMA placement during profiling and execution.
SlideDP is a shared-host runtime for synchronous dataparallel training on a multi-GPU node. It coordinates 𝑅 GPU ranks through one authoritative host state and a cross-rank layer pipeline. Host memory holds the FP32 master parameters and Adam states; each rank processes its local data using temporary BF16 layer copies in a bounded, reusable GPU window. The runtime jointly schedules parameter delivery, gradient aggregation, buffer reuse, and layer updates to overlap host work with GPU computation while preserving synchronous DP semantics. Figure 3 connects these mechanisms to their policy choices. Shared state and communication routes define what work is performed (Section 3.2); the cross-rank pipeline coordinates its execution and overlap (Section 3.3). The model then explains how resource demands and dependencies shape
3.3
Cross-Rank Pipelining and Chunked Overlap
SlideDP overlaps work at two scales: completed layers release updates while earlier layers continue backward, and chunks overlap preparation, movement, and update within a layer. Six logical lanes coordinate parameter preparation and H2D, activation offload and reload, GPU compute, gradient reduction, gradient D2H, and CPU LayerAdam. These lanes share DRAM and host-device paths. The runtime coordinates 5
Data-Parallel GPU Ranks GPU 0
Each rank: BF16 layer window + local activations / gradients
Cross-Rank Layer Pipeline cast+H2D P[i]
GPU compute GPU reduction Grad D2H · rank 0
AG
A[i]
cast+H2D P[i−1]
AG
cast+H2D P[i−2]
A[i−1]
BWD L[i+1] G[i+2]
CPU Adam · rank 0
GPU R−1
BWD L[i] G[i]
L[i+2]
G[i]
L[i+1]
L[i]
Shared Host Training State FP32 master parameters + Adam m/v One authoritative copy across ranks
Reusable staging buffers BF16 parameters / gradient chunks
Probe routes + chunks
AG
BWD L[i−1]
G[i+1]
Staged AutoPolicy Workload ·topology GPU / host memory budgets
A[i−2]
G[i+1] G[i+2]
···
Backward example: sharded dispatch + GPU reduction schematic time
AG: all-gather · ticks: transfer / update chunks
Param dispatch Activation reload
GPU 1
Rank-0 update worker Ordered, chunked LayerAdam
Profile + calibrate MILP + layer placement Check memory feasibility
Selected configuration Dispatch: full / shard Reduction: CPU / GPU Chunks: transfer / update Per-layer activation layout
Figure 3. Overview of SlideDP: shared host state, cross-rank pipelining, and AutoPolicy. rank readiness and buffer completion so that each layer advances through these lanes with one aggregated gradient and one host update. Layer order and rank readiness. During forward, parameter preparation for the next layer overlaps with currentlayer computation. Backward visits layers in reverse order, reloading parameters and offloaded activation state before use. Figure 3 shows the GPU-reduction path for each layer:
starts after the complete local shard is ready. On gradient return, whole-layer reduction precedes chunked D2H; each completed copy releases its Adam chunk while later copies continue. These two overlap paths shorten the delay before dependent work can start. Smaller chunks also incur more launches and queue operations, so AutoPolicy measures the tradeoff under the chosen routes. Exposed waiting. Work is hidden when it finishes before its consumer would otherwise be ready. For a layer update, consumers include buffer reuse and the next parameter dispatch, which may follow computation in the next iteration. Thus the available overlap window follows each operation’s next consumer across the layer pipeline. Section 3.4 models these dependencies together with shared-resource demand. Section 4.4 evaluates chunking with communication routes and activation layout held fixed.
{BWD𝑙,𝑟 }𝑟𝑅−1 =0 −→ Reduce𝑙 −→ D2H𝑙,𝑘 −→ Adam𝑙,𝑘 , where 𝑘 indexes gradient chunks. Once reduction combines all ranks’ contributions, rank 0 returns the aggregated gradient and releases its chunks to the host update worker; the other ranks participate in completion synchronization. GPU computation continues on earlier layers while the completed layer is updated. This moves CPU optimizer work out of the step-end serial tail while preserving one update per layer. Completion and buffer reuse. A GPU slot becomes reusable after its outstanding computation, collective, and copy operations complete. Host gradient staging likewise waits for its previous update consumer. One update worker consumes layers in queue order and their chunks as they become ready. Completion signals release buffers and publish the updated layer version. At iteration 𝑡 + 1, every rank’s dispatch of layer 𝑙 waits for that layer’s update from iteration 𝑡. This per-layer gate preserves the same parameter view across ranks while allowing independent operations to proceed as their inputs and buffers become ready. Together, version gates and buffer completion bound in-flight storage and preserve synchronous DP throughout the overlapped execution. When buffers fill, backpressure exposes the pending transfer or update work. Overlap within a layer. On parameter delivery, the CPU casts chunk 𝑘 + 1 to BF16 while DMA transfers chunk 𝑘. Aligned AVX-512 staging uses non-temporal stores to reduce cache pollution from buffers consumed by DMA. All-gather
3.4
Modeling Shared-Host Scaling
The same logical payload can incur different delays depending on its physical route, concurrent work, and readiness constraints. We model these effects through resource service demands and operation dependencies. Inputs are the workload, communication routes, chunk sizes, activation layout, and resource service costs; the constraints follow the runtime in Sections 3.2–3.3. Resource demand. Let 𝑊 𝑗 be the work per iteration assigned to resource 𝑗 and 𝛽 max an upper bound on its service 𝑗 rate. Capacity gives the steady-state lower bound 𝑇step ≥ max 𝐷 min 𝑗 ,
𝐷 min = 𝑊 𝑗 /𝛽 max . 𝑗 𝑗
𝑗
Resources include GPU compute and copy engines, GPU interconnects, the CPU update worker, host DRAM, and shared PCIe paths. Demands from ranks using a shared path are summed against that path’s capacity; independent GPU demands remain separate. Measured effective rates give emb𝑗 = 𝑊 𝑗 /𝛽b𝑗 . Concurrent casting, DMA, and pirical estimates 𝐷 6
Adam affect these rates. Combined probes such as cast+H2D contribute one joint service cost. Execution dependencies. For a fixed order of compute, collective, and transfer/update operations, completion obeys
3.5
Elastic Checkpointing
Elastic Checkpointing defines the activation-side policy space evaluated by AutoPolicy. It uses memory beyond the bounded GPU parameter window to reduce activation movement or replay computation, combining checkpoint granularity with activation placement, as explored in hybrid-parallel training [36]. A layout 𝑝 ckpt = (𝜋1, . . . , 𝜋𝐿 ) specifies each layer’s checkpoint granularity, activation retention, and offload ratio. Checkpoint families and boundary masks or operator presets determine which tensors are saved and what is replayed. Retaining previously offloaded activations on GPU removes their D2H/H2D demand and reload dependencies; recomputation lowers memory demand but adds computation. Saving additional intermediates reduces replay. Candidate policies trade GPU memory usage against transfer demand, recomputation cost, and exposed execution work. Placement and granularity. Figure 5 separates checkpointing families. Each combines with an offload ratio 𝑟 ∈ [0, 1] that controls activation placement, targeting a fraction 𝑟 of saved bytes moved to CPU at tensor granularity. The ratio trades GPU residency for CPU–GPU transfer demand without changing the chosen recomputation policy, and applies across full-layer, segment-level, operator-level, and no-checkpoint policies. Full-layer checkpointing reconstructs intermediates from saved layer inputs [1]. Segment-level checkpointing uses seven-bit checkpoint-boundary masks to define finer replay boundaries, retaining segment inputs and cross-boundary state needed to reconstruct intermediates. A compact mask set exposes meaningful boundaries in the fused decoder. Both use non-reentrant checkpointing, stopping replay once required tensors are recovered [22]. Selective activation checkpointing (SAC) caches chosen operator outputs and recomputes other operations [22]. Segment masks and SAC offer complementary choices over execution boundaries and operator outputs, varying retained memory, replay scope, and transient tensor lifetimes. No-checkpoint policies use ordinary autograd without replay. Measured candidates. AutoPolicy instantiates a finite set of masks, operator presets, and offload ratios, merging ratios that select the same saved tensors within a checkpoint configuration. Figure 4 shows measured single-layer memory– compute tradeoffs for two models, identifying Pareto-efficient choices at different retained-memory budgets. AutoPolicy selects candidates based on measured pipeline effects rather than memory reduction alone. It calibrates frontier and reference candidates in the runtime, accounting for transient replay peaks before allocating layers. Layout across layers. Static tensor specifications identify saved tensors moved to pinned host buffers, with crosslayer reload prefetch. The planner groups offloaded layers contiguously and places GPU-retained policies toward the decoder tail, ordered from lighter to heavier memory use.
𝑓 (𝑣) = 𝑑 𝑣 + max {𝑎 𝑣 } ∪ {𝑓 (𝑢) : 𝑢 ∈ P (𝑣)} . Here 𝑎 𝑣 is the release time, 𝑑 𝑣 excludes predecessor waits, and P (𝑣) = P𝐺 (𝑣) ∪ P𝑅 (𝑣) ∪ P𝐵 (𝑣). The three sets encode data and parameter-version dependencies, ordering on serialized streams or workers, and buffer-release dependencies, respectively. Collective readiness includes all participating ranks. Shared bandwidth changes service times rather than imposing one serial order on all CPU and DMA operations. For example, a layer’s backward computation waits for its parameters, required activations, and the preceding GPU computation. Sharded delivery can advance parameter readiness, but only after all-gather reconstructs the layer. This advances backward execution only while parameter readiness is the latest predecessor; activation reload or GPU computation can instead determine its start. Buffer availability and the prior update further constrain when delivery can begin. The model thus identifies which readiness constraint a route change must advance to accelerate backward execution. The graph extends across iterations: the next dispatch of a layer waits for its update. Steady-state step time is the interval between successive iteration boundaries after warmup. The final update drain is a separate completion interval, rather than a cost charged in full to every step. Regimes and scaling. When compute and host/transfer service dominate, their interaction has the coarse approximation 𝑇step ≈ max{𝐶, 𝐻 } + 𝐸. 𝐶 includes forward, backward, and replay computation; 𝐻 is the bottleneck in host and transfer service. 𝐸 accounts for additional delay caused by execution dependencies and pipeline boundaries. The dependency model captures buffer waiting and rank readiness. In the host-bound regime (𝐻 > 𝐶), reducing 𝐻 by 𝛿𝐻 ≥ 0 saves min(𝛿𝐻, 𝐻 −𝐶) within this approximation when 𝐶 and 𝐸 remain fixed. Further host savings become hidden once compute dominates. Reducing replay instead lowers 𝐶 and can help GPU-bound execution. A route or activation change can affect both demands and readiness; AutoPolicy therefore measures its effect in the pipeline. Strong scaling gives 𝐶 (𝑅) = 𝐶 (𝐵 global /𝑅, 𝑆, 𝑝 ckpt ), where 𝑆 is sequence length and 𝑝 ckpt the activation layout. If compute efficiency is stable, 𝐻 nearly constant, and boundary costs small, 𝐶 (𝑅) ≈ 𝐶 (1)/𝑅 gives a crossover 𝑅cross ≈ 𝐶 (1)/𝐻 . Under weak scaling, sharding can shorten per-rank H2D service at fixed local batch. Aggregate throughput can then grow superlinearly while H2D limits the step, until shared host capacity or all-gather takes over. Section 4.5 examines these implications through GPU-count scaling and route comparisons across computation windows and topologies. 7
Decoder computation
Forward + backward (ms)
(a) Qwen3-14B, BS128
(b) Llama-3.1-8B, BS128
Full
390
layer
360
Segment
560
x Norm
330
SAC
Norm
480
x
None
440 10
20
0
Retained activations (GiB) SAC
Segment ckpt.
8
16
full ckpt
Norm
no replay
Retained activations (GiB) no ckpt
Norm
Gate/ up
Act × Down + res
Q/K/V
Attn
O+ res
Norm
Gate/ up
Act × Down + res
Q/K/V
Attn
O+ res
Norm
Gate/ up
Act × Down + res
reuse cached outputs; replay other operators
300 0
O+ res
saved boundary states; replay within each segment Norm
operators
Attn
replay layer internals from saved input x
x
boundaries
520
Q/K/V
Q/K/V
Attn
O+ res
Norm
Gate/ up
Act × Down + res
save the tensors required by ordinary autograd
Pareto
boundary state
saved tensors
recompute as needed
Figure 4. Single-layer checkpointing Pareto frontiers for both models (S1024; H100 in (a)).
Figure 5. Checkpoint granularities and recomputation within one decoder layer.
Backward reaches this tail first, releasing large retained states early. Selected layouts must preserve pipeline feasibility; final placement and memory requirements are validated in the complete runtime.
Algorithm 1 Measurement-guided AutoPolicy in SlideDP
3.6
Require: Workload W; topology T ; memory budgets 𝐵 Ensure: Memory-validated policy Π 1: (𝑝 param, 𝑝 red ) ← ProbeRoutes(W, T ) 2: (𝑐 H2D, 𝑐 D2H ) ← TuneChunks(W, 𝑝 param, 𝑝 red ) 3: 𝜌 ← (𝑝 param, 𝑝 red, 𝑐 H2D, 𝑐 D2H ) 4: C0 ← ProfileCkpt(W) 5: C ← ShortlistAndProbe(C0, 𝜌) 6: 𝑛 ← SolveMILP(C, 𝐵) 7: (C, 𝐵) ← CalibrateSelectedSpans(C, 𝑛, 𝐵) 8: 𝑝 ckpt ← PlaceLayers(SolveMILP(C, 𝐵)) 9: Π ← ValidateAndRefine(𝜌, 𝑝 ckpt, C, 𝐵) 10: return Π or a feasible conservative plan
Measurement-Guided Policy Selection
AutoPolicy selects Π = (𝑝 param, 𝑝 red, 𝑐 H2D, 𝑐 D2H, 𝑝 ckpt ) for a workload and topology. The analytical model identifies how these choices affect demand and readiness; execution measurements supply the selector’s costs. Routes and chunks are selected first, then activation choices are calibrated under that configuration. This staging controls selection overhead and captures their interaction with parameter and gradient transfers. Algorithm 1 summarizes the stages. Measure pipeline effects. Route probes compare parameter delivery and gradient return; chunk probes measure their preparation/update overlap. Single-layer profiling records computation time, retained activations, offload traffic, and peak-memory metadata. A reduced real training pipeline then replaces 𝑘 𝑗 reference layers with each shortlisted candidate 𝑗. Let 𝑇 𝑗 (𝑘 𝑗 ) and 𝑇0 denote the maximum per-rank mean steady-state step times of the candidate and matching reference pipelines. With common reference-layer cost 𝜏0 , the calibrated coefficient is 𝑇 𝑗 (𝑘 𝑗 ) − 𝑇0 𝜏 𝑗 = 𝜏0 + , 𝑘𝑗
and transient headroom. For selected candidates J+ = { 𝑗 : 𝑛 𝑗 > 0}, the GPU constraint is ∑︁ 𝑛 𝑗 𝑚 𝑗 + max 𝑝 𝑗 + max ℎ 𝑗 ≤ 𝐵 avail . 𝑗
𝑗 ∈ J+
𝑗 ∈ J+
𝐵 avail subtracts the reference peak reserved memory and a safety margin from GPU capacity. Shared pools and transient headroom are thus reserved once at their maximum requirements. Any configured host-memory and activation-transfer budgets add constraints; candidate count limits bound extrapolation from measured spans. The placement rule above converts counts into 𝑝 ckpt . The objective is a calibrated cost surrogate under the selected routes and chunks. Validation and selection overhead. Short full-model runs validate memory feasibility and compare step times. Reserved-memory feedback adjusts the budget for a bounded number of re-solves, retaining the fastest validated feasible configuration. Formal throughput runs then assess the selected policy’s performance. Transfer and activation measurements can be cached with matching hardware, workload, and implementation identities. With selection overhead 𝑊 and per-step saving Δ𝑡 > 0, the break-even count is ⌈𝑊 /Δ𝑡⌉.
The measured difference captures marginal pipeline cost (overlap and contention). Since total layer count is fixed, 𝜏0 does not affect the allocation ranking. Larger-span probes refine selected candidates before final allocation. Memory calibration separates per-layer residency, shared prefetch pools, and transient headroom, with guards across ranks. Í Allocate and place. The MILP minimizes 𝑗 𝑛Í 𝜏 𝑗 𝑗 , where 𝑛 𝑗 is the integer layer count of candidate 𝑗 and 𝑗 𝑛 𝑗 = 𝐿. Relative to the same measured reference, 𝑚 𝑗 , 𝑝 𝑗 , and ℎ 𝑗 denote increments in per-layer residency, shared-pool memory, 8
SlideDP-fixed SlideDP OOM
6000
4000
2000
0 250 200 150 100 50 0
15000
60
10000 5000 0
4
8
16
32
64
1200
RoundPipe ZeRO-Offload MegaTrain SlideDP OOM
75
20000
RoundPipe ZeRO-Offload
MegaTrain FSDP2
Megatron SlideFormer
SlideDP-fixed SlideDP
OOM
GPU Memory (GiB)
RoundPipe ZeRO-Offload MegaTrain SlideFormer
Throughput (tokens/s)
8000
Host PSS (GiB)
Host PSS (GiB)
Throughput (tokens/s)
10000
800
H100 limit (80 GiB)
45
30 RTX 4090 limit (24 GiB)
15
400 0
16
32
64
128
256
Per-GPU Batch Size
Per-GPU Batch Size
0
8
16
32
64
128
256
Per-GPU Batch Size
Figure 6. Batch-size scaling and CPU memory Figure 7. Batch-size scaling and CPU memory Figure 8. GPU memory vs. for Qwen3-8B on 4×RTX 4090. for Qwen3-14B on 4×H100. batch size for Qwen3-14B. Section 4.6 reports measured selection overhead and amortization alongside the gains over the fixed policy.
4
with four-way tensor parallelism (TP4) [29]. Fixed-workload comparisons match global batch; TP4 realizes it through resident microbatch and accumulation settings and plots global batch divided by four. Model-size sweeps use each system’s largest evaluated feasible batch, annotated above each bar. Measurement protocol. Batch-scaling runs of SlideDP and SlideDP-fixed normally use 15 iterations, discarding the first eight as warmup. Each such point reports the arithmetic mean of two runs. Throughput is measured after policy selection, whose overhead is reported separately in Section 4.6. Workload-specific measurement windows, repetition counts, and launch configurations accompany the plotting data. Variants and metrics. In SlideDP-fixed, the runtime uses full-layer checkpointing, activation offload, and fixed chunk sizes. SlideDP uses a selected policy in the paired comparisons. Throughput is global tokens per step divided by steadystate step time. Model FLOPS equal this rate times the training FLOPs per token, computed from actual projection dimensions with two operations per multiply–accumulate. The count includes forward, backward, and sequence-dependent attention, but excludes checkpoint replay, optimizer work, and transfers. GPU memory is CUDA peak reserved memory unless specified. Host memory is proportional set size (PSS). For SlideDP and SlideFormer, we sum sampled per-rank peaks; PSS apportions shared mappings across ranks.
Evaluation
We first evaluate throughput and capacity, then isolate runtime mechanisms, examine shared-host scaling, and assess policy gains and selection cost. MoE workloads and loss comparisons assess model coverage and correctness. 4.1
Experimental Setup
Platforms. We evaluate SlideDP on three multi-GPU platforms. The PCIe-only workstation has four RTX 4090 24GB GPUs, two Xeon Gold 6338N CPUs, and 1 TiB DDR4 memory. The NVLink-bridged server has eight A800 80GB PCIe GPUs, two Xeon Platinum 8358P CPUs, and 2 TiB DDR4 memory. The NVSwitch server has eight H100 80GB SXM GPUs, two Xeon Platinum 8468 CPUs, and 2 TiB DDR5 memory; we use up to four H100s. These platforms cover the three interconnect topologies in Figure 1(a–c). Implementations use PyTorch Distributed, NCCL, Transformers [32], FlashAttention2 [2], and compatible fused kernels from Liger Kernel [8] and Transformer Engine [20]. Workloads and baselines. We fine-tune all parameters with mixed precision [17], using BF16 computation and FP32 Adam states. Dense workloads span Qwen3-1.7B through 32B [33] and Qwen2.5-72B [23]; the host-memory sweep also includes Mistral-Small-24B [18]. Sequence length defaults to 1024; batches are per GPU unless marked global. Longsequence runs use full-length inputs, all attention positions enabled, and a position limit of 262,144. Baselines are SlideFormer [35], ZeRO-Offload [28], MegaTrain [37], and RoundPipe [15]. SlideFormer’s public multi-GPU implementation uses replicated parameter delivery, CPU gradient reduction, and full-layer checkpointing with activation offload; singleGPU points use its public single-GPU path. ZeRO-Offload uses DeepSpeed ZeRO-3 with CPU parameter and optimizer offload; RoundPipe uses synchronous mode. Baselines use compatible fused kernels where available. H100 comparisons also include GPU-resident FSDP2 [21, 39] and Megatron-LM
4.2
End-to-End Throughput
Batch-size scaling. We compare systems at matched batches (Figures 6 and 7). On four RTX 4090 GPUs, SlideDP achieves geometric-mean speedups of 1.83× over ZeRO-Offload and 1.48× over MegaTrain across batches 4–32. At batch 64, it reaches 9.81K tokens/s, while these baselines and RoundPipe exhaust GPU memory. On four H100s, the corresponding gains are 1.46× over ZeRO-Offload (batches 16–64) and 1.59× over MegaTrain (batches 16–128). At batch 256, SlideDP reaches 23.2K tokens/s and 1,048,576 tokens per step without gradient accumulation. Its throughput is 1.41× and 1.47× the baselines’ respective peaks in this sweep. 9
0
1.7
4
8
72
500
0
b128
b64
RoundPipe ZeRO-Offload MegaTrain
4
b128
b128
b32
b32 b32
b128
b64 b64
b64
b64 b128
b256
b32
b256 b256
b64 b128 b128
b256
b128
b256 b256
b128
b256
b64
b16
32
b128
b16
14
1000
b64
b32
b32
SlideFormer SlideDP Compute-only ref.
1500
FSDP2 Megatron SlideFormer
8
Model Size (B)
SlideDP Compute-only ref. OOM
14
32
Throughput (TFLOPS)
RoundPipe ZeRO-Offload MegaTrain
b8
100
Throughput (TFLOPS)
b32
b64 b64 b32
b32 b32 b32
b16
200
b64 b64
b64 b64
b128 b128 b64
b32 b32
b64 b32
300
b64
Throughput (TFLOPS)
400
2000
b256
2500 500
FSDP2 Megatron
SlideFormer SlideDP-fixed
SlideDP OOM
128K
256K
1500 1000 500 0
72
ZeRO-Offload MegaTrain
2000
Model Size (B)
16K
32K
64K
Sequence length (BS1 per GPU)
Figure 9. Model-size scaling throughput Figure 10. Model-size scaling throughput Figure 11. Sequence-length scaling for for Qwen models on 4×RTX 4090. for Qwen models on 4×H100. Qwen3-14B on 4×H100. 100
4.3
512K
RoundPipe ZeRO-Offload MegaTrain SlideDP Measured Estimated
75
50
Sequence length (BS1)
Model size (B)
Model-size scaling. Figures 9 and 10 compare model throughput at each system’s largest evaluated feasible batch. On RTX 4090, SlideDP sustains 429–470 TFLOPS across 1.7B–32B models; Qwen3-32B achieves 4.24× ZeRO-Offload’s throughput at batches 32 and 8, respectively. Qwen2.5-72B reaches 240 TFLOPS, 1.56× SlideFormer’s throughput; the other three offloading baselines are marked OOM. On H100, SlideDP sustains 1.8–2.0 PFLOPS across 4B–72B models, reaching 1.88× ZeRO-Offload’s throughput at 72B with a fourfold larger batch. At the plotted batches, it matches FSDP2 within 1% for 4B and 8B and exceeds it by 11.2% for 14B (batches 256 and 32, respectively). The dashed H100 reference projects single-GPU layer/stage compute measurements to four GPUs; FSDP2 and TP4 are measured end-to-end GPUresident comparisons. These results show that layer streaming expands the feasible workload range while maintaining competitive throughput against GPU-resident systems. Long sequences. Long sequences increase activation pressure and reduce available memory headroom, making memory placement and pipeline coordination more challenging. Figure 11 varies full attention length for Qwen3-14B at one sequence per GPU. At 64K, SlideDP reaches 2.13 PFLOPS versus 1.64 for ZeRO-Offload and 1.34 for MegaTrain (1.30× and 1.60×). At 128K, it reaches 1.95 PFLOPS versus MegaTrain’s 1.41 PFLOPS; it also completes 256K at 1.86 PFLOPS. FSDP2 and TP4 are marked infeasible at 64K. Crosses denote measured OOMs or longer sequences beyond a measured same-configuration OOM boundary. Together, these sweeps demonstrate matched-workload throughput gains and a broader feasible workload range.
25
0
0
500
1000
Host PSS (GiB)
256K 128K 64K 32K 16K
4
8
14
32
72
Model size (B)
Figure 12. Host memory and estimated sequence capacity. Figures 6 and 7 report host PSS of 106–109 GiB for SlideDP, 183 GiB for ZeRO-Offload, and 211–212 GiB for MegaTrain on RTX 4090 with Qwen3-8B at batches 4–32. On H100 with Qwen3-14B at batches 16–64, SlideDP uses about 202 GiB versus ZeRO-Offload’s 342 GiB. Shared master weights and optimizer states bound persistent host storage, while GPUretained activations avoid additional host copies. Figure 12(a) fits measured BS1 host footprints through 32B: slopes are 14.2 GiB per billion parameters for SlideDP and 19.8–22.9 GiB for other systems. Solid lines connect measurements; dashed lines extrapolate these slopes. The lower growth rate supports larger models within a fixed host-memory budget. Panel (b) estimates sequence capacity from measured working sets at one sequence per GPU, with 78 GiB per GPU and 1984 GiB host budgets. It scales activation storage with token count and accounts for RoundPipe’s pipeline microbatches. MegaTrain uses sampled GPU memory peaks; other systems use CUDA reserved peaks. Projected 14B capacities are 276K tokens for SlideDP, 113K for ZeRO-Offload, and 144K for MegaTrain; Figure 11 reports SlideDP’s successful 256K training run.
Memory Efficiency and Capacity
With full-layer checkpointing and activation offload, SlideDP reduces Qwen3-14B peak reserved memory by 51–59% versus ZeRO-Offload and 29–42% versus MegaTrain at common feasible batches (Figure 8). Batch 256 uses 72.4 GiB and is four times ZeRO-Offload’s largest feasible batch and twice MegaTrain’s in this sweep. The bounded window frees HBM for larger workloads and for retaining activations to reduce transfers and replay (Section 4.6).
4.4
Shared-Host Runtime Mechanisms
Reducing shared-host demand. The fixed-policy comparisons in Figures 6 and 7 evaluate the shared-host datapath. On RTX 4090, both paths use replicated parameter delivery and full-layer checkpointing with activation offload. SlideDP 10
Table 3. AutoPolicy ablation across platforms and workloads. We report the change from SlideDP-fixed to SlideDP. F-r1 and F-r0 denote full-layer checkpointing with and without activation offload. S-r0 and SAC-r0 denote segment-level and selective activation checkpointing without offload; N denotes no checkpointing. Unlisted layers use F-r1. Platform / workload
AutoPolicy layout
Chunk (MiB)
Throughput (Token/s)
GPU mem. (GiB)
CPU mem. (GiB)
4 × RTX 4090, Qwen3-8B 16 4 × RTX 4090, Qwen3-8B 32 4 × RTX 4090, Qwen3-8B 64
Batch Base layout 36 F-r1 36 F-r1 36 F-r1
36 F-r0 30 F-r0 + 3 SAC-r0 31 F-r1 + 5 SAC-r0
64/64 → 8/64 64/64 → 8/16 64/64 → 8/16
6345 → 7198 (+13.4%) 9161 → 9415 (+2.8%) 9684 → 9740 (+0.6%)
6.0 → 10.2 8.8 → 22.1 14.5 → 22.0
124.0 → 106.0 (-14.5%) 142.2 → 109.2 (-23.2%) 178.0 → 168.3 (-5.5%)
4 × H100, Qwen3-14B 4 × H100, Qwen3-14B 4 × H100, Qwen3-14B
32 64 128
40 F-r1 40 F-r1 40 F-r1
31 F-r0 + 9 N 64/64 → 128/128 17370 → 19530 (+12.4%) 37 F-r0 + 3 SAC-r0 64/64 → 32/128 21658 → 22635 (+4.5%) 15 F-r1 + 25 F-r0 64/64 → 8/128 22033 → 22657 (+2.8%)
15.0 → 76.6 22.8 → 76.4 39.4 → 75.9
261.9 → 200.6 (-23.4%) 341.6 → 200.6 (-41.3%) 501.6 → 313.1 (-37.6%)
4 × A800, Qwen3-4B 4 × A800, Qwen3-8B 4 × A800, Qwen3-32B
64 64 64
36 F-r1 36 F-r1 64 F-r1
11.1 → 77.3 14.9 → 69.0 26.5 → 71.8
131.7 → 94.9 (-28.0%) 179.6 → 167.3 (-6.8%) 638.7 → 585.2 (-8.4%)
10 F-r0 + 10 N 29 F-r1 + 7 N 13 F-r0 + 3 N
64/64 → 32/128 64/64 → 128/64 64/64 → 32/128
aggregates gradients on GPU before their return, replacing per-rank D2H transfers and CPU aggregation with one aggregated gradient per layer. On H100, the fixed runtime also uses sharded delivery over NVSwitch. These changes reduce shared-host service demand while retaining the same activation policy. At batches 4–16 on RTX 4090, SlideDP-fixed achieves 2.32–2.54× SlideFormer’s throughput; at batches 16–32 on H100, it achieves 4.03–4.18×. Larger batches provide more GPU computation to hide host work, narrowing the gap as described by the exposure model (Section 3.4).
26805 → 28964 (+8.1%) 15737 → 16281 (+3.5%) 3848 → 3923 (+1.9%)
prefetching and CPU allocation fixed. The four-GPU gain demonstrates that chunk-level overlap improves end-to-end execution with multiple ranks sharing host resources. 4.5
Understanding Shared-Host Scaling
The computation-window effect also appears in the sequence sweep (Figure 11): longer sequences narrow the gap between SlideDP-fixed and SlideFormer. We next examine the model in Section 3.4 by varying GPU count and delivery routes. (a) Qwen3-32B, 64 per GPU
136ms
Time (ms)
1.0×
100
92ms
1.3×
50
51ms Mono (pipeline) Chunked (pipeline) Isolated (no contention)
0
Convert + H2D (fp32→bf16 + DMA)
D2H + Update (DMA + Adam)
2.5
2.6× 2.3×
2.0 1.5
1.3×
1.3×
1.0 0.5 0.0
Mono Chunked bs=8 bs=8
Mono Chunked bs=32 bs=32
10000
Throughput (tokens/s)
1.3×
2.6× 128ms
3.0
Full Model (Qwen3-8B, bs=8) end-to-end training
Throughput (tokens/s)
150
H2D PCIe Interference (lower = better)
H2D Slowdown vs Isolated
Pipeline Component Breakdown (single layer, 64 MiB chunks)
4000 +7.7%
3000 2000
+33.6%
1000 0
Mono Chunked
Mono Chunked
1×4090
4×4090
ZeRO-Offload MegaTrain SlideFormer SlideDP OOM
20000
7500
15000
5000
10000
2500
0
5000
1
2
4
Number of GPUs
Figure 13. Chunked overlap: layer microbenchmarks and Qwen3-8B training on one and four RTX 4090 GPUs.
(b) Qwen3-14B, global batch 256
ZeRO-Offload MegaTrain SlideDP
8
0
1
2
4
8
Number of GPUs
Figure 14. A800 scaling: (a) Qwen3-32B, 64 sequences per GPU; (b) Qwen3-14B, global batch 256.
Chunked pipeline overlap. After reducing shared-host demand, SlideDP uses chunking to overlap the remaining transfers and updates (Section 3.3). Figure 13 evaluates this overlap within a layer. In the single-layer experiment with 64 MiB chunks, conversion plus H2D takes 50.8 ms under pipeline contention, down from 127.8 ms with whole-layer transfers. D2H plus CPU update falls from 135.9 to 91.7 ms. Across the two computation windows in the middle panel, chunking reduces H2D slowdown relative to isolated execution from 2.3–2.6× to approximately 1.3×. Shorter service paths advance parameter and update readiness, reducing exposed waits in the layer pipeline (Section 3.4). The right panel evaluates Qwen3-8B at batch eight per GPU, using full-length inputs on four GPUs and padded variable-length inputs on one. Chunking increases throughput from 3,218 to 3,465 tokens/s on four RTX 4090s (+7.7%) and from 1,386 to 1,852 tokens/s on one (+33.6%), averaging two runs per configuration. Each pair keeps layer-wise
GPU count and computation windows. Figure 14 separates weak and strong scaling. For Qwen3-32B at 64 sequences per GPU, SlideDP reaches 7,963 tokens/s on eight A800s, versus 5,478 for ZeRO-Offload and 5,396 for MegaTrain. Parallel efficiency of 95–98% across two, four, and eight GPUs indicates near-linear weak scaling in the GPU-bound regime. For Qwen3-14B at global batch 256, SlideDP-fixed achieves 1.94× scaling from four to eight GPUs (8.84K to 17.18K tokens/s), versus 1.04× for SlideFormer. Figure 1(e) reports excess step time over the matched compute reference as an aggregate indicator of exposed execution cost. At eight A800s and global batch 256, the excess is 0.90 s for SlideDP, 28.57 s for SlideFormer, and 9.77 s for MegaTrain. Low exposed cost lets the fixed runtime sustain scaling as each rank’s computation window shrinks. At fixed local batch, SlideDP keeps excess time at roughly 1–2 s across one to eight GPUs (Figure 1(d)). 11
(a) GPU-bound (per-GPU batch 64/256)
18000
full delivery
shard +0.8%
shard +1.5%
full +1.0%
full +0.3%
17000
16000
b64 NVLink
b64 PCIe-only
Table 4. MoE throughput on four H100 GPUs (tokens/s).
(b) host-bound (per-GPU batch 8/16)
sharded delivery
b256 NVLink
b256 PCIe-only
8×A800, per-GPU batch × interconnect
Throughput (tokens/s)
Throughput (tokens/s)
19000
shard +16.8%
8000 6000 4000
Model / per-GPU 𝐵 × 𝑆
full +26.0%
shard +32.7%
Qwen3.6-35B-A3B 32 × 4096 Gemma4-26B-A4B 64 × 1024
full +19.2%
2000 0
b8 NVLink
b8 PCIe-only
b16 NVLink
SlideDP
ZeRO-Offload
MegaTrain
44,715
13,931
26,102
27,347
13,822
18,993
b16 PCIe-only
4×A800, per-GPU batch × interconnect
Figure 15. Qwen3-14B route sensitivity on A800 with and without NVLink bridges. at batch 64, five SAC-r0 layers retain selected intermediates alongside 31 F-r1 layers, yielding a 0.6% gain and 5.5% less host PSS. On A800 at batch 64, selected configurations gain 1.9–8.1% throughput with 6.8–28.0% less host PSS under fixed sharded delivery and GPU reduction. Activation placement. Each checkpoint family combines with an offload ratio (Section 3.5). The selected layouts offload compact full-layer inputs while retaining larger intermediate sets from finer-grained policies on GPU. Retaining these intermediates reduces replay without adding activation traffic to the PCIe paths used by parameters and gradients. Selection overhead and amortization. On four H100s with Qwen3-14B and 32 sequences per GPU, a separate selection run reuses a cached single-layer activation profile and takes 1218 s (20.3 min) for route/chunk probes, pipeline calibration, and final validation. It selects 40 F-r0 layers and 32/128 MiB H2D/D2H chunks. With routes and chunks fixed, its final validation reduces step time from 7.797 s with F-r1 to 6.828 s with the selected layout. The measured 0.970 s saving per step yields an estimated break-even of 1,257 steps (2.72 hours of baseline training). This one-time cost would account for 1.4% of a 24-hour training budget.
Topology and communication routes. Figure 15 compares replicated and sharded parameter delivery for Qwen314B on A800. In host-bound four-GPU workloads, sharded delivery improves throughput by 16.8–32.7% with NVLink bridges; after their removal, replicated delivery improves it by 19.2–26.0%. Sharding trades repeated H2D for GPU-side reconstruction, whereas replication avoids reconstruction on a slower GPU path. In GPU-bound eight-GPU workloads, all route/topology combinations remain within 2.5% at each batch. Route choice therefore depends jointly on topology and the computation available to hide communication. 4.6
Measurement-Guided Policy Selection
We compare SlideDP with SlideDP-fixed under the same shared-host runtime to measure the combined benefit of selecting routes, chunk sizes, and activation layouts (Section 3.6). Routes and chunks are probed first; shortlisted activation candidates are then calibrated in a reduced pipeline before budgeted allocation, validation, and bounded refinement. Table 3 compares selected chunks and layouts with matched fixed controls in a separate measurement series from the batch-scaling curves. Selected layouts and gains. In Figure 7, policy selection raises Qwen3-14B throughput on H100 at batch 16 from 10,231 to 12,832 tokens/s (+25.4%). At 64K, policy selection gains 21.7% over SlideDP-fixed (Figure 11). Table 3 illustrates how the selected layouts use HBM to reduce different costs. On RTX 4090 at batch 16, retaining all 36 layer inputs removes activation offload/reload traffic; together with 8/64 MiB H2D/D2H chunks, the configuration gains 13.4% throughput and reduces host PSS by 14.5%. On H100 at batch 32, combining 31 F-r0 layers with nine no-checkpoint layers eliminates replay in the latter. This configuration uses 76.6 GiB versus 15.0 GiB for the fixed policy and gains 12.4% throughput, turning the bounded window’s memory headroom into faster execution. Larger batches use mixed layouts as activation demand grows. On H100 at batch 128, 25 F-r0 layers retain their inputs while 15 F-r1 layers offload them; the configuration gains 2.8% throughput and reduces host PSS by 37.6%. On RTX 4090
4.7
MoE Training and Correctness
The shared-host DP runtime also supports MoE layers. Table 4 compares Qwen3.6-35B-A3B [24] and Gemma4-26BA4B [7] on four H100 GPUs at matched batch sizes and sequence lengths across systems. SlideDP reaches 44.7K and 27.3K tokens/s, respectively. It delivers 1.98–3.21× ZeROOffload’s and 1.44–1.71× MegaTrain’s throughput. Correctness. We assess dense-model correctness by comparing SlideDP with GPU-resident FSDP2 over 200 training steps of Qwen3-14B on four H100s, using a per-GPU batch size of 16 and sequence length 1024. We match the initial weights, C4 samples [25] and their order, and optimizer settings. All ranks complete with finite loss; the maximum absolute differences in global-batch mean loss are 5.2 × 10−4 for F-r1 and 6.3 × 10−4 for the selected layout; the loss curves followed consistent trajectories. The close agreement with FSDP2 in BF16 training loss supports the correctness of SlideDP’s synchronous execution (Section 3.3). 12
5
[12] Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, and Soumith Chintala. 2020. PyTorch distributed: experiences on accelerating data parallel training. Proc. VLDB Endow. 13, 12 (Aug. 2020), 3005–3018. doi:10.14778/3415478.3415530 [13] Changyue Liao, Mo Sun, Zihan Yang, Jun Xie, Kaiqi Chen, Binhang Yuan, Fei Wu, and Zeke Wang. 2024. LoHan: Low-Cost HighPerformance Framework to Fine-Tune 100B Model on a Consumer GPU. arXiv:2403.06504 [cs.DC] doi:10.48550/arXiv.2403.06504 [14] Qijun Luo, Hengxu Yu, and Xiao Li. 2024. BAdam: A Memory Efficient Full Parameter Optimization Method for Large Language Models. In Advances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Eds.), Vol. 37. Curran Associates, Inc., 24926– 24958. https://proceedings.neurips.cc/paper_files/paper/2024/file/ 2c570b0f9938c7a58a612e5b00af9cc0-Paper-Conference.pdf [15] Yibin Luo, Shiwei Gao, Huichuan Zheng, Youyou Lu, and Jiwu Shu. 2026. Efficient Training on Multiple Consumer GPUs with RoundPipe. arXiv:2604.27085 [cs.DC] doi:10.48550/arXiv.2604.27085 [16] Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, Sayak Paul, and Benjamin Bossan. 2022. PEFT: State-of-theart Parameter-Efficient Fine-Tuning methods. https://github.com/ huggingface/peft. [17] Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. 2018. Mixed Precision Training. In International Conference on Learning Representations. https://openreview.net/forum?id=r1gs9JgRZ [18] Mistral AI. 2025. Mistral-Small-24B-Instruct-2501. https:// huggingface.co/mistralai/Mistral-Small-24B-Instruct-2501 [19] Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R Devanur, Gregory R Ganger, Phillip B Gibbons, and Matei Zaharia. 2019. PipeDream: Generalized pipeline parallelism for DNN training. In Proceedings of the 27th ACM symposium on operating systems principles. 1–15. [20] NVIDIA. 2024. Transformer Engine: A Library for Accelerating Transformer Models on NVIDIA GPUs. https://github.com/NVIDIA/ TransformerEngine. Accessed on 2026-09-25. [21] PyTorch Contributors. 2026. torch.distributed.fsdp.fully_shard: PyTorch 2.11 Documentation. https://docs.pytorch.org/docs/2.11/ distributed.fsdp.fully_shard.html Accessed: 2026-09-24. [22] PyTorch Contributors. 2026. torch.utils.checkpoint — PyTorch Documentation. https://docs.pytorch.org/docs/stable/checkpoint.html. Accessed: 2026-06-08. [23] Qwen Team. 2025. Qwen2.5 Technical Report. arXiv:2412.15115 [cs.CL] doi:10.48550/arXiv.2412.15115 [24] Qwen Team. 2026. Qwen3.6-35B-A3B Model Card. https:// huggingface.co/Qwen/Qwen3.6-35B-A3B Accessed: 2026-09-24. [25] Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research 21, 140 (2020), 1–67. https://www.jmlr.org/papers/v21/20-074.html [26] Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 1–16. [27] Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He. 2021. ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learning. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (St. Louis, Missouri) (SC ’21). Association for Computing Machinery, New York, NY, USA, Article 59, 14 pages. doi:10.1145/ 3458817.3476205
Conclusion
This paper presented SlideDP, a shared-host runtime for full-parameter LLM fine-tuning with synchronous data parallelism. Shared state and cross-rank pipelining reduce redundant host work and overlap the remaining transfers and updates with GPU computation. The analytical model explains when reducing shared-resource demand shortens a step; measurements guide communication and activation policies for each workload and topology. The broader design lesson is to manage communication and GPU memory according to their effects on exposed pipeline work. This coordination translates host-memory capacity into larger feasible training workloads while sustaining throughput across commodity workstations and multi-GPU servers.
References [1] Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016. Training deep nets with sublinear memory cost. arXiv preprint arXiv:1604.06174 (2016). [2] Tri Dao. 2023. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning. arXiv:2307.08691 [cs.LG] doi:10.48550/ arXiv.2307.08691 [3] Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. Advances in neural information processing systems 35 (2022), 16344–16359. [4] Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. QLoRA: Efficient Finetuning of Quantized LLMs. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS ’23). Curran Associates Inc., Red Hook, NY, USA, Article 441, 28 pages. [5] Jiarui Fang, Zilin Zhu, Shenggui Li, Hui Su, Yang Yu, Jie Zhou, and Yang You. 2023. Parallel Training of Pre-Trained Models via ChunkBased Dynamic Memory Management. IEEE Transactions on Parallel and Distributed Systems 34, 1 (2023), 304–315. doi:10.1109/TPDS.2022. 3219819 [6] Yangyang Feng, Minhui Xie, Zijie Tian, Shuo Wang, Youyou Lu, and Jiwu Shu. 2023. Mobius: Fine Tuning Large-Scale Models on Commodity GPU Servers. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 (ASPLOS ’23). ACM, 489–501. doi:10.1145/3575693.3575703 [7] Google DeepMind. 2026. Gemma 4 26B A4B Model Card. https: //huggingface.co/google/gemma-4-26B-A4B Accessed: 2026-09-24. [8] Pin-Lun Hsu, Yun Dai, Vignesh Kothapalli, Qingquan Song, Shao Tang, Siyu Zhu, Steven Shimizu, Shivam Sahni, Haowen Ning, and Yanning Chen. 2024. Liger Kernel: Efficient Triton Kernels for LLM Training. arXiv preprint arXiv:2410.10989 (2024). arXiv:2410.10989 [cs.LG] doi:10. 48550/arXiv.2410.10989 [9] Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: LowRank Adaptation of Large Language Models. In International Conference on Learning Representations. https://openreview.net/forum?id= nZeVKeeFYf9 [10] Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al. 2019. GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism. Advances in neural information processing systems 32 (2019). [11] Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014). 13
[28] Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. 2021. ZeRO-Offload: Democratizing Billion-Scale Model Training. In 2021 USENIX Annual Technical Conference (USENIX ATC 21). USENIX Association, 551–564. https://www.usenix.org/conference/ atc21/presentation/ren-jie [29] Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism. arXiv preprint arXiv:1909.08053 (2019). [30] Xiaoyang Sun, Wei Wang, Shenghao Qiu, Renyu Yang, Songfang Huang, Jie Xu, and Zheng Wang. 2022. STRONGHOLD: Fast and Affordable Billion-Scale Deep Learning Model Training. In SC22: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 1–17. doi:10.1109/SC41404.2022.00076 [31] Philippe Tillet, H. T. Kung, and David Cox. 2019. Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations. In Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages (MAPL 2019). Association for Computing Machinery, New York, NY, USA, 10–19. doi:10.1145/3315508.3329973 [32] Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Association for Computational Linguistics, Online, 38–45. https: //www.aclweb.org/anthology/2020.emnlp-demos.6 [33] An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, Le Yu, Lianghao Deng, Mei Li, Mingfeng Xue, Mingze Li, Pei Zhang, Peng Wang, Qin Zhu, Rui Men, Ruize Gao, Shixuan Liu, Shuang Luo, Tianhao Li, Tianyi Tang, Wenbiao Yin, Xingzhang Ren, Xinyu Wang, Xinyu Zhang, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yinger Zhang, Yu Wan, Yuqiong Liu, Zekun Wang, Zeyu Cui, Zhenru Zhang, Zhipeng Zhou, and Zihan Qiu. 2025. Qwen3 Technical Report. arXiv:2505.09388 [cs.CL] https: //arxiv.org/abs/2505.09388 [34] Hanmei Yang, Jin Zhou, Yao Fu, Xiaoqun Wang, Ramine Roane, Hui Guan, and Tongping Liu. 2026. ProTrain: Efficient LLM Training via Automatic Memory Management. In Proceedings of Machine Learning and Systems, A. Chowdhery and Z. Jia (Eds.), Vol. 8. MLSys, Bellevue, WA, USA, 1757– 1770. https://proceedings.mlsys.org/paper_files/paper/2026/file/ 56280ad5fb53967fd55e4ba1b1cfc418-Paper-Conference.pdf [35] Ruijia Yang and Zeyi Wen. 2026. An Efficient Heterogeneous CoDesign for Fine-Tuning on a Single GPU. arXiv:2603.16428 [cs.DC] doi:10.48550/arXiv.2603.16428 [36] Tailing Yuan, Yuliang Liu, Xucheng Ye, Shenglong Zhang, Jianchao Tan, Bin Chen, Chengru Song, and Di Zhang. 2024. Accelerating the Training of Large Language Models using Efficient Activation Rematerialization and Optimal Hybrid Parallelism. In 2024 USENIX Annual Technical Conference (USENIX ATC 24). USENIX Association, Santa Clara, CA, 545–561. https://www.usenix.org/conference/atc24/ presentation/yuan [37] Zhengqing Yuan, Hanchi Sun, Lichao Sun, and Yanfang Ye. 2026. MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU. arXiv:2604.05091 [cs.CL] doi:10.48550/arXiv.
2604.05091 [38] Jiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang, Anima Anandkumar, and Yuandong Tian. 2024. GaLore: Memory-efficient LLM Training by Gradient Low-rank Projection. arXiv preprint arXiv:2403.03507 (2024). [39] Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews, and Shen Li. 2023. PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel. Proc. VLDB Endow. 16, 12 (Aug. 2023), 3848–3860. doi:10.14778/3611540. 3611569
14