The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures
arXiv:2605.11999v1 [cs.DC] 12 May 2026
Bole Ma1[0009−0006−6536−1044] , Ayesha Afzal1[0000−0002−8632−3681] , Jan Eitzinger1[0009−0000−3350−3841] , and Gerhard Wellein2[0000−0001−7371−3026] 1
Erlangen National High Performance Computing Center, Erlangen, Germany 2 Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany {bole.ma, ayesha.afzal, jan.eitzinger, gerhard.wellein}@fau.de
Abstract. Power capping is the standard GPU energy lever in LLM serving, and it appears to work: throughput drops, power readings fall, and energy budgets are met. We show the appearance is illusory for the phase that dominates production serving: autoregressive decode. Across four attention paradigms—GQA, MLA, Gated DeltaNet, and Mamba2— on NVIDIA H200, decode draws only 137–300 W on a 700 W GPU; no cap ever triggers, because memory-bound decode saturates HBM bandwidth rather than compute and leaves power headroom untouched. Firmwareinitiated clock throttling compounds the illusion: these deviations can corrupt any throughput measurement that attributes them to the cap. SM clock locking dissolves both confounds. By targeting the lever that is actually on the critical path, clock locking Pareto-dominates power capping universally, recovering up to 32% of decode energy at minimal throughput loss. We identify three architecture-dependent DVFS behavioural classes and characterise a common energy pattern across novel attention replacements: a heavy prefill cost recouped by efficient decode, eventually halving total request energy relative to GQA at production batch sizes.
Keywords: LLM inference · power capping · SM clock locking · GPU energy efficiency · attention mechanisms · roofline model · DVFS
1
Introduction
Data centres increasingly rely on GPU power capping to manage the energy footprint of LLM inference [9,18]. The implicit model is straightforward: set a board-level watt ceiling, and the driver throttles the GPU to stay within it— trading some throughput for guaranteed power savings. This model is sound for compute-bound workloads that push the GPU near its thermal design power (TDP). We show it fails for the phase that dominates production LLM serving: autoregressive decode. The failure is structural, not incidental. Across four attention paradigms— GQA [2], MLA [6], Gated DeltaNet [21], and Mamba2 [5]—decode power draw
2
B. Ma et al.
on NVIDIA H200 ranges from 137 to 300 W, never approaching even our lowest 280 W cap on a 700 W-TDP GPU. The cap is inert: the driver holds ≈1830 MHz regardless of the configured limit, because memory-bound decode saturates HBM bandwidth, not compute, leaving the GPU’s power headroom untouched. Compounding this, the nvidia-smi --lock-gpu-clocks command silently clamps any requested lock ≥1830 MHz to ≈1830 MHz, distinct from the free-running boost that holds 1980 MHz indefinitely—and even this residual 240 MHz gap above 1590 MHz produces zero throughput gain at 7–13% more power, confirming that decode is entirely memory-paced. SM clock locking dissolves both confounds. By directly controlling the frequency lever that is on the critical path, static clock locking recovers up to 32% of decode energy at less than 1% throughput loss, Pareto-dominating power capping at every matched operating point. We present the first systematic energy characterisation under both mechanisms for these four paradigms—spanning dense attention, compressed-KV, linear-recurrent, and state-space designs—served out-of-the-box via vLLM on H200 SXM. A controlled GQA↔MLA ablation via TransMLA [12] on shared Minitron-4B [14] weights isolates the attention mechanism from confounding model differences. Our contributions: 1. The power-capping illusion. We demonstrate that power capping is structurally ineffective for memory-bound LLM decode regimes that dominate production serving: the GPU never reaches the cap, the driver ignores it, and clock throttling creates a confounding artefact. Table 1 exposes the gap between configured and actual GPU behaviour. 2. SM clock locking as the correct lever. Energy control must target the critical-path resource—for memory-bound decode, SM clock rather than aggregate power. Static clock locking Pareto-dominates power capping universally. We identify three architecture-dependent behavioural classes (batch-invariant, batch-sensitive, compute-light) and provide deployable per-architecture clock policies. 3. Cross-architecture energy landscape. Data-centre energy controls are misaligned with LLM inference because decode is memory-bound—a claim we validate across five architectures. Novel architectures (GDN, Mamba2, MLA) share a common pattern: heavy prefill cost recouped by efficient decode at long context and large batch. MLA’s KV compression crosses below GQA only beyond a batch-size-dependent context threshold; recurrent models cross after ∼1,000 output tokens.
2
Background and Related Work
2.1
Attention Architectures for LLM Inference
The four paradigms differ in what they store per token and how they compute each decode step. GQA caches full key–value pairs per group of heads and remains the dominant mechanism in mainstream open-weight transformers. MLA instead caches a compressed latent (576 dims vs. 2,048) and reconstructs KV via
The Illusion of Power Capping in LLM Decode
3
projections at decode time; introduced in DeepSeek-V2 [6], it has since propagated to models such as GLM and Mistral variants. GDN replaces attention with a linear recurrence over a fixed-size state; it is the attention mechanism in the Qwen3.5 series [21], a family known for competitive small-model releases. Mamba2 dispenses with attention entirely via SSM layers, achieving O(1) per-step decode cost; NVIDIA deploys Mamba2-hybrid architectures in its Nemotron line [5]. All models are ≈4B parameters. GQA-ctrl (Minitron-4B [14]) uses the same GQA mechanism as Qwen3-4B but serves as the controlled baseline for MLA: it and the TransMLA [12]-converted variant share base weights, differing only in the attention mechanism. The energy consequences follow from a simple question: which kernel class dominates, and is it compute- or memory-bound? 2.2
GPU Clock Scaling and Power Capping
NVIDIA GPUs expose two static energy levers via NVML. SM clock locking fixes the compute-core frequency while keeping HBM at its rated clock—the operator chooses the exact trade-off between speed and power. In principle, memory clock can also be adjusted, but querying the clock after setting it reveals no change: the driver silently ignores the request, leaving HBM locked at its rated frequency regardless. Power capping sets a board-level power ceiling and lets the driver select clocks dynamically—simpler but less predictable. Crucially, a power cap only constrains the GPU when actual power draw exceeds the cap; if the workload is light enough that the GPU stays below the limit, the cap is inert and the driver runs at its default clock—a distinction that will prove decisive (Section 5.2). We evaluate both; all settings are applied statically before serving. Prior work has characterised these levers for HPC kernels [3] and training [23], where the rule of thumb is simple: compute-bound kernels scale with clock, memory-bound kernels do not. But modern LLM inference mixes kernel types in architecture-dependent ways that this rule alone cannot predict. 2.3
Gap in Prior Work
The phenomenon underlying our result, low GPU power draw during decode, is not new; what is new is asking whether it invalidates power capping as an energy lever. POLCA [16] observed a two-stage power profile in LLM inference (high during prefill, low during decode) and concluded that the resulting headroom enables power oversubscription in inference clusters—a capacity-planning insight, not a critique of power capping itself. DVFS-based approaches [11,20] apply clock scaling to decode and demonstrate energy savings, but benchmark against default governor baselines rather than power capping, leaving open the question of whether the two mechanisms are even comparable for memory-bound decode. No prior work explicitly poses the question: is power capping effective for memorybound LLM decode? —and none gives the negative answer. Equally absent is a cross-architecture energy comparison. Existing inference studies focus on throughput and latency [8,10,13,4,19], not energy. The few studies that do report inference energy or apply DVFS restrict themselves to
4
B. Ma et al.
standard GQA/MHA transformers [16,11,20]; none examines novel attention replacements—MLA’s compressed KV path, linear-attention hybrids such as GDN, or SSM-based models such as Mamba2—where qualitatively different kernel types may alter the DVFS response entirely. No prior work asks how the interaction between SM clock and attention kernel type shapes per-architecture energy profiles, or whether MLA’s KV compression actually saves energy on real hardware.
3
Experimental Setup
3.1
Deployment Baseline and Hardware
Every model is served via vLLM [8] in BF16, taken directly from HuggingFace with no custom kernels—the scenario most practitioners actually face. We run on a single NVIDIA H200 SXM (HBM3e, 4.8 TB/s bandwidth, 989 TFLOPS BF16 dense peak, 700 W TDP). This single-card setup directly mirrors the decode pool model widely adopted in industry: disaggregated serving systems [17,25] route prefill and decode requests to dedicated GPU pools, so each decode-pool card sees a decode-only workload—exactly what we measure. Energy is measured via NVML power sampling at 50 ms intervals, integrated with the trapezoidal rule; for operations shorter than 100 ms (≈44% of prefill configs) we fall back to the product of snapshot power and wall-clock latency. Results are cross-validated against NVML hardware energy counters, which agree to within 2% for operations ≥200 ms but have millijoule-level granularity that makes them unreliable for short prefills. Nsight Compute (NCU) provides per-kernel roofline diagnostics. 3.2
Experimental Design
Our primary metric is energy per token (mJ/tok), measured separately for prefill and decode. For prefill, the denominator is the number of input tokens processed (i.e. prompt length × batch size); for decode it is the number of output tokens generated. We sweep five SM clock levels from 390 to 1980 MHz (with HBM held at rated speed), five power cap levels from 280 to 700 W, batch sizes from 1 to 32, and sequence lengths from 1K to 64K tokens. Each configuration is repeated 10–20 times; we report medians. Three warmup iterations precede every measurement run. Figures showing energy-per-token vs. sequence length include ±1 s.d. shaded bands. 3.3
Models and Controls
All models are ≈4B parameters. The most important design choice is a controlled pair: GQA-ctrl and MLA share the same Minitron-4B base weights [14], differing only in the attention mechanism. TransMLA [12] converts the checkpoint to activate vLLM’s compressed-cache path, caching a 576-dim latent per token instead of GQA-ctrl’s 2,048 dims—a 3.6× compression. This is the only controlled
The Illusion of Power Capping in LLM Decode H200 Roofline -- Decode (all architectures) Colour = architecture Shape = kernel class Area proportional to GPU time Architecture GQA (Qwen3-4B) GQA-ctrl (Minitron-4B) MLA (Minitron-4B) GDN (Qwen3.5-4B) Mamba2 (Nemotron-Nano)
100
H200 Roofline -- Prefill (all architectures) Colour = architecture Shape = kernel class Area proportional to GPU time
390 MHz
10
1 BW
1e-01
1e-02
=
T 4.8
s B/
Kernel class (shape) GEMM Attention SSM Elementwise Other
1e-01
SM clock ceiling 1980 MHz 1590 MHz 1185 MHz 780 MHz 390 MHz
1
Architecture GQA (Qwen3-4B) GQA-ctrl (Minitron-4B) MLA (Minitron-4B) GDN (Qwen3.5-4B) Mamba2 (Nemotron-Nano)
1000
1980 1590 MHz MHz 1185 MHz 780 MHz
10
100
1000
Achieved Performance (TFLOPS)
Achieved Performance (TFLOPS)
1000
5
100
1980 1590 MHz MHz 1185 MHz 780 MHz 390 MHz
10 /s
1 BW
1e-01
1e-02
=
4.8
TB
Kernel class (shape) GEMM Attention SSM Elementwise Other
1e-01
Arithmetic Intensity (FLOPs / byte)
SM clock ceiling 1980 MHz 1590 MHz 1185 MHz 780 MHz 390 MHz
1
10
100
1000
Arithmetic Intensity (FLOPs / byte)
Fig. 1. H200 roofline for all four paradigms (plus GQA-ctrl, the MLA ablation control): decode (left, BS = 1, seq = 1024) and prefill (right, BS = 1, seq = 4096). In decode, every kernel across all architectures clusters deep in the memory-bound region, orders of magnitude below the ridge (206 FLOPs/byte) and nowhere near the compute-bound ceiling— confirming that no decode workload approaches the condition under which a power cap would engage. In prefill, GDN’s elementwise and Mamba2’s SSM kernels remain memory-bound despite compute-bound GEMMs.
GQA↔MLA ablation in the literature; without it, any MLA comparison confounds attention mechanism with model weights, vocabulary, and training data. We formalise six testable hypotheses; four are confirmed and two require qualification (notably MLA only saves decode energy beyond a batch-size-dependent context threshold).
4
The Hardware Substrate
Before examining architecture-specific effects, we need to understand what the GPU is actually doing during inference. The H200’s roofline ridge sits at ≈206 FLOPs/byte: kernels below this threshold are waiting on memory and do not benefit from faster compute; kernels above it are compute-limited and scale with clock. Prefill processes all tokens in parallel via large GEMMs—solidly compute-bound. Decode at BS = 1 reduces each multiply to a matrix-vector operation—solidly memory-bound. This single fact determines everything that follows (Figure 1).
4.1
Decode Is Universally Memory-Bound
During decode, the GPU spends most of its time loading model weights from HBM. Tensor cores sit over 88% idle. The SM clock is simply not on the critical path, and reducing it saves roughly a quarter of the energy with negligible throughput loss—consistently across every architecture and batch size we tested (Figure 3).
6
4.2
B. Ma et al.
Batch Size Drives the Regime Boundary
Batching has a far larger effect than any DVFS or architecture choice: increasing BS from 1 to 32 reduces energy-per-token by over 20× by amortising the cost of loading weights. But batching also shifts the compute/memory balance in architecture-dependent ways, revealing three DVFS classes (Figure 2). Batchinvariant architectures (GQA, GQA-ctrl) stay memory-bound even at BS = 32 (a representative production serving regime); a single low clock works at all batch sizes. Batch-sensitive architectures (MLA, Mamba2) contain enough additional per-step work—KV decompression data movement for MLA, SSM scan compute for Mamba2—that large batches push them toward the compute-bound regime; the optimal clock must rise with batch size. Compute-light GDN (65% elementwise kernels, 1.8% tensor-core utilisation) has so little arithmetic intensity that even BS = 32 cannot shift the balance—it tolerates the most aggressive underclocking unconditionally. A practical ceiling on DVFS savings comes from the H200’s idle power floor (≈75 W): a 5× clock reduction yields only ∼1.5× power reduction, because DVFS controls only the dynamic component.
5
Power Capping vs. Clock Locking by Architecture
5.1
The DVFS Spectrum
Figure 2 maps the decode landscape. The headline result is uniformity: every architecture saves roughly a quarter of its decode energy by underclocking, because the dominant kernels in all decode paths are memory-bound. The subtlety emerges in the Pareto frontiers (Figure 3), which plot absolute throughput (tok/s) against tokens per joule (tok/J) for every clock and power-cap setting. At every configuration, SM clock locking dominates power capping: it reaches lower energy at comparable throughput. The power-cap curves appear degenerate—all five cap settings cluster at nearly identical throughput and energy—while clock locking traces a clean frontier. Both the cap’s inertness and a firmware-imposed clock throttle deserve detailed examination (Section 5.2). Among architectures, the pattern is intuitive: the less compute a decode path uses, the more energy underclocking saves. GDN—whose decode is two-thirds elementwise operations —benefits the most: clocking from 1980 to 780 MHz saves 49 W (30%) at zero throughput loss, reducing the GPU to only 117 W total draw. Mamba2 and GQA fall in between; MLA saves the least (though still 47 W / 24%), not because of extra GEMM compute—GEMM counts are nearly identical to GQA-ctrl—but because KV decompression emits hundreds of small cat/copy/reshape kernels per step that are insensitive to SM clock. 5.2
Why Power Capping Fails for Memory-bound Decode
Power capping is the standard energy management tool in data centres: set a board-level watt ceiling and let the driver manage clocks. For LLM decode, it is entirely ineffective.
The Illusion of Power Capping in LLM Decode
7
Decode DVFS: Optimal Clock, Savings, and Absolute Energy vs Batch Size (Nseq = 4096)
780
1185
1185
1980
GDN
390
390
390
780
780
1785
Mamba2
390
780
390
780
780
GQA-ctrl
780
1185
1185
1185
1185
1387 982
MLA
390
780
780
780
1185
1
4
8
16
32
Batch Size
GQA GDN MHz
780
Mamba2 GQA-ctrl
585 0
MLA
23%
23%
21%
18%
7%
±0.5%
±1.3%
±0.1%
±2.4%
±0.2%
27%
29%
26%
27%
25%
±1.4%
±0.8%
±0.3%
±1.1%
±0.5%
25%
24%
23%
25%
23%
±0.1%
±0.2%
±0.2%
±3.3%
±0.2%
20%
21%
18%
14%
8%
±0.1%
±0.2%
±0.1%
±0.2%
±0.2%
24%
23%
22%
24%
22%
±0.3%
±0.2%
±0.2%
±0.2%
±0.2%
1
4
8
16
32
Batch Size
Energy/Token at Optimal Clock (mJ) 40
GQA
1882
504
277
161
30
GDN
2964
806
428
227
129
20
Mamba2
1762
468
252
139
85.8
10
GQA-ctrl
1729
459
258
143
96.0
0
MLA
2088
542
280
148
83.6
1
4
8
16
32
Batch Size
107 10 3 mJ/tok
Energy Savings vs Best Power Cap (%)
780
%
Optimal SM Clock (MHz) GQA
10 2
Decode DVFS: Optimal Clock, Savings, and Absolute Energy vs Batch Size (Nseq = 16384)
1185
1185
1185
1980
GDN
390
390
390
780
780
1785
Mamba2
780
390
780
780
1185
GQA-ctrl
1185
1185
1185
1185
1185
1387 982
MLA
780
780
1185
1185
1185
1
4
8
16
32
Batch Size
585 0
GQA GDN MHz
1185
Mamba2 GQA-ctrl MLA
23%
18%
6%
7%
1%
±0.0%
±0.0%
±0.1%
±0.0%
±0.1%
27%
29%
27%
26%
19%
±0.0%
±0.0%
±0.8%
±0.0%
±0.7%
24%
25%
21%
24%
11%
±0.0%
±0.0%
±0.2%
±0.0%
±0.1%
20%
10%
9%
3%
-1%
±0.1%
±0.2%
±0.1%
±0.3%
±0.3%
19%
22%
22%
18%
7%
±0.1%
±0.1%
±0.1%
±0.2%
±0.1%
1
4
8
16
32
Batch Size
Energy/Token at Optimal Clock (mJ) 40
GQA
2017
629
415
289
30
GDN
3015
844
460
260
163
20
Mamba2
1793
482
265
152
99.8
10
GQA-ctrl
1815
575
358
269
222
0
MLA
2195
579
318
183
119
1
4
8
16
32
Batch Size
242 10 3 mJ/tok
Energy Savings vs Best Power Cap (%)
780
%
Optimal SM Clock (MHz) GQA
10 2
Fig. 2. Decode DVFS heatmaps: energy-optimal SM clock (left), SM clock-down supremacy over optimal power capping across all examined decode configurations (centre), and absolute energy per token (right). All energy saving results are rock-stable across repeated runs (max stddev ≤3%, typically <0.5%). The absolute energy per token (right) grows with sequence length for all architectures, as each decode step must stream an increasingly large KV cache from HBM; the penalty is strongly architecturedependent. GQA more than doubles (107 → 242 mJ/tok, 2.26×) from 4K to 16K, consistent with its O(L) KV bandwidth cost. MLA grows more modestly (1.42×) owing to compressed KV representations, while Mamba2—unburdened by any KV cache— rises only 1.16× (86 → 100 mJ/tok). At 16K sequence length—exceeding the median real-world conversation context by over 100× and 4× beyond the 4K threshold used to define “long-prompt” workloads in standard inference benchmarks [7]—where KV cache traffic saturates HBM bandwidth, GQA/GQA-ctrl savings collapse to −1–9% for large batch sizes, while Mamba2 surpasses MLA and achieves 99.8 mJ/tok at BS = 32.
Mismatch between configured power cap vs. actual power draw. Table 1 exposes the gap between what an operator configures and what the GPU does. Under every power-cap setting from 280 W to 700 W, the NVML-reported actual SM clock remains ≈1830 MHz and the actual power draw stays in the range 160– 300 W. The reason is elementary: the cap is a ceiling, not a target. Memory-bound decode never pushes the GPU hard enough to reach even our lowest cap, so the driver ignores it. The resulting throughput spread across caps (0.3–2.8%) exceeds per-setting noise (by 1.2–27×), but corresponds to only a few watts of incidental variation—operationally meaningless. The disguise of requested vs. actual SM clock. Even when we bypass power capping and directly lock the SM clock via nvidia-smi --lock-gpu-clocks, the H200 firmware imposes its own limit. Requesting 1980 MHz yields only ≈1830 MHz sustained; all settings ≤1590 MHz are honoured exactly. This is not thermal throttling (GDN at BS = 1 draws only 167 W at 42 ◦ C yet is still
8
B. Ma et al. Decode: Throughput vs Energy Efficiency 1K 8K 0.55
780 MHz 780 MHz 780 MHz 1185 MHz
390 MHz 390 MHz 390 MHz
0.50
tok / J
BS=1 (latency-sensitive)
GQA
0.45
390 MHz
1980 MHz 19801980 MHzMHz 1980 MHz
0.40 0.35 70
75
80
0.35
39.2
39.5
tok / J
BS=4 (online serving)
1185 MHz 1980 MHz 1185 MHz
390 MHz
390 MHz
1.0 200
250
300
Throughput (tok/s)
Throughput (tok/s)
1.2 390 MHz
1980 MHz 1980 MHz 1185 MHz
1.0
1980 MHz
0.8
tok / J
BS=8
390 MHz 780 MHz 1980 1185MHz MHz
390 MHz 390 MHz
tok / J
BS=16 (throughput / offline)
500
600
Throughput (tok/s)
142
2.25
1980 MHz
1980 MHz
1.50
700
284
286
1980 MHz
390 MHz
1000
390 MHz
1590 MHz 1250
555
GQA 10.0
8
1980 MHz
7
1980 MHz
390 MHz 780 MHz
1000
1500
565
570
1980 MHz
1590 MHz 2000 2500
Throughput (tok/s)
1590 MHz
6
600
620
390 MHz
390 MHz
1120
4
8
780 MHz 780 MHz 780 MHz 1980 MHz 1980 MHz 1980 MHz
900
1000
1100
1200
Throughput (tok/s)
390 MHz 390 MHz 390 MHz
Throughput (tok/s)
1500
1.2
780 MHz 1185 MHz 1980 MHz
1980 MHz 1980 MHz
Throughput (tok/s)
2500
1980 MHz 1980MHz MHz 1980 1980 MHz
260
Throughput (tok/s)
1980 MHz 1980 MHz
390 MHz
1590 MHz 800
600
6 390 MHz
2
390 MHz 390 MHz
500
1000
390 MHz
2.5
1980 MHz
400
450
500
550
Throughput (tok/s)
1185 MHz 1980 MHz
1980 MHz 1590 MHz 1590 MHz 1500
1185 MHz
6
3
390 MHz
1980 MHz 1980 MHz
1590 MHz 390 MHz 1185 MHz 1980 MHz 1590 MHz 1590 MHz 1000 2000 3000 1980 MHz
Throughput (tok/s)
1980 MHz
1980 1185MHz MHz 1980 MHz
390 MHz
800
1000
Throughput (tok/s)
1200
MLA
390 MHz 780 MHz
390 MHz
780 MHz 1980 MHz 1185 MHz
390 MHz
5 4
600 390 MHz
7
1980 MHz
10 5
1980MHz MHz 1185 MHz 1980 1980 MHz
390 MHz
2.0
Throughput (tok/s)
15
1185 MHz 1185 MHz
390 MHz
3.0
MLA 390 MHz1185 MHz 1980 MHz
8
300 390 780 MHzMHz
3.5
Throughput (tok/s)
4
280
1185 MHz
1980 MHz
2000
1980 MHz
240
390 MHz
400
780 MHz 780 MHz
MLA
390 MHz
1980 MHz 1185 MHz
390 MHz
1160
390 MHz 780 MHz 390 MHz
1.4
76
390 MHz
390 MHz
1.6
1980 MHz
2
74
Throughput (tok/s)
GQA-ctrl 1185 MHz 1185 MHz 1185 MHz
72
1.8
1980 MHz 1980 MHz 1185 MHz
400
390 MHz
3
1300
6 1140
300
4
660
780 1980 MHzMHz
390 MHz
5
70
GQA-ctrl
390 MHz 390 MHz
6
10
1980 MHz
1100
640
Throughput (tok/s)
12
390 MHz 7801980MHz MHz 1980 MHz
5
390 MHz
Mamba2
390 MHz
780 MHz
7.5 5.0
1980 MHz 1980 MHz 1980 MHz 1980 MHz
GDN
390 MHz
12.5
560
Throughput (tok/s)
780 MHz 780 MHz
3.0
7
780 MHz
390 MHz 1980 MHz 1980 MHz 1980 MHz 1980 MHz
3
1.5
Mamba2
390 MHz
4
1980 MHz
1980 1590MHz MHz 1980 MHz
390 MHz
750
290
780 MHz 1980 MHz 1185 MHz
GQA-ctrl
390 MHz 390 MHz 780 MHz
780 MHz
390 MHz 390 MHz 390 MHz
Mamba2
390 MHz
1185 MHz
390 MHz
2
288
Throughput (tok/s)
68
MLA
Throughput (tok/s)
390 MHz
2.5
1980MHz MHz 1980 1980 MHz 1980 MHz
0.35
110
Throughput (tok/s)
1.0
330
390 MHz
1980 MHz 1980 MHz 1980 MHz 1980 MHz
282
Throughput (tok/s)
tok / J
3.5
0.40
100
Throughput (tok/s)
GDN
6
500
BS=32 (industry batch)
4.0
390 MHz
780 MHz
90
2.0
1980 1980 MHz MHz 19801980 MHzMHz
325
11851980 MHz MHz
390 MHz
7801185 MHzMHz 780 780 MHzMHz
390 MHz 390 390 MHzMHz
GQA-ctrl 390 MHz
1.6
0.45
1980 MHz 1980 MHz 1980 MHz
MHz 390 MHz 780390 MHz 390 MHz
2.0
390 MHz 780 MHz
390 MHz
2.00
80
Mamba2
320
390 MHz
1.75
85
Throughput (tok/s)
144
390 MHz 780 MHz
4
84
2.2
Throughput (tok/s)
GQA
8
1980 MHz 1980 MHz 1980 MHz
1980 MHz
1980 MHz 1185 MHz
1590 MHz
400
83
780 MHz 1185 MHz MHz 1185
0.4
1.8
350
390 MHz
300
0.40
GDN
4
2
780 MHz
MLA
390 MHz 390 MHz 390 MHz 390 MHz
19801980 MHzMHz 1980 MHz 1980 MHz
40
390 MHz 390 MHz390 MHz 780 MHz
GQA 3
39.8
GQA-ctrl 0.6 0.5
0.45 19801980 MHzMHz 1980 MHz
GDN
390 MHz
390 MHz 390 MHz 390 MHz 390 MHz
780 MHz 780 MHz
0.55
1980 MHz
GQA 2.0
Pareto-5% Optimal
0.50
39
390 MHz 780 MHz
Min Freq Max Freq
390 MHz
390 MHz 390 MHz 390 MHz
0.25
85
SM Clock Power Cap
Mamba2
0.30
Throughput (tok/s)
1.5
16K 32K
GDN
390 MHz
1185 MHz
12.5 10.0 7.5 5.0
1980 MHz 780 MHz
390 MHz
1980 MHz 1185 MHz
390 MHz
1980 MHz
780 MHz
1500
1750
1185 1980MHz MHz 2000
2250
Throughput (tok/s)
Fig. 3. Decode DVFS Pareto frontier. Power-cap points cluster in a degenerate blob—all five cap settings produce nearly identical throughput and energy because the GPU draws less than 300 W, below even the 280 W cap. Mamba2 and GDN are worth a brief note on presentation. Their traces under the clock-cap axis can appear erratic compared with GQA and MLA, but the irregularity is a scale artefact: the entire throughput range spanned by varying the SM clock is only 2–10 tok/s for these two architectures, versus 5–2,000 tok/s for GQA and MLA. In absolute terms, aggressive underclocking costs Mamba2 and GDN almost nothing in throughput, so the visual noise is meaningless.
clamped) and not a silicon fmax limit (free-running GPU Boost reaches exactly 1980 MHz indefinitely on the same die). The --lock-gpu-clocks command itself enforces a conservative sustained-frequency policy that clamps any requested lock ≥1830 MHz to the documented base clock of 1830 MHz—a side effect of the lock mechanism rather than a hardware limit. Crucially, the 240 MHz gap (1830 vs. 1590) is wasted: across all four paradigms (and the GQA-ctrl ablation control), sequence lengths from 1K to 65K, and batch sizes from 1 to 32, the median throughput difference between 1590 and 1980 MHz is <0.1%, with half of all configurations showing a slight inversion (1590 MHz faster). Decode throughput is completely insensitive to SM clock above ≈1590 MHz because the memory subsystem sets the pace; the extra clock cycles only waste power (+7–13%). Watt savings from clock locking. At 780 MHz and seq = 1024—the regime where decode is most firmly memory-bound—every architecture saves 47–90 W (24–
The Illusion of Power Capping in LLM Decode
9
Table 1. Power cap vs. actual GPU behaviour during decode (BS = 1, seq = 1024, median over 10 reps). Despite a 2.5× range in the configured cap, the actual SM clock and power draw are identical—the cap never triggers. Actual SM clock (MHz) Actual power (W) Cap (W) GQA GDN
MLA GQA GDN
280 420 500 600 700
1830 1830 1830 1830 1830
1830 1590 1830 1830 1830
1830 1830 1830 1830 1830
207 200 207 207 207
167 167 167 167 167
MLA 231 230 231 231 231
32%) with <1% throughput loss. GDN benefits the most (30% at BS = 1, 32% at BS = 32) because its decode path is almost entirely elementwise operations; the SM clock reduction translates directly into dynamic power savings without contending with any compute bottleneck. At BS = 32, the absolute savings grow to 60–90 W because higher utilisation amplifies the dynamic power component. At longer contexts (seq ≥16K) with large batch sizes, the workload shifts toward compute-bound and the optimal clock rises to 1185–1590 MHz, reducing achievable savings to 5–15%; in a handful of extreme configurations no clock below 1980 MHz satisfies the <1% loss budget. The practical conclusion is stark: data-centre operators who rely solely on power capping for LLM inference workloads—the current industry default—gain zero energy savings during decode, which is the dominant phase in production serving. Static SM clock locking costs nothing to implement (a single nvidia-smi call at job start) and Pareto-dominates power capping at every matched energy– throughput budget we tested. Why this is not an under-utilization artifact. A natural concern is that the low power draw during decode is an artifact of under-utilization (e.g., singleGPU execution or small models). This is not the case. Decode is fundamentally memory-bound: each step performs matrix–vector operations with low arithmetic intensity, requiring repeated HBM weight fetches. As shown in Figure 1, all decode kernels lie far below the roofline ridge point (≈ 206 FLOPs/byte), indicating performance is limited by memory bandwidth rather than compute throughput. Accordingly, if decode were compute-bound, increasing SM frequency would improve throughput; instead, throughput is invariant above ≈ 1590 MHz, confirming it is not compute-limited. Increasing utilization via batching or parallelism does not change arithmetic intensity, which is an algorithmic property. While batching improves throughput and energy efficiency by amortizing memory traffic, the workload remains memory-bound and below the GPU power limit. Consequently, power capping does not engage even at high batch sizes (e.g., BS = 32) or sustained load. In tensor- or pipeline-parallel settings, each GPU still executes memory-bound kernels on partitioned weights; parallelism increases aggregate utilization but not per-GPU arithmetic intensity, and therefore does not
10
B. Ma et al.
move decode across the roofline ridge. Thus, decode is inherently memory-bound: arithmetic intensity is unchanged by deployment configuration, and GPU power remains below TDP not due to under-utilization, but due to memory bandwidth bottlenecks.
6
The Compressed and Recurrent Architecture Crossover
Novel attention replacements—MLA’s compressed KV cache, GDN’s linear recurrence, Mamba2’s SSM—all exhibit the same pattern: a heavy prefill cost that efficient decode eventually recoups. 6.1
The Prefill Penalty
Even at their optimal clocks, GDN and Mamba2 consume an order of magnitude more prefill energy per token than the transformers—and the gap widens further at longer sequences. The cause is a double penalty: throughput plateaus because sequential state recurrence cannot be parallelised across sequences, while power draw is higher because the GPU is fully occupied issuing low-intensity elementwise instructions. The result is the worst of both worlds—low throughput at high power. This penalty reflects vLLM’s unfused eager-mode execution (Section 7.2), not an inherent architectural limit; fused kernels could substantially close the gap. MLA pays a smaller but persistent prefill tax. Its non-power-of-2 head dimension (dh = 192) wastes tensor-core lanes, reducing FlashAttention TC utilisation from 58% to 51% and slowing attention kernels by 1.6× relative to GQA-ctrl (dh = 128). However, this tile-alignment penalty is the smaller of two costs: the KV decompression path—concatenation, reshape, and copy operations that reconstruct full heads from compressed latents—adds a data-movement overhead that persists even in decode, where it accounts for 90% of the MLA–GQA gap (Section 6.2). Fixing dh to a power of 2 would help prefill attention but would not address the dominant decompression cost. Because both costs scale with sequence length, the gap widens rather than closes at longer contexts, and DVFS cannot help: MLA and GQA-ctrl respond similarly to clock reduction. 6.2
The Decode Payoff
Yet in decode, the story reverses (Figure 2, decode 16K panel). Mamba2 confirms the O(1) decode promise: its per-step latency is constant regardless of context length. At large batch size and long context, this gives it more than a 2× energy advantage over GQA. As batch size grows and SSM scan compute becomes significant, the optimal clock must rise, placing Mamba2 in the “batch-sensitive” class. MLA’s crossover is more gradual. At short context, compressing the KV cache is pointless: weight loading dominates HBM traffic and the decompression path is pure overhead, making MLA 12–29% worse than GQA-ctrl. Contrary to
The Illusion of Power Capping in LLM Decode
11
the intuition that MLA simply adds a few cheap GEMMs, its GEMM count is nearly identical to GQA-ctrl’s (225 per decode step). The overhead is instead dominated by the data-movement machinery that reconstructs full KV heads from compressed latents—hundreds of small concatenation, reshape, and copy kernels per step that are individually cheap but collectively account for 90% of the gap and are entirely insensitive to SM clock. A fused decompression kernel could eliminate most of this cost. As context grows, however, KV cache traffic grows linearly— steeply for GQA-ctrl, gently for MLA’s compressed latent—and a crossover emerges. It arrives sooner at higher batch sizes: at BS = 32 the crossover is already at 4K tokens; at BS = 1 it never arrives. At the most aggressive configuration (BS = 32, seq = 65K), MLA uses less than half of GQA-ctrl’s decode energy. MLA’s energy advantage is strictly decode-specific; disaggregated serving [1] can exploit this by routing decode to MLA-optimised pools while keeping prefill on GQA hardware. A cautionary example: MiniCPM3-4B [15] advertises MLA but vLLM’s backend routing silently expands its latents to full KV dimensions, negating compression entirely. MLA’s benefits must be verified in the deployment stack, not inferred from the model card.
6.3
Total Request Energy and Deployable Policy
The per-phase results leave a natural question: which architecture wins for a complete request? Figure 4 answers this. At low batch size (BS = 1), architecture barely matters: all transformers cluster together, and GDN is always the most expensive. The picture changes dramatically at production batch sizes (BS = 32). MLA’s decode efficiency makes it cheapest from nearly the first output token. Mamba2 starts expensive (its prefill penalty is visible as a steep initial offset) but crosses below GQA after roughly a thousand output tokens—exactly the regime of agentic coding assistants and document processing. The crossover mechanism is clarified by comparing absolute energy gaps. At BS = 32 and 16K context, Mamba2’s prefill costs ≈35× more energy per token than GQA—but the absolute difference is only ∼10 mJ/tok, because prefill is already extremely cheap. In decode, the ratio is a modest 1.7×, yet the absolute saving exceeds 99 mJ/tok because decode energy per token is orders of magnitude higher. Every output token therefore repays the prefill penalty many times over.
6.4
Deployable Clock Policies
The key insight: batch-invariant architectures (GQA, GQA-ctrl) can use a single low decode clock regardless of load, while batch-sensitive ones (MLA, Mamba2) need to raise it when the serving engine fills large batches. GDN is the simplest to operate—it tolerates aggressive underclocking unconditionally.
12
B. Ma et al. Total Request Energy: Pareto-5% (solid) vs Min-Energy (dashed) GQA
GDN
Mamba2
GQA-ctrl
Total request energy (mJ)
Total request energy (mJ)
Prefill ctx = 4K, BS = 1
MLA
Pareto-5%
Prefill ctx = 16K, BS = 1
Prefill ctx = 32K, BS = 1
10 8
10 8
10 8
10 7
10 7
10 7
10 6
10 6
10 6
Decode output tokens
Decode output tokens
Decode output tokens
Prefill ctx = 4K, BS = 32
Prefill ctx = 16K, BS = 32
Prefill ctx = 32K, BS = 32
10 6
10 5
256
512
Min-energy
1K
2K
4K
Decode output tokens
8K
16K
32K
10 6
10 6
10 5
10 5
256
512
1K
2K
4K
Decode output tokens
8K
16K
32K
256
512
1K
2K
4K
Decode output tokens
8K
16K
32K
Fig. 4. Total request energy vs. decode output length. Solid: Pareto-5% clock; dashed: min-energy clock (the two nearly overlap). Top: BS = 1; bottom: BS = 32. At low batch, architectures cluster; at high batch, MLA and Mamba2 pull ahead as decode length grows, while GDN crosses only at long context.
7
Discussion
7.1
Implications for Data-Centre Power Management
Our results expose a gap between common data-centre practice and LLM inference reality. Power capping is the industry-standard energy knob: cluster schedulers set per-GPU or per-node power limits to meet facility-level power budgets [9]. This works well when GPU workloads operate near TDP—training runs, dense HPC kernels, large-batch prefill—because the cap constrains actual behaviour. But as LLM serving shifts toward decode-dominated workloads (long outputs, agentic multi-turn, streaming), the GPU spends most of its time in a low-power, memory-bound state that never reaches the cap. In our measurements, decode draws 160–300 W on a 700 W GPU; a facility-level 280 W cap achieves precisely nothing. The fix is straightforward: replace power capping with static SM clock locking for decode pools. In disaggregated serving architectures (Splitwise [17], DistServe [25]), where prefill and decode run on separate GPU pools, each pool can be locked at its phase-optimal clock —no dynamic switching required. For colocated serving (e.g. single-GPU vLLM), a conservative decode clock (780 MHz) applied globally saves 47–90 W per GPU at short-to-moderate context with negligible throughput loss. At data-centre scale (tens of thousands of GPUs), this translates to megawatts of savings (e.g., at 50 W savings per GPU across 10,000 GPUs, this corresponds to 0.5 MW of continuous power reduction) that power capping cannot deliver.
The Illusion of Power Capping in LLM Decode
13
A subtler implication concerns monitoring and accounting. If operators track “power cap utilisation” (actual draw / cap) as a proxy for energy efficiency, decode workloads will appear highly efficient (30–40% of cap)—masking the fact that the GPU is simply idle most of the time. Clock locking makes the trade-off explicit: the lower clock directly reduces both instantaneous power and energy per token, and the throughput impact is measurable and bounded. 7.2
Limitations
The software stack matters as much as the architecture: vLLM serves GDN and Mamba2 via unfused eager mode, and the order-of-magnitude prefill gap reflects this rather than an inherent architectural limit; custom fused kernels [22,24] could substantially close it. Beyond this: our measurements cover a single GPU (no multi-GPU communication energy), a single framework (TensorRT-LLM may differ), and dense models only (MoE routing may interact with DVFS differently). The clamp we observe (1980→1830 MHz under --lock-gpu-clocks) is a side effect of the lock command’s conservative sustained-frequency policy rather than a hardware limit: free-running GPU Boost reaches 1980 MHz on the same die. This could be specific to the H200 SXM and the 590.48 driver; other GPU generations, driver versions, or lock mechanisms may behave differently. NVML timeseries samples (50 ms cadence) recorded die temperature alongside clock and power for every run and consistently show 42–45 ◦ C under clamp, ruling out thermal effects at sustained load. Finally, while multi-GPU tensor-parallel decode increases per-GPU utilisation and could in principle push power draw above a cap, our batch-size sweep provides direct evidence that this concern does not apply: at BS = 32—a regime of high request concurrency that approaches the memory-bandwidth saturation point—power draw still remains well below every tested cap level across all five architectures. Arithmetic intensity does not change with batch size or parallelism strategy; the memory-boundedness is structural, not an artefact of low utilisation.
8
Conclusion
Power capping—the standard data-centre energy lever—is structurally ineffective for memory-bound LLM decode. This result is scale-invariant: decode arithmetic intensity is an algorithmic property unchanged by batch size, model size, or tensor parallelism—larger models generate more memory traffic, not less, and tensor parallelism does not move per-GPU kernels above the roofline ridge point. The GPU draws only 137–300 W during memory-bound decoding on an H200 rated at 700 W; no cap we apply ever triggers, and the driver holds ≈1830 MHz regardless. A silent clamp inside --lock-gpu-clocks (1980→1830 MHz under lock, distinct from the free-running boost that holds 1980 MHz) compounds the illusion by introducing clock deviations unrelated to the cap—and the extra 240 MHz above 1590 MHz is entirely wasted, producing zero throughput gain at 7–13% more power. Neither the configured power limit nor the requested clock
14
B. Ma et al.
frequency reflects actual GPU behaviour—a double disguise that makes power capping not merely ineffective but actively misleading for decode workloads. SM clock locking dissolves both confounds. By controlling the frequency lever that is genuinely on the critical path, it recovers up to 32% of decode energy at negligible throughput loss, Pareto-dominating power capping at every matched operating point. The optimal clock is architecture- and batch-size-dependent—we identify three behavioural classes and provide a deployable policy table—but the superiority of clock locking over power capping is universal across all architectures and configurations we tested. Beyond the DVFS mechanism, the cross-architecture characterisation reveals a shared pattern among novel designs: recurrent and compressed-KV architectures pay an upfront prefill cost that their efficient decode recoups within roughly a thousand output tokens at production batch sizes, eventually halving total request energy relative to GQA. Future work should evaluate whether the power-cap ineffectiveness we observe on H200 extends to other GPU generations, whether fused recurrent kernels close the prefill gap, whether KV quantisation shifts the MLA crossover point, and whether adaptive DVFS or workload-aware frequency scaling can outperform static clock locking at runtime. Acknowledgments. This work has been funded by the Free State of Bavaria in the DSgenAI project (Grant Nr.: RMF-SG20-3410-2-18-4). The authors gratefully acknowledge the scientific support and HPC resources provided by the Erlangen National High Performance Computing Center (NHR@FAU) of the Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU). The hardware is funded by the German Research Foundation (DFG).
References 1. Agrawal, A., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B.S., Ramjee, R.: Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills (2023), https://arxiv.org/abs/2308.16369 2. Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., Sanghai, S.: Gqa: Training generalized multi-query transformer models from multi-head checkpoints (2023), https://arxiv.org/abs/2305.13245 3. Bridges, R.A., Imam, N., Mintz, T.M.: Understanding gpu power: A survey of profiling, modeling, and simulation methods. vol. 49. Association for Computing Machinery, New York, NY, USA (Sep 2016). https://doi.org/10.1145/2962131 4. Dao, T., Fu, D.Y., Ermon, S., Rudra, A., Ré, C.: FlashAttention: Fast and memoryefficient exact attention with IO-awareness. In: Advances in Neural Information Processing Systems (NeurIPS) (2022) 5. Dao, T., Gu, A.: Transformers are SSMs: Generalised models and efficient algorithms through structured state space duality. In: Proceedings of the 41st International Conference on Machine Learning (ICML) (2024), https://arxiv.org/abs/2405. 21060 6. DeepSeek-AI, Liu, A., Feng, B., Wang, B., Wang, B., Liu, B., Zhao, C., Dengr, C., Ruan, C., Dai, D., et al.: Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model (2024), https://arxiv.org/abs/2405.04434
The Illusion of Power Capping in LLM Decode
15
7. Kolluru, S.: Comparative analysis of large language model inference serving systems: A performance study of vllm and huggingface tgi (2025), https://arxiv.org/abs/ 2511.17593 8. Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C.H., Gonzalez, J., Zhang, H., Stoica, I.: Efficient memory management for large language model serving with PagedAttention. In: Proceedings of the ACM Symposium on Operating Systems Principles (SOSP) (2023), https://arxiv.org/abs/2309.06180 9. Li, B., Samsi, S., Gadepally, V., Tiwari, D.: Clover: Toward sustainable ai with carbon-aware machine learning inference service. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. SC ’23, Association for Computing Machinery, New York, NY, USA (2023). https: //doi.org/10.1145/3581784.3607034 10. Liu, Z., Yuan, J., Jin, H., Zhong, S., Xu, Z., Braverman, V., Chen, B., Hu, X.: KIVI: A tuning-free asymmetric 2bit quantization for KV cache. arXiv preprint arXiv:2402.02750 (2024), https://arxiv.org/abs/2402.02750 11. Maliakel, P.J., Ilager, S., Brandic, I.: Characterizing llm inference energyperformance tradeoffs across workloads and gpu scaling (2026), https://arxiv. org/abs/2501.08219 12. Meng, F., Tang, P., Tang, X., Yao, Z., Sun, X., Zhang, M.: Transmla: Multi-head latent attention is all you need (2025), https://arxiv.org/abs/2502.07864 13. Micikevicius, P., Stosic, D., Burgess, N., Cornea, M., Dubey, P., Grisenthwaite, R., Ha, S., Heinecke, A., Judd, P., Kamalu, J., Mellempudi, N., Oberman, S., Shoeybi, M., Siu, M., Wu, H.: Fp8 formats for deep learning (2022), https://arxiv.org/ abs/2209.05433 14. Muralidharan, S., Sreenivas, S.T., Joshi, R., Chochowski, M., Patwary, M., Shoeybi, M., Catanzaro, B., Kautz, J., Molchanov, P.: Compact language models via pruning and knowledge distillation. arXiv preprint arXiv:2407.14679 (2024), https://arxiv. org/abs/2407.14679 15. OpenBMB Team: MiniCPM3-4B: A language model with function call, long context, and retrieval augmented generation. Hugging Face model card (2024), https: //huggingface.co/openbmb/MiniCPM3-4B 16. Patel, P., Choukse, E., Zhang, C., Íñigo Goiri, Warrier, B., Mahalingam, N., Bianchini, R.: Polca: Power oversubscription in llm cloud providers (2023), https: //arxiv.org/abs/2308.12908 17. Patel, P., Choukse, E., Zhang, C., Shah, A., Íñigo Goiri, Maleki, S., Bianchini, R.: Splitwise: Efficient generative llm inference using phase splitting (2024), https: //arxiv.org/abs/2311.18677 18. Patterson, D., Gonzalez, J., Hölzle, U., Le, Q., Liang, C., Munguia, L.M., Rothchild, D., So, D.R., Texier, M., Dean, J.: The carbon footprint of machine learning training will plateau, then shrink (Jul 2022). https://doi.org/10.1109/MC.2022.3148714 19. Shah, J., Bikshandi, G., Zhang, Y., Thakkar, V., Ramani, P., Dao, T.: Flashattention3: Fast and accurate attention with asynchrony and low-precision (2024), https: //arxiv.org/abs/2407.08608 20. Shi, T., Wu, Y., Liu, S., Ding, Y.: Greenllm: Disaggregating large language model serving on heterogeneous gpus for lower carbon emissions (2024), https://arxiv. org/abs/2412.20322 21. Yang, S., Kautz, J., Hatamizadeh, A.: Gated delta networks: Improving mamba2 with delta rule. In: Proceedings of ICLR (2025) 22. Yang, S., Zhang, Y.: Fla: A triton-based library for hardware-efficient implementations of linear attention mechanism (Jan 2024), https://github.com/fla-org/ flash-linear-attention
16
B. Ma et al.
23. You, J., Chung, J.W., Chowdhury, M.: Zeus: Understanding and optimizing gpu energy consumption of dnn training (2022), https://arxiv.org/abs/2208.06102 24. Yu, C., Zeng, B., Chen, H., Yang, Z., Zhang, Z., Li, H., Zhou, J.: cula: Cuda linear attention (2026), https://github.com/InclusionAI/cuLA 25. Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., Zhang, H.: Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving (2024), https://arxiv.org/abs/2401.09670