Conceptio › Archive › arXiv CS
arXiv CSopen access

Measured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM Serving

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

Measured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM Serving Muhammad Abdur Rab Siddiqui

arXiv:2609.23085v1 [cs.PF] 19 Sep 2026

[email protected] University of Alberta Edmonton, Alberta, Canada

Daniela Rojas

Chen Yang

[email protected] University of California San Diego La Jolla, California, USA

[email protected] Hong Kong University of Science and Technology (Guangzhou) Guangzhou, China

Wenqi Cui

Yuanyuan Shi

Yize Chen

[email protected] New York University New York, New York, USA

[email protected] University of California San Diego La Jolla, California, USA

[email protected] University of Alberta Edmonton, Alberta, Canada

Abstract

1

Large language models (LLMs) and agentic AI systems are creating rapidly growing inference energy demands as model sizes grow and reasoning trajectories extend. While in practice, many queries do not require the capabilities of the largest available model, and routinely directing such queries to a high-capability model can introduce unnecessary, considerable computation and energy consumption. In this paper, we investigate whether adaptive routing across a heterogeneous pool of LLMs can reduce this energy burden without substantially compromising task performance. We design a language-model-based router that reads in each query and selects an answer model from a fixed candidate pool. The candidate models are first profiled through an offline tournament that records their correctness, latency, power, and GPU energy for each query. Using these measurements, the router is trained through supervised finetuning followed by group relative policy optimization (GRPO) with the tailored paradigms. Results demonstrate that learned routing can selectively allocate expensive model capacity based on query contexts, and effectively improve the accuracy–energy trade-off of multi-LLM serving. Across seven benchmark tasks, we also observe a sharp accuracy–energy phase transition among routers, providing practical insights in terms of striking the energy efficiency while still keeping LLM performance.

Artificial intelligence (AI), particularly large language models (LLMs), already places substantial demands on computing infrastructure and its electricity supply. Inference is becoming the central part of this energy burden. As deployed models with billions or trillions of parameters repeatedly process requests throughout their operational lifetime, so even modest energy costs per response accumulate at service scale [3, 20]. Longer reasoning outputs and agentic workflows can increase this demand further by generating more tokens and invoking models multiple times for a single user request [4]. These workloads make inference efficiency an important concern for sustainable AI serving with regard to cutting down energy spent answering queries while preserving task performance. As current LLMs vary by model sizes, specialized areas, and associated computing demands, model routing offers one approach to this challenge. Instead of assigning every query to the same LLM, in practice, a router can select an answer model from a heterogeneous pool according to the query’s requirements and characteristics [13, 23]. As modern AI models differ to a significant extent in their capabilities across reasoning, coding, and knowledge-intensive tasks, it creates opportunities to allocate computational effort selectively [10]. Always sticking to a high-capability model can spend unnecessary energy on queries that a smaller model could answer correctly, whereas always using a small model can sacrifice accuracy on more demanding inputs. Query-dependent LLM model selection therefore offers a way to exploit model diversity and adjust the quality-cost trade-off [10, 12, 19]. However, making routing effective and tailored for energy reduction presents a few challenges. Indeed, the cost of a model invocation depends on both models and queries. Input and output lengths, hardware, and serving configurations affect energy consumption, so current practices like parameter counts, API prices, and fixed model-level costs do not fully expose the joules consumed by the underlying generation task [3, 27]. This issue becomes especially important in agentic LLM systems, where one user request may trigger multiple model calls for planning, tool use, verification, and revision [28]. Moreover, an energy-centric router shall be able to judge and evaluate whether a candidate will answer correctly before observing its output. Reducing energy by systematically choosing weaker models can undermine the service’s purpose, and the useful

ACM Reference Format: Muhammad Abdur Rab Siddiqui, Daniela Rojas, Chen Yang, Wenqi Cui, Yuanyuan Shi, and Yize Chen. 2027. Measured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM Serving. In Proceedings of Make sure to enter the correct conference title from your rights confirmation email (Preprint, In Submission). ACM, New York, NY, USA, 16 pages. https://doi.org/XXXXXXX.XXXXXXX

Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. Preprint, In Submission, Publishing venue © 2027 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-XXXX-X/2018/06 https://doi.org/XXXXXXX.XXXXXXX

Introduction

Preprint, In Submission, Publication time, Publishing venue

Siddiqui et al.

where LLM models below this region can remain relatively inexpensive, whereas higher accuracy requires substantially more frequent use of capable models. Our proposed policy trained via the mixed KL-GRPO scheme reaches similar accuracy as UniRoute [12] , while consistently using about 23% less mean answer energy (1692 J vs. 2192 J). Ablation studies also indicate interesting observations regarding reward selection, model collapse, and tradeoff between AI performance and energy footprints. We will open-source the proposed router after reviewing period.

1.1 Figure 1: Overview of LLM-based routing learned from measured inference outcomes and RL feedbacks. We collect and construct offline profiling dataset from LLM queries to record correctness, GPU energy, and latency to train a small LLM routing controller through SFT and GRPO; At serving time, the intended controller selects one candidate from the fixed executer pool, which achieves improved tradeoff between LLM serving energy and performance.

routing requires learning when additional model capacity is warranted. Another concern is routing itself consumes computation, while most of recent works on router design may not distinguish the selected model’s consumption from the overhead of making and executing that decision [12, 19]. In this work, we study whether a tailored small language model can learn the unique energy-aware routing policy from measured AI serving outcomes. To achieve this router training, we first conduct an offline tournament in which every candidate answer model processes every training question, producing logs of correctness, latency, power, and GPU energy. These logs define an oracle route and a gated reward, where we propose that incorrect answers receive a fixed penalty, while correct answers receive a reward reduced by the selected model’s within-query energy rank. Built upon recent advancements in agentic post-training paradigms [11, 18] together with this novel profiling data, the routing controller is trained through supervised fine-tuning (SFT) followed by group relative policy optimization (GRPO) [22], including a variant regularized toward the SFT policy. Rather than guided to output complete final answers, based on query contexts, we design the training paradigm to let the router only output the routing decision, which is both lightweight and highly generalizable after finetuning on the small energy-aware datasets. During training and evaluation, model outcomes are replayed from the tournament logs, so answer models do not need to be invoked repeatedly. Figure 1 illustrates our proposed router training and implementation workflow. This work makes three main contributions. First, we formulate LLM model routing around measured GPU energy rather than model size or monetary price alone. Second, we illustrate a practical way of profiling LLM energy and performance during inference, and develop a compact LLM-based routing controller trained through an SFT-to-RL pipeline under a gated accuracy-energy reward. Third, we evaluate proposed and classical model routers on a nine-model pool across a diverse held-out mixture of seven benchmarks. Results indicate a sharp accuracy-energy transition near 60% accuracy,

Related Work

1.1.1 Energy-efficient LLM inference. Due to the surging energy demand associated with AI usage, research start to investigate the power and energy impact of AI training and AI serving [3, 15]. Inference energy costs are alarming as agentic tasks involve complicated reasoning loops, long reasoning chains, and tendency of using larger sizes of LLMs [17]. Researchers have identified hardware selection [27], AI model training [29], LLM model caching and serving strategies [8, 26], preemptive power management [21] as opportunities of energy reduction. Prior work has measured substantial variation in LLM inference energy across tasks, models, and deployment configurations [3, 20]. Such measurement literature shows that energy depends on model size, batching, hardware, precision, and generation length. For instance, most routers shown in Table 2 still optimize accuracy, embedding similarity, or APIstyle cost proxies rather than measured GPU energy on the serving stack. GreenServ [33] is the main exception among our baselines because it explicitly trades accuracy against normalized inference cost in the bandit reward. We evaluate classical and learned routers on the same held-out mix with logged GPU energy, power, and latency, and we study how reward shaping and RL reshape the accuracy–energy frontier over a multi-model local pool. 1.1.2 Routers for AI Serving. A growing line of work considers practical serving systems, and treats model selection as a firstclass inference problem instead of fixing one backend for every query. Most of these routers are designed to optimize efficiency [12], while outputs with least token outputs may not be directly translated to energy consumption. RouterBench formalizes this setting with a large offline benchmark of per-model outcomes and popular routing baselines [10]. Among them, 𝑘-nearest-neighbour routing scores each candidate from the error rates of models on embedding neighbours of the test prompt. UniRoute [12] represents each LLM through its errors on representative prompt clusters, so a model can be scored and routed to even when it was not in the original training pool. Smoothie [7] takes a label-free route. It embeds each candidate’s prompt–output pair and estimates per-sample quality without correctness labels on the routing split. RouteLLM [19] learns a binary strong–weak router from preference data and thresholds a win probability to trade cost against quality. In our design, we adapt that rule to a strong–weak pair from the local candidate pool rather than the paper’s API checkpoints. GreenServ [33] targets energyaware serving. It models routing as a contextual bandit and uses LinUCB [14] is a classical bandit-based approach that maintains a disjoint linear reward model per arm and explores with an uncertainty bonus. GreenServ’s reward mixes accuracy with normalized inference cost, which is close in spirit to our gated accuracy–energy

Measured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM Serving

objective, while our controller is a learned language model rather than a hand-built feature vector. The broader AI serving system’s goal is not a single router knob, posing challenges to router design. Indeed, agentic serving has two coupled controls. One is which model answers a query. The other is how much reasoning that call is allowed to spend. Most routers named above decide only the first. ARES [28] decides the second, and it does so adaptively. It chooses how much reasoning effort an agent should spend on the current query, so harder items can think longer and easier items can stop sooner. It does not assign one fixed thinking budget to every call. That split is part of what motivates this work. A fixed model wastes energy on easy items and still fails hard ones and a fixed thinking budget has the same problem inside one model. ARES is the closest prior statement of that second failure, and it treats reasoning effort as something the system should choose rather than freeze. Our controller attacks the first control with a small language model and a gated accuracy– energy reward on measured GPU logs. Together the two controls point to the same end goal. Agentic inference should spend compute only when the task justifies it. 1.1.3 Agentic RL. Training the router as a policy rather than only as a classifier connects our setup to agentic RL for language models. Recent advances have shown effectiveness of RL in post training which go beyond text imitation usually happening in SFT approaches. AgentLightning [18] is an agent training framework that wraps rollout generation and policy optimization for multi-step workflows. Hugging Face TRL [25] is a separate library that implements GRPO and related on-policy algorithms for language models. We explored both stacks on the same offline tournament reward. Our reported SFT, GRPO, and KL-GRPO numbers use TRL, so training, checkpoints, and held-out evaluation stay in one consistent pipeline. Agent Lightning experiments served as a cross-check on the same reward design rather than as a second training path in the main tables. The distinction between supervised fine-tuning (SFT) and on-policy RL matters here [30]. Prior analyses show that RL can improve exploration and reward shaping beyond imitation of static routing labels, while SFT remains a stable warm start [2]. To that end, in this work we study whether RL on replayed tournament rewards can learn when extra answer-model cost is worth higher accuracy under measured energy constraints.

2 Problem Formulation 2.1 Heterogeneous model pool For modern AI model serving, it is hard to apply one LLM model fit all inputs. This is due to the fact that it becomes challenging to strike a balance between being accurate on hard items and staying efficient on easy queries. For instance, a small, open-weight model generates few tokens and spends little energy, but it can fail queries which need more model and reasoning capacity. Whereas a large, reasoning-based model does the reverse, because they may overthink, and can waste significant amount of joules on items a smaller model would have easily answered [32]. To this end, we study a serving system that selects one answer LLM for each incoming query. Let M = {𝑚 1, . . . , 𝑚𝐾 } denote a fixed pool of 𝐾 candidate LLMs, and let A = {1, . . . , 𝐾 } index

Preprint, In Submission, Publication time, Publishing venue

the available routes. A query context 𝑥 contains the question and any metadata provided to the router. The routing policy 𝜋𝜃 (𝑎 | 𝑥) selects a candidate index 𝑎, after which model 𝑚𝑎 generates an answer 𝑦. The router makes the model-selection decision; the selected candidate performs the task. Each evaluated query involves one routing decision and one answer-model outcome. Multi-turn recovery and tool use could also be applied to our framework with proper extensions. The pool of candidate LLMs is heterogeneous by design, and the model card is illustrated in Table 1. Candidates differ in model family, parameter count, and reasoning style, from sub-billion Qwen2.5 drafts through mid-size models such as Gemma2-2B, Qwen2.5-7B, and Llama 3.1 8B, up to DeepSeek-R1 8B and 32B. That spread is what makes routing meaningful. These differences create opportunities for query-dependent selection, since a model that is effective for one task need not offer the best accuracy–energy trade-off for another. For example, Figure 2 shows a GSM8K query answered correctly by Qwen2.5-0.5B using 82 J and by DeepSeek-R1-8B using 2864 J, while Llama 3.1-8B answers incorrectly using 638 J. For this question, though Llama model has larger size, its overthinking and extra steps around equation modeling causes both inefficient and inaccurate inference. Such an example motivates measuring outcomes for each query-energy-model triplet rather than inferring suitability from model size alone. We assume all answer models are open-weight checkpoints served locally through Ollama on that device, so each candidate can be invoked under the same power logger rather than billed as a separate API price.

2.2

Recorded outcomes and energy measurement

For query 𝑖 and candidate 𝑎, let 𝑦𝑖𝑎 be the recorded answer and 𝑦𝑖★ its reference. We denote correctness by 𝐶𝑖𝑎 = 𝑔𝑏 (𝑖 ) (𝑦𝑖𝑎 , 𝑦𝑖★) ∈ {0, 1},

(1)

where 𝑔𝑏 (𝑖 ) is the evaluator for benchmark 𝑏 (𝑖). This notation covers task-specific answer matching and code-test evaluation without requiring exact and literal string equality. Let 𝐸𝑖𝑎 , 𝐿𝑖𝑎 , and 𝑃¯𝑖𝑎 denote the each answer call 𝑖’s incremental GPU energy in joules, wall-clock duration in seconds, and mean incremental GPU power in watts, respectively. The candidate models are served locally through Ollama on NVIDIA A100-SXM4-80GB GPUs with tensor parallelism. During profiling, the logger polls board power through NVML using pynvml at a 200 ms interval and subtracts an idle-power baseline. We adopt nvidia-ml-py bindings, which read the same device board (𝑡) denote the board power trace counters as nvidia-smi. Let 𝑃𝑖𝑎 idle and 𝑃𝑖𝑎 the baseline associated with the query call 𝑖, the measurement quantities of power, energy, and duration are modeled as inc board idle 𝑃𝑖𝑎 (𝑡) = 𝑃𝑖𝑎 (𝑡) − 𝑃𝑖𝑎 ; ∫ 𝑡 end 𝑖𝑎 inc 𝐸𝑖𝑎 = 𝑃𝑖𝑎 (𝑡) d𝑡;

(2) (3)

start 𝑡𝑖𝑎

𝐸𝑖𝑎 𝑃¯𝑖𝑎 = . (4) 𝐿𝑖𝑎 The energy integral is estimated from the sampled trace. Note these quantities describe idle-subtracted GPU consumption within the end start 𝐿𝑖𝑎 = 𝑡𝑖𝑎 − 𝑡𝑖𝑎 ,

Preprint, In Submission, Publication time, Publishing venue

GSM8K, question 31. Gunter is trying to count the jelly beans in a jar. He asks his friends how many they think are in the jar. One says 80. Another says 20 more than half the first one. A third says 25% more than the first one. What is their average guess? Qwen2.5-0.5B (correct) 82 J | 3.8 s Answer: 80 Response: “First, let’s calculate each friend’s guess: (1) The first friend guesses 80 jelly beans. (2) The second friend guesses 20 more than half of 80, which is (80/2) + 20 = 60. (3) The third friend guesses 25% more than the first one, which is 1.25 × 80 = 100. The average is (80 + 60 + 100)/3 = 80.” Llama 3.1-8B (wrong) 638 J | 7.8 s Answer: 106.67 Response: “Let’s denote Gunter’s friend who said 80 as F1, the second friend as F2, and the third friend as F3. F2 = (1/2) ·F1+20. F3 = 1.25·F1. Now we can find the average of their guesses by adding up all three guesses and dividing by 3: Average = (F1 + ((1/2) * F1 + 20) + (1.25 * F1))/3 Combine like terms. Average = (F1 + (0.5 * F1) + 20 + (1.25 * F1))/3 Simplify the numerator. Average = (3.75 * F1 + 20)/3 Now we can plug in the value of F1, which is 80. Average = (3.75 * 80 + 20)/3 Multiply 3.75 and 80. Average = (300 + 20)/3 Add 300 and 20. Average = 320/3 Divide 320 by 3. Average = 106.67” DeepSeek-R1-8B (correct) 2,864 J | 22.8 s Answer: 80 Response: “The first friend guesses 80 jelly beans. The second friend guesses 20 more than half of the first guess: half of 80 is 40, so 40+20 = 60. The third friend guesses 25% more than the first guess: 25% of 80 is 20, so 80 + 20 = 100. The average is (80 + 60 + 100)/3 = 80.”

Figure 2: Same benchmark query from GSM8K, different profilings of three local models. Responses are truncated tournament logs; energy and duration are measured effective GPU cost for that answer call.

measured query call interval, which is the core energy associated with AI inference. They do not represent total server or data-centerlevel energy consumption. We first profile AI model’s energy and performance through an offline tournament of mild size, which records the outcomes of the candidate models on the collected questions. With given query input 𝑥𝑖 , the per-query record is profiled as   𝐾 D𝑖 = 𝑥𝑖 , {(𝐶𝑖𝑎 , 𝐸𝑖𝑎 , 𝐿𝑖𝑎 , 𝑃¯𝑖𝑎 )}𝑎=1 . (5) These are realized outcomes of the underlying LLM inference generations. Replaying a record reproduces that recorded outcome, rather than estimating the answer model’s stochastic behavior or measurement variation. The complete outcome vector provides feedback for all candidates on a training query, although the learned policy selects only one candidate at inference time.

2.3

Energy-Aware AI serving objectives

In this work, we explore the possibilities of meeting LLM service objectives at low energy cost. On a fixed logged dataset D of 𝑁 queries, the expected accuracy and answer energy of a feasible

Siddiqui et al.

routing policy are 𝑁 ∑︁ 𝐾 ∑︁ bD (𝜃 ) = 1 𝜋𝜃 (𝑎 | 𝑥𝑖 ) 𝐶𝑖𝑎 , 𝐴 𝑁 𝑖=1 𝑎=1

(6a)

𝑁 𝐾 1 ∑︁ ∑︁ 𝜋𝜃 (𝑎 | 𝑥𝑖 ) 𝐸𝑖𝑎 . 𝐸bD (𝜃 ) = 𝑁 𝑖=1 𝑎=1

(6b)

For deterministic routing, these expressions reduce to averages of the selected candidates’ outcomes. For sampled decisions, the corresponding sample averages estimate these expectations. Handling of invalid controller outputs is discussed in Section 3.1. One design to handle the desired accuracy-energy trade-off is bD (𝜃 ) max 𝐴 𝜃

subject to 𝐸bD (𝜃 ) ≤ 𝐵;

(7)

where 𝐵 is an average answer-energy budget in joules per query. This constraint expresses a service goal, not a guarantee made by the training algorithm. The reported method instead optimizes the gated rank-based surrogate defined in Section 3.2 and evaluates its resulting accuracy and measured answer energy. Problem (7) can be treated as a single-step contextual decision problem. A contextual-bandit formulation accommodates either a classical or a text-conditioned neural policy. In addition, our logged training data provide detailed outcomes for every candidate. We therefore distinguish the routing problem, the controller’s representation, and the feedback used to train it. We note with the addition of the router policy 𝜋𝜃 , the complete serving energy also depends on the controller and deployment costs, including model loading and residency under the chosen serving configuration. Those costs is measured use the consistent metric to establish end-to-end savings, and they are not supplied by the lookup of 𝐸𝑖𝑎 alone. More importantly, as LLM’s consumed energy are strongly dependent on input queries and output contexts, it requires dedicated design and investigation of the router. Table 1: Candidate language models used in the evaluation, with ARC Challenge [5] accuracy and mean effective answer energy per query on all 472 collected questions. Model Qwen 2.5 Qwen 2.5 Qwen 2.5 Phi-3 Mini Gemma 2 Qwen 2.5 Llama 3.1 DeepSeek-R1 DeepSeek-R1

Nominal Size

ARC Acc.

Energy (J)

0.5B 1.5B 3B 3.8B 2B 7B 8B 8B 32B

0.314 0.665 0.727 0.803 0.722 0.871 0.803 0.932 0.939

66.1 84.3 131.7 170.5 56.2 140.8 196.1 2142.2 3307.9

3 Agentic Learning to Route 3.1 Router’s environment To fully release the routing controller’s effectiveness, we propose to utilize a lightweight, finetuned LLM in a text-conditioned approach. Its role is to interpret a query and select a candidate from the

Measured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM Serving

Preprint, In Submission, Publication time, Publishing venue

Table 2: Held-out avg. results by routing policies. Policy

Accuracy

Reward

Duration (ms)

Model Energy (J)

Power (W)

always_smallest always_medium always_large K-NN UniRoute Smoothie RouteLLM GreenServ (𝜆=0.5)

0.189 0.791 0.735 0.790 0.618 0.599 0.721 0.567

-

3780 28672 25706 23366 15371 9269 25053 5584

91.8 4127.9 5112.4 3697.6 2192.2 1070.6 4958.6 277.0

23.2 125.9 174.5 121.3 84.8 59.4 170.3 41.4

SFT (𝜆𝑒 =0.5) KL-GRPO (sft 𝜆𝑒 =0.5, rl 𝜆𝑒 =0.4, step 500) KL-GRPO (sft 𝜆𝑒 =0.5, rl 𝜆𝑒 =0.4, step 1500) GRPO (sft 𝜆𝑒 =0.5, rl 𝜆𝑒 =0.4, step 750) GRPO (both 𝜆𝑒 =0.3, step 1000)

0.502 0.591 0.609 0.584 0.573

0.400 0.686 0.720 0.666 0.612

7821 10932 13118 7942 6002

822.8 1339.8 1691.9 811.8 297.0

45.9 47.1 55.2 48.5 46.8

fixed pool, using a smaller LLM than most of the candidate answer models. The controller prompt contains the question like in Figure 9, a fixed legend mapping labels to candidate indices, and an optional difficulty field when available. Recorded correctness and energy are regarded as training feedback rather than inputs supplied to the controller when selecting a model. Let 𝑝𝜃 (𝑜 | 𝑥) denote the probability of a controller completion 𝑜 given query context 𝑥, and let 𝑓 (𝑜) parse its route label into an index in A. The induced route probabilities are 𝜋𝜃 (𝑎 | 𝑥) =

∑︁

𝑝𝜃 (𝑜 | 𝑥).

(8)

𝑜: 𝑓 (𝑜 )=𝑎

This separates probabilities over generated texts from probabilities over candidate models. The reported variants use either a short rationale with a Choice: <letter> field or a choice-only completion. Both formats produce the same model-selection action. An un-parseable output is an invalid action and receives a negative training reward. If multi-turn routing is enabled, the router also sees any prior user-assistant turns on the same item, the current turn number, the model chosen last time, and whether that attempt was correct. The router can output Stop when a satisfactory answer already exists and does not require to route to candidate LLMs. In all reported experiments we use a single routing turn per question, so each item produces one letter choice. Probability on invalid outputs is separate from the valid-route probabilities in Eq. (8). The tournament contains approximately 3, 000 collected questions from MMLU, GSM8K, BBH, ARC Challenge, HellaSwag, MATH500, and HumanEval. The reported training and held-out subsets contain 2050 and 913 questions, respectively; benchmark counts are given in Table 5 in Appendix. The training subset supplies oracle labels and policy rewards. We use Qwen2.5-1.5B-Instruct as a text-conditioned routing controller, and find it achieve desired posttraining behaviors with respect to routing decisions. During policy optimization, the controller generates new completions, but answermodel outcomes are retrieved from the fixed training records. Thus, router sampling is on-policy with respect to the rollout controller, while its environmental feedback is replayed.

3.2

Router reward design

To ensure in the RL post-training, the router can effectively evaluate the input queries’ contexts and make decisions, we come up with a ranking-based reward for energy, which is in the same scale as of performance reward. Such a reward combines task correctness with preferences over inference costs (e.g., energy, task completion time and etc). For query 𝑖 and candidate 𝑎, 𝐶𝑖𝑎 denotes correctness, 𝐿𝑖𝑎 denotes answer latency in seconds, and 𝑃¯𝑖𝑎 denotes mean incremental GPU power in watts. The measured answer energy 𝐸𝑖𝑎 is expressed in joules. We derive a bounded energy penalty from the ascending rank 𝑘𝑖𝑎 of 𝐸𝑖𝑎 among the 𝐾 candidates for the same query: 𝑧𝑖𝑎 =

𝑘𝑖𝑎 − 1 ∈ [0, 1]. 𝐾 −1

(9)

With distinct energies, zero corresponds to the cheapest candidate and one to the most expensive. For example, energies of 85, 95, 280, and 3800 J produce penalties of 0, 1/3, 2/3, and 1. 𝑧𝑖𝑎 denotes rank throughout. It is distinct from energy 𝐸𝑖𝑎 . Ranking bounds the training penalty but discards the magnitude of differences in joules. Let 𝝀 = (𝜆𝑡 , 𝜆𝑒 , 𝜆𝑝 ) denote nonnegative cost coefficients associated with session duration, energy rank, and average GPU power. The general gated reward is ( 𝑅correct − 𝜆𝑡 𝐿𝑖𝑎 − 𝜆𝑒 𝑧𝑖𝑎 − 𝜆𝑝 𝑃¯𝑖𝑎 , 𝐶𝑖𝑎 = 1, gen 𝑟𝑖𝑎 (𝝀) = (10) −𝑅wrong, 𝐶𝑖𝑎 = 0. Latency, energy, and power describe related but different aspects of a session call. The same energy can be consumed over different durations and at different mean power levels. Thus, separate latency and power coefficients can express preferences over those quantities, but an average-power penalty alone does not enforce an instantaneous device or rack power limit. We distinguish this general reward form from the configuration used in the reported experiments as follows. Reported experimental configuration. In our paper, all reported reward configurations set 𝜆𝑡 = 𝜆𝑝 = 0 to focus on energy analysis, with 𝑅correct = 2 and 𝑅wrong = 1. The active training reward

Preprint, In Submission, Publication time, Publishing venue

Siddiqui et al.

therefore reduces to ( gen

𝑟𝑖𝑎 (𝜆𝑒 ) = 𝑟𝑖𝑎 (0, 𝜆𝑒 , 0) =

invalid-action penalty. Denote these completion rewards by 𝑟𝑖𝑔 . Their group statistics and advantages are 2 − 𝜆𝑒 𝑧𝑖𝑎 , 𝐶𝑖𝑎 = 1, −1, 𝐶𝑖𝑎 = 0.

(11)

Latency and mean GPU power are consequently measured and reported but do not contribute to the optimization signal in these runs. The oracle and policy-optimization descriptions below use this specialization. Invalid controller outputs and unavailable outcomes receive a separate negative penalty rather than the valid-route score above. Equivalently, the active reward can be expressed as 𝑟𝑖𝑎 (𝜆𝑒 ) = 3𝐶𝑖𝑎 − 1 − 𝜆𝑒 𝐶𝑖𝑎 𝑧𝑖𝑎 .

(12)

Our reward design distinguishes the costs of correct routes, while it assigns equal reward to cheap and expensive failures. It is an empirical learning surrogate instead of a guarantee of meeting the joule budget in Eq. (7). Evaluation counts measured answer energy for successful and unsuccessful answers alike. Joint optimization with nonzero latency or power weights is not demonstrated by the reported experiments.

3.3

𝑎𝑖★ ∈ arg max 𝑟𝑖𝑎 (𝜆𝑒SFT ), 𝑎∈ A

𝐺

1 ∑︁ 𝑟𝑖𝑔 , 𝐺 𝑔=1

𝐴𝑖𝑔 =

𝑟𝑖𝑔 − 𝜇𝑖 , 𝜎𝑖 + 𝛿

𝜎𝑖2 =

1 ∑︁ (𝑟𝑖𝑔 − 𝜇𝑖 ) 2, 𝐺 𝑔=1

(15)

(16)

where 𝛿 > 0 stabilizes the denominator. A completion receives positive advantage when its reward exceeds the group mean. If all group rewards are equal, all advantages are zero and that group supplies no reward-driven update. At the completion level, the clipped policy surrogate is expressed using importance sampling ratio 𝜌𝑖𝑔 (𝜃 ) =

𝑝𝜃 (𝑜𝑖𝑔 | 𝑥𝑖 ) , 𝑝𝜃 old (𝑜𝑖𝑔 | 𝑥𝑖 )

(17)

with the resulting loss function  Lclip (𝜃 ) = −E𝑖, 𝑜𝑖1:𝐺 ∼𝑝𝜃

old

𝐺  1 ∑︁ min 𝜌𝑖𝑔 (𝜃 )𝐴𝑖𝑔 , 𝐺 𝑔=1



Oracle construction and supervised initialization

clip(𝜌𝑖𝑔 (𝜃 ), 1 − 𝜀, 1 + 𝜀 )𝐴𝑖𝑔

, (18)

For each training query, the oracle target is selected from (13)

where 𝜆𝑒SFT is the coefficient used during target construction. Under the reported positive coefficients 0 < 𝜆𝑒SFT ≤ 1, every correct route receives a larger reward than every incorrect route. The oracle therefore selects a lowest-energy correct candidate whenever one exists. Changing the positive coefficient within this range does not change that hard target if the logs and tie-breaking rule are fixed. For 𝜆𝑒SFT = 0, correct candidates tie, and the draft’s accuracy-only configuration selects the largest correct model. When all candidates are wrong, every valid route ties under the reward. Let 𝑜𝑖★ be the target completion encoding the selected route in the chosen output format. Supervised fine-tuning initializes the controller by minimizing the completion negative log likelihood 1 ∑︁ LSFT (𝜃 ) = − log 𝑝𝜃 (𝑜𝑖★ | 𝑥𝑖 ), (14) 𝑁 tr 𝑖 ∈ Itr

where Itr indexes the 𝑁 tr training queries. We can include two variations of the target during SFT. In the choice-only variant, the target contains the route field; in the rationale variant, it also contains an explanation. The resulting checkpoint is evaluated as an SFT baseline and initializes policy optimization. SFT learns to imitate target completions, whereas the next stage uses the rewards of the controller’s sampled decisions.

3.4

𝐺

𝜇𝑖 =

Group-relative policy optimization

We fine-tune the controller using GRPO [22], which is based on groupwise comparisons instead of using a separate critic model. For each training query 𝑖, a rollout policy 𝑝𝜃 old samples 𝐺 completions {𝑜𝑖𝑔 }𝐺 𝑔=1 . Each completion is parsed and scored using its selected candidate’s logged outcome and coefficient 𝜆𝑒RL , or the

where 𝜀 controls ratio clipping and is distinct from the advantage stabilizer 𝛿. This surrogate favors higher-reward completions while limiting the incentive for large policy-ratio changes. The KL-regularized variant additionally anchors the controller to the SFT reference distribution 𝑝 ref . At the distribution level its regularized objective is   LKL-GRPO (𝜃 ) = Lclip (𝜃 ) + 𝛽 E𝑖 𝐷 KL (𝑝𝜃 (· | 𝑥𝑖 )∥𝑝 ref (· | 𝑥𝑖 )) , (19) with regularization coefficient 𝛽 ≥ 0. This expression concerns controller completions rather than the marginal distribution over model indices in Eq. (8). Note KL regularization discourages departure from the reference policy. It does not by itself guarantee diverse routing or prevent concentration on a single candidate. Some configurations center rewards by subtracting a query-specific mean candidate reward. If the same scalar is subtracted from every reward before the group normalization in Eq. (16), it cancels from the advantages. Such centering is therefore distinct from any additional reward scaling or clipping used by the trainer.

3.5

Training configurations and evaluation

In this paper, we unify and use TRL to conduct main SFT, GRPO, and KL-GRPO post-training [25]. Policy optimization starts from the corresponding SFT checkpoint, and samples multiple controller completions per training prompt. The main reported configuration uses 𝜆𝑒SFT = 0.5 and 𝜆𝑒RL = 0.4 as we find these hyperparameters give the best overall SFT and RL performances. The experiments also examine matched coefficients and alternative checkpoints. These coefficients identify the oracle-construction and policy-reward settings, respectively. Their effects on hard targets, sampled rewards, and learned policies should be distinguished. The output format is also part of each configuration, where choice-only and rationalegenerating controllers have different generation costs, and should be compared with their matching SFT initializations.

Measured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM Serving

Preprint, In Submission, Publication time, Publishing venue

On held-out queries, the router can be treated as the lightweight add-on to inference engines, where it selects a LLM model and the evaluator retrieves that candidate’s stored outcome. We report accuracy, mean answer-model energy, latency, power, and modelselection frequencies. Reward comparisons require a common evaluation coefficient: reward values computed with different 𝜆𝑒 are different metrics even when the policies share an SFT initialization. Accuracy and measured joules provide the primary comparison across configurations and baselines. The replay protocol isolates model-selection behavior on a fixed outcome table. It does not measure live changes in answer generation, model-loading cost, queueing, or concurrent serving. Controller measurements and answer-model measurements must therefore be distinguished when discussing complete serving costs. The evaluation addresses single-turn routing on the profiled pool. Adaptation to new deployment conditions and multi-turn agent workflows would require additional experiments.

reward is (1 − 𝜆)Acc − 𝜆 · energy and is warmed on train and frozen before test. It is not the paper’s online serving stack. RouteLLM is a binary router over two fixed answer models, not a classifier over the full pool. We set the strong arm to DeepSeek-R1-32B and the weak arm to the smallest pool model, train a prompt classifier for 𝑃 (strong wins | 𝑞) on train outcomes, and route to the strong arm when 𝑃 ≥ 𝛼. Fixed-model baselines always route every question to one predetermined answer model (always-smallest, always-medium, or always-large). Table 2’s methods and results can be split into three categories. High-accuracy part is K-NN and always-medium, near 79% at about 3.7–4.1 kJ. Always-large and RouteLLM spend even more and do worse. The mid-accuracy band is UniRoute, Smoothie, and our 𝜆𝑒 =0.5→0.4 GRPO runs, about 58 to 62% at 0.8 to 2.2 kJ. The cheap accuracy band is GreenServ 𝜆=0.5 and matched GRPO 𝜆𝑒 =0.3, about 57% at about 280–300 J. SFT sits below that mid accuracy band. Always-smallest is cheap in energy but achieves poor accuracy.

4 Numerical Studies 4.1 Simulation setup

Reward parameters. We instantiate Eq. (11) with 𝑅correct = 2, 𝑅wrong = 1, 𝜆𝑡 = 𝜆𝑝 = 0, and energy-rank mode for 𝑧𝑖𝑎 in Eq. (9), so the training signal is the gated rank surrogate rather than raw joules. Baselines do not train on this gated reward, so the Reward column is blank for them. For SFT and TRL the Reward column is therefore 𝑟𝑖𝑎 (𝜆𝑒 ) = 2 − 𝜆𝑒 𝑧𝑖𝑎 when 𝐶𝑖𝑎 = 1 and 𝑟𝑖𝑎 (𝜆𝑒 ) = −1 when 𝐶𝑖𝑎 = 0. SFT uses 𝜆𝑒SFT = 0.5. The 0.5→0.4 runs use 𝜆𝑒RL = 0.4. Matched GRPO uses 𝜆𝑒 = 0.3 at both stages. Therefore, the Reward column should only be compared among SFT and RL inside that block. Acc., duration, energy, and power are the columns to compare across the whole table.

We evaluate routing policies on a mixed held-out set drawn from seven standard LLM benchmarks that span knowledge, reasoning, math, and code: MMLU [9], GSM8K [6], BBH [24], ARC Challenge [5], HellaSwag [31], MATH-500 [16], and HumanEval [1]. These tasks stress different answer-model costs, from short multiplechoice items to long generative math and code solutions, which is what makes energy-aware routing nontrivial on the shared ninemodel pool. Power measurements are based on NVML1 . To keep energy and performance comparable across runs, we place all router training and inference on instances of 1x NVIDIA A100-SXM4-80,GB GPU. Answer models are served locally with Ollama on that device. The routing controller is trained and run with Hugging Face Transformers and TRL on the same GPU. Offline tournament collection used a separate machine with the same A100-SXM4-80,GB specification. During each answer-model generation we poll board power at a 200, ms interval via pynvml, subtract a short idle baseline, and record instantaneous effective power 𝑃 (𝑡) in watts. Effective energy for the call is 𝐸 call as defined in Eq. (3), and latency is the wall-clock duration of the same generation. Table 2 reports mean accuracy together with measured duration, energy, and power summaries. All methods in Table 2 are scored on the same 913 held-out test questions covering a spectrum of difficulty and categories. Scoring is an offline lookup from pre-logged tournament outcomes, so no live answer models are called at evaluation time. K-NN estimates per-model error from embedding neighbours on the train split. UniRoute matches the prompt to a train K-means cluster and uses per-cluster error plus a cost term. Smoothie is labelfree and scores candidates from embeddings of [prompt, output]. Unlike the others, Smoothie uses all logged generations on the test item to decide, so we treat it as an output-informed diagnostic rather than a deployable pre-generation router. We still charge only the chosen model’s energy. GreenServ is a local LinUCB adaptation using task, semantic-cluster, and complexity features. Its 1 https://docs.nvidia.com/deploy/nvml-api/

4.2

Results and discussion

Accuracy–energy trade-off among routers. Figure 4 plots mean answer energy against held-out accuracy for the multi-model routers and learned controllers in Table 2. Blue markers are TRL sweeps over 𝜆𝑒 [25]. Among these methods, policies near 60% accuracy generally report higher mean answer energy than policies that remain below that range by relying more on smaller candidate models. The additional energy is associated with more frequent selection of higher-capacity models when the query requires it. This comparison is restricted to the routers and learned controllers under study, and does not describe every fixed model in the candidate pool. Under that scope, KL-GRPO step 1500 nearly matches UniRoute (60.9% vs. 61.8%) at about 23% lower mean answer energy (1.69 vs. 2.19 kJ), and GRPO step 750 is close to Smoothie (58.4% vs. 59.9%) at about 24% lower mean answer energy (0.81 vs. 1.07 kJ). Role of a mid-size fixed arm. A single mid-size model can look strong in aggregate. For instance, Always-Qwen2.5-7B reaches 64.0% at 281.6 J on the same held-out pool (Appendix Table 8), yet it falls to 26.2% on BBH, so locking every query to 7B can be a weak suite-wide policy even when the overall mean is high. Our question is not whether one fixed arm can beat a router on the mean, but whether a controller can allocate capacity across the pool. Figure 3 shows that the learned and classical routers still call 7B when it helps. About one third of GreenServ traffic, about 30% for GRPO step 750, and about 19% for both KL-GRPO step 1500 and

Preprint, In Submission, Publication time, Publishing venue

Siddiqui et al.

Figure 3: Model selection mix on the held-out set. for roughly 52–57% of the total energy for SFT and KL-GRPO. For KL-GRPO step 1500 the mean MATH-500 cost is about 6.6 kJ, while MMLU, GSM8K, ARC, and HellaSwag each sit near 0.15–0.49 kJ. This concentration is why the mid-accuracy rows still appear expensive in the aggregate table. Collapsed cheap policies reduce the MATH-500 spend, but Table 3 shows that their MATH-500 accuracy also drops sharply, from 73.8% under KL-GRPO step 1500 to 36.9% under GRPO 𝜆𝑒 =0.3.

Figure 4: Energy–Accuracy Pareto comparison.

UniRoute use that arm. In that sense 7B is a useful pool member, not a substitute for routing. MATH-500 accounts for most energy. Table 4 breaks energy down by benchmark and shows how uneven the cost is. Although MATH-500 is only about 13% of the held-out questions, it accounts

Accuracy is as uneven as energy. Table 3 is the matching accuracy view of the same methods. The Overall row recovers Table 2, but the benchmark mix is not uniform. K-NN is the only method that stays high on both BBH (69.7%) and HumanEval (93.2%); RouteLLM and always-large look strong in the aggregate table while BBH falls to 27.6%. SFT’s 50.2% overall hides a GSM8K collapse to 38.2%, whereas GRPO step 750 recovers GSM8K to 75.7% without moving into the high-energy band of Table 4. GreenServ 𝜆=0.5 and GRPO 𝜆𝑒 =0.3 remain close overall (56.7% against 57.3%), yet GreenServ keeps MATH-500 at 57.4% while the collapsed GRPO run does not. KL-GRPO step 1500 is the mixed mid-band policy that is even on ARC (78.2%) and MATH-500 (73.8%), weaker on MMLU (45.6%), and still far cheaper than K-NN. Choice of energy coefficient. Our main-table schedule uses 𝜆𝑒 =0.5 at SFT and 𝜆𝑒 =0.4 at RL. The SFT stage is meant to warmstart an energy-aware hybrid policy that still under-routes toward cheap drafts. Softening the RL weight to 0.4 and adding a KL term to that SFT policy is what lets the controller abandon the 0.5B model and move into the mid-accuracy band without jumping to always-large energy. Setting both stages to 𝜆𝑒 =0.3 instead collapses

Measured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM Serving

always smallest

GreenServ (𝜆=0.5) Smoothie

SFT (𝜆𝑒 =0.5)

Preprint, In Submission, Publication time, Publishing venue

GRPO KL-GRPO (step 750) (step 1500)

GRPO (𝜆𝑒 =0.3)

UniRoute

K-NN

always RouteLLM largest

MMLU GSM8K BBH ARC Challenge HellaSwag Math500 HumanEval

14.8 27.1 2.1 35.4 20.4 17.2 10.2

51.7 82.6 15.2 77.6 57.1 57.4 54.2

49.0 86.8 42.8 76.2 53.7 62.3 33.9

40.9 38.2 33.8 74.8 40.8 73.0 57.6

38.9 75.7 40.7 76.2 57.8 67.8 47.5

45.6 61.8 52.4 78.2 58.5 73.8 54.2

50.3 81.2 37.9 78.9 57.8 36.9 50.8

58.4 70.8 19.3 87.1 62.6 78.7 52.5

63.1 89.6 69.7 88.4 72.8 86.1 93.2

65.1 89.6 27.6 91.2 66.7 85.2 94.9

65.1 89.6 27.6 92.5 72.8 86.9 94.9

Overall

18.9

56.7

59.9

50.2

58.4

60.9

57.3

61.8

79.0

72.1

73.5

Table 3: Held-out accuracy (%) by benchmark and method.

Figure 5: SFT versus KL-GRPO effective GPU power on four held-out questions. Each panel shows the controller decode followed by the answer-model draft, with energy for the router, the answer, and their sum. Title deltas are relative to that total. onto Llama 3.1 8B on every scored item, with zero selection entropy and the weak MATH-500 accuracy in Table 3. Under our hybrid energy-rank setup, any positive 𝜆𝑒 yields the same hard oracle labels (cheapest correct), so the SFT tags 0.3/0.5/0.7 share supervision targets. What changes across these runs is the learned policy and, during RL training, the reward scale. In Table 2, KL-GRPO step 1500 is the run among our own policies that reaches above 60% held-out accuracy at the lowest mean answer energy in that band. The fuller sweep is in Table 6.

SFT versus RL. Reinforcement learning improves the controller from 50.2% accuracy under SFT to 60.9% at KL-GRPO step 1500. This gain raises mean answer energy from 0.82 to 1.69 kJ because the policy is more willing to pay for a capable answer model rather than fail cheaply. GRPO step 750 provides a different operating point, improving accuracy to 58.4% while reducing mean answer energy to 0.81 kJ. The four examples in Figure 5 show that this trade-off is question dependent. In HumanEval q149, RL is correct and cheaper overall; in HellaSwag q395 and ARC-Challenge q438 it is also cheaper once router energy is included with the answer; and in MATH-500 q351, RL avoids an expensive failed R1-32B call

Preprint, In Submission, Publication time, Publishing venue

by selecting R1-8B. Thus, the effect of RL is not a uniform increase in model size or energy, but a change in when additional answermodel cost is justified for higher accuracy. The more direct training effect appears in the controller itself, which is the only component updated by SFT and RL. On the heldout evaluation, the majority of SFT controller decisions take about 2.4 s, whereas after KL-GRPO the majority take about 0.8 s, a reduction of more than 50%, which is also visible as the shorter controller band in every panel of Figure 5. Because we train only this decoder, that drop is a change in how the router writes. Indeed, SFT still spends time on a longer reasoning plus choice generation, while RL concentrates the policy on a short, decisive letter, so the routing step itself becomes cheaper even when the chosen answer model is not. The effective controller power in the plotted traces is consistent across more than thirty checkpoints from various policies collected on that A100-SXM4-80 GB machine; a later session on the same GPU class produced a different effective draw for otherwise similar decisions, so controller watts follow machine condition more than the training method. Latency, not power, is therefore the reliable signal of how RL changed the controller in our runs. LLM controller versus bandits and plug-in routers. Figure 3 highlights the difference between an LLM-based router and the baselines visible through their model-selection patterns. K-NN, UniRoute, Smoothie, GreenServ, and RouteLLM never decode a routing token. They instead score embeddings, clusters, or a small contextual vector, and then apply an explicit knob such as 𝜆, 𝛼, or a cost term. For instance, GreenServ 𝜆=0.5 stays cheap while still mixing mid-size models, K-NN stays R1-heavy because neighbours vote for correctness (418 calls to R1-8B and 294 to R1-32B), and RouteLLM mostly sends traffic to its strong arm. In contrast, the LLM controller directly reads the question and generates a routing decision, enabling task-dependent model selection without a hand-built feature map. Under KL-GRPO, this results in structured routing, with MATH-500 favoring R1-8B, ARC favoring Gemma, and GSM8K favoring the 1.5B model, which is the mix behind the even ARC and MATH-500 accuracies in Table 3. KL-GRPO step 1500 distributes traffic across eight models and rarely selects R1-32B, while GRPO step 750 is more 7B-heavy and achieves lower energy at similar accuracy. Therefore, it can match UniRoute without large energy. GRPO step 750 is likewise mixed but more 7B-heavy, and therefore cheaper at similar accuracy. This flexibility can also lead to policy collapse. With GRPO 𝜆𝑒 =0.3, the controller routes every item to Llama 3.1 8B, yielding zero selection entropy despite competitive aggregate energy and accuracy. SFT with 𝜆𝑒 =0.5 exhibits a different failure mode by over-selecting the 0.5B model, resulting in low energy but poor accuracy. These results show that effective learned routing requires maintaining task-dependent model diversity rather than collapsing to a single low-cost decision. Latency and Power Profile. Latency and power follow the same three sections as accuracy. The high accuracy rows draw about 121–175 W and take roughly 23–29 s, the middle rows draw about 48–85 W over 8–15 s, and the cheap rows draw about 23–47 W over 4–6 s. The TRL policies sit with Smoothie and GreenServ on power, at 47–55 W, rather than with K-NN, and KL-GRPO step 1500 is faster than UniRoute while GRPO step 750 is faster than Smoothie.

Siddiqui et al.

As before, the duration column reflects answer model time only and does not include the controller. Controller Energy and Training Cost. Tables report answermodel energy because that is the quantity every baseline shares once a model is chosen. Most competing methods are lookup-like mechanisms rather than generative controllers like the proposed router, so we do not invent a controller cost for them or leave the comparison on a guess. To visualize detailed per-session power trajectories and investigate where the energy savings come from, Figure 5 and Figure 10 (in Appendix) use the same four held-out questions, the same 200 ms NVML protocol, and NVIDIA A100SXM4-80,GB GPUs. Across those sessions, the occupancy of the board actively change. SFT remains the longer decode (about 2.4– 2.8 s versus about 0.8–1.0 s) because the warm-start samples emit more tokens, so SFT controller energy stays higher even when both runs sit on the same GPU class. KL-GRPO controller energy is a few hundred joules in Fig. 5 and about 30 J in Fig. 10. Latency, not watts, is the stable signature of the training paradigm shift. That add-on is why we keep answer model energy in the main table. That cost grows if the controller itself is scaled (larger routers, more candidates, or multi-turn routing) or if the pool moves to still heavier answer models (48B and above), where choosing which call to issue matters more than zeroing the router’s own watts. Depending on steps, total training cost varied from 0.32 to 0.40 kWh, and remains a one-time investment for a reusable local controller.

5

Conclusion and Future Works

In this work, we look into the energy-efficient serving of LLMs, and investigate novel energy-aware model routing over a heterogeneous pool of LLMs via measured GPU energy rank. Instead of using model size or monetary cost as the routing objective, we propose to let the small language-model router evaluate the whole input queries’ contexts and energy profiles and finetune the router model with tailored reward. Such a router firstly undergoes supervised fine-tuning followed by reinforcement learning-based policy optimization, and it is evaluated on a held-out mixture of seven benchmarks using offline tournament logs. Our results show that routing policies can substantially alter the accuracy–energy trade-off by changing how frequently different answer models are selected, leading to a promising energy-efficient LLM serving paradigm while preserving model performance. Compared with router trained via SFT, selected GRPO and KL-GRPO checkpoints achieve higher held-out accuracy while maintaining lower energy than several existing routing baselines at comparable operating points. These results suggest that just with a mild-size finetuning dataset collected, measured-energy-aware routing is effective for realizing heterogeneous LLM serving, which is both flexible and generalizable across task domains. Future work will evaluate the approach in online and multi-turn agent settings and account for the full end-to-end routing overhead.

References [1] Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021). [2] Chu, T., Zhai, Y., Yang, J., Tong, S., Xie, S., Schuurmans, D., Le, Q. V., Levine, S., and Ma, Y. Sft memorizes, rl generalizes: A comparative study of foundation model post-training. arXiv preprint arXiv:2501.17161 (2025).

Measured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM Serving

always smallest

GreenServ (𝜆=0.5) Smoothie

SFT (𝜆𝑒 =0.5)

Preprint, In Submission, Publication time, Publishing venue

GRPO KL-GRPO (step 750) (step 1500)

GRPO (𝜆𝑒 =0.3)

UniRoute

K-NN

always RouteLLM largest

MMLU GSM8K BBH ARC Challenge HellaSwag Math500 HumanEval

116.5 91.1 62.7 72.0 45.5 167.6 110.3

213.6 306.0 122.2 157.3 127.5 854.0 223.7

228.8 1.5k 1.4k 161.8 154.3 3.8k 388.0

146.3 159.5 1.6k 77.6 100.5 3.5k 312.1

123.9 389.2 2.3k 129.7 382.0 2.0k 178.3

154.2 489.1 2.5k 180.1 162.3 6.6k 3.9k

263.4 331.9 305.8 204.4 176.6 571.4 238.0

2.8k 572.6 234.0 1.9k 1.5k 6.3k 3.3k

1.1k 2.3k 3.6k 2.0k 2.9k 9.6k 7.8k

3.9k 2.4k 4.3k 3.2k 3.4k 8.6k 16.5k

3.9k 2.4k 4.4k 3.3k 3.8k 9.1k 16.5k

Overall

91.8

277.0

1.1k

822.8

811.8

1.7k

297.0

2.2k

3.7k

5.0k

5.1k

Table 4: Mean effective answer energy (J) by benchmark and method.

[3] Chung, J.-W., Ma, J. J., Wu, R., Liu, J., Kweon, O. J., Xia, Y., Wu, Z., and Chowdhury, M. The ml. energy benchmark: Toward automated inference energy measurement and optimization. Advances in Neural Information Processing Systems 38 (2026). [4] Chung, J.-W., Wu, R., Ma, J. J., and Chowdhury, M. Where do the joules go? diagnosing inference energy consumption. arXiv preprint arXiv:2601.22076 (2026). [5] Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457 (2018). [6] Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168 (2021). [7] Guha, N., Chen, M. F., Chow, T., Khare, I. S., and Re, C. Smoothie: Label free language model routing. Advances in Neural Information Processing Systems 37 (2024), 127645–127672. [8] He, X., Fang, Z., Lian, J., Tsang, D. H., Zhang, B., and Chen, Y. Freesh: Fair, resource-and energy-efficient scheduling for llm serving on heterogeneous gpus. arXiv preprint arXiv:2511.00807 (2025). [9] Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300 (2020). [10] Hu, Q. J., Bieker, J., Li, X., Jiang, N., Keigwin, B., Ranganath, G., Keutzer, K., and Upadhyay, S. K. Routerbench: A benchmark for multi-llm routing system. arXiv preprint arXiv:2403.12031 (2024). [11] Jiang, D., Lu, Y., Li, Z., Lyu, Z., Nie, P., Wang, H., Su, A., Chen, H., Zou, K., Du, C., et al. Verltool: Towards holistic agentic reinforcement learning with tool use. arXiv preprint arXiv:2509.01055 (2025). [12] Jitkrittum, W., Narasimhan, H., Rawat, A. S., Juneja, J., Wang, C., Wang, Z., Go, A., Lee, C.-Y., Shenoy, P., Panigrahy, R., et al. Universal model routing for efficient llm inference. In International Conference on Learning Representations (2026), vol. 2026, pp. 10169–10218. [13] Li, H., Zhang, Y., Guo, Z., Wang, C., Tang, S., Zhang, Q., Chen, Y., Qi, B., Ye, P., Bai, L., et al. Llmrouterbench: A massive benchmark and unified framework for llm routing. In Findings of the Association for Computational Linguistics: ACL 2026 (2026), pp. 37733–37754. [14] Li, L., Chu, W., Langford, J., and Schapire, R. E. A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th international conference on World wide web (2010), pp. 661–670. [15] Li, Y., Mughees, M., Chen, Y., and Li, Y. R. The unseen ai disruptions for power grids: Llm-induced transients. arXiv preprint arXiv:2409.11416 (2024). [16] Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K. Let’s verify step by step. In International Conference on Learning Representations (2024), vol. 2024, pp. 39578– 39601. [17] Luccioni, S., Jernite, Y., and Strubell, E. Power hungry processing: Watts driving the cost of ai deployment? In Proceedings of the 2024 ACM conference on fairness, accountability, and transparency (2024), pp. 85–99. [18] Luo, X., Zhang, Y., He, Z., Wang, Z., Zhao, S., Li, D., Qiu, L. K., and Yang, Y. Agent lightning: Train any ai agents with reinforcement learning. arXiv preprint arXiv:2508.03680 (2025). [19] Ong, I., Almahairi, A., Wu, V., Chiang, W.-L., Wu, T., Gonzalez, J. E., Kadous, M., and Stoica, I. Routellm: Learning to route llms from preference data. In International Conference on Learning Representations (2025), vol. 2025, pp. 34433– 34448. [20] Oviedo, F., Kazhamiaka, F., Choukse, E., Kim, A., Luers, A., Nakagawa, M., Bianchini, R., and Ferres, J. M. L. Energy use of ai inference, efficiency pathways,

and test-time scaling. Joule (2026). [21] Patel, P., Choukse, E., Zhang, C., Goiri, Í., Warrier, B., Mahalingam, N., and Bianchini, R. Characterizing power management opportunities for llms in the cloud. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 (2024), pp. 207–222. [22] Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., Wu, Y., et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300 (2024). [23] Song, W., Huang, Z., Cheng, C., Gao, W., Xu, B., Zhao, G., Wang, F., and Wu, R. Irt-router: Effective and interpretable multi-llm routing via item response theory. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (2025), pp. 15629–15644. [24] Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q., Chi, E. H., Zhou, D., et al. Challenging big-bench tasks and whether chain-of-thought can solve them. In Findings of the Association for Computational Linguistics: ACL 2023 (2023), pp. 13003–13051. [25] von Werra, L., Belkada, Y., Tunstall, L., Beeching, E., Thrush, T., Lambert, N., Huang, S., Rasul, K., and Gallouédec, Q. TRL: Transformers Reinforcement Learning, 2020. [26] Wang, D., Liu, B., Lu, R., Zhang, Z., and Zhu, S. Storellm: Energy efficient large language model inference with permanently pre-stored attention matrices. In Proceedings of the 16th ACM International Conference on Future and Sustainable Energy Systems (2025), pp. 398–406. [27] Wilkins, G., Keshav, S., and Mortier, R. Hybrid heterogeneous clusters can lower the energy consumption of llm inference workloads. In Proceedings of the 15th ACM international conference on future and sustainable energy systems (2024), pp. 506–513. [28] Yang, J., Hou, B., Wei, W., Bao, Y., and Chang, S. Ares: Adaptive reasoning effort selection for efficient llm agents. arXiv preprint arXiv:2603.07915 (2026). [29] You, J., Chung, J.-W., and Chowdhury, M. Zeus: Understanding and optimizing { GPU } energy consumption of { DNN } training. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) (2023), pp. 119–139. [30] Yu, Q., Zhang, Z., Zhu, R., Yuan, Y., Zuo, X., Yue, Y., Dai, W., Fan, T., Liu, G., Liu, L., et al. Dapo: An open-source llm reinforcement learning system at scale. Advances in Neural Information Processing Systems 38 (2026), 113222–113244. [31] Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th annual meeting of the association for computational linguistics (2019), pp. 4791–4800. [32] Zhao, H., Yan, Y., Shen, Y., Xu, H., Zhang, W., Song, K., Shao, J., Lu, W., Xiao, J., and Zhuang, Y. Let lrms break free from overthinking via self-braking tuning. Advances in Neural Information Processing Systems 38 (2026), 1861–1887. [33] Ziller, T., Ilager, S., Tundo, A., Bartocci, E., Mariani, L., and Brandic, I. Greenserv: Energy-efficient context-aware dynamic routing for multi-model llm inference. arXiv preprint arXiv:2601.17551 (2026).

Preprint, In Submission, Publication time, Publishing venue

A

Siddiqui et al.

Simulation Details

Table 5 lists how many questions we collected per benchmark and how they split into train and held-out. We originally collected 3,000 tournament items. Before building the final evaluation set we dropped 37 questions where at least one pool model failed to load or return a usable log, so those incomplete items never enter train or held-out scoring. The retained mix is 2,050 train and 913 held-out questions. Train prefixes and held-out suffixes are benchmark-specific as in the table. Every retained question was run through all nine answer models to build the offline correctness and energy logs used for oracle labeling, SFT, RL, and replay evaluation. Table 5: Number of questions per benchmark used for training and held-out evaluation. Collected counts are before removing 37 incomplete tournament items.

B

Benchmark

Collected

Train

Held-out

MMLU GSM8K BBH ARC Challenge HellaSwag MATH-500 HumanEval

475 473 472 472 472 472 164

325 325 325 325 325 325 100

149 144 145 147 147 122 59

Total

3,000

2,050

913

Additional Simulation Results

The duration plots explain why KL-GRPO step 1500 is cheaper than UniRoute at almost the same held-out accuracy. Figure 7 shows that 564 of 913 KL-GRPO answers finish in 0–5 s, versus 212 for UniRoute, while UniRoute places 307 items in the 15–30 s band against 114 for KL-GRPO. The stacks are model-specific. KLGRPO fills the short bins with Qwen2.5-1.5B, Gemma2-2B, and Qwen2.5-7B. UniRoute fills the 15–30 s band with DeepSeek-R1-8B and DeepSeek-R1-32B. Neither policy uses many calls above 60 s (27 versus 16). Figure 6 splits the same bins by task. KL-GRPO’s short bin is mostly GSM8K, ARC, and HellaSwag. UniRoute’s 5– 8 s peak is BBH-heavy, and its 15–30 s peak mixes MMLU, ARC, and HellaSwag. MATH-500 and HumanEval dominate the 30 s and longer bands for both policies, which is why a minority of items still drive mean energy. Figure 8 shows two cases where the controller picks a small correct arm instead of a large one. GSM8K q448 goes to Qwen2.51.5B at 87.6 J. MMLU q358 goes to Gemma2-2B at 108.8 J. The traces are illustrative. The written reasons are generic and should not be read as causal explanations of the policy. Figure 9 is the task-conditional view of the same mix. KL-GRPO step 1500 uses Qwen2.5-1.5B on GSM8K, DeepSeek-R1-8B on MATH500, and Gemma2-2B on ARC. GRPO step 750 is 7B-heavy on GSM8K and MATH and Gemma-heavy on ARC. SFT still puts mass on Qwen2.5-0.5B on GSM8K. RouteLLM stays on R1-32B. GRPO 𝜆𝑒 =0.3 locks onto Llama 3.1 8B on every task. That is why aggregate mid-band rows can hide collapse or under-routing.

Table 6 orders the TRL sweep by mean answer energy. Matched 𝜆𝑒 =0.3 is cheap and collapsed (57.3% at 297 J). Matched 𝜆𝑒 =0.7 stays below 44% accuracy. The 0.5→0.4 family occupies the mid band. KL-GRPO step 1500 is the highest-accuracy run in that family at 60.9% and 1691.9 J. GRPO step 750 is the cheaper 0.5→0.4 operating point at 58.4% and 811.8 J. Runs warmed from SFT (𝜆𝑒 =1.0) jump to about 74% accuracy at more than 5 kJ because that warm start is the accuracy-oracle policy, not a hybrid 𝜆𝑒 =1.0 label set. Table 7 makes the same split at SFT. Hybrid warm starts sit near 45–54% and 0.75–0.82 kJ. SFT (𝜆𝑒 =1.0) is sft_acc at 73.8% and 4706.6 J. Table 8 is the pool profile that those policies draw from. On heldout testing, Always-Qwen2.5-7B is 64.0% at 281.6 J, but BBH is only 26.2%, indicating Qwen2.5-7B’s imbalanced performance across benchmarks. DeepSeek-R1-8B is the strongest overall arm (79.1%), but also comes as the expensive always-medium baseline (4127.9 J). DeepSeek-R1-32B is weaker than R1-8B on BBH (27.6% versus 69.7%) despite higher energy. Train and held-out accuracy ranks are similar. Train energy is higher because long MATH generations are more common in the train prefix. Figure 10 repeats the four-question layout of Fig. 5 on a later A100-SXM4-80GB session using the KL-GRPO step 1500 of a different KL beta variation and the matching SFT eval. The KL-GRPO controller is still about 0.8 s and here about 27–32 J. SFT is still about 2.8 s and about 470–560 J. Answer traces for HumanEval q149 and ARC q438 overlap because both policies pick the same model. The large SFT fill is extra controller tokens, not a different GPU. The two figures disagree on absolute watts because board occupancy differed across sessions on the same SKU. Controller energy is therefore a mild add-on to Table 2, not a reason to move the main comparison off answer joules.

Measured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM Serving

Preprint, In Submission, Publication time, Publishing venue

Figure 6: Held-out answer-duration band counts by benchmark for KL-GRPO step 1500 and UniRoute.

Figure 7: Held-out answer-duration band counts aggregated over all benchmarks for KL-GRPO step 1500 and UniRoute.

Example 1. GSM8K, question 448. A salesman bought 48 sneakers for $576. He sold 17 at a Monday flash sale for $20 each and the remaining 31 to a department store for $25 each. What was his profit?

Example 2. MMLU, question 358. Why does the sky appear blue? (multiple choice; correct option D: molecules scatter blue light more effectively than red light.)

Qwen2.5-1.5B (correct) 87.6 J Controller. Choice: B REASONING: The chosen model (qwen2.5:1.5b) has been shown to perform well on tasks involving financial calculations. Answer. Cost per pair $576/48 = $12. Flash sale 17 × $20 = $340. Store 31 × $25 = $775. Profit $340 + $775 − $576 = $539. Final answer: 539

Gemma2-2B (correct) 108.8 J Controller. Choice: E REASONING: The chosen model (gemma2:2b) is known for its large language model capacity. Answer. Blue light is scattered more by air molecules than other colors. Final answer: D

Figure 8: Two held-out routing examples. GSM8K q448 to Qwen2.5-1.5B (87.6 J). MMLU q358 to Gemma2-2B (108.8 J). Controller text is from a free-form run. Answers and joules are tournament logs.

Preprint, In Submission, Publication time, Publishing venue

Siddiqui et al.

Figure 9: Held-out model-selection mix on GSM8K, MATH-500, and ARC Challenge for the non-always policies. KL-GRPO step 1500 is mostly Qwen2.5-1.5B on GSM8K, DeepSeek-R1-8B on MATH-500, and Gemma2-2B on ARC; GRPO step 750 is 7B-heavy on GSM8K/MATH and Gemma-heavy on ARC; SFT still places mass on the 0.5B model on GSM8K. RouteLLM stays on R1-32B and GRPO 𝜆𝑒 =0.3 locks onto Llama 3.1 8B on every task. Table 6: TRL RL checkpoints from Table 2 and the blue-circle sweep in Fig. 4, ordered by increasing mean energy. Choice only means the controller is trained to emit a route letter (Choice: A) without a free-form reasoning line. Checkpoint

Acc. Dur. (ms) Energy (J) Power (W)

GRPO step 1000 (SFT 𝜆𝑒 =0.3, RL 𝜆𝑒 =0.3) GRPO step 750 (SFT 𝜆𝑒 =0.7, RL 𝜆𝑒 =0.7, run2) GRPO step 1000 (SFT 𝜆𝑒 =0.7, RL 𝜆𝑒 =0.7, run3) GRPO step 1500 (SFT 𝜆𝑒 =0.7, RL 𝜆𝑒 =0.7) KL-GRPO step 500 (SFT 𝜆𝑒 =0.7, RL 𝜆𝑒 =0.7) KL-GRPO step 1536 (SFT 𝜆𝑒 =0.3, RL 𝜆𝑒 =0.3) KL-GRPO step 500 (SFT 𝜆𝑒 =0.5, RL 𝜆𝑒 =0.4, choice only) GRPO step 750 (SFT 𝜆𝑒 =0.5, RL 𝜆𝑒 =0.4, choice only) KL-GRPO step 1750 (SFT 𝜆𝑒 =0.5, RL 𝜆𝑒 =0.4, choice only) KL-GRPO step 500 (SFT 𝜆𝑒 =0.3, RL 𝜆𝑒 =0.3) KL-GRPO step 500 (SFT 𝜆𝑒 =0.5, RL 𝜆𝑒 =0.4) KL-GRPO step 500 (SFT 𝜆𝑒 =0.2, RL 𝜆𝑒 =0.2) GRPO step 1500 (SFT 𝜆𝑒 =0.5, RL 𝜆𝑒 =0.4, choice only) GRPO step 1000 (SFT 𝜆𝑒 =0.5, RL 𝜆𝑒 =0.4) KL-GRPO step 1500 (SFT 𝜆𝑒 =0.5, RL 𝜆𝑒 =0.4) GRPO step 1750 (SFT 𝜆𝑒 =0.5, RL 𝜆𝑒 =0.4, choice only) GRPO step 1250 (SFT 𝜆𝑒 =0.2, RL 𝜆𝑒 =0.2) GRPO step 2000 (SFT 𝜆𝑒 =1.0, RL 𝜆𝑒 =0.3) GRPO step 1250 (SFT 𝜆𝑒 =1.0, RL 𝜆𝑒 =0.6) GRPO step 500 (SFT 𝜆𝑒 =1.0, RL 𝜆𝑒 =1.0)

0.573 0.421 0.432 0.384 0.527 0.503 0.524 0.584 0.563 0.549 0.591 0.552 0.581 0.575 0.609 0.589 0.590 0.744 0.736 0.738

6002 5634 6821 6070 7286 7410 7987 7942 8703 9814 10932 11809 11860 12561 13118 11764 14076 25574 25581 25825

297.0 354.7 561.8 569.3 648.7 719.2 786.6 811.8 986.8 1296.8 1339.8 1526.8 1588.5 1599.5 1691.9 1796.2 1875.9 5033.7 5051.1 5126.7

46.8 30.4 31.8 27.9 45.5 44.4 46.0 48.5 50.9 48.6 47.1 49.8 54.8 49.2 55.2 78.9 52.9 170.9 173.5 174.2

Table 7: SFT warm-start checkpoints for the TRL sweep and Table 2. The labelled SFT 𝜆𝑒 =0.5 marker in Fig. 4 matches the 𝜆𝑒 =0.5 row. SFT (𝜆𝑒 =1.0) is the accuracy-oracle warm start sft_acc, not a hybrid 𝜆𝑒 =1.0 label set. Checkpoint

Acc.

Dur. (ms)

Energy (J)

Power (W)

SFT (𝜆𝑒 =0.7) SFT (𝜆𝑒 =0.3) SFT (𝜆𝑒 =0.5, choice only) SFT (𝜆𝑒 =0.5) SFT (𝜆𝑒 =1.0)

0.468 0.456 0.544 0.502 0.738

7517 7588 7912 7821 24414

749.9 776.7 779.9 822.8 4706.6

45.5 45.0 47.0 45.9 160.3

Measured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM Serving

Preprint, In Submission, Publication time, Publishing venue

Table 8: Held-out per-model accuracy (%) and mean effective answer energy (J) on 913 questions. Overall is item-weighted.

Model

MMLU GSM8K BBH

ARC HellaSwag MATH-500 HumanEval Overall

Qwen 2.5 0.5B

Acc. Energy

14.8 116.5

27.1 91.1

2.1 62.7

35.4 72.0

20.4 45.5

17.2 167.6

10.2 110.3

18.9 91.8

Qwen 2.5 1.5B

Acc. Energy

37.6 126.5

59.0 129.7

8.3 84.0

68.7 92.9

44.9 66.1

32.8 226.7

30.5 105.0

41.4 117.1

Qwen 2.5 3B

Acc. Energy

53.7 198.4

64.6 233.5

11.0 195.3

74.8 136.7

59.9 124.4

49.2 497.6

33.9 167.7

51.2 219.6

Phi-3 Mini 3.8B

Acc. Energy

49.0 236.1

50.0 297.0

6.2 233.8

78.9 177.5

55.1 172.3

18.0 471.4

20.3 246.7

42.2 257.8

Gemma 2 2B

Acc. Energy

38.3 108.5

55.6 141.3

10.3 100.6

74.8 63.2

56.5 46.4

20.5 227.1

25.4 92.1

42.2 109.9

Qwen 2.5 7B

Acc. Energy

59.1 219.9

84.0 311.6

26.2 263.4

84.4 149.2

61.9 141.0

67.2 701.0

67.8 221.5

64.0 281.6

Llama 3.1 8B

Acc. Energy

50.3 263.4

81.2 331.9

37.9 305.8

78.9 204.4

57.8 176.6

36.9 571.4

50.8 238.0

57.3 297.0

DeepSeek-R1 8B

Acc. 59.7 Energy 2997.6

92.4 3318.5

69.7 93.2 3626.6 2095.0

70.7 3744.1

86.1 8882.8

89.8 6379.1

79.1 4127.9

DeepSeek-R1 32B Acc. 65.1 Energy 3874.6

89.6 2354.2

27.6 92.5 4403.3 3276.1

72.8 3750.8

86.9 9080.4

94.9 16475.5

73.5 5112.4

Table 9: Train per-model accuracy (%) and mean effective answer energy (J) on 2,050 questions. Overall is item-weighted.

Model

MMLU GSM8K BBH

ARC HellaSwag MATH-500 HumanEval Overall

Qwen 2.5 0.5B

Acc. Energy

20.9 125.8

34.5 300.7

6.8 859.4

29.5 63.4

29.5 57.6

11.7 2063.2

30.0 117.1

22.5 555.8

Qwen 2.5 1.5B

Acc. Energy

42.5 151.3

61.5 888.4

18.8 803.3

65.5 80.4

40.0 72.9

32.0 2000.3

64.0 100.8

44.4 638.5

Qwen 2.5 3B

Acc. Energy

45.8 203.6

52.0 250.4

23.7 698.4

71.7 129.4

56.6 135.1

46.2 1198.4

52.0 172.6

49.5 423.1

Phi-3 Mini 3.8B

Acc. Energy

41.8 249.5

58.5 1806.7

16.9 80.9 4574.7 167.4

55.4 181.6

11.1 2971.3

30.0 238.2

43.4 1589.2

Gemma 2 2B

Acc. Energy

40.9 112.5

51.7 153.5

21.2 92.3

71.1 53.1

53.5 56.2

14.8 1740.8

45.0 111.5

42.3 355.5

Qwen 2.5 7B

Acc. Energy

59.4 273.4

86.2 338.7

24.6 241.0

88.3 137.0

64.6 159.4

59.4 1292.8

70.0 213.4

64.0 397.6

Llama 3.1 8B

Acc. Energy

43.7 336.8

76.6 1242.2

40.6 80.9 1905.0 192.4

55.4 188.0

31.4 16106.0

56.0 223.4

54.8 3176.9

DeepSeek-R1 8B

Acc. 66.8 Energy 3812.9

95.1 3586.1

71.4 93.2 5815.8 2163.6

72.9 4468.3

85.8 13004.1

85.0 8476.1

81.1 5621.5

DeepSeek-R1 32B Acc. 70.8 Energy 4981.8

94.2 2448.9

32.6 94.5 6337.2 3322.3

69.8 3862.0

79.1 12974.5

93.0 13168.3

74.4 6021.0

Preprint, In Submission, Publication time, Publishing venue

Siddiqui et al.

Figure 10: Same four held-out questions as Fig. 5 on a later NVIDIA A100-SXM4-80,GB session. KL-GRPO step 1500 of a different KL beta variation versus SFT. Router energy is about 30 J for KL-GRPO and a few hundred joules for the longer SFT decode. Answer curves reuse tournament logs.

Record · ID 1028667 · SHA-256 2fbaaa089103927c
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.