Conceptio › Archive › arXiv CS
arXiv CSopen access

SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM Inference

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

arXiv:2609.21704v1 [cs.LG] 18 Sep 2026

SpecQuant – Speculative Decoding with MultiParent Quantization for Adaptive LLM Inference Harish KB

Jagadeeswaran M

Pradheep P

Computer Science and Engineering Vellore Institute of Technology Vellore, India [email protected]

Computer Science and Engineering Vellore Institute of Technology Vellore, India [email protected]

Computer Science and Engineering Vellore Institute of Technology Vellore, India [email protected]

Yuvanesh S

Sivakumar T

Electronics and Communication Engineering Vellore Institute of Technology Vellore, India [email protected]

School of Computer Science and Engineering Vellore Institute of Technology Vellore, India [email protected]

Abstract—Running large language models (LLMs) locally continues to be limited by restrictions of compute and memory on consumer hardware. The popular acceleration technologies, such as quantization, speculative decoding, and adaptive inferencing, offer substantial speed boosts but usually necessitate retraining, per architecture tuning, or draft models. SpecQuant is a trainingfree framework, that combines speculative decoding with multiparent quantization to perform adaptive, efficient inference of LLMs. SpecQuant derives multiple quantized variants (INT4, FP8, FP16) from a shared base model, and dynamically routes queries based on predicted complexity; lightweight variants are used for simple or factual tasks, and full-precision models are used for complex reasoning tasks or long-context inputs. The shared-weight design of SpecQuant ensures sufficient token acceptance for speculative decoding without compatibility issues using separate draft parent models. We evaluate SpecQuant on Qwen2.5 based models on the MMLU, AlpacaEval, and GSM8K datasets, or benchmarks, demonstrating 35–43% speedups without degrading accuracy greater than 2%, substantial within the LLM community. SpecQuant enables practical on-device LLM deployment across diverse hardware without special infrastructure or expertise. Index Terms—Speculative Decoding, Large Language Models, Quantization, Adaptive Inference, Edge Computing

I. I NTRODUCTION There is great interest in the local deployment of Large Language Models (LLMs), driven by privacy and the ability to run models locally (without internet or cloud) for convenience. However, the hardware for consumer devices is often a significant barrier to practical deployment. Current LLMs employ a large memorized footprint the RAG model with 7 billion parameters often requires a real-time 14 GB or more Published in the 2026 Fifth International Conference on Power, Control and Computing Technologies (ICPC2T), IEEE, Raipur, India, 11-13 March 2026, pp. 371-375. DOI: 10.1109/ICPC2T68221.2026.11646348. © 2026 IEEE. Personal use of this material is permitted. Other uses require IEEE permission, including republication, promotional use, creating collective works, resale or redistribution, or reuse of copyrighted components.

memory, when processing in fact, this is generally much more than most consumer hardware can provide. Several methodologies have been developed to mitigate these challenges. Speculative decoding frameworks such as EAGLE [1] and Medusa [2] employ draft models with reduced parameter counts to generate token candidates, which are subsequently validated by a larger parent model. These approaches attempt to accelerate inference by parallelizing draft model execution while awaiting verification from the parent model. Adherence to the tenets of speculative decoding [3]. The approach of quantization enables smaller memory footprints by employing low precision representations of weights and activations, consequently reducing memory usage and computational costs. Conditional stopping techniques in adaptive inference, such as early exit [4] and layer skipping [5], can be used for simpler input samples. Despite their potential, current solutions have inherent drawbacks. In many cases, adaptive inference techniques need the model to be retrained or the architecture to be changed, thus making the deployment complicated. Speculative decoding methods have to keep a separate draft and parent models for that, which leads to difficulties in compatibility and increased memory consumption. Besides that, the majority of the present implementations do not have hardware-aware adaptation mechanisms that could consider the heterogeneous memory and computational characteristics of consumer devices. This void is a potential that can be realized by developing practical solutions that would make it possible to locally deploy efficient LLMs on different hardware configurations without the need for having a certain level of expertise or a great amount of computational resources. II. L ITERATURE R EVIEW A. Speculative Decoding Methods Speculative decoding has emerged as a viable method to enhance LLM inference speed without sacrificing output qual-

ity. Methods such as EAGLE [1] utilize smaller draft models to make predictions about token candidates, and subsequently learn which of the token candidates it should select given token candidates from draft predictions that are consistent with the output of a larger parent model. While draft predictions are consistent with the parent model output, the increased search space of the draft model allows EAGLE [1] and Medusa [2] to accept multiple tokens at once, thus accelerating inference time. However, the speculative decoding process often involves training a draft model, which can be expensive and may not be fully aligned with the parent model. B. Model Compression and Quantization Quantization [6] methods are an effective strategy for reducing memory needs for large models. Research shows it is possible to lower memory from high precision 16-bit all the way down to 2-bits [7] while still achieving manageable performance improvements. Additionally, post-training quantization [6] does not require re-training of the model which allows for faster deployment forms models during the inference process. Yet, much of the existing and newer work does not leverage multiple quantization levels to work concurrently in optimizing performance while still getting core updates in the single-precision level. C. Adaptive Inference Techniques Adaptive methods, such as early exit [4] and layer skipping [5], enable models to cease computation when reasoning is not warranted according to the input occuring. AdaInfer [8] and similar methods reason about the computation required based on the difficulty of the input dynamically, based on input difficulty. Though effective at reducing average inference time, these adaptations for early computation termination typically impose architecture modifications or retraining, and thus may not be appropriate for end-users who wish to use pre-trained models. D. Resource-Aware Model Selection FrugalGPT [9] introduced the process of routing queries to different model sizes based on complexity of the prompt, focusing on the cloud inference setting. Their work showed that there is some ability to route based on cognition, but the implication never considered constraints of local deployments or the memory limitations of consumer devices.

1) Low-Precision Parent (Q4): As the name suggests, INT4 architecture or Q4 uses minimum space memory and produces tokens at high-speed pace. The conversion process compresses each 16-bit parameter of the parent architecture into 4 bits in this case. Thus, our model representation would be: Q4 = f (W16 , 4) Where W16 denotes the parent FP16 weights. 2) Medium-Precision Parent (Q8): FP8 architecture or Q8 architecture stands out from the rest, as it represents something in-between. While being lighter than the parent version, the FP8 architecture captures numerical detail better than Q4. Thus, the model would look as follows: Q8 = f (W16 , 8) It performs exceptionally well in moderately complex tasks. 3) Full-Precision Parent (Q16): Our initial FP16 model known as Q16, which serves as a baseline for other models and offers the maximum level of precision: Q16 = W16 When decoding speculatively, the router decides which of those three parents is to be chosen according to the input complexity. In the draft model, the following guess set is provided: ( if parent agrees ytdraft , yt = f (x1:t−1 ), otherwise And what those terms actually represent: • yt : the resulting token at step t. draft : a guess generated by the draft model. • yt • f (x1:t−1 ): a token that will be predicted by the parent using all previous tokens. • If the parent agrees with the draft guess, it gets to stick around. • If not, the parent independently predicts the correct token. The model generates quick suggestions as parents validate their validity or otherwise simultaneously. In cases where the parent feels satisfied with the suggestions, then things flow fast and smoothly. However, in cases where the parent is unsatisfied with the suggestions, the model resorts to predictions made by itself. The advantage with this technique is that since both ver- sions of the models use the same weights, they agree quite often.

III. P ROPOSED F RAMEWORK B. Routing Scheme A. Multi-Parent Quantization Post-training quantization enables us to create two lighter versions using one architecture. To start off with, we use the full-precision 16-bit floating-point model, after which we create the 4-bit and 8-bit versions. What they share is their similar base weights, but they differ in precision and speed.

We have managed to create a routing mechanism which determines the right quantized versions to be run depending on a specific prompt. This means that the trade-off between speed and accuracy can easily be achieved. Our routing technique uses three versions of the models according to the degree of qua- Q16 (FP16).

1) Prompt Complexity Estimator: Before executing the prompt using the inference process, we use a prompt complexity estimator which considers three components for prompt complexity estimation, namely • Prompt Length: This refers to the number of tokens in the input prompt. The longer the length of a prompt, the higher the degree of complexity. • Syntactic Complexity: The syntax structure of the input. For this factor, we consider both the depth of grammatical structures and sentence structures. • Named Entity Density: Refers to the density of names (people, places, institutions), locations, organizations, or terminologies used. It can be seen as the complexity of the prompt relative to the number of words. 2) Routing Classification: With the above criteria, the system classifies each prompt and routes it accordingly: • Low Complexity: Brief prompts with basic syntax are routed to Q4. The model is extremely fast and has low memory requirements, thus making it suitable for easy tasks. • Moderate Complexity: Moderately lengthy prompts or those with more complex sentence structures are routed to Q8. It provides superior accuracy compared to Q4 yet remains resource-efficient. • High Complexity: Lengthy, intricate, or highly complex prompts are directly routed to Q16 to ensure maximum accuracy where needed. C. Hardware-Aware Deployment Strategy The proposed framework integrates a dynamic device placement strategy.A system effectively allocates model components to either theSelection of processing unit, either a Graphics Processing Unit (GPU) or a Central Processing Unit (CPU), is contingent on the availability of relevant hardware resources.There is an assurance that The attainment of optimal performance metrics remains contingent upon the operational capabilities of the system.”The design ensures broad compatibility across conventional systems, obviating the need for specialized hardware modifications.”dedicated GPUs. 1) GPU-Accelerated Mode: When first initialized, the system checks to see if an available GPU is present. When a capable GPU is detected, the framework will run at its highest performance mode. • Optimal Placement: If VRAM is available, both the parent model and draft model are loaded directly into the GPU. This minimizes latency since both components are on the same device, which allows for the highest throughput. • Constrained VRAM Placement: When the VRAM on offer is insufficient to hold both models, a hybrid placement strategy is used by the system. To verify the speculative cycle which is the most critical and frequent step, the parent model thus remains on the GPU. The draft model, therefore, is moved to system RAM and executed on the CPU. Here, the draft model is producing speculative

tokens on the CPU and the parent model is verifying the tokens that are sent to the GPU. Hence, the data transfer overhead between the CPU and GPU is minimized. 2) CPU-Only Fallback Mode: In the absence of a detectable or available GPU, the system automatically defaults to a CPU-only fallback mode. This ensures the framework remains functional on any standard computer, regardless of its graphics hardware. • CPU Execution: Both the parent and draft models reside within the system’s RAM. The whole computational chain, including prompt analysis, draft generation, and parent verification, is carried out by the CPU. Although this operation is slower in terms of inference time as compared to GPU acceleration, the main benefit of this mode is that it is accessible to everyone, thus enabling local LLM inference for a wide range of users without the need for a specially designed machine. This two-part approach helps our system perform well when resources are available, while still functioning effectively when they are not. D. Execution Pipeline The Execution Pipeline represents the pathway a system takes for each user Request in order to generate text in both an efficient and accurate manner. It is divided into four stages, which consist of the following: complexity assessment, draft generation, parent verification, and token acceptance or token correction. 1) Complexity Assessment: The Complexity Assessment is undertaken by the router responding to a user’s request. The router checks different characteristics of a prompt to understand how difficult it will be to calculate. OEMs determine complexity through three (3) characteristics of a prompt; prompt length (measured in tokens), syntactic difficulty to parse, and the quantity of entities related to the prompt. Based on its analysis of these characteristics, the complexity of a prompt will be either low, medium, or high. If the complexity is found to be low or medium, it will pass to the second level of Development (Draft Creation). If it is found to be complex, the Custom Process will be performed directly on the parent model at 100 2) Draft Generation: Drafts that are generated from lowcomplexity input can be produced in much less time and require significantly fewer resources than drafts produced from mid-complexity input; thus drafts that are generated from lowcomplexity input can generate numerous speculative tokens within a short period of time and with minimal resource requirements while drafts generated from mid-complexity input will produce draft tokens that are more accurate than draft tokens produced from low-complexity input. Drafts that are generated from each of these draft variations can be produced with a wide range of draft candidate tokens (i.e., all candidate tokens generated in one single time period) in order to maximize the potential for increased parallelization of draft generation across multiple draft candidate tokens.

3) Parent Verification: The parent model that was selected in the fp16 model is also being tested against the same input that was used for the draft tokens. In order for both the draft and parent model to have the same tokens in their input, both the parent and draft tokens must match their potential for input of the same generated tokens as they were produced from the draft. Each quantized version of a model will use a version of the original model’s base weight, so all parent models (and their drafts) should behave very similarly and therefore should produce high acceptance rates for drafts that have produced accurate results. 4) Token Acceptance or Correction: If all of the tokens in the speculative block have been approved by the parent model, then all of the tokens are accepted as a part of the completed data and added to the final output. Then, the pipeline feeds all the accepted tokens back into the parent model so that they can be used as new contextual tokens for creating the next draft cycle. If a token in the draft is rejected, then the parent model provides the correct token sequence. The router will remove the speculative block from the draft and recommence the speculative decoding process starting with the first token that was rejected, thus guaranteeing that the output is correct. 5) Direct Parent Inference: High complexity prompts do not go through the speculative stages. The parent model with High Precision (Q16) will continue processing the prompt in an autoregressive manner until all tokens have been generated. IV. R ESULTS SpecQuant was assessed using Qwen/Qwen2.5-7B- Instruct as the parent model, and its quantized models (INT4, FP8, and FP16) were produced using post-training quantization. The performance was tested on three different benchmarks pertaining to varying types of reasoning: MMLU (Benchmark 1 – factual reasoning), Alpaca Eval Subset (Benchmark 2 – instruction following), and GSM8K (Benchmark 3 – mathematical reasoning). TABLE I AVERAGE INFERENCE TIME AND SPEEDUP ACROSS BENCHMARKS . Benchmark

Baseline (s)

SpecQuant (s)

Speedup (%)

MMLU Alpaca Eval GSM8K

4.07 6.42 8.83

3.01 4.50 6.43

35.2 42.6 37.3

On all metrics, SpecQuant attains an average reduction of 38.4% in generation latency compared to normal decoding. This is due to the speculatively validated quantization by SpecQuant that eliminates unnecessary evaluation of the Parent model. In cases where the Parent rejects a larger percentage of draft tokens, speculative decoding does not yield any performance improvements. To verify our complexity-based routing method, we analyzed the complexity category breakdown of test prompt divisions by each benchmark. Complexity categorization provided us with good indicators of model selection because: based on benchmark averages, Low Complexity (avg 28%) achieved

Fig. 1. SpecQuant achieves 35–43% speedup across factual, instructionfollowing, and mathematical reasoning tasks. TABLE II ACCURACY AND T OKEN ACCEPTANCE R ATES Benchmark MMLU Alpaca Eval GSM8K

Normal Acc. (%) 72.5 75.0 70.0

SpecQuant Acc. (%) 72.3 74.9 69.9

∆ Acc. -0.2 -0.1 -0.1

Acceptance Rate (%) 66.1 60.9 57.9

35.2% speedup with Q4; Medium Complexity (avg 52%) achieved 42.6% speedup with Q8; and high complexity (avg 20%) executed directly with Q16. The distribution across three benchmarks (MMLU, Alpaca Eval and GSM8K) followed the same pattern confirming successful selection of prompts from lighter quantised models by the routing mechanism. In addition, acceptance percentage rates for tokens (57.9%66.1%) were good across all complexity types; supporting that lightweight quantised models could be integrated into speculative decoding without the expectation of loss of acceptance quality V. C ONCLUSION SpecQuant shows a significant performance gain of inference acceleration by combining speculative decoding with a verification of the quantized model. The framework, tested across three different benchmarks, MMLU, Alpaca Eval Subset, and GSM8K, achieves a speedup of about 38.4% on average with less than 2% accuracy drop, thus, a very good trade, off between efficiency and performance is realized. The token acceptance rates vary between 57.9 and 66.1% which reflects a stable alignment between draft and parent during the speculative verification. It is worth noting that this alignment is almost always significantly higher when the draft model is a quantized version of the parent model. Since the two models share the same weights and have similar architectural characteristics, their token distributions differ less during decoding, thus higher acceptance rates can be observed even at lower precisions. The latter indicates that quantization does not merely reduce the weight of the draft model but also enhances the compatibility with the parent model, hence a quantized, parent pair is a particularly efficient one for speculative decoding pipelines.

The routing mechanism based on the complexity factor effectively enables distribution of prompts among the quantized variants, with 80% of prompts used in tests being successfully decoded using SpecQuant using either Q4 or Q8 allocation. SpecQuant highlights that routing complexity is essential to enable maximum speed-up with minimal loss of accuracy. SpecQuant appears to be a useful framework for implementing the use of LLMs in situations where limited hardware resources are available. SpecQuant allows for achieving the same quality of output results, but with significantly reduced computational burden. In this way, LLM inference can be done efficiently without the need for any kind of specialized hardware or lengthy model retraining, thus, making it available for deployment scenarios of a consumer grade type. R EFERENCES [1] Chen, Tianle, et al. ”EAGLE: Speculative sampling requires rethinking generation quality and efficiency.” arXiv preprint arXiv:2401.08294, 2024. [2] Cai, Tianyu, et al. ”Medusa: Simple LLM inference acceleration framework with multiple decoding heads.” arXiv preprint arXiv:2401.10774, 2024. [3] Leviathan, Yaniv, Matan Kalman, and Yossi Matias. ”Fast inference from transformers via speculative decoding.” International Conference on Machine Learning (ICML). PMLR, 2023. [4] Schuster, Tal, et al. ”Confident adaptive language modeling.” Advances in Neural Information Processing Systems (NeurIPS), 2022. [5] Bae, Young Jin, et al. ”Layer skipping for multi-exit neural networks.” IEEE Access, vol. 9, pp. 94958–94970, 2021. [6] Frantar, Elias, et al. ”Gptq: Accurate post-training quantization for generative pre-trained transformers.” arXiv preprint arXiv:2210.17323 (2022). [7] Chee, Jerry, et al. ”Quip: 2-bit quantization of large language models with guarantees.” Advances in Neural Information Processing Systems 36 (2023): 4396-4429. [8] Shen, Li, et al. ”AdaInfer: Adaptive Inference for Large Language Models via Early Exit and Confidence Estimation.” arXiv preprint arXiv:2404.01221, 2024. [9] Zhou, Xingyu, et al. ”FrugalGPT: How to use large language models while reducing cost and improving performance.” arXiv preprint arXiv:2305.05176, 2023. [10] Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J. ”Accelerating large language model decoding with speculative sampling.” arXiv preprint arXiv:2302.01318, 2023a. [11] Kumar, Avinash, et al. ”HELIOS: Adaptive Model And EarlyExit Selection for Efficient LLM Inference Serving.” arXiv preprint arXiv:2504.10724 (2025). [12] Xia, H., Ge, T., Wang, P., Chen, S.-Q., Wei, F., and Sui, Z. ”Speculative decoding: Exploiting speculative execution for accelerating seq2seq generation.” In Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 3909–3925, 2023. [13] Shazeer, N. ”Fast transformer decoding: One write-head is all you need.” arXiv preprint arXiv:1911.02150, 2019. [14] Zhang, J., Wang, J., Li, H., Shou, L., Chen, K., Chen, G., and Mehrotra, S. ”Draft & verify: Lossless large language model acceleration via selfspeculative decoding.” arXiv preprint arXiv:2309.08168, 2023. [15] Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. ”Language models are few-shot learners.” Advances in Neural Information Processing Systems, 33: 1877–1901, 2020. [16] Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. ”PaLM: Scaling language modeling with pathways.” arXiv preprint arXiv:2204.02311, 2022. [17] Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. ”LLM.int8(): 8-bit matrix multiplication for transformers at scale.” arXiv preprint arXiv:2208.07339, 2022.

[18] Hewitt, J., Manning, C. D., and Liang, P. ”Truncation sampling as language model desmoothing.” October 2022. doi: 10.48550/ARXIV.2210.15191. [19] Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S. ”AWQ: Activation-aware weight quantization for LLM compression and acceleration.” arXiv preprint arXiv:2306.00978, 2023. [20] Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. ”Training language models to follow instructions with human feedback.” arXiv preprint arXiv:2203.02155, 2022. [21] Zhou, Y., Lyu, K., Rawat, A. S., Menon, A. K., Rostamizadeh, A., Kumar, S., Kagy, J.-F., and Agarwal, R. ”DistillSpec: Improving speculative decoding via knowledge distillation.” arXiv preprint arXiv:2310.08461, 2023. [22] Gale, T., Elsen, E., and Hooker, S. ”The state of sparsity in deep neural networks.” arXiv preprint cs.LG/1902.09574, 2019. [23] Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S. ”SmoothQuant: Accurate and efficient post-training quantization for large language models.” In International Conference on Machine Learning, pp. 38087–38099. PMLR, 2023a. [24] Dao, Tri, et al. ”FlashAttention-2: Faster attention with better parallelism and work partitioning.” arXiv preprint arXiv:2307.08691, 2023. [25] Zhou, Ming, et al. ”DejaVu: Dynamic early exiting for efficient large language model inference.” arXiv preprint arXiv:2402.09668, 2024.

A PPENDIX The complete implementation of SpecQuant, including all experimental code and configurations, is publicly available at https://github.com/HyperKuvid-Labs/SpecQuant.

Record · ID 1006872 · SHA-256 4ca2663bd9e1ccc5
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.