Conceptio › Archive › arXiv CS
arXiv CSopen access

Benchmarking Post-Quantum Cryptography in Lightweight Virtualization Environments on Embedded Hardware

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

Benchmarking Post-Quantum Cryptography in Lightweight Virtualization Environments on Embedded Hardware Nikolai Puch✉ Technical University of Munich Munich, Germany Fraunhofer AISEC Garching, Germany [email protected]

Chi Hieu Ta

Moritz Beckel

CarByte Engineering GmbH Rülzheim, Germany [email protected] [email protected]

Technical University of Munich Munich, Germany [email protected]

arXiv:2609.23902v1 [cs.CR] 20 Sep 2026

Abstract Post-Quantum Cryptography (PQC) is being deployed at the same time as embedded systems increasingly adopt lightweight virtualization for workload isolation and security. Both trends change performance characteristics, yet their interaction is not well understood. To address this research gap, we present a measurement study of PQC primitives on embedded-class ARM hardware under three execution environments with a shared software stack: native execution, a Docker container, and a Unikraft unikernel running under QEMU. We benchmark five signature and five key encapsulation mechanism families, and two classical algorithms each, for comparison. We evaluate them using different parameter sets for a total of around 70 configurations, measuring execution time, memory, and energy per operation. To better gauge the impact on applications, we evaluated 13 TLS 1.3 cipher combinations as well. We find that container overhead is negligible for primitive computation, whereas unikernel overhead depends on the algorithm. For CROSS, FrodoKEM, Classic McEliece, and the NIST-standardized post-quantum families ML-KEM, ML-DSA, and the SHAKE variants of SPHINCS+, the overhead is near-native. A moderate overhead (1.28-1.53 ×) arises in the alternates BIKE, HQC, and MAYO, and, above all, in Falcon signing (17.8-19.2 ×), while the SHA-2 variants of SPHINCS+ run faster than native (0.54-0.67 ×). Per-operation energy closely tracks execution time in all environments. For TLS handshakes, container and unikernel clients need more handshake time and energy per handshake when the cryptographic cost is low, while all three environments converge once expensive postquantum algorithms dominate the handshake. In these cases algorithm choice affects per-handshake energy by up to three orders of magnitude, far outweighing the environment. Overall, virtualization cost is inversely related to cryptographic cost: environment choice matters most for computationally cheap, standardized algorithms, while for expensive schemes, algorithm choice alone dominates performance.

Keywords Post-Quantum Cryptography, Unikernel, Container, TLS, Embedded Systems, Benchmarking

1

Introduction

A large-scale Cryptographically Relevant Quantum Computers (CRQC) would break the public-key cryptography that secures today’s Internet [37]. Even before a CRQC exists, traffic can be recorded and decrypted retroactively once such machines exist, in so-called Harvest-Now, Decrypt-Later attacks. Recent advances

[23] and such scenarios have accelerated the transition, with the European Union (EU) aiming for a transition by the end of 2030 for high-risk use cases [1]. Thus, the migration to Post-Quantum Cryptography (PQC) is already underway: National Institute of Standards and Technology (NIST) has standardized ML-KEM [30], MLDSA [29], and SLH-DSA [31], selected HQC in its fourth round [17], and announced Falcon for future standardization [16]. Transport Layer Security (TLS) is the most immediate deployment target, as both its key exchange and its authentication must be replaced with Post-Quantum (PQ) algorithms, at the cost of larger keys and signatures. At the same time, embedded and edge systems increasingly rely on software-based isolation for easier deployment and security. Containers are today’s dominant mechanism, but they require a full host operating system. Unikernels aim to become an alternative and promise smaller images, a reduced attack surface, and near-native performance on top of a hypervisor. To achieve this, unikernels compile an application together with only the OS components it needs [28]. This is especially interesting for constrained embedded domains, like the automotive domain, where compute, memory and energy are at a premium. PQC and lightweight virtualization each distinctly impact the system’s performance. Their combination is largely unexplored: post-quantum algorithms stress memory, entropy sources, and floating-point units in ways classical cryptography does not, and unikernels virtualize exactly these resources differently than containers do. Operators tasked with choosing between deploying an application utilizing PQC in a container or in a unikernel on embedded hardware currently face a lack of benchmark data to guide their decision. This paper addresses this gap and answers the following research questions: RQ1 How do PQC primitives perform in unikernels versus other virtualization solutions? RQ2 How does this impact secure communication via TLS? RQ3 What is the energy cost of PQC under the different virtualization options? To answer them, we build a benchmarking framework that runs an identical cryptographic stack, based on liboqs 0.12.0, natively, in a Docker container, and in a Unikraft unikernel on a Raspberry Pi 4B, and we measure primitive execution time, memory, and per-operation energy as well as TLS 1.3 handshake throughput and energy. Contributions.

Nikolai Puch, Chi Hieu Ta, and Moritz Beckel

• Automated benchmarking framework covering three execution environments on ARM64 embedded hardware with an identical crypto stack (Section 5). • A primitive-level comparison of 60 PQC configurations across the three environments, covering speed, memory, and per-operation energy (Section 6). • A TLS 1.3 handshake study over 13 signature/KEM combinations, including energy measurements (Section 6).

Client

Server TCP SYN TCP SYN-ACK

KeyGen

ClientHello + supported_group + signature_algorithms + signature_algorithms_cert + key_share

2 Background 2.1 Post-Quantum Cryptography Post-quantum algorithms are commonly grouped by function into Digital Signature Algorithms (DSAs) and Key Encapsulation Mechanisms (KEMs), and by the underlying hardness assumption. Table 1 gives an overview of the current NIST algorithms and their status. Following prior embedded evaluations, we use the NISTcompetition names of the algorithms and their liboqs implementations rather than the standardized names, because the implementations we benchmark predate final standard release. Signatures: Dilithium, standardized as ML-DSA in FIPS 204 [29], is a module-lattice scheme and the primary NIST signature choice. Falcon is a lattice scheme selected for standardization as FIPS 206 [16]. Its current implementations depend on double-precision floatingpoint arithmetic, which is missing on many embedded platforms. However, our evaluation platform provides this. SPHINCS+, standardized as SLH-DSA in FIPS 205 [31], is a stateless hash-based scheme with conservative security assumptions but comparatively expensive signing. KEMs: Kyber a lattice-based KEM, standardized as ML-KEM in FIPS 203 [30], is the primary KEM by NIST. HQC, a code-based scheme, was selected in NIST’s fourth round [17] as an alternative. Table 1: PQC algorithms and their publication status. Hash-based ciphers SPHINCS+ stateless DSA Lattice-based ciphers CRYSTALS-Kyber KEM CRYSTALS-Dilithium DSA Falcon DSA Code-based ciphers HQC KEM

FIPS 205 [31] FIPS 203 [30] FIPS 204 [29] FIPS 206 (WiP [18]) Round 4 selection [17]

PQC in TLS. Multiple approaches exist for making TLS quantum resistant [25]. The currently favoured method is a PQ TLS 1.3 handshake, which replaces both the ephemeral key exchange and the authentication via DSA. KEMs are comparatively fast and small on embedded devices and are only done once per handshake. The required changes are displayed in Figure 1. However, the DSA is applied once per link of the certificate chain, so signature verification cost and certificate size dominate the handshake overhead for many schemes [25, 34, 39].

2.2

Virtualization on Embedded Systems

Virtualization partitions the resources of a physical machine into isolated instances. Two families are relevant for resource-constrained

Encaps

ServerHello key_share + {EncryptedExtensions} {CertificateRequest*} {Certificate} {CertificateVerify} {Finished}

Decaps Verify Sign

Sign

{Certificate*} {CertificateVerify*} {Finished}

Verify

[Application Data]

Figure 1: Shows the PQ TLS 1.3 handshake. Orange shows the necessary changes to the DSA and blue to the KEM.

systems. OS-level virtualization (containers, e.g., Docker) isolates processes using kernel features such as namespaces and cgroups; it adds little runtime overhead but shares the host kernel, so its isolation boundary is the kernel interface itself. It is worth mentioning that OS-level virtualization is not a security feature by default and measures are necessary to secure it. Hypervisor-based virtualization, on the other hand, runs guests on a virtual machine monitor. Depending on the use-case and underlying hardware, a complete operating system, or images with smaller footprints, such as unikernels and MicroVMs are virtualized. The latter applies this model to single-purpose workloads instead of providing operating system capabilities. The security property is inherently given by the isolation of the hypervisor between host and guest. A visual comparison between containers, virtual machines, MicroVMs, and unikernels is given in Figure 2. This study omits MicroVMs and compares containers with unikernels. Extending the comparison to MicroVMs is left to future work. Unikernels. A unikernel, also known as library operating system, compiles the application together with only the OS components it uses into a single-address-space image that boots directly on a hypervisor. One of the advantages, in light of required performance, is the absence of a dedicated user and kernel space. With one singular address space no context switches must be performed, thus increasing the overall performance. Ordinarily, removing the separation between user and kernel space would introduce security issues on its own. However, with unikernels this is intentional. This issue is resolved by the design of unikernels, as they are compiled and used as single-purpose and read-only binary files. Hence, unikernels are immutable. [28]

PQC in Lightweight Virtualization

Architectural difference matters for cryptographic workloads: in our configuration the unikernel guest has no hardware entropy source of its own. It receives a random seed on the kernel command line at boot (drawn from the host’s /dev/urandom) and expands it internally, whereas native and containerized processes obtain entropy from the host kernel (e.g., via getrandom) throughout their lifetime. VM-1

VM-2

App-1

App-1

App-2

App-2

OS Lib Kernel

OS Lib Kernel

Cont-1

Cont-2

App Lib OS

App Lib OS

Container Engine

VM-1

VM-2

App Lib OS

App Lib OS

Uk-1

Uk-2

KVM

App Lib

App Lib

KVM

Hypervisor

OS Lib Kernel

OS Lib Kernel

Hardware

Hardware

Hardware

Hardware

Virtual Machine

Container

MicroVM

Unikernel

Hypervisor

infrastructure can likewise be used to balance such power tradeoffs [36]. Unikernel performance. Madhavapeddy and Scott were the first to introduce unikernels also known as library operating systems [28]. Meanwhile, many other instances of unikernel frameworks have been introduced, such as rumprun which is based on the rump kernel, and OSv [8, 14]. In addition to rumprun and OSv, Unikraft is introduced as another framework which reports near-native or better-than-Linux-guest performance for server workloads such as nginx and Redis. According to the github page of rumprun the last commit was performed six years ago at the time of our work. Further, the work of Kuenzer et al. show that their Unikraft unikernels report a 1.7–2.7× performance improvement over Linux guests for common server applications, image sizes around 1 MB, and boot times in the millisecond range. In addition, it outperforms rumprun and OSv in the same use-cases, hence we opted for Unikraft for our tests due to still ongoing development compared to rumprun and better performance compared to rumprun and OSv. [27]

Figure 2: Visual comparison of virtualization architectures

3

Related Work

PQC on embedded systems. Evaluating post-quantum algorithms on constrained hardware has been an integral part of the NIST standardization process, which included benchmarks on the ARM Cortex-M4 microcontroller [16]. PQClean [10] collects clean, standalone C implementations of these algorithms. In 2026 it was superseded by the PQ Code Package [9], a collection of open-source implementations maintained within the Linux Foundation. Targeting Microcontroller (µC) platforms more specifically, pqm4 [4] collects optimized implementations of the standardized algorithms or recent standardization candidates for the ARM Cortex-M4. liboqs [5] integrates many of these implementations and forms the basis for the OQS-provider [6], which brings PQC to OpenSSL [7]. Based on these libraries, several works benchmark the thenmost-recent PQ algorithms on embedded platforms within specific application domains, including Internet of Things (IoT) [26], industrial control [35], and automotive systems [20, 21]. Among these domains, TLS is the most prominently researched application for PQC on embedded devices. Most research in this field likewise builds on pqm4 or liboqs and targets the ARM Cortex-M4 and Cortex-A53. Results show that DSA selection and network conditions, owing to the increased key and signature sizes, dominate the performance of PQ TLS [34]. Mixed certificate chains [32, 36] allow tuning the key/signature size vs. performance tradeoff to the specific use case, while KEMTLS [25] offers an alternative approach. Beyond the Cortex-M4/A53 focus of this body of work, Bürstinghaus-Steinbach et al. [22] benchmarked Kyber and SPHINCS+ inside mbedTLS across a broader range of embedded platforms. Energy consumption, by contrast, is evaluated far less frequently: Tasopoulos et al. measure PQ TLS 1.3 handshake time on resource-constrained devices [39] and, in follow-up work, its energy consumption [38]. They find that although Falcon completes its handshake faster, its power consumption exceeds that of Dilithium due to the larger compute share it requires. The aforementioned mixed-certificate

Unikernel vs Container. Unikernel and container have been subjected to comparative benchmarking across various experimental setups. Goethals et al., for instance, compare unikernels against containers in the context of microservices. They conducted their tests on an x86_64 machine with 4GB of RAM and 160GB HDD which is used as a server, while client tasks were done on a separate machine to avoid distorting the measurements. While containers were spawned on Ubuntu 18.04, they opted for a XenServer 7.5 for their unikernel setup. Furthermore, Goethals et al. divided their test cases into single threaded and multithreaded tests. Also, they emphasised that only one container or virtual machine is active during tests. [24] In [33] a comparison between container and multiple unikernel frameworks is performed. The test cases comprise an HTTP server and key-value store to represent workloads usually found in cloud applications. In this regard, Plauth et al. use an x86_64 architecture CPU including 2x8GB of RAM, and KVM as their hypervisor for their test environment. Similar to Plauth et al., Acharya et al. benchmark containers and multiple unikernel frameworks in the context of network functions virtualization. They opt for two server hardware variants to test their application. I.e., one with 8 core x86_64 CPU and 32GB of DDR4 and another platform with a 64bit ARMv8 and 48 cores with 128GB of DDR4 RAM. Operating system-wise, Acharya et al. use Ubuntu 16.04.3 LTS in conjunction with KVM as hypervisor and QEMU 2.5.0. The performed test cases contain regular CPU performance, as well as memory and network bandwidth benchmarks. [15] Although not strictly a head-on comparison of two technologies against each other, Bartolomeo et al. propose a hybrid usage of containers and unikernels for edge computing. The goal is to reduce CPU load by orchestrating the usage of either unikernels or containers depending on the use-case in which either performs best. For this, they test both technologies under different circumstances. Bartolomeo et al. target systems involving two server set-ups with x86_64 and a Raspberry Pi 4 with the ARM Cortex-A72 and 8GB of DDR4 RAM. They use KVM as hypervisor but use an extended version of Oakestra to spawn and orchestrate unikernels and containers in parallel. [19] As we can see, unikernels are mainly used and benchmarked within

Nikolai Puch, Chi Hieu Ta, and Moritz Beckel

a server comparable set-up, i.e., using x86_64 CPUs with a reasonable amount of RAM. While Acharya et al. use a dedicated ARMv8 test environment, it is a server-sized setup, and they specifically stated that unikernel tests on it are excluded due to the lack of stable ARM support of rumprun and OSv [15]. Only Bartolomeo et al. with the Raspberry Pi setup offer a comparable test environment and results to our research, as they use an ARMv8 single board computer and the Unikraft framework to build their unikernels. However, their results still differ in the ultimate goal, as post-quantum algorithms or in general applications involving the usage and performance of encrypted communication is not one of their evaluation criteria. Accordingly, cryptographic workloads, and post-quantum algorithms in particular, have not been part of prior evaluations and research. Hence, to the best of our knowledge, no prior work measures PQ primitives or PQ TLS inside lightweight virtualization, neither containers nor unikernels, on embedded hardware, nor compares the environments against each other under an identical cryptographic stack. The present paper bridges this gap.

4

Methodology

Our goal is to isolate the effects of the transition to post-quantum cryptographic on the performance in different execution environments. We therefore keep hardware, cryptographic library, libc, and compiler toolchain constant and vary only the environment: Native, Container, and unikernel. For these environments we measure three quantities: speed, memory, and power. We evaluate the isolated primitives alone as well as in the context of a full TLS 1.3 handshake. All benchmarks are orchestrated by a single Python driver to make runs reproducible.

4.1

Metrics

Primitive speed. For each algorithm and operation (key generation, signing, verification for DSAs and key generation, encapsulation, decapsulation for KEMs) we execute the operation in a loop and report the mean time per operation together with its population standard deviation. We let each operation run for at least a minimum wall-clock time as well as a minimum number of operations, and then collect the final wall-clock time as well as the number of completed operations. The speed evaluation uses wallclock-time via clock_gettime(CLOCK_REALTIME), while the clock for the timestamps of the power benchmark uses perf_time_ns. The primitive dataset evaluated in this paper uses a 20 s window per operation. We benchmark 40 signature and 20 KEM configurations, covering all NIST security levels available in liboqs 0.12.0 for Dilithium, Falcon (including padded variants), SPHINCS+, (SHA-2 and SHAKE, fast and small variants), MAYO, CROSS, BIKE, Classic McEliece, HQC, Kyber, and FrodoKEM (AES and SHAKE variants). The TLS datasets stem from an earlier measurement campaign, primitive and TLS results are therefore not from the same run. However, it follows overall a similar approach: We measure the number of completed TLS 1.3 connections in a 30 s window using openssl s_time against openssl s_server. The server always runs natively on the Raspberry Pi 4 (Pi4), while the client runs in the environment under test. This choice isolates the client-side handshake cost but implies that each environment reaches the server

over a different network path: loopback for native, Docker’s NAT for the container, and a QEMU bridge network for the unikernel. We discuss this in more detail in Section 8. We report initial (fullhandshake) connections. Certificates are generated per signature algorithm with a single self-signed Certificate Authority (CA) and one server certificate. We evaluate 13 signature/KEM combinations: Dilithium and Falcon paired with Kyber at matching security levels, SPHINCS+-128s with Kyber512, and Dilithium2 paired with BIKE-L1, HQC-128, and FrodoKEM-640 (AES and SHAKE), with RSA-2048+ECDHE and ECDSA+ECDHE as classical baselines. The subset follows prior post-quantum TLS studies on embedded hardware [38] to ease cross-comparison. Memory. For the primitive campaign the benchmark binary tagged with a special pid flag which is passed to the measurement loop to monitor memory. The monitor reads /proc/pid/statm and /proc/pid/status at 200 Hz and 10 Hz, respectively. From statm it collects instantaneous virtual-memory (VM) and resident-set-size (RSS) page counts, which are converted to bytes using the system page size. From status it tracks the kernel-maintained high-water marks VmPeak and VmHWM. The resulting dataset contains, per operation: mean RSS and mean VM over all 200 Hz samples, the samplemaximum of each, and the kernel-reported VmHWM and VmPeak values. For the TLS memory stage the driver spawns the client process under the same pid wrapper and then collects memory readings with top -b -p pid -d 0.5. Mean VIRT and RES values are averaged across all samples. For the unikernel, the monitored process in both cases is QEMU, so the reported figures reflect the memory footprint of the whole guest as seen by the host, including all guest RAM actually touched. This is the deployment-relevant quantity for an operator placing workloads. For the container, the same pid mechanism targets containerd-shim-v2, a host-side management process. We discuss the resulting comparability implications for the container vs. unikernel comparison in Section 8. Power and energy. The earlier TLS campaign used a FNIRSI FNB58 USB power meter placed inline in the device’s USB-C supply. Segment boundaries are detected by current-draw change. The FNB58 runs at 100 Sa/s and is logged to a separate device. According to the data sheet the power meter has a resolution of 0.000 01 volt and ±0.02% + 2 digits [13]. Although the manufacturer’s technical specifications claim high static resolution, a formalized error propagation and statistical uncertainty analysis were omitted for this hardware, due to the absence of a certified chain of traceability to primary international metrology standards. Hence, we decided to opt for professional hardware to perform the primitive campaign. For the primitive campaign, the device is powered by a Rohde & Schwarz HMC8043 programmable power supply [12] controlled over SCPI/TCP. The R&S HMC8043 has a resolution of 1 mV and 0.1 mA, if 𝐼 < 1 A and logging < 100 Sa/s. The reading accuracy is < 0.05%+2 mV and < 0.05%+2 mA. At 100 Sa/s and the resulting resolution of 10 ms might cause aliasing, as load changes in the Pi4 can be faster than 50 Hz. However, the long runtime of 20 s ensures that this error is acceptable. The built-in logging function is used to create a trace over the whole campaign. The trace is then aligned with the per-operation start/stop timestamps recorded by the benchmark driver and alignment is manually verified. This is then used to split the trace per operation. We report

PQC in Lightweight Virtualization

energy per operation (𝐸𝑜𝑝 ), i.e., the integrated segment power divided by the number of iterations executed in the segment. For TLS we derive energy per completed handshake by dividing the segment energy by the connection count recorded in the same run. In terms of power per operation (𝑃𝑜𝑝 ) an additional step must be performed to acquire the value. For this, we start our test with the respective environment and leave it in idle state for 10 s. With the R&S HMC8043 and < 100 Sa/s we receive ≈ 1000 power samples which in turn is averaged over the time. This power draw serves as the baseline. Subsequently, the operation is started and measured over 20 s. The resulting power is consequently subtracted by the baseline yielding 𝑃𝑜𝑝 . We report this baseline-subtracted value rather than raw power because the idle baseline drifts by up to ≈600 mW between measurement windows independent of which environment is running. Comparatively, this is large in relation to the many operations’ own power draw above the idle state which is ≈0.7–1.4 W. So, if not subtracted the raw energy would be dominated by this drift for a substantial share of comparisons. Formal measurement-error. We conduct a formal measurementerror estimation according to JCGM GUM-1:2023[2]. Type A is calculated for memory and energy, while type B is only calculated for the energy measurements by the R&S HMC8043. For the type A evaluation we conducted 25 measurement campaigns for the algorithms ML-DSA, ML-KEM, ECDSA, and X25519. The other parameters were kept the same as in the evaluation runs. We calculate the standard deviation 𝑠 (𝑥) over the 25 runs to calculate the uncertainty 𝑢𝐴 (𝑥) (see Equation 2). For the memory we calculate an uncertainty of 𝑢𝐴 (𝑟𝑠𝑠) = 0.25 − 20.89 KiB, for baseline-subtracted power of 𝑢𝐴 (𝑃) = 7.693 − 121.267 mW and for baseline-subtracted energy of 𝑢𝐴 (𝐸) = 0.0008 − 5.830 mJ. We estimate the type B uncertainty of the energy with equation 1. 𝑢𝐸 (𝐸) is the uncertainty of the energy measurement, 𝑢𝑡 (𝐸) is the uncertainty of the time base, and 𝑢 𝑗𝑢𝑚𝑝 (𝐸) is the uncertainty of the energy at an edge in the powertrace.

√︃ 𝑢𝐵 (𝐸) = 𝑢𝐸2 (𝐸) + 𝑢𝑡2 (𝐸) + 𝑢 2𝑗𝑢𝑚𝑝 (𝐸)

4.2

Controls

All binaries are built for ARM64 and statically linked against musl 1.2.3. The native binary is built inside the same Alpine container as the containerized one and exported, so that libc differences cannot confound the comparison. Benchmark processes were bound to a single core with os.sched_setaffinity, respectively the – cpuset-cpus flag in Docker. For the primitives campaign after each operation the temperature is checked, and the campaign is paused until the temperature drops below 50 ◦ C to prevent thermal throttling. The highest momentary peaks we measured was up to 53.0 ◦ C, which remains well below the Pi4 default thermal-throttling threshold, so throttling can reasonably be excluded for the current primitive datasets.

5 Experimental Setup 5.1 Hardware and Host Software All experiments are conducted on a Raspberry Pi 4 Model B (Broadcom BCM2711, quad-core ARM Cortex-A72, 64-bit). The platform was chosen as a widely available microprocessor-class stand-in for automotive or embedded hardware, whose processors are architecturally comparable [3]. The Raspberry Pi 4 has 8 GB of RAM, runs Debian 12 (bookworm) at kernel version 6.12.34+rpt-rpiv8 with QEMU 7.2.17, Docker 28.3.1, and Python 3.11.2 for instrumentation. Depending on the campaign the Pi is powered by a Rohde & Schwarz HMC8043 programmable power supply or through a FNIRSI FNB58 USB-C meter inline with a power support plug. For a detailed description see subsection 4.1. Figure 3 gives an overview of the measurement setup. Raspberry Pi 4B runner Network

tls Ntv.

power trace

5.12340V 0.31200A

primitives Cntr.

logs

Uk.

HMC8043

Powerout

(1)

The following values are the basis for the calculation: From the datasheet of R&S HMC8043 [12] we obtained the relative error 𝑎𝑈 = 𝑎𝐼 = 0.05%, and reading error 𝑏𝑈 = 2 mV, and 𝑏 𝐼 = 2 mA for voltage and current respectively. From measurement of the clock the Pi4 𝑎𝑡 = 630 ns. From the empirical evaluation of our powertrace, we determine: |Δ𝑃𝑠 | · Δ𝑡 = 23.371 mJ and |Δ𝑃𝑒 | · Δ𝑡 = 23.893 mJ for the start and end respectively. Evaluating equation 6 and 7 over all 537 primitive operations in the power dataset, we calculate the relative uncertainty of voltage and current as a function of energy as 𝑢𝑈 (𝐸) = 51.9 mJ and 𝑢𝐼 (𝐸) = 189.0 mJ. Hence, formulas 3, 4, and 5 are estimated to be 𝑢𝐸 (𝐸) = 0.19 joule, 𝑢𝑡 (𝐸) = 1.82 mJ, and 𝑢 𝑗𝑢𝑚𝑝 (𝐸) = 9.648 mJ. Consequently, equation 1 is evaluated as 𝑢𝐵 (𝐸) = 0.19 joule, 𝑢 𝑗𝑢𝑚𝑝 (𝐸) is constant across operations and contributes only 5% of the uncertainty of 𝑢𝐵 (𝐸) for a typical 20 s segment. 𝑢𝑡 (𝐸) is negligible (≈ 10−3 % relative) regardless of operation.

Figure 3: Shows a schematic overview of the measurement setup. Power supply is based on the primitives campaign. In the Raspberry Pi the runner with the benchmark binaries and the different environments is outlined.

5.2

Cryptographic Stack

All three environments use OpenSSL 3.4.1 with liboqs 0.12.0, the native and container variants additionally load oqs-provider 0.8.0 to expose liboqs algorithms through OpenSSL’s provider interface, while the unikernel links liboqs statically into the image. A single benchmark source tree is shared by all three environments. It is derived from the liboqs speed tests and calls post-quantum algorithms through the native liboqs API. We extended the benchmark macro to include a minimum number of iterations, as well as a minimum time, added CSV export for offline analysis, and fixed an unrelated bug in RDTSC. We evaluate all PQ algorithm families

Nikolai Puch, Chi Hieu Ta, and Moritz Beckel

supported by liboqs 0.12.0, however, we had to exclude some of the larger parameter sets of CROSS-small and Classic-McEliece. As a baseline it also implements the classical ECDSA, ECDHE, X25519, and RSA-2048 through OpenSSL’s EVP interface. For ECDSA and ECDHE the prime256v1 curve is chosen. All binaries are compiled for ARM64 and statically linked against musl 1.2.3.

5.3

The Three Environments

Native. The baseline runs the benchmark and OpenSSL binaries directly on the host OS. To keep the libc identical across environments, the native binaries are built inside the same Alpine container used for the container variant and exported to the host. Container. The container variant runs the identical binaries inside a Docker container based on Alpine Linux. The TLS client inside the container reaches the server via Docker’s host gateway. Unikernel. Unikraft 0.18.0 ("Helene") is used to build unikernels [27]. Unikraft provides a native port of OpenSSL, however, the version does not support the provider architecture of newer OpenSSL versions. Hence, we compile OpenSSL and liboqs as static libraries into a unikernel using Unikraft. We opted for static library builds, as compared to rewriting the Makefile, this approach yields less configuration overhead. Three application images are built: a primitive benchmark, an OpenSSL command-line image restricted to the s_server, s_client, and s_time tools, and a liboqs self-test image used to validate correctness of the port. Porting required patches to musl and to auxiliary-vector handling on ARM64. Images run under qemu-system-aarch64 with -machine virt-cpu max, a bridged virtual network (172.44.0.0/24), and a 9pfs share for exchanging files with the host. To reduce the overhead for the unikernel we build and sign certificates for the TLS tests on the host system and bootstrap the unikernel via 9pfs. The launcher enables KVM acceleration, and the benchmark driver verifies this precondition before every unikernel run and aborts otherwise. Randomness is provided as a boot-time seed on the kernel command line, drawn from the host’s /dev/urandom. The guest memory allocation was raised from 64 MiB after out-of-memory failures during porting.

5.4

Benchmark Orchestration

Each measurement campaign is orchestrated for reproducibility with a Python driver. The driver executes all stages and metrics over all environments and algorithm configurations. It records peroperation statistics as JSON, and stores per-algorithm timestamps for aligning the external power trace.

6

Evaluation

We first evaluate the primitives in isolation (RQ1), then TLS handshakes (RQ2), and finally power and energy for both (RQ3). Throughout, Ntv., Cnt., and Uk. denote the native, container, and unikernel environments. This section showcases the results, while Section 7 will go into more detail on potential root causes.

6.1

Primitive Performance

We present a comparative assessment of the mean time per operation for the approximately 70 algorithm variants and parameter sets. Figure 4 plots the results for the different environments for

ML-DSA, Falcon, SPHINCS+ at the different security levels plus classical algorithms as a baseline. In line with expectations, the time increases at higher security levels. Matching previous work, only ML-DSA and Falcon verification are competitive with ECDSA. Similarly, Figure 5 shows that only ML-KEM is on par with X25519. A representative overview of level-1 DSA and KEM as well as a comparison between the environments is shown in Tables 2 and 3 (a complete set can be found in Tables 9 and 10). Table 4 illustrates an aggregated overview per algorithm family to identify more quickly algorithms of interest. When comparing container against native the overhead ranges from 0.96 to 1.03 for the DSA, except for the SPHINCS+-SHA2 family (min 0.895) with the effect being more dominant for the short version and the higher security levels. For the baseline classical algorithms an overhead of (1.155) is restricted solely to the RSA key generation. For KEMs the overhead ranges from 0.986 to 1.038. The sole outliers are Classic-McEliece-460896 key generation (0.870) and ML-KEM-1024 encapsulation (0.781). With the exception of isolated outliers, the data indicates that the container overhead remains negligible at our benchmark resolution. The native benchmarks reveal identical trends. Specifically, DSA ML-DSA offers a balanced performance profile, whereas Falcon serves as a viable alternative for verification-intensive applications. SPHINCS+ is an alternative solely due to its conservative security assumptions, whereas CROSS is limited to providing fast key generation. For the unikernel, a more nuanced picture emerges. ML-DSA, the SHAKE variants of SPHINCS+, and CROSS variants performed near native (0.909-1.088). For the KEMs the SHAKE variant of FrodoKEM, and largely ML-KEM perform similar to the native run (0.986-1.091). ML-KEM-1024 encapsulation (0.787) is the outlier here too, with almost exactly the same value as in the container environment. MAYO (1.534), BIKE (1.280), and HQC (1.282) show a moderate overhead. Falcon key generation (2.304) and verification (1.427) are in a similar range, however, signing is significantly slower with an overhead ratio from 17.748-19.2 ×. The padded variants are also consistently slower. Conversely, the SHA-2 variants of SPHINCS+ run consistently faster in the unikernel than natively (0.542-0.671). Notably, the lattice-based algorithm families NIST has standardized, ML-KEM and ML-DSA, are among the ones that virtualize the best. Similarly, SHAKE variants of SPHINCS+ and FrodoKEM perform even slightly faster. Memory. Memory-wise, it is important to understand that the metrics are not directly comparable, as the memory layout is different for each virtualization technology. Instead, they provide insights into how much memory has to be allocated for each technology in practice. Table 5 shows the mean resident memory during primitive execution for the security level 1 variants of the DSAs and KEMs. The native processes use on average 4.200 MiB of RAM, the containerized process image 13.715 MiB, and the unikernel 57.655 MiB when computing the PQ algorithms. This shows the expected tradeoff for the statically compiled unikernel. While the classical algorithms require similar amounts of memory in the container, and even slightly less in the native environment, the unikernel variant requires almost 25 % more, when running the OpenSSL-implemented classical baselines. This is again expected based on the increased dependencies. Within one environment,

PQC in Lightweight Virtualization

Native (keygen/sign/verify) Docker (keygen/sign/verify) Unikraft (keygen/sign/verify) keygen sign verify

Execution time (us)

105

104

103

48

A

RS

EC

A-2 0

DS

ple

SP

HIN

CS

+-

SH

AK E

-25

6f-

sim

ple sim 2f-19 AK E SH +CS

HIN SP

SP

HIN

CS

+-

SH

AK E

-12

Fal c

8f-

on -1

sim

ple

02 4

12 on -5 Fal c

A-8 7 ML -DS

A-6 5 ML -DS

ML -DS

A-4 4

102

Figure 4: Mean time per operation for the signature primitives (NIST level 1 parameter sets and classical baselines) across the native, container, and unikernel environments. Table 2: Signature primitive performance (NIST level 1 parameter sets and classical baselines). Ntv. is the native mean time per operation in µs, over a 20 s window, Cnt.× and Uk.× are container and unikernel ratios compared to native.

Algorithm ML-DSA-44 Falcon-512 SPHINCS+-SHA2-128f-simple SPHINCS+-SHA2-128s-simple SPHINCS+-SHAKE-128f-simple SPHINCS+-SHAKE-128s-simple MAYO-1 cross-rsdp-128-fast ECDSA RSA-2048

Ntv.

Keygen Cnt.×

Uk.×

Ntv.

Sign Cnt.×

Uk.×

Ntv.

Verify Cnt.×

Uk.×

208.4 17237.5 5366.6 345295.9 5439.8 352488.2 1070.2 56.8 42.7 319292.4

1.00 1.03 0.97 0.97 1.00 1.00 1.00 1.00 1.01 1.16

1.00 2.38 0.59 0.59 1.00 0.95 1.60 0.98 0.96 1.11

914.7 633.0 125512.5 2629678.9 127668.4 2677162.1 2380.7 1821.3 96.5 4967.0

1.01 0.99 0.98 0.97 1.00 1.00 1.00 1.01 1.00 1.00

1.02 17.75 0.59 0.58 0.98 0.95 1.20 1.04 0.99 1.00

223.2 90.3 7238.9 2640.6 7426.7 2590.8 728.7 1037.1 291.4 133.1

1.00 1.00 1.01 0.92 1.01 1.00 1.02 1.00 1.00 1.01

1.01 1.41 0.61 0.60 0.99 0.91 2.00 1.03 0.99 1.01

differences between PQC algorithms are small. The maximum difference between signature algorithms is 0.999 MiB and 1.591 MiB for KEMs. Both this difference and the standard deviation are negligible between all environments, with only containers showing a slightly less spread (𝜎𝑐𝑛𝑡 = 0.145 compared to 𝜎𝑛𝑎𝑡 = 0.296 and 𝜎𝑢𝑘 = 0.301). In the native environment the SHAKE variant of SPHINCS+ requires the least memory. However, it has some of the outliers with the highest memory demand in the container environment. For unikernel it requires only slightly less memory than the

mean. MAYO tends to require the most memory between all environments. For KEMs Classic-McEliece requires more than average memory. BIKE, ML-KEM, and FrodoKEM tend to require less than average memory across all environments and operations. However, all have a few outliers, often in decapsulation in the container environment, where these algorithms require more memory.

6.2

TLS Handshake Throughput

Table 6 reports our results for completed initial TLS 1.3 handshakes in a 30 s window. As discussed in Section 4, the client environments

Nikolai Puch, Chi Hieu Ta, and Moritz Beckel

Native (keygen/encaps/decaps) Docker (keygen/encaps/decaps) Unikraft (keygen/encaps/decaps) keygen encaps decaps

105

Execution time (us)

104

103

19

ML -KE

X2

M10

55

24

8 M76

ML -KE

ML -KE

M51

2

56 C-2 HQ

92 C-1 HQ

HQ

C-1

28

102

Figure 5: Mean time per operation for the KEM primitives (NIST level 1 parameter sets and classical baseline) across the native, container, and unikernel environments. Table 3: KEM primitive performance (NIST level 1 parameter sets, classical baseline, and hybrid). Ntv. is the native mean time per operation in µs, over a 20 s window, Cnt.× and Uk.× are container and unikernel ratios compared to native.

Algorithm ML-KEM-512 BIKE-L1 HQC-128 FrodoKEM-640-AES FrodoKEM-640-SHAKE Classic-McEliece-348864 ECDHE x25519_Kyber512

Ntv.

Keygen Cnt.×

Uk.×

Ntv.

Encaps Cnt.×

Uk.×

Ntv.

Decaps Cnt.×

Uk.×

79.0 42923.7 5216.3 17010.6 5810.3 344534.1 62.3 132.4

1.01 1.00 1.00 1.00 1.00 1.04 0.99 1.00

1.09 1.28 1.29 0.92 0.99 1.06 1.00 1.23

94.4 2176.9 10495.6 17226.8 6573.4 168.8 245.3 303.8

1.01 1.00 1.00 1.00 0.99 1.00 1.00 1.00

1.02 1.29 1.28 0.92 0.99 1.51 1.01 1.14

114.8 35255.4 16640.0 17300.3 6483.2 56628.3 — 294.6

1.00 1.00 1.00 0.99 1.00 1.00 — 1.00

1.02 1.24 1.22 0.91 1.00 0.97 — 1.22

reach the always native server over different network paths inside the Pi4, so the numbers reflect deployment-level differences including transport, not cryptographic overhead alone. They also stem from an earlier measurement campaign than the primitive results. The lattice-based combinations (Dilithium/Falcon + Kyber) perform the best among all tested combinations and throughout all environments. For the container the slowdown for these combinations is 25-27 % and for the unikernel environment 28-36 %. For our setup Dilithium2+Kyber512 performed the best among all combinations, with 5.0 seconds per connection in the native environment and 6.3 and 6.6 in container and unikernel. That means that this pairing is on par with classical ECDSA+ECDHE in every environment, which matches prior observations that lattice-based PQC is competitive with classical elliptic-curve cryptography on handshake cost [39]. When looking at the slower combinations, it becomes

obvious that expensive cryptography hides the environment. This is especially obvious for SPHINCS+ where the expensive verification operation limits all environments to 11 connections in all three environments (2.7 seconds per connection). At this scale the perconnection transport and virtualization cost is negligible relative to the cryptographic cost. Interestingly, for BIKE-L1 and HQC-128 the container environment outperforms the native one. However, the difference is within plausible single-run variation, which is a plausible explanation as we did not see such behavior in our primitive runs.

6.3

Power and Energy

We conducted a comprehensive evaluation of power and energy, as it is often a limiting factor for embedded devices.

PQC in Lightweight Virtualization

Table 4: Per-family execution-time overhead relative to native, across all parameter sets and operations of each family (𝑛 = number of operation). DSAs above, KEMs below.

Family

𝑛

Container/native med. range

Unikernel/native med. range

ML-DSA Falcon SPHINCS+ (SHA-2) SPHINCS+ (SHAKE) MAYO CROSS ECDSA RSA-2048

9 12 18 18 12 45 3 3

1.003 1.000 0.964 1.003 0.999 1.001 1.001 1.006

1.000–1.014 0.964–1.031 0.895–1.012 0.994–1.017 0.971–1.020 0.981–1.014 0.995–1.011 1.003–1.155

1.007 2.307 0.588 0.959 1.534 1.017 0.993 1.007

0.993–1.020 1.404–19.200 0.542–0.671 0.909–1.024 1.181–2.240 0.965–1.088 0.961–0.994 1.003–1.114

ML-KEM BIKE Classic McEliece HQC FrodoKEM (AES) FrodoKEM (SHAKE) ECDHE

9 9 12 9 9 9 2

1.001 0.999 1.001 1.000 0.999 0.995 0.998

0.781–1.009 0.995–1.001 0.870–1.038 0.986–1.002 0.995–1.003 0.994–1.007 0.993–1.004

1.016 1.280 1.052 1.282 0.918 0.993 1.002

0.787–1.091 1.232–1.294 0.907–1.655 1.220–1.292 0.910–0.923 0.984–0.998 0.999–1.006

Table 5: Mean resident memory (MiB) during primitive execution (DSAs above, KEMs below), averaged over all operations of each algorithm. Algorithm

Native

Container

Unikernel

ML-DSA-44 Falcon-512 SPHINCS+-SHA2-128f-simple SPHINCS+-SHA2-128s-simple SPHINCS+-SHAKE-128f-simple SPHINCS+-SHAKE-128s-simple MAYO-1 cross-rsdp-128-fast ECDSA RSA-2048

4.1 4.1 4.1 4.0 3.9 3.9 4.5 4.0 4.0 3.9

13.7 13.6 13.6 13.7 13.7 13.7 13.7 13.6 13.7 13.6

57.4 57.5 57.5 57.6 57.5 57.6 57.8 57.5 76.1 78.5

ML-KEM-512 BIKE-L1 HQC-128 FrodoKEM-640-AES FrodoKEM-640-SHAKE Classic-McEliece-348864 ECDHE

4.0 3.9 4.1 3.9 3.9 4.5 4.0

13.7 13.7 13.7 13.9 13.7 13.6 13.7

57.4 57.4 57.5 57.5 57.5 58.0 76.4

Primitives. Table 7 reports the power and energy per operation for each algorithm family, averaged across all of the family’s parameter sets, integrated from the HMC8043 power trace, with the device’s idle baseline power subtracted out (see Section 4). Raw values are listed in Table 13 and 14. Mean power during PQC primitive execution ranges from 0.97-1.57 W across algorithms and environments. However, execution time is still the driving factor so that the overall baseline-subtracted energy results still closely track the time results. ML-DSA shows good performance overall. However, it does not achieve the very low energy requirements of ECDSA. Falcon requires the least energy (𝐸𝑛𝑎𝑡 = 0.166 mJ) for verification and CROSS

Table 6: Mean time per completed initial TLS 1.3 handshake in ms, divided by the number of handshakes completed in the 30 s measurement window. Server native environment in all cases. Ratios to native in parentheses. Signature

KEM

Dilithium2 Falcon-512 Dilithium3 Falcon-1024 Dilithium3 Falcon-1024 SPHINCS+-128s Dilithium2 Dilithium2 Dilithium2 Dilithium2 RSA-2048 ECDSA

Kyber512 Kyber512 Kyber768 Kyber768 Kyber1024 Kyber1024 Kyber512 BIKE-L1 HQC-128 Frodo-640-AES Frodo-640-SHAKE ECDHE ECDHE

Native

Container

Unikernel

5.0 6.3 (1.25) 6.6 (1.31) 5.1 6.4 (1.26) 7.0 (1.36) 5.7 7.2 (1.27) 7.2 (1.28) 5.9 7.4 (1.26) 7.7 (1.30) 5.8 7.3 (1.25) 7.4 (1.28) 6.1 7.6 (1.25) 7.9 (1.30) 2727 2727 (1.00) 2727 (1.00) 156 143 (0.91) 158 (1.01) 50.2 48.2 (0.96) 51.7 (1.03) 71.1 74.3 (1.04) 70.4 (0.99) 27.9 30.3 (1.08) 30.0 (1.07) 9.5 11.0 (1.16) 11.3 (1.19) 5.1 6.5 (1.29) 6.9 (1.36)

(𝐸𝑛𝑎𝑡 = 0.100 mJ) the least for key generation. SPHINCS+, on the other hand, requires three to four orders of magnitude more than the lattice schemes. E.g., SPHINCS+ SHA-2 signing requires 1594 × more energy than ML-DSA in the native environment. Similarly, for the KEMs ML-KEM has an average power draw of 1.21 W in the native environment. Its fast execution drops the required energy to 1.93 mJ, which is on equal footing with ECDHE. In contrast, HQC requires more power in all but the unikernel environment than ML-KEM, which confounds the effect of the longer execution time per operation. This means it requires two orders of magnitude more energy in the end. The environments seem to have only a marginal impact on the required power overall. The average container to native ratio over all PQ DSAs is 1.021 × and the unikernel to native ratio is 0.980 ×. For the KEMs the values are 1.023 × and 0.958 ×. While still small in absolute numbers, the unikernel thus reduces the required power. When comparing energy between environments, execution time again becomes the major contributor. Even though Falcon’s power draw drops from 1.1 W to 0.97 W when shifting from native to unikernel execution, its energy rises by a factor of 16.8 × to 17.82 mJ. This is consistent with the overhead during speed measurements of 17.75 × for Falcon-512 signing. Similarly, the SHA-2 variant of SPHINCS+ is able to reduce its average above baseline power cost from 1.28 W in the native environment to 1.24 W in the unikernel (factor: 0.974 ×). Paired with the previously observed speed-up this results in a reduction of the required energy by a factor of 0.60 ×. In conclusion the measurements show that energy per operation tracks execution time. TLS: energy per handshake follows throughput. Table 8 shows mean system power and derived energy per completed handshake for the TLS campaign. Mean system power varies little for the PQC algorithms (3.21-3.46 W). Energy per handshake, however, follows throughput: for the fast lattice combinations the system energy cost per handshake is 34-57 mJ, outperforming RSA and in some combinations even ECDSA with ECDHE. However, switching the KEM from the fast Kyber512 to the second fastest Frodo-640-SHAKE already increases the required system energy by a factor of 5.67 ×. For the DSAs, Falcon is only slightly more energy intensive, on the

Nikolai Puch, Chi Hieu Ta, and Moritz Beckel

Table 7: Mean power 𝑃 (W) and energy 𝐸 (mJ) per operation above the idle baseline (DSA and KEM primitive families and classical baselines), averaged across all of each family’s NIST parameter sets. Derived from the HMC8043 power trace aligned with per-operation timestamps.

Family

Op

Native Container Unikernel 𝑃 𝐸 𝑃 𝐸 𝑃 𝐸 (W) (mJ) (W) (mJ) (W) (mJ)

Keygen Sign Verify Keygen Falcon Sign Verify Keygen SPHINCS+ (SHA-2) Sign Verify Keygen SPHINCS+ (SHAKE) Sign Verify Keygen MAYO Sign Verify Keygen CROSS Sign Verify Keygen ECDSA Sign Verify Keygen RSA-2048 Sign Verify

1.23 1.22 1.18 1.07 1.10 1.22 1.27 1.33 1.23 1.25 1.21 1.18 1.28 1.22 1.57 1.21 1.32 1.30 1.10 0.93 0.89 0.99 0.98 0.75

0.485 1.83 0.475 33.78 1.06 0.166 323 2917 10.55 311 2330 9.93 5.58 11.10 4.37 0.100 8.79 4.97 0.047 0.089 0.260 363 4.98 0.102

1.17 1.17 1.25 1.15 1.21 1.17 1.19 1.31 1.26 1.20 1.33 1.34 1.29 1.22 1.44 1.22 1.32 1.28 1.29 1.12 1.06 0.73 0.68 0.69

0.467 1.72 0.519 40.29 1.21 0.165 288 2726 10.33 287 2784 11.29 6.21 12.74 4.01 0.110 9.29 5.16 0.063 0.114 0.323 261 3.65 0.101

1.20 1.31 1.30 1.02 0.97 1.07 1.20 1.29 1.24 1.23 1.31 1.22 1.12 1.13 1.15 1.23 1.25 1.27 1.20 1.09 0.95 1.01 1.01 0.97

0.446 1.92 0.518 84.50 17.82 0.211 183 1620 6.04 285 2652 9.49 7.87 13.21 6.48 0.105 8.58 4.88 0.050 0.107 0.281 349 5.20 0.133

Keygen Encaps Decaps Keygen BIKE Encaps Decaps Keygen Classic McEliece Encaps Decaps Keygen HQC Encaps Decaps Keygen FrodoKEM (AES) Encaps Decaps Keygen FrodoKEM (SHAKE) Encaps Decaps Keygen ECDHE Encaps Decaps

1.14 1.23 1.27 1.44 1.40 1.42 1.40 1.24 1.07 1.38 1.33 1.42 1.12 1.14 1.11 1.25 1.25 1.39 1.19 1.23 —

0.157 0.188 0.233 247 12.12 205 914 0.346 78.48 22.56 42.26 74.12 46.86 52.93 51.54 17.38 19.23 21.91 0.075 0.300 —

1.25 1.30 1.35 1.47 1.23 1.30 1.49 1.41 1.07 1.43 1.46 1.36 1.06 1.22 1.14 1.23 1.27 1.25 1.02 0.92 —

0.190 0.212 0.271 255 11.67 207 921 0.399 81.12 25.01 50.27 76.50 51.46 57.38 52.67 19.41 22.29 21.45 0.066 0.230 —

1.21 1.25 1.18 1.25 1.11 1.04 1.48 1.25 1.07 1.14 1.10 1.04 1.12 1.25 1.25 1.30 1.25 1.23 1.21 1.05 —

0.178 0.203 0.229 250 13.63 175 743 0.534 79.48 26.34 49.98 76.11 45.91 51.31 52.88 18.53 19.80 19.77 0.076 0.266 —

ML-DSA

ML-KEM

other hand, SPHINCS+ requires 567 × more system energy. When comparing environments, container costs more system energy per handshake than native by a factor of 1.022 ×. Unikernels require even more energy, raising the factor to 1.039 ×. However, summarizing, algorithm choice dwarfs environment choice: a SPHINCS+ signed handshake costs three orders of magnitude more energy than a Dilithium signed one in every environment.

7

Discussion

When looking at the unikernel, we consistently observe at the primitives level that the standardized lattice-based algorithms run at essentially native speed in the unikernel. However, MAYO, BIKE,

Table 8: Mean system power draw 𝑃 (W) and energy per completed handshake 𝐸ℎ𝑠 (mJ) during the TLS power campaign. Measured inline at the device supply with the FNIRSI FNB58. Connection counts and energy stem from the same run.

Signature

KEM

Dilithium2 Falcon-512 Dilithium3 Falcon-1024 Dilithium3 Falcon-1024 SPHINCS+-128s Dilithium2 Dilithium2 Dilithium2 Dilithium2 RSA-2048 ECDSA

Kyber512 Kyber512 Kyber768 Kyber768 Kyber1024 Kyber1024 Kyber512 BIKE-L1 HQC-128 Frodo-640-AES Frodo-640-SHAKE ECDHE ECDHE

Native 𝑃 𝐸ℎ𝑠

Container 𝑃 𝐸ℎ𝑠

Unikernel 𝑃 𝐸ℎ𝑠

3.26 34 3.27 44 3.42 48 3.24 35 3.25 45 3.37 50 3.29 39 3.29 50 3.44 53 3.25 41 3.29 52 3.41 56 3.28 40 3.29 51 3.46 55 3.28 42 3.28 53 3.43 57 3.32 19281 3.33 19710 3.41 19933 3.25 1053 3.36 997 3.25 1073 3.22 334 3.34 334 3.25 350 3.21 472 3.22 499 3.26 476 3.34 193 3.35 217 3.39 212 3.17 63 3.21 75 3.29 79 3.23 35 3.24 45 3.32 49

HQC, and above all Falcon’s signing routine, which we will discuss in more detail in the following paragraph, carry a measurable unikernel penalty. At the TLS level containers and unikernels each still give up between one seventh and one quarter of native throughput for fast handshake combinations, including the standardized lattice pairing Dilithium2+Kyber512, while for BIKE, HQC, and FrodoKEM the environments are harder to tell apart and for SPHINCS+ indistinguishable. This can be explained by the larger portion of the static costs associated with virtualization, like the different network paths in the environments (see Section 8). So at least part of this TLS-level gap plausibly reflects the unikernel’s connection path rather than the cryptography running over it. The energy results reinforce this observation from a different angle. Mean power draw differs only marginally across environments, both for primitives (0.96-1.02 × native, averaged over all PQ families) and for TLS handshakes (system energy per handshake at 1.022 × native in the container and 1.039 × in the unikernel). In the cases where environment influences energy, it does so mainly by moving time rather than power. This shows that virtualization cost is inversely related to cryptographic cost and algorithm choice remains the dominant factor. Only when the algorithms are sufficiently fast, the choice of environment begins to influence the performance. As discussed, Falcon signing is the single largest outlier (17.7519.20 ×). Since Falcon is the only benchmarked scheme whose signing procedure depends on double-precision floating-point arithmetic [16], a plausible cause for this behavior includes a different floating-point/math-library configuration in the Unikraft build or emulated floating-point paths in the guest. This is reinforced by the fact that verification, which is integer-dominated, shows only moderate overhead. To exclude emulated execution as the cause we enforced that KVM acceleration was active during the recorded campaigns. Notably, the anomaly is confined to latency, not power draw. Falcon signing’s mean power is in fact slightly lower in the unikernel than natively (0.97 W vs. 1.10 W), a behavior more consistent with the guest stalling, e.g. on an emulated or absent hardware FPU path, than with it doing additional work. However, we were

PQC in Lightweight Virtualization

not able to trace the root cause to e.g. a missing floating-point configuration or math library. Until root-caused, Falcon signing inside Unikraft on QEMU should be considered impractical on this class of hardware. The SPHINCS+ split is equally notable: the speed-up is specific to the SHA-2 parameter sets (0.54-0.67 ×), while the SHAKE parameter sets perform near-native. Since the two variants differ exactly in their internal hash function, this points at the SHA-2 code path rather than at the signature scheme itself. The power trace corroborates that this is a genuine effect rather than measurement noise. Mean power above baseline for the SHA-2 variant is also slightly lower in the unikernel than natively (1.24 W vs. 1.28 W, factor 0.974 ×), which combined with the speed-up compounds to a 0.60 × reduction in energy per operation. This is caused by an indirection during compilation. The unikernel uses the liboqs SHA-2 implementation instead of the one by OpenSSL, as OPENSSL_cpuid_setup normally crashes the unikernel via its SIGILL probe, and this choice also reduces build complexity. However, this also pulls in an ARMv8 Cryptography-Extension-accelerated SHA-2, compiled with -mcpu = cortex-a53+crypto. This setup is most likely faster than the OpenSSL version and thus the main contributor to the speed-up. The unikernel receives its randomness as a one-time boot seed and expands it internally, while native and containerized processes call into the host kernel for entropy during operation. This may also contribute to the divergent behavior to a lesser extent, as algorithms that request entropy frequently pay a syscall-path cost natively and in containers that the unikernel avoids, whereas the unikernel’s software expansion adds cost elsewhere. This effect could be isolated by an experiment fixing the entropy source across environments. Security-wise, boot-time-only seeding is a doubleedged sword: it removes a runtime dependency, but the guest’s entire entropy pool derives from one seed exposed on the QEMU command line. This has to be considered during deployment of cryptographic unikernel workloads, as this flag is visible to hostside observers. The liboqs 0.12.0 release ships both the round-3 submissions Dilithium/Kyber and the FIPS-final ML-DSA/ML-KEM used throughout the previous sections. Still, comparing these versions shows some interesting results. On this Pi4 target, oqsconfig.h compiles in an AArch64-optimized backend for every Dilithium and Kyber parameter set, but no such backend exists for ML-DSA or ML-KEM in this liboqs version, which therefore always run the portable reference implementation. This leads to a measurable effect. The ratios between native and container for Dilithium, ML-DSA, Kyber, and ML-KEM and the ratio for ML-DSA and ML-KEM in the unikernel are all approximately 1.0 ×. However, it rises, on average, to 1.568 × for Dilithium and 2.220 × for Kyber in the unikernel environment. This is not caused by a slowdown of the FIPS versions, but instead of a faster runtime of the round-3 variants in both the native and container environment. For example, Dilithium2 signing takes 466.618 µs, but ML-DSA-44 requires 914.715 µs in the native environment. Dilithium2 needs only around 0.5 × the time. While the container has the same 50 % difference, in the unikernel Dilithium needs 937.512 µs and ML-DSA 933.227 µs, basically the same amount of time. This shows an environment-specific slowdown inside the Unikraft guest: even though the optimized

_aarch64 code is compiled into that same guest image, it is evidently not being selected at runtime there. Plausible mechanisms for this behavior include liboqs’s runtime CPU-feature probe reporting negatively, or being unavailable, inside Unikraft’s minimal runtime despite the feature being physically present, or a mismatch between the vCPU model QEMU presents to the guest and the host CPU used for native/container. We evaluated with X25519-Kyber also a hybrid KEM. However, as expected, its performance closely follows its individual component algorithms with an overhead of ≈ 1.2 % in the native environment. The environment speed overhead lies between the individual overheads with no apparent additional influence. Memory is additive for native, but only up to ≈ 56 % of the individual parts. Due to the larger size, this effect is marginal for container and unikernel. Energy is dominated by the execution time, leading to the same results as for the speed performance overhead. Due to its composition, it follows Kyber and not ML-KEM, which leads to slightly different splits. When comparing container with unikernel, for pure computation the container is essentially free while the unikernel is not, the opposite of what one might expect given that a unikernel eliminates system-call and scheduling overhead. For power, the situation reverses and the unikernel’s mean power draw during primitive execution is, if anything, slightly below native, so its computational cost shows up in execution time and only marginally in power. For TLS, the two are close, but especially for the fast algorithms the container consistently beats the unikernel environment. The distinct path to the server for each environment may confound this observation. Memory tells another story at a different layer: the unikernel image adds about 5 × the amount of resident memory over native compared to the container environment as seen from the host. The unikernel figure includes the QEMU process itself. The way we measure the memory of the container makes it, however, largely independent of the selected algorithm (see Section 8). On the other hand native memory correlates moderately with the unikernel’s (Spearman 𝜌 ≈ 0.44–0.49, 𝑝 < 0.01 for signatures and KEMs alike), so some of each algorithm’s native memory footprint survives underneath the unikernel’s much larger fixed cost. However, as the other discussions in this section show, these effects are only marginal compared to when differences arise in the choice of implementation, either during build or runtime. Finally, our study measures performance only and does not discuss the direct security considerations. The isolation guarantees of containers and unikernels differ qualitatively [27, 28], and a deployment decision must weigh both.

8

Limitations

The underlying virtualization method has a large impact on the performance of a unikernel. The unikernel launcher derives its -enable-kvm flag from whether /dev/kvm exists and falls back silently. To control this behavior the benchmark driver precedes every unikernel run with an explicit check that aborts the whole campaign if /dev/kvm is missing or not read/write accessible. Thus, all three primitive-campaign runs underlying this paper’s numbers confirm KVM was active, allowing us to rule out TCG-emulation as a possible cause for outliers.

Nikolai Puch, Chi Hieu Ta, and Moritz Beckel

The TLS server always runs natively on the Pi4. The clients connect over different stacks, depending on the environment: loopback (native), Docker’s host gateway (container), or a QEMU bridge (unikernel). The measured differences on fast combinations therefore conflate handshake computation with network-path cost, thus the absolute throughput ratios should not be read as pure cryptographic overhead. To isolate this effect, we ran the TLS benchmark with session resumption enabled. With resumption, native and container throughput rise substantially for the fast combinations (Dilithium2+Kyber512: 1.6 × and 1.5 × respectively). The unikernel stays relatively consistent (1.1 ×), indicating that it is limited by communication and not cryptographic computation. Slower operations using e.g. SPHINCS+ still change drastically (≈ 500 ×), showing the impact of the cryptography in these cases. However, this difference in network stacks would also affect real-world applications, and, moreover, the primitive benchmarks are unaffected. All results stem from one Pi4 and from a single 20 s per primitive and 30 s TLS operation measurement window. Within-window standard deviations are recorded for primitives, but there are no independent repetitions from which cross-run confidence intervals could be derived. The primitive dataset was collected with a more mature toolchain, whereas the TLS speed and TLS power datasets stem from an earlier campaign. However, most changes impact robustness, accuracy, and context, e.g., switching from the FNB58 USB meter to the HMC8043 power supply. Thus, we can use the primitive runs as cross-comparisons between primitive-level and TLS-level results to verify and analyze qualitative findings observed inside the TLS run. For quantitative analysis, however, the primitive results should be favored. For the primitive campaign, memory is sampled by polling proc information. Which process gets tracked is a choice, and for the container and the unikernel there are multiple candidates that could each be measured. For the container, the whole Docker/containerd daemon, the containerd-shim process, and the actual containerized benchmark process are three differently sized quantities. To still include some overhead from the container environment we decided to measure the shim’s footprint at the cost of washing out more of the cryptographic algorithm’s cost. For the unikernel, the whole host-side qemu-system-aarch64 process and the guest-internal heap footprint are likewise potential quantities. The driver tracks the former, against a fixed 1 GiB guest allocation for the primitive benchmarks. Peak usage and allocation behavior inside either the container or the unikernel guest are therefore not visible, and the container figures in particular should not be read as purely reflecting the benchmarked algorithms. Future work should extend the analysis onto all potential measurands to gain a more differentiated picture. Both instruments used for power measurements measure wholesystem input power to understand the power draw of the whole system during cryptographic operations. The primitive campaign uses a lab-grade programmable supply (HMC8043) with trace alignment, which improves on the earlier TLS setup. The HMC8043 trace additionally records an idle baseline immediately before each primitive operation, which drifts by up to ≈600 mW between measurement windows independent of which environment is running. This drift is large relative to many operations’ own power draw above idle, so we report energy with this baseline subtracted rather

than raw energy. The TLS campaign’s FNB58 trace does not record a comparable per-handshake baseline, so its energy-per-handshake figures are not baseline-subtracted. This has to be considered when deriving conclusions from the results. Finally, we measure performance, memory, and power only. We do not evaluate isolation strength, side channels, boot time, image size, or build/toolchain effort. Additionally, our port is a research prototype, which could be further tuned. However, we decided against it as its setup already required significantly higher effort than the container version. Still, its performance may not represent a tuned production unikernel.

9

Conclusion

We presented the first measurement study of post-quantum cryptography under lightweight virtualization on embedded hardware, comparing native execution, Docker containers, and Unikraft unikernels with an identical OpenSSL/liboqs stack on a Pi4. We find that containers are essentially free for post-quantum computation, whereas unikernel overhead is strongly algorithm-dependent. For ECDSA, RSA, CROSS, FrodoKEM, Classic McEliece, and, notably, the NIST-standardized ML-KEM, ML-DSA, and the SHAKE variants of SPHINCS+, the overhead is near-native, while it concentrates in the alternates (1.28–1.53× for BIKE, HQC, and MAYO) and, above all, in Falcon signing (17.8–19.2×). The SHA-2 variants of SPHINCS+ run faster than native (0.54–0.67×) in our environment. Separately, we identify a build artifact: the pre-standardization Kyber and Dilithium implementations run 1.57–2.22× slower in the unikernel than the ML-KEM and ML-DSA that superseded them, even though the same AArch64-optimized backend is compiled into both, evidently going unused at runtime in the unikernel guest. Per-operation energy tracks execution time in every environment. In TLS, both environments cost 14–27% of native handshake throughput while the cryptography is cheap, and nothing once it is expensive. System power differs by at most a few percent throughout. This shows that algorithm selection matters most, but the careful construction and verification of which implementation and cryptographic libraries are used can have a significant impact as well. Future work includes root-causing the Falcon signing anomaly, verifying that the SPHINCS+ SHA-2 outlier is caused by the linked cryptographic libraries, isolating the entropy-handling effect experimentally, and adding boot-time and image-size measurements to complete the deployment trade-off. The comparison could also be broadened by also including MicroVMs. Furthermore, migrating the current implementation to Unikraft v0.21.0 is also on the list of future work. Version 0.21.0 introduces support for reseeding the Cryptographically Secure Pseudo-Random Number Generator [11], which would be of interest for evaluating the entropy-handling. A dedicated evaluation of Unikraft’s native CSPRNG in terms of security and performance might be relevant independent of its impact on PQC.

References [1] 2025. A Coordinated Implementation Roadmap for the Transition to Post-Quantum Cryptography | Shaping Europe’s Digital Future. https://digital-strategy.ec.europa.eu/en/library/coordinated-implementationroadmap-transition-post-quantum-cryptography

PQC in Lightweight Virtualization

[2] 2023. Guide to the Expression of Uncertainty in Measurement — Part 1: Introduction. doi:10.59161/JCGMGUM-1-2023 [3] i.MX8 MPU Platforms 2026. I.MX8 MPU Platforms. i.MX8 MPU Platforms. https://www.nxp.com/design/design-center/development-boardsand-designs/automotive-development-platforms/i-mx-mpu-platforms:AUTOiMX8-PLATFORMS [4] mupq 2025. Mupq/Pqm4. mupq. https://github.com/mupq/pqm4 [5] Open Quantum Safe 2025. Open-Quantum-Safe/Liboqs. Open Quantum Safe. https://github.com/open-quantum-safe/liboqs [6] Open Quantum Safe 2026. Open-Quantum-Safe/Oqs-Provider. Open Quantum Safe. https://github.com/open-quantum-safe/oqs-provider [7] Open Quantum Safe 2026. OpenQuantumSafe TLS. Open Quantum Safe. https: //openquantumsafe.org/applications/tls.html [8] 2026. OSv - the Operating System Designed for the Cloud. https://osv.io/ [9] Linux Foundation 2026. PQ Code Package. Linux Foundation. https://github. com/pq-code-package [10] PQClean 2025. PQClean/PQClean. PQClean. https://github.com/PQClean/ PQClean [11] GitHub [n. d.]. Releases · Unikraft/Unikraft. GitHub. https://github.com/unikraft/ unikraft/releases#release-RELEASE-0.21.0 [12] 2021. R&S HMC804x Power Supply User Manual / Benutzerhandbuch. https://scdn.rohde-schwarz.com/ur/pws/dl_downloads/dl_common_library/ dl_manuals/dl_user_manual/HMC804x_UserManual_de_en_05.pdf [13] 2026. FNIRSI FNB58 USB Fast Charge Tester. https://www.fnirsi.com/products/ fnb58 [14] 2026. rumpkernel/rumprun. https://github.com/rumpkernel/rumprun originaldate: 2015-02-20T17:55:46Z. [15] Ashijeet Acharya, Jérémy Fanguède, Michele Paolino, and Daniel Raho. 2018. A Performance Benchmarking Analysis of Hypervisors Containers and Unikernels on ARMv8 and X86 CPUs. In 2018 European Conference on Networks and Communications (EuCNC) (2018-06). 282–287. doi:10.1109/EuCNC.2018.8443248 [16] Gorjan Alagic, Daniel Apon, David Cooper, Quynh Dang, Thinh Dang, John Kelsey, Jacob Lichtinger, Yi-Kai Liu, Carl Miller, Dustin Moody, Rene Peralta, Ray Perlner, Angela Robinson, and Daniel Smith-Tone. 2022. Status Report on the Third Round of the NIST Post-Quantum Cryptography Standardization Process. NIST IR 8413-upd1 pages. doi:10.6028/NIST.IR.8413-upd1 [17] Gorjan Alagic, Maxime Bros, Pierre Ciadoux, David Cooper, Quynh Dang, Thinh Dang, John Kelsey, Jacob Lichtinger, Yi-Kai Liu, Carl Miller, Dustin Moody, Rene Peralta, Ray Perlner, Angela Robinson, Hamilton Silberg, Daniel SmithTone, and Noah Waller. 2025. Status Report on the Fourth Round of the NIST Post-Quantum Cryptography Standardization Process. NIST IR 8545 pages. doi:10.6028/NIST.IR.8545 [18] Gorjan Alagic, Dustin Moody, Maxime Bros, Pierre Ciadoux, Quynh Dang, Thinh Dang, John Kelsey, Jacob Lichtinger, Yi-Kai Liu, Carl Miller, Rene Peralta, Ray Perlner, Angela Robinson, Hamilton Silberg, Daniel Smith-Tone, and Noah Waller. 2026. Status Report on the Second Round of the Additional Digital Signature Schemes for the NIST Post-Quantum Cryptography Standardization Process. NIST IR 8610 pages. doi:10.6028/NIST.IR.8610 [19] Giovanni Bartolomeo, Patrick Sabanic, Nitinder Mohan, and Jorg Ott. 2025. Supporting Hybrid Virtualization Orchestration for Edge Computing. In Proceedings of the 8th International Workshop on Edge Systems, Analytics and Networking (New York, NY, USA, 2025-03-31) (EdgeSys ’25). Association for Computing Machinery, 19–24. doi:10.1145/3721888.3722093 [20] Joppe W. Bos, Alexander Dima, Alexander Kiening, and Joost Renes. 2023. PostQuantum Secure Over-the-Air Update of Automotive Systems. escar Europe 21st (2023). doi:10.13154/294-10380 [21] Joppe W. Bos, Joost Renes, and Amber Sprenkels. 2022. Dilithium for Memory Constrained Devices. In Progress in Cryptology - AFRICACRYPT 2022, Lejla Batina and Joan Daemen (Eds.). Springer Nature Switzerland, Cham, 217–235. [22] Kevin Bürstinghaus-Steinbach, Christoph Krauß, Ruben Niederhagen, and Michael Schneider. 2020. Post-Quantum TLS on Embedded Systems: Integrating and Evaluating Kyber and SPHINCS+ with Mbed TLS. In Proceedings of the 15th ACM Asia Conference on Computer and Communications Security (New York, NY, USA, 2020-10-05) (ASIA CCS ’20). Association for Computing Machinery, 841–852. doi:10.1145/3320269.3384725 [23] Madelyn Cain, Qian Xu, Robbie King, Lewis R. B. Picard, Harry Levine, Manuel Endres, John Preskill, Hsin-Yuan Huang, and Dolev Bluvstein. 2026. Shor’s Algorithm Is Possible with as Few as 10,000 Reconfigurable Atomic Qubits. arXiv:2603.28627 [quant-ph] doi:10.48550/arXiv.2603.28627 [24] Tom Goethals, Merlijn Sebrechts, Ankita Atrey, Bruno Volckaert, and Filip De Turck. 2018. Unikernels vs Containers: An In-Depth Benchmarking Study in the Context of Microservice Applications. In 2018 IEEE 8th International Symposium on Cloud and Service Computing (SC2) (2018-11). 1–8. doi:10.1109/SC2.2018. 00008 [25] Ruben Gonzalez and Thom Wiggers. 2022. KEMTLS vs. Post-quantum TLS: Performance on Embedded Systems. In Security, Privacy, and Applied Cryptography Engineering (Cham, 2022), Lejla Batina, Stjepan Picek, and Mainack Mondal (Eds.). Springer Nature Switzerland, 99–117. doi:10.1007/978-3-031-22829-2_6

[26] Yacoub Hanna, Jessica Bozhko, Samet Tonyali, Ricardo Harrilal-Parchment, Mumin Cebe, and Kemal Akkaya. 2025. A comprehensive and realistic performance evaluation of post-quantum security for consumer IoT devices. Internet of Things 33 (2025), 101650. [27] Simon Kuenzer, Vlad-Andrei Bădoiu, Hugo Lefeuvre, Sharan Santhanam, Alexander Jung, Gaulthier Gain, Cyril Soldani, Costin Lupu, Ştefan Teodorescu, Costi Răducanu, Cristian Banu, Laurent Mathy, Răzvan Deaconescu, Costin Raiciu, and Felipe Huici. 2021. Unikraft: Fast, Specialized Unikernels the Easy Way. In Proceedings of the Sixteenth European Conference on Computer Systems (New York, NY, USA, 2021-04-21) (EuroSys ’21). Association for Computing Machinery, 376–394. doi:10.1145/3447786.3456248 [28] Anil Madhavapeddy and David J. Scott. 2014. Unikernels: The Rise of the Virtual Library Operating System. 57, 1 (2014), 61–69. doi:10.1145/2541883.2541895 [29] National Institute of Standards and Technology (US). 2024. Module-Lattice-Based Digital Signature Standard. NIST FIPS 204 pages. doi:10.6028/NIST.FIPS.204 [30] National Institute of Standards and Technology (US). 2024. Module-Lattice-Based Key-Encapsulation Mechanism Standard. NIST FIPS 203 pages. doi:10.6028/NIST. FIPS.203 [31] National Institute of Standards and Technology (US). 2024. Stateless Hash-Based Digital Signature Standard. NIST FIPS 205 pages. doi:10.6028/NIST.FIPS.205 [32] Sebastian Paul. 2022. On the Transition to Post-Quantum Cryptography in the Industrial Internet of Things. Ph. D. Dissertation. Technische Universität Darmstadt, Darmstadt. doi:10.26083/tuprints-00021368 [33] Max Plauth, Lena Feinbube, and Andreas Polze. 2017. A Performance Survey of Lightweight Virtualization Techniques. In Service-Oriented and Cloud Computing (Cham, 2017), Flavio De Paoli, Stefan Schulte, and Einar Broch Johnsen (Eds.). Springer International Publishing, 34–48. doi:10.1007/978-3-319-67262-5_3 [34] Nikolai Puch, Maximilian Pursche, Sebastian N. Peters, and Michael P. Heinl. 2026. SoK: The Engineer’s Guide to Post-Quantum Cryptography for Embedded Devices. In 2026 IEEE 11th European Symposium on Security and Privacy (EuroS&P) (2026-07). 23–42. doi:10.1109/EuroSP68448.2026.00014 [35] Leonie Reichert, Nicolas Coppik, and Soeren Finster. 2025. Performance Evaluation of Quantum-Resistant Algorithms on Industrial Embedded Systems. Availability, Reliability and Security (ARES) (2025), 127–148. doi:10.1007/978-3-03200630-1_8 [36] Maximilian Schöffel, Frederik Lauer, Carl C. Rheinländer, and Norbert Wehn. 2022. Secure IoT in the Era of Quantum Computers-Where Are the Bottlenecks? Sensors 22, 7 (2022). doi:10.3390/s22072484 [37] Peter W. Shor. 1997. Polynomial-Time Algorithms for Prime Factorization and Discrete Logarithms on a Quantum Computer. 26, 5 (1997), 1484–1509. arXiv:https://doi.org/10.1137/S0097539795293172 doi:10.1137/ S0097539795293172 [38] George Tasopoulos, Charis Dimopoulos, Apostolos P. Fournaris, Raymond K. Zhao, Amin Sakzad, and Ron Steinfeld. 2023. Energy Consumption Evaluation of Post-Quantum TLS 1.3 for Resource-Constrained Embedded Devices. In Proceedings of the 20th ACM International Conference on Computing Frontiers (New York, NY, USA, 2023-08-04) (CF ’23). Association for Computing Machinery, 366–374. doi:10.1145/3587135.3592821 [39] George Tasopoulos, Jinhui Li, Apostolos P. Fournaris, Raymond K. Zhao, Amin Sakzad, and Ron Steinfeld. 2022. Performance Evaluation of Post-Quantum TLS 1.3 on Resource-Constrained Embedded Systems. In Information Security Practice and Experience, Chunhua Su, Dimitris Gritzalis, and Vincenzo Piuri (Eds.). Vol. 13620. Springer International Publishing, 432–451. doi:10.1007/978-3-03121280-2_24

A

Ethical Considerations

As far as we are aware, this work does not raise ethical concerns.

B

Use of AI

In this work AI has been used for: • Coding assistance for creating the benchmark and analysis scripts. • Drafting, proofreading, grammar and spelling correction of the article itself. The authors remain solely responsible for the content and results of this work.

C

Full Primitive Results

Tables 9 and 10 list the mean per-operation times for all 60 benchmarked algorithm configurations.

Nikolai Puch, Chi Hieu Ta, and Moritz Beckel

Tables 11 and 12 list the mean per-operation resident memory for the same 60 configurations. Tables 13 and 14 list the mean per-operation energy for the same 60 configurations.

𝑢𝑡 (𝐸) = 𝐸

Error Propagation and Statistical Uncertainty Analysis Type A: Statistical Evaluation of Uncertainty. 𝑠 (𝑥) 𝑢𝐴 (𝑥) = √ 𝑛 Type B: Evaluation of Uncertainty Formulas. √︃ 𝑢𝐸 (𝐸) = 𝑢𝑈2 (𝐸) + 𝑢𝐼2 (𝐸)

(2)

|Δ𝑃𝑠 | · Δ𝑡 2 |Δ𝑃𝑒 | · Δ𝑡 2 ) +( √ ) √ 12 12 Δ𝑡 ∑︁ 𝐼𝑖 (𝑎𝑈 𝑈𝑖 + 𝑏𝑈 ) 𝑢𝑈 (𝐸) = √ 3 𝑖 Δ𝑡 ∑︁ 𝑢𝐼 (𝐸) = √ 𝑈𝑖 (𝑎𝐼 𝐼𝑖 + 𝑏 𝐼 ) 3 𝑖

Received 22 August 2026

(3)

(4)

√︄

𝑢 𝑗𝑢𝑚𝑝 (𝐸) =

D

𝑎𝑡 √ 𝑇𝑜𝑝 3

(

(5) (6)

(7)

PQC in Lightweight Virtualization

Table 9: Full signature-primitive results. Mean time per operation in µs over a 20 s window per operation. Ntv. is the native value; Cnt. and Uk. give the container/unikernel value, with the ratio to native in parentheses.

Algorithm Dilithium2 Dilithium3 Dilithium5 ML-DSA-44 ML-DSA-65 ML-DSA-87 Falcon-512 Falcon-1024 Falcon-padded-512 Falcon-padded-1024 SPHINCS+-SHA2-128f-simple SPHINCS+-SHA2-128s-simple SPHINCS+-SHA2-192f-simple SPHINCS+-SHA2-192s-simple SPHINCS+-SHA2-256f-simple SPHINCS+-SHA2-256s-simple SPHINCS+-SHAKE-128f-simple SPHINCS+-SHAKE-128s-simple SPHINCS+-SHAKE-192f-simple SPHINCS+-SHAKE-192s-simple SPHINCS+-SHAKE-256f-simple SPHINCS+-SHAKE-256s-simple MAYO-1 MAYO-2 MAYO-3 MAYO-5 cross-rsdp-128-balanced cross-rsdp-128-fast cross-rsdp-128-small cross-rsdp-192-balanced cross-rsdp-192-fast cross-rsdp-256-balanced cross-rsdp-256-fast cross-rsdpg-128-balanced cross-rsdpg-128-fast cross-rsdpg-128-small cross-rsdpg-192-balanced cross-rsdpg-192-fast cross-rsdpg-192-small cross-rsdpg-256-balanced cross-rsdpg-256-fast ECDSA RSA-2048

Ntv.

Keygen Cnt.

Uk.

Ntv.

Sign Cnt.

Uk.

Ntv.

158 293 449 208 369 566 17238 50310 17639 49660 5367 345296 7937 510142 22561 375745 5440 352488 8011 516730 22132 328001 1070 2374 3953 10335 57 57 57 121 121 212 213 29 29 29 53 53 53 81 81 43 319292

158 (1.00) 294 (1.00) 450 (1.00) 209 (1.00) 370 (1.00) 565 (1.00) 17771 (1.03) 48501 (0.96) 17788 (1.01) 48761 (0.98) 5201 (0.97) 334043 (0.97) 7705 (0.97) 489995 (0.96) 20186 (0.89) 338981 (0.90) 5465 (1.00) 353836 (1.00) 8003 (1.00) 517272 (1.00) 22187 (1.00) 328805 (1.00) 1071 (1.00) 2368 (1.00) 3951 (1.00) 10341 (1.00) 57 (1.00) 57 (1.00) 57 (1.01) 122 (1.00) 121 (1.00) 213 (1.00) 212 (1.00) 29 (1.01) 29 (1.01) 29 (1.01) 53 (1.00) 53 (1.00) 53 (1.01) 82 (1.01) 81 (1.01) 43 (1.01) 368907 (1.16)

208 (1.32) 371 (1.27) 563 (1.25) 209 (1.00) 369 (1.00) 562 (0.99) 40995 (2.38) 117906 (2.34) 40056 (2.27) 110284 (2.22) 3161 (0.59) 202330 (0.59) 4987 (0.63) 310677 (0.61) 12742 (0.56) 203649 (0.54) 5445 (1.00) 333623 (0.95) 7768 (0.97) 492093 (0.95) 20826 (0.94) 335812 (1.02) 1709 (1.60) 3061 (1.29) 6113 (1.55) 15722 (1.52) 56 (0.97) 56 (0.98) 56 (0.98) 119 (0.98) 118 (0.98) 205 (0.97) 205 (0.96) 29 (1.01) 29 (1.00) 29 (1.01) 56 (1.06) 56 (1.05) 56 (1.06) 88 (1.08) 88 (1.09) 41 (0.96) 355618 (1.11)

467 736 950 915 1492 1817 633 1285 627 1277 125513 2629679 209350 4698340 462626 4637549 127668 2677162 208003 4645198 446433 4039135 2381 3327 8894 23076 3331 1821 12435 7864 4443 15213 9353 2541 1297 9098 3808 3002 14005 6545 5033 97 4967

465 (1.00) 738 (1.00) 937 (0.99) 926 (1.01) 1513 (1.01) 1823 (1.00) 629 (0.99) 1280 (1.00) 628 (1.00) 1280 (1.00) 123323 (0.98) 2545501 (0.97) 202809 (0.97) 4513138 (0.96) 419039 (0.91) 4221643 (0.91) 128252 (1.00) 2681182 (1.00) 208158 (1.00) 4650068 (1.00) 448006 (1.00) 4044271 (1.00) 2377 (1.00) 3324 (1.00) 8914 (1.00) 23065 (1.00) 3306 (0.99) 1831 (1.01) 12429 (1.00) 7873 (1.00) 4426 (1.00) 15179 (1.00) 9312 (1.00) 2551 (1.00) 1297 (1.00) 9135 (1.00) 3816 (1.00) 2976 (0.99) 14057 (1.00) 6500 (0.99) 5027 (1.00) 97 (1.00) 4980 (1.00)

938 (2.01) 1526 (2.07) 1824 (1.92) 933 (1.02) 1518 (1.02) 1816 (1.00) 11234 (17.75) 24506 (19.07) 11230 (17.90) 24515 (19.20) 73635 (0.59) 1530117 (0.58) 128942 (0.62) 2865978 (0.61) 261087 (0.56) 2522469 (0.54) 125275 (0.98) 2535920 (0.95) 200594 (0.96) 4425673 (0.95) 421752 (0.94) 4038059 (1.00) 2852 (1.20) 3929 (1.18) 10555 (1.19) 27522 (1.19) 3477 (1.04) 1903 (1.04) 13056 (1.05) 7862 (1.00) 4464 (1.00) 15972 (1.05) 9769 (1.04) 2591 (1.02) 1353 (1.04) 9290 (1.02) 3834 (1.01) 2963 (0.99) 13921 (0.99) 6636 (1.01) 5237 (1.04) 96 (0.99) 4980 (1.00)

152 253 435 223 360 588 90 177 91 177 7239 2641 11245 3829 11943 6280 7427 2591 10983 3871 11625 5803 729 792 2601 6564 1931 1037 7407 4307 2566 7338 5200 1545 787 5511 2340 1865 8796 3836 3129 291 133

Verify Cnt.

Uk.

152 (1.00) 225 (1.48) 254 (1.00) 362 (1.43) 435 (1.00) 591 (1.36) 224 (1.00) 226 (1.01) 360 (1.00) 363 (1.01) 588 (1.00) 592 (1.01) 90 (1.00) 128 (1.41) 177 (1.00) 255 (1.44) 90 (1.00) 127 (1.40) 177 (1.00) 255 (1.45) 7329 (1.01) 4449 (0.61) 2421 (0.92) 1590 (0.60) 11002 (0.98) 6888 (0.61) 3849 (1.01) 2567 (0.67) 11217 (0.94) 6928 (0.58) 5792 (0.92) 3431 (0.55) 7504 (1.01) 7324 (0.99) 2600 (1.00) 2356 (0.91) 11012 (1.00) 11047 (1.01) 3846 (0.99) 3666 (0.95) 11822 (1.02) 11292 (0.97) 5851 (1.01) 5487 (0.95) 743 (1.02) 1459 (2.00) 795 (1.00) 1775 (2.24) 2590 (1.00) 5082 (1.95) 6374 (0.97) 12324 (1.88) 1928 (1.00) 1987 (1.03) 1034 (1.00) 1072 (1.03) 7388 (1.00) 7531 (1.02) 4293 (1.00) 4201 (0.98) 2516 (0.98) 2527 (0.98) 7334 (1.00) 7301 (0.99) 5185 (1.00) 5494 (1.06) 1562 (1.01) 1607 (1.04) 788 (1.00) 810 (1.03) 5519 (1.00) 5703 (1.03) 2347 (1.00) 2340 (1.00) 1864 (1.00) 1865 (1.00) 8765 (1.00) 8734 (0.99) 3830 (1.00) 3969 (1.03) 3123 (1.00) 3219 (1.03) 290 (1.00) 290 (0.99) 134 (1.01) 134 (1.01)

Nikolai Puch, Chi Hieu Ta, and Moritz Beckel

Table 10: Full KEM-primitive results (Structure as in Table 9).

Algorithm

Ntv.

Keygen Cnt.

Uk.

Ntv.

Encaps Cnt.

Uk.

Ntv.

Decaps Cnt.

Uk.

BIKE-L1 42924 42878 (1.00) 54991 (1.28) 2177 2178 (1.00) 2810 (1.29) 35255 35090 (1.00) 43744 (1.24) BIKE-L3 133308 133185 (1.00) 170274 (1.28) 6682 6683 (1.00) 8605 (1.29) 110868 110283 (0.99) 136565 (1.23) BIKE-L5 334433 333915 (1.00) 427914 (1.28) 16743 16756 (1.00) 21663 (1.29) 276644 276171 (1.00) 343380 (1.24) Classic-McEliece-348864 344534 357498 (1.04) 364574 (1.06) 169 169 (1.00) 254 (1.51) 56628 56681 (1.00) 55146 (0.97) Classic-McEliece-348864f 160245 158878 (0.99) 154329 (0.96) 168 167 (0.99) 254 (1.51) 56650 56647 (1.00) 55141 (0.97) Classic-McEliece-460896 1389270 1209027 (0.87) 1288431 (0.93) 350 354 (1.01) 571 (1.63) 84538 84569 (1.00) 88917 (1.05) Classic-McEliece-460896f 536332 547725 (1.02) 486712 (0.91) 343 353 (1.03) 568 (1.65) 84500 84554 (1.00) 88946 (1.05) HQC-128 5216 5204 (1.00) 6732 (1.29) 10496 10513 (1.00) 13422 (1.28) 16640 16597 (1.00) 20296 (1.22) HQC-192 15877 15879 (1.00) 20509 (1.29) 31860 31854 (1.00) 40983 (1.29) 49512 48832 (0.99) 61736 (1.25) HQC-256 29152 29128 (1.00) 37499 (1.29) 58484 58570 (1.00) 74982 (1.28) 90338 90515 (1.00) 112967 (1.25) Kyber512 46 47 (1.02) 86 (1.86) 52 52 (1.00) 104 (1.98) 43 43 (1.00) 117 (2.73) Kyber768 68 68 (1.00) 136 (2.01) 79 79 (1.01) 164 (2.08) 69 69 (1.00) 182 (2.63) Kyber1024 100 100 (1.01) 205 (2.06) 115 115 (1.01) 239 (2.08) 104 103 (0.99) 265 (2.54) ML-KEM-512 79 79 (1.01) 86 (1.09) 94 95 (1.01) 96 (1.02) 115 114 (1.00) 117 (1.02) ML-KEM-768 130 131 (1.01) 137 (1.06) 151 150 (1.00) 151 (1.00) 179 179 (1.00) 181 (1.01) ML-KEM-1024 200 200 (1.00) 205 (1.03) 286 223 (0.78) 225 (0.79) 262 261 (1.00) 262 (1.00) FrodoKEM-640-AES 17011 16934 (1.00) 15680 (0.92) 17227 17253 (1.00) 15908 (0.92) 17300 17212 (0.99) 15758 (0.91) FrodoKEM-640-SHAKE 5810 5839 (1.00) 5759 (0.99) 6573 6539 (0.99) 6508 (0.99) 6483 6490 (1.00) 6453 (1.00) FrodoKEM-976-AES 39204 39142 (1.00) 35957 (0.92) 39906 39811 (1.00) 36642 (0.92) 39704 39703 (1.00) 36204 (0.91) FrodoKEM-976-SHAKE 13127 13042 (0.99) 12941 (0.99) 14527 14452 (0.99) 14477 (1.00) 14389 14495 (1.01) 14358 (1.00) FrodoKEM-1344-AES 74155 74077 (1.00) 67467 (0.91) 75064 75293 (1.00) 69026 (0.92) 75071 75245 (1.00) 68942 (0.92) FrodoKEM-1344-SHAKE 23771 23699 (1.00) 23399 (0.98) 26549 26386 (0.99) 26352 (0.99) 26365 26198 (0.99) 26317 (1.00) ECDHE 62 62 (0.99) 62 (1.00) 245 246 (1.00) 247 (1.01) — — — X25519 83 84 (1.00) 73 (0.88) 249 249 (1.00) 240 (0.96) 248 248 (1.00) 240 (0.96) x25519_Kyber512 132 133 (1.00) 163 (1.23) 304 304 (1.00) 346 (1.14) 295 295 (1.00) 360 (1.22) x25519_Kyber768 153 154 (1.00) 214 (1.39) 331 332 (1.00) 405 (1.22) 320 321 (1.00) 425 (1.33) x25519_Kyber1024 185 186 (1.00) 282 (1.52) 367 367 (1.00) 479 (1.31) 356 355 (1.00) 507 (1.42)

PQC in Lightweight Virtualization

Table 11: Full signature-primitive resident memory results. Mean per-operation resident memory in MiB for the native, container, and unikernel environments.

Algorithm

Keygen Sign Ntv. Cnt. Uk. Ntv. Cnt.

Verify Uk. Ntv. Cnt. Uk.

Dilithium2 Dilithium3 Dilithium5 ML-DSA-44 ML-DSA-65 ML-DSA-87 Falcon-512 Falcon-1024 Falcon-padded-512 Falcon-padded-1024 SPHINCS+-SHA2-128f-simple SPHINCS+-SHA2-128s-simple SPHINCS+-SHA2-192f-simple SPHINCS+-SHA2-192s-simple SPHINCS+-SHA2-256f-simple SPHINCS+-SHA2-256s-simple SPHINCS+-SHAKE-128f-simple SPHINCS+-SHAKE-128s-simple SPHINCS+-SHAKE-192f-simple SPHINCS+-SHAKE-192s-simple SPHINCS+-SHAKE-256f-simple SPHINCS+-SHAKE-256s-simple MAYO-1 MAYO-2 MAYO-3 MAYO-5 cross-rsdp-128-balanced cross-rsdp-128-fast cross-rsdp-128-small cross-rsdp-192-balanced cross-rsdp-192-fast cross-rsdp-256-balanced cross-rsdp-256-fast cross-rsdpg-128-balanced cross-rsdpg-128-fast cross-rsdpg-128-small cross-rsdpg-192-balanced cross-rsdpg-192-fast cross-rsdpg-192-small cross-rsdpg-256-balanced cross-rsdpg-256-fast ECDSA RSA-2048

4.3 4.3 4.3 4.1 4.1 4.1 4.1 4.1 4.0 4.3 4.1 4.0 4.1 4.1 4.1 4.1 3.9 3.9 3.9 3.9 4.0 3.9 4.5 4.4 4.4 4.9 4.3 4.0 4.8 4.4 4.1 4.8 4.3 4.0 4.0 4.5 4.1 4.1 4.6 4.5 4.3 4.0 3.9

57.5 57.5 57.5 57.5 57.4 57.6 57.4 57.5 57.5 57.5 57.6 57.8 57.5 57.8 57.6 57.7 57.5 57.6 57.5 57.9 57.6 57.7 57.8 57.7 57.8 58.2 57.6 57.5 58.1 57.8 57.7 58.2 57.8 57.5 57.6 57.8 57.6 57.7 58.3 57.8 57.7 76.1 78.3

13.6 13.6 13.8 13.8 13.6 13.6 13.6 13.7 13.8 13.6 13.6 13.5 13.5 13.9 13.8 13.7 13.6 13.7 13.9 13.5 13.9 13.7 13.7 13.5 13.6 13.8 13.7 13.7 13.6 13.8 13.6 13.9 13.9 13.8 13.8 13.8 13.7 13.4 13.9 13.4 13.9 13.7 13.5

57.5 57.4 57.6 57.5 57.4 57.6 57.5 57.6 57.5 57.5 57.5 57.4 57.6 57.5 57.5 57.5 57.5 57.5 57.5 57.5 57.5 57.6 57.8 57.7 57.8 58.3 57.6 57.5 58.2 57.9 57.8 58.2 57.9 57.5 57.5 57.9 57.6 57.6 58.3 57.8 57.7 76.1 78.8

4.3 4.3 4.3 4.1 4.1 4.1 4.1 4.1 4.0 4.3 4.1 4.0 4.1 4.1 4.1 4.1 3.9 3.9 3.9 3.9 4.0 3.9 4.5 4.4 4.4 4.9 4.3 4.0 4.8 4.4 4.1 4.8 4.3 4.0 4.0 4.5 4.1 4.1 4.6 4.5 4.3 4.0 3.9

13.6 13.7 13.6 13.6 14.1 13.6 13.6 13.8 13.8 13.6 13.5 13.8 13.6 14.0 13.6 13.8 13.8 13.7 13.7 13.7 13.8 13.9 13.6 13.8 13.7 13.6 13.9 13.5 13.8 13.7 13.8 13.8 13.7 13.8 13.6 13.8 13.9 13.8 13.6 13.7 13.9 13.8 13.5

4.3 4.3 4.3 4.1 4.1 4.1 4.1 4.1 4.0 4.3 4.1 4.0 4.1 4.1 4.1 4.1 3.9 3.9 3.9 3.9 4.0 3.9 4.5 4.3 4.4 4.9 4.3 4.0 4.8 4.4 4.1 4.8 4.3 4.0 4.0 4.5 4.1 4.1 4.6 4.5 4.3 4.0 3.9

13.8 13.6 13.7 13.7 13.7 13.6 13.8 13.8 13.9 13.8 13.6 13.8 13.6 13.9 13.7 13.7 13.7 13.7 13.8 13.6 13.5 13.8 13.8 13.6 13.6 13.8 13.6 13.5 13.9 13.8 13.7 13.8 13.6 13.6 13.6 13.8 13.9 13.5 13.8 13.7 13.5 13.6 13.8

57.5 57.5 57.6 57.4 57.4 57.6 57.5 57.4 57.5 57.5 57.5 57.4 57.5 57.5 57.5 57.6 57.5 57.6 57.5 57.5 57.5 57.6 57.7 57.6 57.8 58.3 57.6 57.5 58.2 57.8 57.7 58.2 57.8 57.5 57.4 57.8 57.6 57.8 58.2 57.8 57.7 76.1 78.4

Nikolai Puch, Chi Hieu Ta, and Moritz Beckel

Table 12: Full KEM-primitive resident memory results. Mean per-operation resident memory in MiB for the native, container, and unikernel environments.

Algorithm

Keygen Encaps Decaps Ntv. Cnt. Uk. Ntv. Cnt. Uk. Ntv. Cnt. Uk.

BIKE-L1 BIKE-L3 BIKE-L5 Classic-McEliece-348864 Classic-McEliece-348864f Classic-McEliece-460896 Classic-McEliece-460896f HQC-128 HQC-192 HQC-256 Kyber512 Kyber768 Kyber1024 ML-KEM-512 ML-KEM-768 ML-KEM-1024 FrodoKEM-640-AES FrodoKEM-640-SHAKE FrodoKEM-976-AES FrodoKEM-976-SHAKE FrodoKEM-1344-AES FrodoKEM-1344-SHAKE ECDHE X25519 x25519_Kyber512 x25519_Kyber768 x25519_Kyber1024

3.9 4.0 4.0 4.5 4.5 5.3 5.3 4.0 4.0 4.0 4.1 4.0 4.0 4.0 4.0 4.0 3.9 3.9 3.9 4.0 4.1 4.0 3.9 3.8 4.3 4.3 4.4

13.5 13.5 13.8 13.6 14.0 13.9 13.9 13.6 13.8 13.9 13.8 13.8 13.6 13.7 14.0 13.6 13.9 13.9 13.5 13.6 13.7 13.6 13.5 13.8 13.6 13.7 13.6

57.4 57.4 57.7 58.1 58.1 59.0 58.9 57.4 57.4 57.5 57.4 57.4 57.4 57.4 57.4 57.5 57.5 57.5 57.6 57.5 57.5 57.5 75.9 75.7 75.6 75.5 75.6

3.9 4.0 4.0 4.5 4.5 5.2 5.2 4.1 4.1 4.3 4.1 4.0 4.0 4.0 4.0 4.0 3.9 3.9 4.0 4.0 4.1 4.1 4.1 3.8 4.4 4.3 4.4

13.8 13.2 13.6 13.5 13.9 13.9 13.6 13.8 13.8 13.8 13.7 14.1 13.8 13.6 13.7 13.9 13.7 13.9 13.6 13.5 13.8 13.9 13.8 13.6 13.8 13.5 13.6

57.4 57.4 57.6 58.0 58.0 58.7 58.7 57.4 57.5 57.6 57.4 57.5 57.4 57.4 57.4 57.4 57.5 57.5 57.5 57.5 57.6 57.6 77.0 75.9 75.8 75.7 75.7

4.0 4.0 4.3 4.5 4.5 5.2 5.2 4.1 4.1 4.3 4.1 4.0 4.0 4.0 4.0 4.0 3.9 3.9 4.0 4.0 4.3 4.1 — 3.8 4.4 4.3 4.4

13.7 13.9 13.2 13.8 13.7 13.9 13.7 13.7 13.8 13.6 13.9 13.8 13.5 13.7 13.9 13.6 13.9 13.4 13.7 13.7 13.6 13.8 — 13.6 13.8 13.5 13.8

57.4 57.5 57.8 58.0 58.0 58.7 58.7 57.6 57.5 57.7 57.4 57.5 57.4 57.4 57.4 57.4 57.5 57.5 57.6 57.7 57.6 57.6 — 75.7 75.7 75.8 75.7

PQC in Lightweight Virtualization

Table 13: Full signature-primitive power and energy results above the idle baseline. 𝑃 (W) is the mean power draw during the operation, 𝐸 (mJ) is the energy per operation. Derived from the HMC8043 power trace aligned with per-operation timestamps.

Algorithm

Ntv. P E

Keygen Cnt. P E

Uk. P E

Ntv. P E

Sign Cnt. P E

Uk. P E

Ntv. P E

Verify Cnt. P E

Uk. P E

Dilithium2 0.94 0.149 1.34 0.222 1.35 0.284 1.33 0.616 1.16 0.558 1.18 1.11 1.06 0.161 1.15 0.180 1.28 0.291 Dilithium3 1.32 0.388 1.35 0.415 1.07 0.401 1.34 0.975 1.34 1.03 1.39 2.12 1.36 0.344 1.36 0.359 1.02 0.375 Dilithium5 1.21 0.542 1.36 0.644 1.28 0.729 1.38 1.30 1.35 1.32 1.29 2.41 1.40 0.612 1.36 0.618 1.31 0.785 ML-DSA-44 1.03 0.214 1.29 0.310 1.27 0.268 1.11 1.28 1.33 1.28 1.30 1.24 1.11 0.249 1.35 0.315 1.27 0.292 ML-DSA-65 1.35 0.500 1.07 0.413 1.29 0.484 1.22 1.82 1.09 1.68 1.30 2.02 1.12 0.405 1.31 0.493 1.35 0.495 ML-DSA-87 1.31 0.742 1.14 0.678 1.04 0.587 1.33 2.39 1.09 2.19 1.34 2.50 1.31 0.771 1.11 0.749 1.28 0.767 Falcon-512 1.12 19.48 1.12 20.57 0.89 36.29 1.22 0.771 1.20 0.799 0.84 9.61 1.18 0.108 1.02 0.097 0.91 0.119 Falcon-1024 1.09 53.51 1.11 57.29 1.05 126 1.21 1.56 1.21 1.62 0.89 22.47 1.20 0.213 1.22 0.227 0.98 0.257 Falcon-padded-512 1.23 21.39 1.14 21.10 1.08 45.06 0.97 0.612 1.23 0.812 1.08 12.46 1.19 0.109 1.23 0.115 1.19 0.155 Falcon-padded-1024 0.83 40.74 1.22 62.19 1.05 130 1.01 1.30 1.20 1.63 1.06 26.74 1.30 0.233 1.20 0.222 1.19 0.313 SPHINCS+-SHA2-128f-simple 1.27 6.90 1.03 5.64 0.78 2.54 1.28 163 1.33 172 1.19 89.56 1.06 8.11 1.27 9.69 1.28 5.91 SPHINCS+-SHA2-128s-simple 1.27 513 1.25 503 1.19 283 1.40 3847 1.33 3473 1.23 1931 1.31 4.41 1.15 3.75 1.23 2.28 SPHINCS+-SHA2-192f-simple 1.31 10.71 1.34 11.00 1.27 6.38 1.39 298 1.32 287 1.27 169 1.42 16.35 1.26 14.70 1.30 9.08 SPHINCS+-SHA2-192s-simple 1.28 777 1.04 625 1.24 469 1.31 6297 1.31 6119 1.29 3772 1.04 5.88 1.29 7.63 1.02 3.37 SPHINCS+-SHA2-256f-simple 1.24 29.46 1.23 27.47 1.48 19.81 1.29 613 1.26 553 1.44 389 1.26 16.23 1.30 15.65 1.35 9.92 SPHINCS+-SHA2-256s-simple 1.26 600 1.27 559 1.22 316 1.32 6284 1.33 5750 1.31 3372 1.31 12.31 1.26 10.58 1.29 5.68 SPHINCS+-SHAKE-128f-simple 1.27 6.98 1.36 7.76 1.43 7.85 1.28 167 1.47 201 1.41 185 1.29 10.18 1.48 12.14 1.27 9.45 SPHINCS+-SHAKE-128s-simple 1.31 542 1.28 542 1.14 450 1.34 3684 1.34 3676 1.27 3302 1.06 3.62 1.36 4.54 1.36 4.43 SPHINCS+-SHAKE-192f-simple 1.10 9.06 1.27 10.78 1.25 10.18 1.22 260 1.24 286 1.26 261 1.22 13.93 1.26 14.66 1.25 13.89 SPHINCS+-SHAKE-192s-simple 1.19 735 1.05 662 1.31 787 1.04 4957 1.35 6454 1.32 6060 1.12 6.59 1.29 8.14 1.24 6.80 SPHINCS+-SHAKE-256f-simple 1.30 30.08 1.18 27.96 1.25 27.65 1.34 614 1.28 595 1.27 548 1.22 15.38 1.27 16.78 1.27 14.64 SPHINCS+-SHAKE-256s-simple 1.30 541 1.08 473 1.00 426 1.04 4296 1.33 5492 1.34 5558 1.17 9.90 1.39 11.46 0.95 7.71 MAYO-1 1.28 1.38 1.28 1.43 0.93 1.61 1.22 2.90 1.19 2.96 1.01 2.94 1.46 1.08 1.46 1.12 1.09 1.61 MAYO-2 1.39 3.29 1.22 3.02 1.19 3.68 1.18 3.91 1.16 4.08 1.18 4.69 1.45 1.15 1.42 1.18 1.03 1.84 MAYO-3 1.20 4.73 1.25 5.17 1.19 7.35 1.38 12.30 1.16 10.89 1.14 12.27 1.66 4.27 1.45 3.91 1.25 6.50 MAYO-5 1.25 12.91 1.41 15.22 1.18 18.84 1.09 25.30 1.37 33.02 1.18 32.96 1.69 10.98 1.44 9.83 1.25 15.97 cross-rsdp-128-balanced 1.20 0.068 1.36 0.081 1.32 0.075 1.28 4.23 1.49 5.16 1.25 4.38 1.24 2.39 1.42 2.86 1.20 2.43 cross-rsdp-128-fast 1.18 0.068 1.18 0.077 1.13 0.064 1.28 2.33 1.22 2.58 1.17 2.24 1.27 1.32 1.10 1.21 1.48 1.62 cross-rsdp-128-small 1.20 0.069 1.20 0.072 1.19 0.067 1.31 16.42 1.31 17.17 1.13 14.90 1.25 9.28 0.96 7.44 1.40 10.70 cross-rsdp-192-balanced 1.45 0.177 1.22 0.156 1.17 0.142 1.50 11.79 1.30 10.74 1.30 10.38 1.24 5.36 1.25 5.61 1.24 5.27 cross-rsdp-192-fast 1.04 0.126 1.39 0.178 1.18 0.143 1.51 6.71 1.29 5.97 1.27 5.72 1.50 3.85 1.30 3.43 1.26 3.22 cross-rsdp-256-balanced 1.21 0.258 1.32 0.294 1.24 0.259 1.26 19.58 1.42 22.56 1.20 19.64 1.47 10.79 1.42 10.64 1.21 8.94 cross-rsdp-256-fast 1.21 0.258 1.21 0.272 1.38 0.288 1.25 11.62 1.50 14.45 1.39 13.73 1.23 6.40 1.46 7.83 1.20 6.72 cross-rsdpg-128-balanced 1.21 0.035 1.20 0.037 1.37 0.041 1.28 3.26 1.27 3.40 1.41 3.70 1.26 1.97 1.32 2.14 1.39 2.28 cross-rsdpg-128-fast 1.38 0.040 1.21 0.037 1.21 0.036 1.29 1.68 1.28 1.73 1.24 1.70 1.29 1.02 1.29 1.06 1.48 1.21 cross-rsdpg-128-small 1.38 0.040 0.91 0.028 1.21 0.036 1.42 12.93 1.28 12.20 1.27 12.07 1.17 6.44 1.26 7.32 1.26 7.26 cross-rsdpg-192-balanced 1.24 0.066 1.06 0.059 1.19 0.067 1.43 5.43 1.15 4.57 1.24 4.85 1.42 3.31 0.97 2.38 1.23 2.92 cross-rsdpg-192-fast 1.22 0.065 1.39 0.078 1.08 0.061 1.26 3.76 1.43 4.37 0.96 2.89 1.26 2.35 1.40 2.76 1.27 2.39 cross-rsdpg-192-small 1.24 0.066 1.22 0.068 1.35 0.077 1.30 18.19 1.27 18.70 1.15 16.13 1.27 11.13 1.44 13.17 1.09 9.75 cross-rsdpg-256-balanced 0.95 0.077 1.25 0.106 1.21 0.108 1.30 8.52 1.30 8.94 1.47 9.84 1.30 4.99 1.32 5.33 1.09 4.39 cross-rsdpg-256-fast 1.10 0.089 1.22 0.104 1.18 0.106 1.08 5.43 1.28 6.76 1.25 6.60 1.26 3.95 1.26 4.20 1.25 4.11 ECDSA 1.10 0.047 1.29 0.063 1.20 0.050 0.93 0.089 1.12 0.114 1.09 0.107 0.89 0.260 1.06 0.323 0.95 0.281 RSA-2048 0.99 363 0.73 261 1.01 349 0.98 4.98 0.68 3.65 1.01 5.20 0.75 0.102 0.69 0.101 0.97 0.133

Nikolai Puch, Chi Hieu Ta, and Moritz Beckel

Table 14: Full KEM-primitive power and energy results (Structure as in Table 13).

Algorithm

Ntv. P E

Keygen Cnt. P E

Uk. P E

Ntv. P E

Encaps Cnt. P E

Uk. P E

Ntv. P E

Decaps Cnt. P E

Uk. P E

BIKE-L1 1.42 61.29 1.55 69.78 1.38 76.82 1.39 3.04 1.19 2.72 0.84 2.40 1.39 48.92 0.92 33.34 1.17 52.05 BIKE-L3 1.46 195 1.41 199 1.38 239 1.39 9.40 1.10 7.55 1.28 11.25 1.39 156 1.57 181 1.00 140 BIKE-L5 1.44 486 1.44 497 1.00 434 1.40 23.93 1.39 24.73 1.21 27.22 1.46 411 1.41 407 0.95 332 Classic-McEliece-348864 1.23 397 1.55 548 1.67 562 1.31 0.220 1.35 0.242 1.14 0.297 1.05 60.91 1.07 64.67 0.97 54.76 Classic-McEliece-348864f 1.29 207 1.39 234 1.42 223 1.01 0.171 1.34 0.238 1.29 0.334 1.06 60.86 1.05 63.43 1.24 70.16 Classic-McEliece-460896 1.64 2279 1.56 2096 1.33 1344 1.34 0.520 1.41 0.537 1.28 0.749 1.10 99.90 0.82 75.68 1.04 96.49 Classic-McEliece-460896f 1.45 772 1.48 805 1.51 842 1.32 0.474 1.54 0.580 1.28 0.756 1.06 92.27 1.32 121 1.04 96.49 HQC-128 1.42 7.37 1.44 7.84 1.00 6.85 1.42 14.86 1.54 16.91 0.97 13.18 1.42 23.60 1.25 21.24 0.84 17.28 HQC-192 1.44 22.99 1.45 23.94 1.29 27.43 1.45 46.18 1.41 47.22 1.23 51.37 1.44 70.29 1.41 72.25 1.01 63.03 HQC-256 1.28 37.32 1.41 43.25 1.14 44.74 1.12 65.75 1.41 86.66 1.12 85.38 1.41 128 1.43 136 1.28 148 Kyber512 1.43 0.067 1.00 0.049 1.24 0.108 1.17 0.061 1.29 0.072 1.26 0.132 1.17 0.050 1.32 0.060 1.27 0.151 Kyber768 1.32 0.089 1.17 0.081 1.24 0.171 1.53 0.121 1.11 0.093 1.25 0.208 0.91 0.063 1.29 0.095 1.26 0.232 Kyber1024 1.34 0.134 1.56 0.162 1.07 0.222 1.09 0.126 1.22 0.143 0.95 0.230 1.36 0.142 1.19 0.127 1.28 0.344 ML-KEM-512 0.97 0.077 1.26 0.105 1.11 0.097 1.29 0.122 1.29 0.129 1.09 0.106 1.29 0.149 1.44 0.214 1.08 0.128 ML-KEM-768 1.31 0.170 1.21 0.199 1.24 0.172 1.31 0.199 1.35 0.210 1.39 0.215 1.32 0.236 1.29 0.241 1.17 0.214 ML-KEM-1024 1.12 0.223 1.28 0.264 1.27 0.265 1.10 0.244 1.27 0.297 1.26 0.289 1.21 0.314 1.30 0.357 1.30 0.347 FrodoKEM-640-AES 1.20 20.65 0.87 15.57 1.14 18.26 1.01 17.62 1.15 20.84 1.20 19.21 0.94 16.11 1.23 22.58 1.20 19.30 FrodoKEM-640-SHAKE 1.31 7.58 1.04 6.34 1.22 7.14 1.33 8.76 1.14 7.76 1.34 8.88 1.44 9.36 1.09 7.32 1.29 8.47 FrodoKEM-976-AES 1.16 45.71 1.15 47.56 1.05 38.33 1.20 48.08 1.32 55.13 1.33 49.07 1.19 47.22 0.99 40.04 1.19 43.97 FrodoKEM-976-SHAKE 1.28 16.78 1.30 17.93 1.45 18.90 1.32 19.32 1.32 20.09 1.13 16.41 1.35 20.17 1.31 19.73 1.08 15.71 FrodoKEM-1344-AES 1.00 74.23 1.17 91.27 1.18 81.16 1.23 93.10 1.21 96.17 1.22 85.65 1.20 91.28 1.21 95.38 1.36 95.38 FrodoKEM-1344-SHAKE 1.16 27.77 1.34 33.97 1.24 29.56 1.10 29.61 1.37 39.01 1.27 34.09 1.37 36.20 1.35 37.30 1.31 35.11 ECDHE 1.19 0.075 1.02 0.066 1.21 0.076 1.23 0.300 0.92 0.230 1.05 0.266 — — — — — — X25519 1.22 0.102 1.24 0.108 1.27 0.095 1.02 0.261 1.00 0.260 1.00 0.244 — — — — — — x25519_Kyber512 1.08 0.142 1.35 0.187 1.27 0.212 0.91 0.280 1.18 0.381 1.09 0.387 1.17 0.345 1.05 0.324 1.08 0.396 x25519_Kyber768 1.28 0.197 1.39 0.223 1.25 0.272 1.09 0.362 1.18 0.411 1.09 0.449 1.13 0.365 1.17 0.396 1.11 0.480 x25519_Kyber1024 1.28 0.238 1.29 0.249 1.13 0.322 1.09 0.404 1.22 0.469 1.26 0.615 1.09 0.394 0.94 0.348 1.14 0.584

Record · ID 1028605 · SHA-256 3f67cd533f117b28
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.