Conceptio › Archive › arXiv CS
arXiv CSopen access

Efficient TCitH-Based Alternatives to SLH-DSA: Cross-Layer ASIC Design of Mirath

Hiandra Tomasi et al. · arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

Efficient TCitH-Based Alternatives to SLH-DSA: Cross-Layer ASIC Design of Mirath Hiandra Tomasi, Maximilian Schöffel, Johannes Feldmann, and Norbert Wehn Microelectronic Systems Design Research Group RPTU Kaiserslautern-Landau Kaiserslautern, Germany {tomasi, m.schoeffel, j.feldmann, norbert.wehn}@rptu.de

arXiv:2609.35053v1 [cs.CR] 28 Sep 2026

Abstract—To address the security risks posed by quantum computers, the U.S. National Institute of Standards and Technology (NIST) has standardized the post-quantum signature schemes ML-DSA, FN-DSA, and SLH-DSA. While ML-DSA and FN-DSA are lattice-based, SLH-DSA relies on hash-based assumptions. To support cryptographic agility against future vulnerabilities, NIST is evaluating non-lattice candidates as alternatives to SLH-DSA. Among these, TCitH-based schemes are particularly promising due to their compact keys and small signatures. However, their high computational complexity and memory footprint pose significant challenges for efficient implementations on resource-constrained embedded platforms. They remain largely unexplored in this context, particularly in ASIC implementations. To address this gap, we use a cross-layer methodology combining algorithmic and hardware layers to present, to the best of our knowledge, the first ASIC implementation of Mirath, a TCitH-based signature scheme, in a RISC-V-based system. The design is implemented in a 22 nm FD-SOI technology node. Compared with an SLH-DSA ASIC implemented in the same technology node, the proposed architecture achieves 17.8× lower signing latency while requiring 58 % less total cell area, showing the potential of TCitH-based signatures as efficient non-lattice alternatives from an implementation perspective. Index Terms—TCitH, MPCitH, PQC, Mirath, RISC-V, ASIC

I. I NTRODUCTION The threat of quantum computers (QCs) [1] has accelerated the transition to Post-Quantum Cryptography (PQC). The National Institute of Standards and Technology (NIST) selected ML-DSA, FN-DSA, and SLH-DSA for standardization [2]– [4] and subsequently launched an additional digital-signature standardization process [5] to strengthen crypto-agility, i.e., the ability to replace algorithms if vulnerabilities emerge. To justify adoption, non-lattice candidates are required to provide a significant performance advantage over SPHINCS+, the scheme underlying SLH-DSA [5]. Multi-Party Computation-in-the-Head (MPCitH) signatures [6] are promising candidates in this process. In particular, the Threshold-Computation-in-the-Head (TCitH) framework [7], adopted by Mirath [8], RYDE [9], and MQOM [10], significantly reduces signature sizes relative to earlier corresponding MPCitH-based designs. These schemes combine small public keys and compact signatures with security assumptions distinct from those of the standardized schemes. This work was partly funded by the German Federal Ministry of Research, Technology and Space as part of the project “PoQ-KiKi” under grant number 16KIS2064.

However, their reference implementations are computationally and memory intensive, leading to high latency, memory usage, and energy consumption on resource-constrained devices. Efficient implementations of MPCitH-based schemes remain comparatively under-explored, and some candidates were eliminated before optimized implementation baselines were established [11]. The state-of-the-art (SoA) consists mainly of software optimizations for off-the-shelf embedded devices [12]–[14] and FPGA implementations [15]–[17], while ASIC implementations of schemes from this process remain largely unexplored. In contrast, ASIC implementations of standardized PQC schemes are well documented [18]–[22]. In this context, hardware/software (HW/SW) co-design based on the RISC-V Instruction Set Architecture (ISA) [23] is an established approach for embedded PQC [18]–[22], combining software programmability with dedicated acceleration of computational bottlenecks. Following this approach, we present a RISC-V-based HW/SW co-design of Mirath to challenge the perception that MPCitH-based signatures are inherently inefficient [24]. We use Mirath as a representative TCitH-based scheme to investigate their computational and memory bottlenecks. The main contributions of this work are: 1) We present, to the best of our knowledge, the first ASIC implementation of a TCitH-based signature scheme. Our design is implemented in 22 nm FD-SOI technology. 2) We profile Mirath to identify computational and memory bottlenecks, develop dedicated accelerators for dominant computational kernels, and perform a HW/SW designspace exploration to quantify the individual and combined impact of the proposed accelerators on execution time. 3) From a crypto-agility perspective, we demonstrate TCitHbased signatures as an efficient non-lattice alternative, achieving lower signing latency than all considered SoA SPHINCS+/SLH-DSA ASIC implementations and execution times comparable to SoA Dilithium/ML-DSA and Falcon/FN-DSA implementations. This paper is organized as follows: Section II provides the theoretical background. Section III describes the design objectives and cross-layer methodology. Section IV details the algorithmic optimizations, while Section V presents the hardware implementation. Experimental results and comparisons

© 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.

to SoA are provided in Sections VI and VII, respectively, followed by the conclusion in Section VIII.

III. D ESIGN O BJECTIVES Throughout this work, we consider the Mirath-1a-fast parameter set and use Mirath’s reference C implementation (denoted as PRef in the following) using test vectors with 33byte message size as baseline. Our hardware platform is based on our proprietary RISC-V core that implements the RV64I with the M, C, Zba, Zbb, Zicsr, and Zifencei extensions and uses a three-stage pipeline, an instruction-prefetch depth of four, and support for four outstanding read requests.

Verify

K1: Sample and Expand seeds Sample and expand seedsk , seedpk to obtain S, C ′ and H ′ .

S1: Commit witnesses Derive randomized witness shares, and compute commitments with digest hcom .

V1: Reconstruct Parse σ, recover the opened shares, rebuild commitments, and recompute hsh .

K2: Compute syndrome y Build E = (S | SC ′ ), set e = vec(E), and compute y.

S2: Compute proof Derive hsh and Γ, compute the polynomial-proof (e) (e) values αmid , αbase , and hpiop .

K3: Output keys pk = (seedpk , y) sk = (seedsk , seedpk )

pk, msg, σ

Sign

pk, sk

II. BACKGROUND A digital signature algorithm (DSA) consists of the operations KeyGen, Sign, and Verify. MPC-in-the-Head (MPCitH) signatures locally simulate an MPC protocol among virtual parties and commit to their views. Threshold Computation-in-the-Head (TCitH) replaces additive secret sharing with packed Shamir secret sharing [25] to reduce the opening data. Mirath is a TCitH-based scheme. Although it did not advance to the third round of the NIST process due to similarities with other candidates, its security claims remain unaffected. Fig. 1 summarizes its computation flow. Mirath is based on the MinRank Syndrome Problem [26]. Let q = ps be a prime power, with p prime and s ∈ N>0 , and let m, n, k, r, µ, ρ, τ, N, λ ∈ N>0 denote the MinRank dimensions, extension degree, number of parallel polynomial checks, protocol repetitions, virtual parties per repetition, and seed security parameter, respectively. Given H = [I mn−k ∥ (mn−k)×mn (mn−k)×k H ′ ] ∈ Fq , H ′ ∈ Fq , and y ∈ Fmn−k , q m×n the witness is E ∈ Fq satisfying H vec(E) = y and rank(E) ≤ r. It is represented as E = S[I r ∥ C ′ ], with r×(n−r) . S ∈ Fm×r and C ′ ∈ Fq q During KeyGen, seedsk , seedpk ∈ {0, 1}λ are expanded to derive (S, C ′ ) and H ′ , from which y is computed. During Sign, for e ∈ [1, τ ] and i ∈ [1, N ], a Goldreich–Goldwasser–Micali (GGM) tree derives leaf seeds (e) (e) seedi ∈ {0, 1}λ . Each seed is committed as comi , with all commitments bound by the digest hcom , and expanded (e) ′(e) r×(n−r) (e) , and v rnd,i ∈ Fρ×1 , C rnd,i ∈ Fq into S rnd,i ∈ Fm×r qµ . q ρ×(mn−k) , these values are combined into the Using Γ ∈ Fqµ (e) (e) polynomial-proof values αmid , αbase ∈ Fρ×1 q µ . The BAVC opening πBAVC reveals sufficient seed-tree information to reconstruct all non-challenged parties while keeping the challenged leaf seeds hidden. Verify reconstructs these parties and recomputes the corresponding commitments, shares, and proof values. At the implementation level, the block cipher Advanced Encryption Standard (AES) [27] is used for GGM-tree expansion, pseudorandom-share generation, and, in this work, leafseed commitment generation. Federal Information Processing Standard (FIPS) 202 [28], instantiated with the Secure Hash Algorithm 3 SHA3-256 hash function and the SHAKE128 extendable-output function (XOF), is used for matrix and seed expansion, transcript hashing and challenge generation.

KeyGen

S3: Open and output Compute πBAVC and output the signature σ.

V2: Recompute proof Derive Γ and recompute the evaluation/base proof terms and h′piop .

V3: Accept / reject Accept iff h′piop = hpiop and vgrinding = 0.

Fig. 1. Overview of Mirath’s KeyGen, Sign, and Verify operations. TABLE I L ATENCY AND MEMORY COMPARISON OF THE REFERENCE AND ALGORITHMICALLY OPTIMIZED IMPLEMENTATIONS AT 500 MH Z . Impl.

Cycles [M]

Latency [s]

Data [kB]

Code [kB]

Sign

PRef P0 (Alg. Opt.)

14,471.64 54,019.66

28.94 108.04

377 35

28 27

Verify

PRef P0 (Alg. Opt.)

21,476.88 66,321.67

42.95 132.64

380 32

27 25

Function

The PRef results in Table I reveal two main limitations. First, Sign and Verify require substantial data memory, restricting their applicability to resource-constrained embedded systems. Second, their long execution times result in a prohibitive latency for general applications. Throughout this work, latency denotes the computation time of one cryptographic operation. These observations define the two main design objectives of this work: (i) to reduce the memory requirements of Mirath through algorithmic optimization, and (ii) to reduce the resulting execution latency by accelerating the dominant computational bottlenecks. We therefore adopt a cross-layer optimization strategy consisting of an algorithmic layer for memory reduction, followed by an implementation layer based on HW/SW co-design for hardware acceleration. IV. A LGORITHMIC O PTIMIZATION The main memory bottleneck is the pre-computation of round inputs. During Sign, the complete GGM tree and leaf commitments each require approximately 139 kB, while duplicating the 4352 leaf seeds for hcom adds 70 kB to the 377 kB data-memory footprint. We instead derive leaves and commitments on demand from the root along its corresponding path (Fig. 2), immediately use and discard them, and incrementally absorb commitments into hcom . Verify applies the

Level 0 1 node

0 0

Level 1 2 nodes

2

0

Level 2 4 nodes

Level 11 2048 nodes

Level 12 256 internal + 3840 leaves

Level 13 512 leaves

1

1

1

3

4

5

6

levels 3–10 .. . branch bits omitted

.. .

.. .

.. .

···

2047

0

1

4095

4096

0

1

8191

0

1

···

0

1

0

4349

4350

4351

0

8192

···

8701

···

2175

2174

0

1

4352

4094

···

8189

1

8190

1

8702

Fig. 2. Compact layout of the Mirath-1a-fast GGM seed tree. The 4352 leaves do not form a perfect binary tree and therefore span two depths. The orange nodes and arrows show one exact root-to-leaf path, ending at node 8702. At each internal node, the 16-byte parent seed is used as the AES-128 key. Each branch uses a distinct 16-byte plaintext block derived from the salt, node index, and branch bit, and the 16-byte ciphertext becomes the child seed. TABLE II P ROFILING RESULTS FOR K E Y G E N , S I G N , AND V E R I F Y . VALUES ARE REPORTED AS PERCENTAGES OF THE TOTAL EXECUTION TIME . Computational Kernel AES FIPS 202 GF Arithmetic Others

KeyGen (%)

Sign (%)

Verify (%)

0.00 1.59 98.35 0.06

98.65 0.01 1.32 0.03

99.36 0.01 0.62 0.02

Fig. 3. RISC-V architecture and TCitH accelerator unit after memory optimization for the P0 –P12 configurations.

accelerators to reduce data movement. A memory-mapped interface is used instead of custom RISC-V instructions because the targeted kernels involve multiple wide operands. The accelerator modules implement progressively coarser offload granularities. While the standalone AES module processes one block per software request and the GGM walker derives a single requested tree node, the Accumulator evaluates all parties of one protocol repetition, and the MPC-Emulation module executes complete algebraic kernels. A. AES IP

same strategy to visible leaves reconstructed from the sibling seeds in πBAVC . This introduces a memory–computation trade-off: the optimization replaces the full tree and commitment arrays with one 224 B path and one 32 B commitment, reducing Sign/Verify data memory by approximately 90 %. Repeated tree traversals, however, increase AES evaluations and latency by up to 273 %, motivating hardware acceleration. The comparison between PRef and the algorithmically optimized version, denoted as P0 , is shown in Table I. V. HW/SW C O -D ESIGN We profiled the optimized software implementation to identify computational bottlenecks and guide HW/SW partitioning. As shown in Table II, KeyGen is dominated by GF arithmetic, whereas Sign and Verify are dominated by AES operations associated with GGM-tree traversal and commitment generation. Based on these results, we developed a suite of hardware accelerators. The overall system architecture, comprising the processor and the TCitH accelerator, is illustrated in Fig. 3. The processor communicates with the accelerator through an AXI4-Lite interface. Local storage is provided by two 8 kB memories to hold operands, intermediate values, and results. Software controls data transfers, accelerator configuration, execution, and result retrieval. Intermediate values are retained locally and, where possible, transferred directly between

The system integrates an open-source AES IP core [29]. Although the core supports AES-128/256 encryption and decryption, only AES-128 encryption is utilized for the targeted parameter set. The remaining modes are preserved to maintain versatility for broader application contexts. The AES core is unused during KeyGen. During Sign and Verify, the core accelerates GGM seed derivation, leaf-seed commitments, and pseudorandom-share generation. B. GGM Tree Accelerator In the GGM Tree module, the AES IP is scheduled internally, eliminating the need for one processor command per AES block. The accelerator receives a root seed, salt, and node or leaf index, determines the corresponding left/right directions of the path, and applies the AES-based child derivation function only along the selected path. It neither generates both children nor materializes the complete tree. The local memories retain the preceding tree path, allowing consecutive traversals with a common prefix to resume from the first differing level. This accelerator is not used during KeyGen. During Sign, it performs the root-to-leaf traversals and derives the siblingnode and hidden-party seeds required for πBAVC . During Verify, reconstruction starts from the sibling-path seeds disclosed in the signature rather than from the original root. This mode is integrated only when the Accumulator is used.

Without it, software traverses from the disclosed seeds by invoking the AES IP for the required tree edges. C. FIPS 202 Accelerator The FIPS 202 hardware module processes aligned 64-bit portions of the absorb and squeeze operations and executes the K ECCAK permutations. Therefore, software retains control of the incremental context, unaligned input and output fragments, padding, byte extraction, and the protocol-specific construction of each hash input. Additionally, when combined with the Accumulator, 32-byte party commitments are transferred directly to the K ECCAK input through a 64-bit ready/valid interface. This avoids returning every commitment to the processor and subsequently writing it back to the hash accelerator. The FIPS 202 accelerator is used for three classes of operations: expansion of public and secret matrices; computation of commitment and transcript digests; and derivation of the protocol challenges. D. Accumulator The Accumulator is always combined with the GGM Tree module. It is not used during KeyGen. During Sign, software loads the salt, root seed, and secret matrices and issues one command per repetition e. The accelerator derives the party seeds and commitments, expands the shares, and computes the accumulations, party-weighted bases, and auxiliary values. AES seed generation and share processing are overlapped such that one AES block can be accumulated while the next is generated. Extension-field arithmetic uses eight parallel byte-wide F28 multipliers, processing one 64-bit word per arithmetic step. Base-field elements in F16 use the same embedded representation as PRef . The resulting auxiliary and base values remain in the local memories and can be consumed directly by the MPC-Emulation accelerator when it is also enabled. During Verify, software parses πBAVC into a path table containing the disclosed sibling seeds and the leaf ranges covered by their subtrees. For each visible party, the accelerator selects the corresponding entry and completes the remaining GGM traversal from the disclosed seed. It then reconstructs the visible parties, expands their shares, and regenerates their commitments. For the hidden parties, the auxiliary values and commitment are obtained from the signature. When direct commitment streaming is enabled, the hidden commitments are inserted at the corresponding position in the stream to the FIPS 202 accelerator. E. MPC-Emulation Accelerator The MPC-Emulation accelerator performs the algebraic proof computations after share accumulation. It is implemented as a multi-phase state machine and reuses the same eight F28 multiplier lanes as the Accumulator. An early-exit mode reuses the matrix-product datapath to compute y. This mode is used for syndrome computation during KeyGen and in Sign during a reconstruction of y as an implementationspecific prerequisite for hpiop . The full mode computes the

TABLE III HW/SW PARTITIONS AND THEIR CORRESPONDING HARDWARE ACCELERATORS . Partition PRef P0 P1 P2 P3 P4 P5 P6 P7 P8 P9 P10 P11 P12

Hardware Accelerators None (Reference Implementation) None (Algorithmically Optimized Software) AES IP FIPS 202 MPC-Emulation AES IP, GGM Tree FIPS 202, MPC-Emulation AES IP, GGM Tree, FIPS 202 AES IP, GGM Tree, Accumulator AES IP, GGM Tree, MPC-Emulation AES IP, GGM Tree, FIPS 202, Accumulator AES IP, GGM Tree, FIPS 202, MPC-Emulation AES IP, GGM Tree, Accumulator, MPC-Emulation AES IP, GGM Tree, FIPS 202, Accumulator, MPC-Emulation

polynomial-proof values during Sign and Verify. It is not used during KeyGen. F. Partitionings The previously described accelerator functionalities can be selectively implemented in software or hardware, resulting in multiple HW/SW partitioning configurations. This work considers 14 distinct partitions, shown in Table III. Since KeyGen was not algorithmically optimized, PRef and P0 are the same, and only P2 , P3 , and P5 are applicable. For Sign, all partitions are supported, whereas for Verify, all partitions except P4 are applicable, as discussed in Section V-B. Subsequent references to P12 also include KeyGen, which uses only its supported accelerators. VI. R ESULTS The ASIC results are based on the GlobalFoundries 22 nm FD-SOI technology. Timing is evaluated under worst-case PVT conditions at 125 ◦ C and 0.72 V, whereas power is evaluated under nominal conditions at 25 ◦ C and 0.8 V. The design flow employed Synopsys DesignCompiler, IC-Compiler, and the INVECAS Memory Compiler, with power calculated via back-annotated wiring data. A. Impact of the different partitions on the overall runtime The impact of the HW/SW partitions on execution time and memory requirements is summarized in Table IV for Sign and Verify. For KeyGen, latency decreases from 3.07 ms in P0 to 2.67 ms, 0.60 ms, and 0.20 ms in P2 , P3 , and P5 , respectively. All partitions were validated against PRef . The results show that the effectiveness of each accelerator strongly depends on the targeted operation and on the remaining software bottlenecks. For Sign and Verify, accelerating FIPS 202 or MPCEmulation alone provides negligible benefit, leaving the system still 197–273 % slower than PRef . In contrast, introducing the standalone AES IP with P1 reduces the latency of P0 by 99.61 % and 99.70 % for Sign and Verify, respectively,

TABLE IV ASIC S IGNING AND VERIFICATION LATENCY AND NUMBER OF CLOCK CYCLES OF M IRATH FOR THE CONSIDERED HW/SW PARTITIONS . Sign

Area (µm2 )

Area (%)

RISC-V Interconnect Memories ⊢ Main Memories ⊢ Accelerator Memories TCitH Accelerator ⊢ Control Logic ⊢ FIPS 202 ⊢ Combined Modules ⊢ GGM Tree ⊢ AES IP ⊢ Accumulator ⊢ MPC-Emulation Others

10,741.25 6,186.55 158,841.05 119,712.39 39,128.66 42,686.92 2,513.97 9,795.44 30,377.52 8,486.47 11,172.56 7,040.92 3,677.57 16,635.54

4.57 2.63 67.57 50.92 16.64 18.16 1.07 4.17 12.92 3.61 4.75 3.00 1.56 7.08

P12

235,091.33

100.00

Verify

Latency Cycles Latency (ms) (M) (ms) 28,943.28 21,476.88 42,953.76 108,039.32 66,321.67 132,643.34 One Accelerator P1 212.56 425.12 200.35 400.69 P2 54,017.33 108,034.66 63,824.19 127,648.37 P3 51,912.17 103,824.33 66,306.51 132,613.01 Two Accelerators P4 165.58 331.16 Same as P1 ‡ P5 51,909.78 103,819.56 63,750.99 127,501.97 Three Accelerators P6 163.20 326.39 194.35 388.70 P7 79.59 159.18 44.10 88.20 P8 86.39 172.77 155.18 310.36 Four Accelerators P9 72.44 144.87 34.70 69.40 P10 84.05 168.10 148.96 297.92 P11 8.65 17.30 10.80 21.60 All Accelerators P12 1.50 2.99 1.39 2.78 ‡ For Verify, the GGM Tree accelerator is inactive without the Accumulator, so P4 reduces to P1. Impl. PRef P0

TABLE V ASIC AREA BREAKDOWN .

Cycles (M) 14,471.64 54,019.66

AES in area, its impact is minor in intermediate partitions. However, once the other dominant workloads are accelerated, FIPS 202 becomes the primary remaining bottleneck. Adding it to P11 to obtain the fully accelerated P12 reduces signing and verification cycles by a further 82.7 % and 87.1 %, respectively. A similar dependency is observed for KeyGen. B. Area, Power, and Energy

Fig. 4. Layout of the fully accelerated P12 ASIC implementation. TCitH accelerator and local memories in green, RISC-V in blue, memory interconnect in yellow, main memories in red, and others in purple.

confirming AES as the dominant bottleneck. This improvement is achieved at low hardware cost, with the AES IP accounting for only 4.75 % of the total P12 cell area. Adding the GGM Tree accelerator in P4 further reduces signing latency by approximately 22.1 % relative to P1 . Although its marginal gain is smaller than that of the standalone AES IP, it removes another major component of the seedtree computation and enables larger improvements from the Accumulator and MPC-Emulation accelerators. Additionally, the results show that the FIPS 202 module becomes relevant only after the dominant bottlenecks have been addressed. Although it is the second-largest accelerator after

For the implementation, we considered the two representative endpoints of the design space: PRef and the fully accelerated P12 configuration. This enables a direct comparison of the area and energy characteristics of the proposed architecture P12 relative to PRef . To run PRef , the memory subsystem requires one 32 kB SRAM macro for code and six 64 kB SRAM macros for data. The resulting ASIC occupies a core area of 0.61 mm2 with a utilization of 55.74 %, an aspect ratio of 2.02, a total cell area of 0.51 mm2 , and a maximum frequency of 630 MHz. Given the memory reduction that resulted from the algorithmic optimization, the required SRAM macros were significantly reduced from six to one 64 kB macro for P12 . The ASIC occupies a core area of 0.33 mm2 with a utilization of 45.92 %, an aspect ratio of 1.0, a total cell area of 0.24 mm2 , and a maximum frequency of 650 MHz, limited by the main memory access time (Fig. 4). The breakdown of the ASIC area is shown in Table V, where only the total cell area is considered. Both designs were implemented and simulated with a 500 MHz operating frequency. Beyond reducing latency, the proposed P12 implementation also substantially reduces energy consumption. As summarized in Table VII, P12 exhibits a slightly higher power consumption than PRef , which can be attributed to the additional accelerator logic despite the reduced memory footprint. This increase, however, is outweighed by the substantial reduction in execution time, particularly for Sign and Verify. Consequently, the energy consumption of KeyGen, Sign,

TABLE VI C OMPARISON OF OUR WORK AND RELATED HW/SW C O -D ESIGN BASED ASIC IMPLEMENTATIONS OF NIST STANDARDIZED DSA S . A LL IMPLEMENTATIONS CORRESPOND TO THE SMALLEST PARAMETER SET OF THE RESPECTIVE ALGORITHMS . N . R . DENOTES NOT REPORTED . Work

Technology [nm]

Area

Freq. [MHz]

KeyGen

Sign

Mcycles

Lat. [ms]

0.10

Verify

Mcycles

Lat. [ms]

Mcycles

Lat. [ms]

0.20

1.50

2.99

1.39

2.78

2.16 0.71 11.91

42.60 4.90 45.34

53.25 19.62 283.34

2.46 0.44 2.93

3.07 1.76 18.31

0.74 0.27 2.86

1.91 0.91 1.54

2.38 0.72 9.61

0.65 0.36 0.59

0.81 0.29 3.70

65 n.r. 115.80 160 117.79 736.16 48.61 303.81 *: kGE values follow the memory accounting of the respective works. ◦: only accelerator area reported (excludes processor). †: post-place-and-route data (post-synthesis for the others). ⋄: reported area does not include the full memory.

0.26

1.61

[mm2 ]

[kGE]*

This work (P12 )†

22

0.24

381.86

Karl et al. [18] Saarinen [19] Dolmeta et al. [20]◦

22 45 65

0.56 n.r. n.r.

403.30 73.08 115.80

500

SPHINCS+/SLH-DSA 800 250 160

1.73 0.18 1.91

Dilithium/ML-DSA Karl et al. [21]†

22 22 65

Carril et al. [22]⋄ Dolmeta et al. [20]◦

0.46 2.74 n.r.

244.00 13712.00 115.80

800 1200 160

0.59 0.34 0.46

Falcon Dolmeta et al. [20]◦

TABLE VII P OWER AND ENERGY CONSUMPTION OF THE PROPOSED M IRATH IMPLEMENTATIONS AND S OA COMPARISON . N . R . DENOTES NOT REPORTED . Implementation

This work (PRef ) This work (P12 ) Karl et al. [18]

KeyGen

Sign

Verify

Power [mW]

Energy [µJ]

Power [mW]

Energy [µJ]

Power [mW]

Energy [µJ]

16.50 18.20 n.r.

50.66 3.60 n.r.

16.50 26.90 n.r.

477,560 80.54 n.r.

17.50 25.60 48.66

751,695 71.12 9,567

and Verify is reduced by 14×, 5,929×, and 10,569×, respectively. VII. C OMPARISON TO S OA Table VI compares P12 with SoA ASIC implementations of standardized NIST DSAs and their pre-standardization versions. Direct area comparisons are limited by differences in technology nodes and implementation stages. The PVT conditions of the compared works are not reported. As shown, our work P12 achieves 6.6× to 95× lower signing latencies compared with the SPHINCS+/SLH-DSA implementations in [18]–[20]. Furthermore, our design also reduces the verification latency by approximately 1.1× and 6.6× compared with [18] and [20], respectively. The work in [19] retains a 1.6× advantage in verification and reports a smaller logic area. However, the area values are not directly comparable, due to differences in technology libraries and implementation stages. The most comparable implementation [18], which uses the same technology node and accounts for the full system (memory, processor, and accelerator), occupies 2.4× the area of P12 . In summary, with our implementation, Mirath meets the NIST-induced requirement for a non-lattice candidate in the

additional DSA standardization process to outperform SLHDSA with respect to latency when jointly considering the sign and verify operation. Even though Mirath did not advance to the third round, this is an important finding for TCitH schemes with similar computational structures. Energy comparisons are limited because [18] is the only considered work reporting power and energy for an individual cryptographic operation, and only for Verify. For this operation, P12 requires approximately 1.9× lower power and 135× lower energy. P12 achieves latencies comparable to those of SoA Dilithium/ML-DSA and Falcon implementations, supporting the feasibility of TCitH-based signatures as viable alternatives from an implementation perspective. It outperforms the Dilithium/ML-DSA implementation in [20] in all three operations and operates in the same latency range as [21] and [22], while reporting 2× and 11.7× lower area, respectively. VIII. C ONCLUSION In this work, we investigated the implementation characteristics of Mirath, a TCitH-based post-quantum DSA, for resource-constrained systems. We presented, to the best of our knowledge, the first ASIC implementation of a TCitHbased signature scheme and combined algorithmic optimizations with hardware accelerators to substantially reduce the energy consumption of signing and verification relative to the reference software implementation. Our design achieves lower signing latency than all considered SoA SPHINCS+/SLHDSA ASIC implementations. At the same technology node, it also requires less area and achieves substantially lower verification energy. Additionally, compared to SoA implementations of lattice-based algorithms Dilithium/ML-DSA and Falcon, we have shown that, from an implementation perspective, a TCitH-based algorithm is a viable alternative in embedded applications in terms of computation time and required area.

R EFERENCES [1] M. Mosca and M. Piani, “Quantum Threat Timeline Report 2025,” Global Risk Institute, Toronto, ON, Canada, Tech. Rep., Mar. 2026. [Online]. Available: https://globalriskinstitute.org/publication/ quantum-threat-timeline-report-2025b/ [2] National Institute of Standards and Technology, “Module-LatticeBased Digital Signature Standard,” National Institute of Standards and Technology, Gaithersburg, MD, Federal Information Processing Standards Publication 204, Aug. 2024. [Online]. Available: https: //doi.org/10.6028/NIST.FIPS.204 [3] P.-A. Fouque, J. Hoffstein, P. Kirchner, V. Lyubashevsky, T. Pornin, T. Prest, T. Ricosset, G. Seiler, W. Whyte, and Z. Zhang, “Falcon: Fast-Fourier Lattice-Based Compact Signatures over NTRU,” falconsign.info, Technical report, 2022, Supporting documentation. [Online]. Available: https://falcon-sign.info/falcon.pdf [4] National Institute of Standards and Technology, “Stateless HashBased Digital Signature Standard,” National Institute of Standards and Technology, Gaithersburg, MD, Federal Information Processing Standards Publication 205, Aug. 2024. [Online]. Available: https: //doi.org/10.6028/NIST.FIPS.205 [5] ——, “Call for Additional Digital Signature Schemes for the PostQuantum Cryptography Standardization Process,” https://csrc.nist.gov/ Projects/post-quantum-cryptography/additional-signatures-2023, Sep. 2023. [6] Y. Ishai, E. Kushilevitz, R. Ostrovsky, and A. Sahai, “Zero-Knowledge from Secure Multiparty Computation,” in Proceedings of the ThirtyNinth Annual ACM Symposium on Theory of Computing. ACM, Jun. 2007, pp. 21–30. [Online]. Available: https://doi.org/10.1145/1250790. 1250794 [7] T. Feneuil and M. Rivain, “Threshold Computation in the Head: Improved Framework for Post-Quantum Signatures and ZeroKnowledge Arguments,” Journal of Cryptology, vol. 38, no. 3, p. 28, Jul. 2025. [Online]. Available: https://doi.org/10.1007/s00145-025-09543-8 [8] G. Adj, N. Aragon, S. Barbero, M. Bardet, E. Bellini, L. Bidoux, J.-J. Chi-Domı́nguez, V. Dyseryn, A. Esser, T. Feneuil, P. Gaborit, R. Neveu, M. Rivain, L. Rivera-Zamarripa, C. Sanna, J.-P. Tillich, J. Verbel, and F. Zweydinger, “Mirath,” Mirath Team, Algorithm specification, Feb. 2025, Version 2.0; merger of MIRA and MiRitH. [Online]. Available: https://csrc.nist.gov/csrc/media/Projects/ pqc-dig-sig/documents/round-2/spec-files/mirath-spec-round2-web.pdf [9] N. Aragon, M. Bardet, L. Bidoux, J.-J. Chi-Domı́nguez, V. Dyseryn, T. Feneuil, P. Gaborit, A. Joux, R. Neveu, M. Rivain, J.-P. Tillich, and A. Vinçotte, “RYDE Signature Scheme,” RYDE Team, Algorithm specification, Sep. 2025, Version 2.1.0. [Online]. Available: https://pqc-ryde.org/assets/downloads/ryde specification v2.1.0.pdf [10] R. Benadjila, C. Bouillaguet, T. Feneuil, and M. Rivain, “MQOM: MQ on my Mind: Algorithm Specifications and Supporting Documentation,” MQOM Team, Algorithm specification, Sep. 2025, Version 2.1. [Online]. Available: https://mqom.org/docs/mqom-v2.1.pdf [11] G. Alagic, M. Bros, P. Ciadoux, Q. Dang, T. H. Dang, J. Kelsey, J. Lichtinger, Y.-K. Liu, C. Miller, D. Moody, R. Peralta, R. Perlner, A. Robinson, H. Silberg, D. Smith-Tone, and N. Waller, “Status Report on the Second Round of the Additional Digital Signature Schemes for the NIST Post-Quantum Cryptography Standardization Process,” National Institute of Standards and Technology, Gaithersburg, MD, NIST Internal Report 8610, May 2026. [Online]. Available: https://doi.org/10.6028/NIST.IR.8610 [12] R. Benadjila and T. Feneuil, “Breaking the Myth of MPCitH Inefficiency: Optimizing MQOM for Embedded Platforms,” IACR Transactions on Cryptographic Hardware and Embedded Systems, vol. 2026, no. 3, pp. 279–305, Jul. 2026. [Online]. Available: https://doi.org/10.46586/tches.v2026.i3.279-305 [13] D. F. Aranha, J. Degn, J. Eilath, K. Nielsen, and P. Scholl, “FAEST for Memory-Constrained Devices with Side-Channel Protections,” Cryptology ePrint Archive, Paper 2025/1261, 2025. [Online]. Available: https://eprint.iacr.org/2025/1261 [14] S. Bettaieb, L. Bidoux, A. Budroni, M. Palumbi, and L. P. Perin, “Enabling PERK and Other MPC-in-the-Head Signatures on ResourceConstrained Devices,” IACR Transactions on Cryptographic Hardware and Embedded Systems, vol. 2024, no. 4, pp. 84–109, 2024. [Online]. Available: https://doi.org/10.46586/tches.v2024.i4.84-109 [15] S. Deshpande, J. Howe, J. Szefer, and D. Yue, “SDitH in Hardware,” IACR Transactions on Cryptographic Hardware and Embedded

Systems, vol. 2024, no. 2, pp. 215–251, 2024. [Online]. Available: https://doi.org/10.46586/tches.v2024.i2.215-251 [16] B. Funk, T. Bao, L. Bidoux, and J. Xie, “HAKE: Efficient Hardware Accelerator for Key Generation of Post-Quantum Signature Scheme PERK,” Cryptology ePrint Archive, Paper 2026/841, 2026. [Online]. Available: https://eprint.iacr.org/2026/841 [17] M. Schöffel, H. Tomasi, and N. Wehn, “HW/SW Implementation of MiRitH on Embedded Platforms,” in 2025 IEEE 16th Latin America Symposium on Circuits and Systems (LASCAS). IEEE, 2025, pp. 1–5. [Online]. Available: https://doi.org/10.1109/LASCAS64004.2025. 10966273 [18] P. Karl, J. Schupp, and G. Sigl, “Performance and Communication Cost of Hardware Accelerators for Hashing in Post-Quantum Cryptography,” ACM Transactions on Embedded Computing Systems, vol. 24, no. 5, pp. 66:1–66:31, Sep. 2025. [Online]. Available: https://doi.org/10.1145/3676965 [19] M.-J. O. Saarinen, “Accelerating SLH-DSA by Two Orders of Magnitude with a Single Hash Unit,” in Advances in Cryptology – CRYPTO 2024, ser. Lecture Notes in Computer Science, vol. 14920. Springer, 2024, pp. 276–304. [Online]. Available: https: //doi.org/10.1007/978-3-031-68376-3 9 [20] A. Dolmeta, V. Piscopo, M. Hutter, M. Martina, and G. Masera, “HORCRUX: A Complete PQC RISC-V eXtension Architecture,” arXiv preprint arXiv:2607.13939, 2026. [Online]. Available: https: //arxiv.org/abs/2607.13939 [21] P. Karl, J. Schupp, T. Fritzmann, and G. Sigl, “Post-Quantum Signatures on RISC-V with Hardware Acceleration,” ACM Transactions on Embedded Computing Systems, vol. 23, no. 2, pp. 30:1–30:23, Mar. 2024. [Online]. Available: https://doi.org/10.1145/3579092 [22] X. Carril, A. M. Pasoot, E. Parisi, O. Farràs, C. A. Lara-Niño, and M. Moretó, “PQCUARK: A Scalar RISC-V ISA Extension for ML-KEM and ML-DSA,” in 2026 Design, Automation & Test in Europe Conference (DATE). IEEE, 2026, pp. 1–7. [Online]. Available: https://doi.org/10.23919/DATE69613.2026.11539512 [23] RISC-V International, “The RISC-V Instruction Set Manual, Volume I: Unprivileged Architecture,” https://docs.riscv.org/reference/isa/ v20260120/index.html, 2026, Version 20260120. [24] M. J. Kannwischer, M. Krausz, R. Petri, and S.-Y. Yang, “pqm4: Benchmarking NIST Additional Post-Quantum Signature Schemes on Microcontrollers,” Cryptology ePrint Archive, Paper 2024/112, 2024. [Online]. Available: https://eprint.iacr.org/2024/112 [25] M. K. Franklin and M. Yung, “Communication Complexity of Secure Computation (Extended Abstract),” in Proceedings of the Twenty-Fourth Annual ACM Symposium on Theory of Computing. ACM, 1992, pp. 699–710. [Online]. Available: https://doi.org/10.1145/129712.129780 [26] L. Bidoux, T. Feneuil, P. Gaborit, R. Neveu, and M. Rivain, “Dual Support Decomposition in the Head: Shorter Signatures from Rank SD and MinRank,” in Advances in Cryptology – ASIACRYPT 2024, ser. Lecture Notes in Computer Science, K.-M. Chung and Y. Sasaki, Eds., vol. 15485. Springer, 2024, pp. 38–69. [Online]. Available: https://doi.org/10.1007/978-981-96-0888-1 2 [27] National Institute of Standards and Technology, “Advanced Encryption Standard (AES),” National Institute of Standards and Technology, Gaithersburg, MD, Federal Information Processing Standards Publication 197-upd1, May 2023. [Online]. Available: https://doi.org/10.6028/NIST.FIPS.197-upd1 [28] ——, “SHA-3 Standard: Permutation-Based Hash and ExtendableOutput Functions,” National Institute of Standards and Technology, Gaithersburg, MD, Federal Information Processing Standards Publication 202, Aug. 2015. [Online]. Available: https://doi.org/10.6028/NIST.FIPS.202 [29] R. Swann and J. E. Stine, “A Reconfigurable Architecture for Improvement and Optimization of Advanced Encryption Standard Hardware,” in 2021 55th Asilomar Conference on Signals, Systems, and Computers. IEEE, 2021, pp. 1181–1185. [Online]. Available: https://doi.org/10.1109/IEEECONF53345.2021.9723104

Record · ID 1108622 · SHA-256 72b434b05e8d9e26
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.