Conceptio › Archive › arXiv CS
arXiv CSopen access

Horizontal SCA Attacks on Binary kP Algorithms using Chevallier-Mames Atomic Blocks

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

Author's version, accepted for the workshop “Test Methods and Reliability of Circuits and Systems” (TuZ-2026)

Horizontal SCA Attacks on Binary kP Algorithms using Chevallier-Mames Atomic Blocks Gerald Isheanesu Matungamire1,3, Alkistis Aikaterini Sigourou1, Gerrit Schrock1,2, Zoya Dyka1,3, Peter Langendoerfer3 and Ievgen Kabin1 1 IHP – Leibniz Institute for High Performance Microelectronics, Frankfurt (Oder), Germany 2 Nordhausen University of Applied Sciences, Nordhausen, Germany 3 BTU Cottbus-Senftenberg, Cottbus, Germany {matungamire, sigourou, schrock, dyka, langendoerfer, kabin}@ihp-microelectronics.com Abstract— Scalar multiplication kP is the operation most frequently targeted in Elliptic Curve (EC) cryptosystems. To protect against single-trace Side-Channel Analysis (SCA) attacks, the atomicity principle and various atomic block patterns have been proposed in the past. In this work we use our software and hardware implementations to demonstrate that binary right-toleft and left-to-right kP algorithms, when implemented with Chevallier-Mames atomic block patterns, are still vulnerable to single-trace SCA attacks. The vulnerability remains true for the left-to-right kP algorithm with projective coordinate randomization. Keywords— Elliptic Curve (EC), kP, Side-Channel Analysis (SCA) attacks, atomic block, atomic patterns, Simple SCA.

I. INTRODUCTION Elliptic curves (EC) are the basis of cryptographic protocols for the exchange of secret keys, authentication and digital signature approaches. The most time and energy consuming operation in EC cryptosystems is the EC point-scalar multiplication, denoted as kP operation (where k is a scalar and P is a point on an EC). This operation is mostly attacked with the goal of revealing the value of the secret scalar k. Sidechannel analysis (SCA) attacks exploit side-channel effects that can be measured during the targeted kP execution, for example, current drawn from the power supply or electromagnetic (EM) radiation emitted by the attacked chip. Analysing the power or the EM trace, attackers can reveal the scalar k. The resistance of kP algorithms to single-trace attacks remains an issue despite many countermeasures proposed in the past. The atomicity principle is a countermeasure strategy based on the representation of kP operations as a series of small blocks of mathematical operations. The notation “atomic block” was introduced in [1], proposing decomposing the kP operation into many small atomic blocks: each block performs the same sequence of operations causing similar energy consumption over the same execution time, i.e. the shapes of the atomic blocks are very similar to each other. Attackers do not know which block refers to the processing of ‘0’ bit value in the binary representation of k, and which to the ‘1’. A similar idea was proposed in [2], before [1]. In [3], [4], [5], the vulnerability of kP algorithms implemented using different atomic block patterns to single-trace SCA attacks was described. The SCA leakages were caused by the data value processed (data-bit) [3] or data flow (address-bit) [4] or by their combination [5]. The last two vulnerabilities were observed for hardware implementations of the atomic patterns [6] and [7]. In this paper, we demonstrate that not only a hardware implementation of

binary kP algorithms using Chevallier-Mames’ atomic patterns [1] can be vulnerable to simple SCA, but also software implementation for microcontrollers, even when implemented using an open source crypto library with time-constant field operations. II. IMPLEMENTATION DETAILS AND MEASUREMENTS We implemented two binary kP double-and-add algorithms: the left-to-right and the right-to-left algorithms using atomic block patterns for EC point doubling (further denoted as PD, or 2Q) and EC point addition (further denoted as PA, or Q+P) operations as proposed in [1]. Each atomic block consists of: multiplication, addition, negation, and addition, known as the MANA atomic block. The atomic patterns for a PD and a PA consist of 10 and 16 MANA atomic blocks, respectively. The sequences of field operations in our hardware and software implementations are given in [8]. For our software implementation, we selected a TI LAUNCHXL-F28379D LaunchPad board [9] with the TMS320F28379D C2000 32-bit dual-core microcontroller as the target device and the opensource cryptographic library FLECC in C [10] due to the availability of the constant-time modular arithmetic functions. The implementation was done for the NIST EC P-256 [11]. For the left-to-right kP algorithms, we implemented two versions: without any additional countermeasures (i.e., only atomic blocks were the countermeasure) and a version with randomised projective coordinates of the point P corresponding to [12]. We did not apply the randomisation of the point coordinates in the right-to-left kP algorithm. In our hardware implementation, we realised only the left-to-right kP algorithm using the previous IHP design [5] as a basis and changing the data flow corresponding to [1]. All operations were implemented serially and described in VHDL using AMD Vivado Design Suite v2024.2.1 (64-bit). The design was ported for a Digilent Arty Z7-20 board [13] with a Zynq XC7Z020-1CLG400C FPGA. Measurements of EM emanation during a kP operation were performed using a near-field microprobe MFA-R 0.2-75 [14] manufactured by Langer. An integrated circuit scanner Langer ICS 105 [15] was used to guarantee a precise position and orientation of the probe over the board. We measured EM traces on the power supply capacitor C49 located on the back side of the attacked board, see Fig. 1. The EM trace of a kP execution on the attacked FPGA was measured also on a power supply capacitor on the back side of the attacked board. The placement and orientation of the EM probe are shown in Fig. 2. All

algorithms were executed using the affine coordinate of the base point G of the EC P-256 as inputs. We used a 256-bit-long scalar k for FPGA execution andonly the 5-bit-long binary scalar k=111112 for all executions on the microcontroller. We used a Teledyne LeCroy WavePro 604HD oscilloscope to capture the EM trace. The FPGA design runs at 10 MHz, and the trace was measured with the oscilloscope’s sampling rate of 1 GS/s. The kP execution on the microcontroller (in RAM) was clocked at 100 MHz and measured with 2.5 GS/s.

Reduction is performed in each clock cycle. If both or even only one of the multiplicands is equal to “0” or “1”, the partial products such as 0‧0, 0‧1, 1‧1, or x‧1 are calculated consuming much less energy than processing 64 bit long multiplicands x‧y. These processes determine the shape of the multiplications and their distinguishability, see Fig. 4.

Fig. 4. Atoms Δ1, … , Δ10 of the Q+P operation from Fig. 3, zoomed in.

Fig. 5 shows the traces measured on the microcontroller. On each trace all PDs and PAs are shown.

Fig. 1. Measurements of EM trace on the back side of the microcontroller’s board: position and orientation of the EM probe are shown, zoomed in.

Fig. 2. Measurements of EM trace on the back side of the attacked FPGA board: position and orientation of the EM probe are shown, zoomed in.

III. ANALYSIS A. Vulnerability to simple SCA Fig. 3 shows a part of the oscilloscope’s waveform for FPGA measurement. The shown part corresponds to the execution of a PA and a PD. Each PA consists of 16 atomic blocks Δ1, …, Δ16; each PD consists of 10 atomic blocks Δ1, … , Δ10. To avoid the successful simple SCA attack, the shapes of all atomic blocks have to be very similar, but it is clearly visible that the shapes of the atomic blocks Δ1-Δ4 and Δ9 in the Q+P differ significantly from those of the other atomic blocks.

Fig. 3. A part of the kP trace measured on the FPGA demonstrating the distinguishability of the EC point addition and the EC point doubling patterns

Fig. 4 shows the first ten atoms of the Q+P operation from Fig. 3, zoomed in. Each atom starts with a field multiplication, which is the most consuming energy field operation. In our implementation, we used the field multiplier based on the iterative 4-segment Karatsuba multiplication method adapted for the multiplication of prime finite field elements as explained in [16]. Corresponding to this method, both 256 bit long operands are segmented into 4 parts, each 64 bit long. To calculate the field product, nine partial products are calculated and accumulated in the output register of the field multiplier.

Fig. 5. EM traces measured on the microcontroller for kP executions implementing: a) – the left-to-right algorithm; b) – the left-to-right kP algorithm with randomization of projective coordinates; c) – the right-to-left algorithm.

In Fig. 5-a) the atomic blocks for a PD and a PA are shown also zoomed in, as an example demonstrating their distinguishability in the left-to-right algorithm. It is clearly visible that the atomic patterns Δ1 and Δ3 in Q+P differ significantly from those of all other atomic blocks. For the rightto-left algorithm (see Fig. 5 - c), the first PD and PA operations differ from all other patterns, too. The reasons for the visible distinguishability of atomic blocks are primarily due to operanddependent energy consumption in field multiplication processing small operand values such as 1 and -3, corresponding to [1]. Only the left-to-right algorithm with randomization of the projective coordinates does not have an easily observable distinguisher (see Fig. 5-b), i.e. the shape of all atomic blocks appears visually similar in the EM trace. Thus, the well-known randomisation of elliptic curve point coordinates proposed by [12] as a countermeasure against vertical attacks can also effectively prevent the simple data-bit SCA attacks by replacing the special operand values ‘1’ with long binary numbers. B. Atoms are distinguishable Additionally, we investigated the distinguishability of atomic block shapes for the left-to-right kP algorithm with randomization of projective coordinates (see trace in Fig. 5-b) at

the beginning of each atom to prove our assumption that the data flow, i.e. the use of different variables for passing data to/from a function, can cause distinguishability of the atoms. This kind of vulnerability is also known as address-bit side-channel leakage phenomena, where the information leakage comes from the addresses of registers or memory locations that are accessed during execution. Regular binary kP algorithms, such as the Montgomery ladder 1 , are vulnerable to this kind of attacks analysing many kP execution traces [22] as well as a single trace [23], at least when implementing kP operation in hardware. Atomic patterns [6] are also distinguishable due to the addressbit phenomenon [24]. Different address randomisation methods were proposed in the past [25], [26], [27], [28] and successfully broken by analysing a single trace [29], [30]. Some methods increasing the resistance of hardware implementations to address-bit SCA are proposed in [31].

Fig. 6. A part of each atom Δ1, from all four point doubling operations, aligned: traces of the atoms are marked in red; template is marked in black.

The nature of this vulnerability and suitable countermeasures have not investigated yet deeply for embedded processors and microcontrollers. Only a very few works [24], [32], [33] describe such investigations. In our investigation case, although the same operations are performed in each atom, different variables and data-flow paths lead to different address accesses. This can cause observable but very short – only a few clock cycles long – SCA leakage at the start of field operation functions, which allows distinguishing between atoms. Thus, we selected only 24 clock cycles at the beginning of the first atom in the first PD as a template for the comparison of all other atoms. Fig. 6 shows a part of each atom Δ1 in each PD (red lines). The traces are aligned. The part of the first atom Δ1, which we selected as template, is marked in black. All atoms Δ1 have a very similar shape at their beginning, i.e. a high correlation with the template is expected. We compared the template with the corresponding parts of other atoms in PDs. A similar shape was observed not for all of them but only for atoms Δ2, Δ4 and Δ7. Fig. 7 shows a part of each atom Δ3 in each PDs (red lines). The traces are aligned and the template is shown as black line (the same template as in Fig. 7). All atoms Δ3 have a very similar shape at their beginning, but it differs from the template in a few clock cycles significantly. Similar differences are observable for other atoms in PDs.

Fig. 7. A part of each atom Δ3 (red lines), from all four PD operations, aligned; template is marked in black.

Based on these observations, we decided to use the template for the calculation of Pearson coefficients for each 24 clock cycles long part of the measured trace to determine all very 1 It is the notation for kP algorithms based on [17]. Please note that the Montgomery ladder is known as resistant to timing and simple SCA attacks

similar shapes. Thus, we calculated Pearson’s coefficients for the whole trace, see Fig. 8-a). The calculated coefficients are mostly in the range of -0.5…0.7, and only a few of them are close to 1.0. We marked the coefficients higher than 0.9 with red dotted lines.

Fig. 8. Distinguishability of PD and PA operations in the left-to-right kP algorithm with randomization of projective coordinates: a) - Pearson coefficients calculated for template and each 24 clock cycles long part of the measured trace; red dotted lines show the coefficients higher than 0.9; b) – labeled measured trace.

The distinguishability of PD and PA operations using the marked correlation coefficients is demonstrated in Fig. 8-b). Each Δ1, Δ2, Δ4, and Δ7 exhibits a high correlation with the template during PD, while in PA, only atom Δ1 maintains this correlation. Conversely, atoms Δ2, Δ4, and Δ7 do not show the same relationship. Additionally, the duration of Q+P is longer than 2Q. These facts can be successfully exploited to distinguish between the atomic patterns and extract the processed scalar. In our investigations, we focused on a scalar contatining all ‘1’, implying that a Q+P operation follows every 2Q operation. For other scalars, this is not the case; processing the bit value ‘0’ requires only a 2Q operation and processing the bit value ‘1’ requires the sequence of a 2Q and a Q+P operations. IV. CONCLUSION In this paper we evaluated the distinguishability of Chevallier-Mames atomic block patterns in binary kP algorithms. Both our hardware and software implementations are vulnerable to single-trace SCA attacks. The field multiplication of special small operands is the reason for databit vulnerability, which can be easily observed and exploited for key extraction when left-to right algorithm is selected for the implementation. Additionally, we demonstrated that the projective coordinate randomization proposed by Coron [12] can successfully countermeasure this vulnerability but can not prevent the address-bit SCA leakage. We utilize portions of the measured traces related to data transmission in field multiplications as templates to recognize atomic patterns associated with PD and PA operations. Addressing of different registers is the reason for this vulnerability and remains exploitable even when using projective coordinate randomization. Please note that not only field multiplications but also field additions can be potentially such SCA leakage sources, too. Using the proposed here template, the results published in [24], [32], [33] investigating other atomic patterns, [18], [19], [20], assuming that key-dependent addressing of different registers is indistinguishable from SCA point of view [21].

[6], [7], [34] can be improved. Additionally, atomic patterns [35], can be very promising and have to be investigated as well. REFERENCES

Attacks Against ECSM with the Montgomery Ladder’, in Constructive Side-Channel Analysis and Secure Design, vol. 11421, I. Polian and M. Stöttinger, Eds., in LNCS, vol. 11421. , Cham: Springer International Publishing, 2019, pp. 25–42. doi: 10.1007/978-3-030-16350-1_3.

B. Chevallier-Mames, M. Ciet, and M. Joye, ‘Low-cost solutions for preventing simple side-channel analysis: side-channel atomicity’, IEEE

Trans.Comput. ,vol.53, no.6, pp.760–768,Jun.2004,doi: 10.1109/TC.2004.13.

[20] J. Samotyja and K. Lemke-Rust, ‘Practical Results of ECC Side Channel Countermeasures on an ARM Cortex M3 Processor’, in Proceedings of the 2016 ACM Workshop on Theory of Implementation Security, Vienna Austria: ACM, Oct. 2016, pp. 27–35. doi: 10.1145/2996366.2996371.

C. H. Gebotys, ‘Security-driven exploration of cryptography in DSP cores’, in Proceedings of the 15th international symposium on System Synthesis, in ISSS ’02. New York, NY, USA: Association for Computing

[21] M. Joye and S.-M. Yen, ‘The Montgomery Powering Ladder’, in CHES 2002, B. S. Kaliski, çetin K. Koç, and C. Paar, Eds., Berlin, Heidelberg: Springer, 2003, pp. 291–302. doi: 10.1007/3-540-36400-5_22.

[3]

A. Bauer, et al ‘Horizontal Collision Correlation Attack on Elliptic Curves’, in Selected Areas in Cryptography -- SAC 2013, vol. 8282, T. Lange, K. Lauter, and P. Lisoněk, Eds., in LNCS, vol. 8282, Springer 2014, pp. 553–570. doi: 10.1007/978-3-662-43414-7_28.

[22] K. Itoh, T. Izu, and M. Takenaka, ‘Address-Bit Differential Power Analysis of Cryptographic Schemes OK-ECDH and OK-ECDSA’, in CHES 2002, vol. 2523, B. S. Kaliski, çetin K. Koç, and C. Paar, Eds., in LNCS, vol. 2523. , Springer Berlin Heidelberg, 2003, pp. 129–143. doi: 10.1007/3-540-36400-5_11.

[4]

I. Kabin, et al. , ‘Atomicity and Regularity Principles Do Not Ensure Full Resistance of ECC Designs against Single-Trace Attacks’, Sensors, vol. 22, no. 8, Art. no. 8, Jan. 2022, doi: 10.3390/s22083083.

[23] I. Kabin, Z. Dyka, D. Kreiser, and P. Langendoerfer, ‘Horizontal addressbit DPA against montgomery kP implementation’, in 2017 ReConFig, Cancun:IEEE, Dec.2017, pp.1–8.doi: 10.1109/RECONFIG.2017.8279800.

[5]

A. A. Sigourou, Z. Dyka, P. Langendoerfer, and I. Kabin, ‘Atomic Patterns: Field Operation Distinguishability on Cryptographic ASICs’, in 2025 IEEE CSR, Chania, Crete, Greece: IEEE, Aug. 2025, pp. 990–995.

[24] A. A. Sigourou, I. Kabin, P. Langendörfer, N. Sklavos, and Z. Dyka, ‘Successful Simple Side Channel Analysis: Vulnerability of an atomic pattern kP algorithm implemented with a constant time crypto library to simple electromagnetic analysis attacks’, in 2023 MECO, Jun. 2023, pp. 1–6. doi: 10.1109/MECO58584.2023.10154940.

[1]

[2]

Machinery, Oct. 2002, pp. 80–85. doi: 10.1145/581199.581218.

doi: 10.1109/CSR64739.2025.11130154.

[6]

F. Rondepierre, ‘Revisiting Atomic Patterns for Scalar Multiplications on Elliptic Curves’, in Smart Card Research and Advanced Applications, A.Francillon and P.Rohatgi, Eds.,in LNCS. Cham: Springer International Publishing, 2014, pp. 171–186. doi: 10.1007/978-3-319-08302-5_12.

[7]

P. Longa, ‘Accelerating the Scalar Multiplication on Elliptic Curve Cryptosystems Over Prime Fields’, 2007.

[8]

G. Schrock, et al., ‘Distinguishability of EC Point Doublings and Additions in Binary kP Implementations using Chevallier-Mames Atomic Blocks’, in 2025 IEEE EWDTS, Tbilisi, Georgia: IEEE, Dec. 2025, pp. 1–6. doi: 10.1109/EWDTS67441.2025.11303708.

[9]

‘LAUNCHXL-F28379D Development kit | TI.com’. Available: https://www.ti.com/tool/LAUNCHXL-F28379D

[10] Flecc in C. (Sep. 07, 2022). Scilab. Institute of Information Security. [Online]. Available: https://github.com/isec-tugraz/flecc_in_c [11] ‘Digital Signature Standard (DSS)’, National Institute of Standards and Technology (U.S.), Washington, D.C., NIST FIPS 186-5, Feb. 2023. doi: 10.6028/NIST.FIPS.186-5. [12] J.-S. Coron, ‘Resistance Against Differential Power Analysis For Elliptic Curve Cryptosystems’, in CHES, Ç. K. Koç and C. Paar, Eds., Berlin, Heidelberg: Springer,1999,pp.292–302.doi: 10.1007/3-540-48059-5_25. [13] ‘Arty Z7: Zynq-7000 SoC Development Board’, Digilent. Available: https://digilent.com/shop/arty-z7-zynq-7000-soc-development-board/

[25] K. Itoh, T. Izu, and M. Takenaka, ‘A Practical Countermeasure against Address-Bit Differential Power Analysis’, in CHES 2003, C. D. Walter, Ç. K. Koç, and C. Paar, Eds., Berlin, Heidelberg: Springer, 2003, pp. 382–396. doi: 10.1007/978-3-540-45238-6_30. [26] M. Izumi, et al., ‘Improved countermeasure against Address-bit DPA for ECC scalar multiplication’, inDATE 2010, Dresden: IEEE, Mar. 2010, pp. 981–984. doi: 10.1109/DATE.2010.5456907. [27] L. Batina, et al., ‘Side-channel evaluation of FPGA implementations of binary Edwards curves’, in 2010 17th IEEE ICECS, Athens, Greece: IEEE, Dec. 2010, pp. 1248–1251. doi: 10.1109/ICECS.2010.5724745. [28] N. Pirotte, J. Vliegen, L. Batina, and N. Mentens, ‘Design of a Fully Balanced ASIC Coprocessor Implementing Complete Addition Formulas on Weierstrass Elliptic Curves’, in DSD 2018, Prague: IEEE, Aug. 2018, pp. 545–552. doi: 10.1109/DSD.2018.00095. [29] I. Kabin, Z. Dyka, and P. Langendoerfer, ‘Randomized Addressing Countermeasures are Inefficient Against Address-Bit SCA’, in IEEE CSR 2023 , Venice, Italy: IEEE, Jul. 2023, pp. 580–585. doi: 10.1109/CSR57506.2023.10224968. [30] I. Kabin, et al., ‘Breaking a fully Balanced ASIC Coprocessor Implementing Complete Addition Formulas on Weierstrass Elliptic Curves’, in DSD 2020, Kranj, Slovenia: IEEE, Aug. 2020, pp. 270–276. doi: 10.1109/DSD51259.2020.00051.

[14] ‘Langer EMV - MFA-R 0.2-75, Near-Field Micro Probe 1 MHz up to 1 GHz’.Available: https://www.langer-emv.de/en/product/mfa-active-1mhz-

[31] I. Kabin, ‘Horizontal address-bit SCA attacks against ECC and appropriate countermeasures’, BTU Cottbus - Senftenberg, 2023. doi: 10.26127/BTUOpen-6397.

[15] ‘Langer EMV - ICS 105 set, IC Scanner 4-Axis Positioning System’.

[32] Sze Hei Li, ‘Distinguishability investigation on Longa’s atomic patterns when used as a basis for implementing elliptic curve scalar multiplication algorithms, Master’s Thesis’, BTU Cottbus - Senftenberg, 2024. doi: 10.26127/BTUOpen-6836.

up-to-6-ghz/32/mfa-r-0-2-75-near-field-micro-probe-1-mhz-up-to-1-ghz/854

Available: https://www.langer-emv.com/en/product/langer-scanner/41/ics105-set-ic-scanner-4-axis-positioning-system/144

[16] I. Kabin, et al., ‘Unified field multiplier for ECC: Inherent resistance against horizontal SCA attacks’, in 2018 13th DTIS, Taormina: IEEE, Apr. 2018, pp. 1–4. doi: 10.1109/DTIS.2018.8368560. [17] P. L. Montgomery, ‘Speeding the Pollard and elliptic curve methods of factorization’, Math. Comput., vol. 48, no. 177, pp. 243–264, 1987, doi: 10.1090/S0025-5718-1987-0866113-7. [18] J. Fan, et al., ‘State-of-the-art of secure ECC implementations: a survey on known side-channel attacks and countermeasures’, in 2010 HOST, Anaheim, CA, USA: IEEE, Jun. 2010, pp. 76–87. doi: 10.1109/HST.2010.5513110. [19] M. Azouaoui, et al., ‘Fast Side-Channel Security Evaluation of ECC Implementations: Shortcut Formulas for Horizontal Side-Channel

[33] P. L. Doku, Investigation of the Distinguishability of Giraud-Verneuil Atomic Blocks, Master’s thesis. BTU Cottbus-Senftenberg, 2025. doi: https://opus4.kobv.de/opus4-btu/frontdoor/index/index/docId/7140 [34] C. Giraud and V. Verneuil, ‘Atomicity Improvement for Elliptic Curve Scalar Multiplication’, in Smart Card Research and Advanced Application, vol. 6035, D. Gollmann, J.-L. Lanet, and J. Iguchi-Cartigny, Eds., in LNCS, vol. 6035. Springer Berlin Heidelberg, 2010, pp. 80–101. doi: 10.1007/978-3-642-12510-2_7. [35] P. Longa and A. Miri, ‘Fast and Flexible Elliptic Curve Point Arithmetic over Prime Fields’, IEEE Trans. Comput., vol. 57, no. 3, pp. 289–302, Mar. 2008, doi: 10.1109/TC.2007.70815.

Record · ID 134503 · SHA-256 3fb71c6a1ee0ad91
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.