Conceptio › Archive › arXiv CS
arXiv CSopen access

Design of Memristive Lightweight Encryption For In-Memory Image Steganography

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

Design of Memristive Lightweight Encryption For In-Memory Image Steganography ∗ Seyed Erfan Fatemieh1,†, Reza Shahdi Alizadeh2,‡, and Esmail Zarezadeh 3,§ May 6, 2026

1 Department of Computer Architecture, Faculty of Computer Engineering, University of Isfahan, Isfahan 8174673441, Iran

arXiv:2605.03494v1 [cs.CR] 5 May 2026

2 Department of Computer Engineering, Yadegar-e-Imam Khomeini (RAH) Shahre Rey Branch, Islamic Azad University,

Tehran, Iran, Tehran, Iran 3 Department of Electrical Engineering, AmirKabir University of Technology, Tehran, Iran

Abstract With the expansion of data-intensive applications and increasing data volumes, providing an efficient solution to address growing energy consumption and performance degradation caused by the transfer of large amounts of data between the processor and the main memory has become a severe challenge. The frequent transfer of large amounts of data between internal chip units, memories, and their interconnections exacerbates the vulnerability of the data being accessed. Employing a memristive Computation In-Memory-Array (CIM-A) architecture limits data transfer, thereby addressing both challenges. Furthermore, by integrating lightweight cryptography, developed to secure data in hardware-constrained devices, with CIM-A architectures, the security of data in transit, especially across interconnections, can be ensured. This paper implements two standard lightweight stream ciphers, Trivium and Grain-128a, for CIM using stateful material implication (IMPLY) logic to address these combined security and performance challenges. In addition to redesigning the cryptographic structures, we reduce the hardware complexity of conventional IMPLY-based implementations by proposing an efficient method for shifting data within the shift registers. Applying the proposed data-shifting method to the registers of these ciphers reduces the number of computational steps and decreases energy consumption by up to 42% and 44%, respectively, compared to conventional implementations. Finally, the performance of the proposed circuits is evaluated in a steganography application, demonstrating their practical efficiency.

Keywords Computation In-Memory (CIM), Memristor, Lightweight Cryptography, IMPLY Logic, Steganography, Emerging Technologies.

1

Introduction

Today, data is one of the most important components of computation, and its fast, energy-efficient processing has become a challenge for computer architects more than ever. The Internet of Things (IoT) and Artificial Intelligence (AI) are among the most important processing applications, expanding in various dimensions every day. From the perspective of an architect of a high-volume data-processing system, the secure transmission of data, alongside fast, and energy-efficient processing, is of particular importance. Maintaining data security and integrity poses numerous challenges, and using appropriate cryptographic mechanisms for secure data transmission, especially on edge devices and in IoT-related processing applications, is of high importance. The constant data movement ∗

Preprint Submitted to arXiv Corresponding author: Seyed Erfan Fatemieh ([email protected]) ‡ Reza Shahdi Alizadeh ([email protected]) § Esmail Zarezadeh ([email protected]) †

1

between the processor and memory in the conventional von Neumann architecture is a major concern for designers and researchers, both in terms of resource consumption and circuit evaluation criteria, and in terms of the security of data accessed during inter-unit transfers, especially in interconnections (Yang u. a., 2021; Fatemieh u. a., 2025b). By the middle of the first decade of the new century, technological improvements and reductions in chip transistor dimensions had significantly improved performance and addressed various processing needs (Hennessy und Patterson, 2017). With the invalidation of Dennard’s scaling principle, the mismatch between Gordon Moore’s prediction and manufacturing processes, and the power wall problem, the continuous trend of improving processor efficiency faced a sharp decline due to numerous problems, such as increased leakage currents, which were overcome by applying alternative architectures and the use of multi-core processors for several years (Hennessy und Patterson, 2017; Fatemieh und Reshadinezhad, 2026). One of the main solutions proposed by researchers to overcome the aforementioned problems has been the use of emerging technologies such as Carbon Nanotube Field-Effect Transistors (CNFETs), Quantum-dot Cellular Automata (QCA), and memristors (Fatemieh und Reshadinezhad, 2026; Fatemieh u. a., 2025c; Bagheralmoosavi u. a., 2025; Farahani u. a., 2024). In addition to emerging technologies, numerous computational methods have been proposed across various scientific communities to replace conventional methods. One of the most promising of these methods is Computation In-Memory (CIM). In the conventional von Neumann architecture, the main memory in the memory hierarchy is located far from the central processor. The significant difference between processor speed and memory bandwidth, combined with the massive volume of data transferred for processing and computation in this architecture, has become a major bottleneck in the design, known as the von Neumann (memory) bottleneck (Carvalho, 2002). The problem of data transfer in modern processing applications is so serious that, according to estimates from companies such as Google, from 63% to 90% of total processing energy consumption is spent solely on data movement, without performing any actual computation (TaheriNejad, 2024). The use of CIM architectures that leverage emerging memristive technologies, which enable simultaneous data storage and logic/arithmetic operations, is a promising approach to overcoming the problems posed by the memory wall. Several methods have been proposed for designing arithmetic and logic circuits using memristors, due to their favorable characteristics. In general, these methods can be classified into two categories: stateful and non-stateful (Alhaj Ali, 2020). Stateful methods are methods that can be used to store and process data completely within the memristive array (Fatemieh u. a., 2025b; Fatemieh und Reshadinezhad, 2026; Alhaj Ali, 2020). This means that stateful methods such as material implication (IMPLY) (Kvatinsky u. a., 2013) and Memristor Aided loGIC (MAGIC) (Kvatinsky u. a., 2014) are memristive processing methods that can be used in the CIM-Array (CIM-A) architecture to significantly reduce data transfer volume (Nguyen u. a., 2020). Maintaining information security and privacy across various processing applications related to IoT and AI, where large volumes of data are generated, collected, and transmitted by edge devices, has become a major concern (Yang u. a., 2021). The use of cryptographic mechanisms in these areas is essential. Memristors can be an efficient option for implementing cryptographic platforms due to their small area and power consumption, their resistance to side-channel attacks, and their potential integration into CIM-A architectures to eliminate redundant data transfers (Yang u. a., 2021; Fatemieh u. a., 2025b; Nguyen u. a., 2020; Siddiqi u. a., 2023). Excessive data transfer between the processor and main memory in conventional cryptographic units, which is required to generate and store ciphertext, also increases the risk of attack; this vulnerability can be minimized by applying a CIM-A architecture (Yang u. a., 2021). In (Kaya, 2020; Rai u. a., 2018, 2021), several methods for generating random and pseudorandom numbers using memristors have been investigated. Implementation and acceleration of hash functions using memristors for processing within and near the memory array have also been discussed in (Oved u. a., 2022). In addition to generating random numbers and hash functions, designers have also considered methods for implementing cryptographic mechanisms that specifically utilize memristors. In (Yang u. a., 2021), the characteristics of the process variations during transistor manufacturing in a single chip are applied to generate random numbers. The non-stateful method of One Transistor-One Resistor (1T-1R) is also employed to XOR the input value stored in the memristors with the aforementioned random number to generate the ciphertext (Yang u. a., 2021). In 2024, the lightweight block cipher GIFT, based on memristors, was used to generate ciphertext on edge devices (Siddiqi u. a., 2023). In this study, the XOR structure is located in the peripheral circuits of the memristive crossbar array (sense amplifiers) (Siddiqi u. a., 2023). Hence, the architecture considered by the authors was categorized as a CIM-Periphery (CIM-P) architecture (Nguyen u. a., 2020). Based on the results presented in (Siddiqi u. a., 2023; Zinabu u. a., 2025; Naser und Naif, 2022; Manifavas u. a., 2016; Soto-Cruz u. a., 2024) and an analysis of various lightweight ciphers, it can be concluded that lightweight stream ciphers are more suitable for implementation in devices with limited hardware resources. Using stateful 2

methods in the design of memristor-based circuits not only significantly reduces data transfer and the risk of unauthorized data access during inter-unit transfers, but is also an efficient way to improve circuit metrics. Considering the limitations imposed by power and memory walls, the importance of applying cryptographic mechanisms in inter-unit communications, and the need to minimize data transfer between processing units and memory on a chip, the authors propose the design of IMPLY-based encryption/decryption units for integration into CIM-A architectures, ensuring high compatibility with conventional memristive crossbar arrays. The lightweight stream ciphers evaluated in this article were selected based on global evaluations and standards (Naser und Naif, 2022; Manifavas u. a., 2016; Paar u. a., 2024; Hell u. a., 2019). The selected ciphers have been redesigned and improved based on a reliable stateful method, IMPLY (Seiler, 2025), to align seamlessly with the structure of conventional memristive crossbar arrays. Furthermore, the functionality of the proposed designs has been evaluated behaviorally in a steganography application. The results confirmed the practical applicability of the proposed designs in this domain. The main contributions of this paper are as follows: 1. Presenting a systematic methodology for porting the classic implementation algorithms of the Trivium (Paar u. a., 2024; De Cannière und Preneel, 2008) and Grain-128a (Ågren u. a., 2011) lightweight stream ciphers to the structure of CIM-A architectures based on the stateful IMPLY design method; 2. Developing an algorithm for reducing the computational steps and energy consumption of Linear Feedback Shift Registers (LFSRs) and Nonlinear Feedback Shift Registers (NFSRs) used in the Trivium and Grain-128a lightweight stream ciphers; 3. Redesigning and evaluating the functional correctness of all the basic logic blocks required in the IMPLYbased improved implementation algorithms of the Trivium and Grain-128a lightweight stream ciphers; and 4. Examining the functionality and image quality metrics of the proposed memristive encryption and decryption units in a steganography application. The remainder of this paper is organized into four sections. In Section 2, the basic concepts, definitions, and related work are reviewed. In Section 3, the proposed algorithm for implementing LFSRs and NFSRs, as well as the step-by-step implementation method for redesigning the Trivium and Grain-128a lightweight stream ciphers based on the IMPLY design method for CIM, are presented. The circuit-level and application-level simulation results are reported in Section 4. Finally, the paper’s conclusion is provided in Section 5.

2

Background and related work

2.1

Lightweight stream ciphers

Symmetric ciphers are divided into two categories: stream ciphers and block ciphers (Paar u. a., 2024). In stream ciphers, the message is encrypted by adding it (modulo-2 addition) to a pseudorandom keystream by applying XOR operations. Stream ciphers are often faster than block ciphers and require fewer hardware resources, making them suitable for resource-constrained devices (Paar u. a., 2024; Psomiadis, 2020). In general, block ciphers are more susceptible to error propagation when noise corrupts data, and stream ciphers are recommended when the message size is unknown, or the messages are transmitted in continuous streams (Psomiadis, 2020). Stream ciphers were developed based on the One-Time Pad (OTP) cipher, but because of the OTP’s limitations, they are used as an alternative (Zinabu u. a., 2025). Pseudorandom Number Generators (PRNGs) are a main pillar of stream ciphers, which use various structures, including LFSRs. The LFSR structure, in combination with nonlinear logic gates, acquires nonlinear characteristics and is applied in stream ciphers (Paar u. a., 2024). Over the last three decades, numerous symmetric ciphers have been proposed by various companies and research institutions and evaluated across aspects such as information security, complexity, and the resources required for evaluation (Zinabu u. a., 2025; Naser und Naif, 2022; Manifavas u. a., 2016; Soto-Cruz u. a., 2024; Paar u. a., 2024; Hell u. a., 2019; Ågren u. a., 2011; Psomiadis, 2020). Among the most important of these processes are the eSTREAM project and the evaluations of the US National Institute of Standards and Technology (NIST) (Zinabu u. a., 2025; Naser und Naif, 2022; Manifavas u. a., 2016; Soto-Cruz u. a., 2024; Paar u. a., 2024; Hell u. a., 2019; Ågren u. a., 2011; Psomiadis, 2020). Several stream ciphers have also been examined in these processes. Cryptography is widely applied in processing environments such as the IoT, but many devices in these and similar 3

domains are embedded devices with limited resources (Naser und Naif, 2022; Soto-Cruz u. a., 2024; Psomiadis, 2020). Therefore, relying on conventional cryptographic algorithms with high computational complexity in these applications is associated with significant limitations. Lightweight cryptography is a different approach from conventional algorithms that enables the maintenance of confidentiality and data integrity on devices with limited resources. Most efficient lightweight ciphers are symmetric ciphers, and in addition to evaluating their security, assessing their circuit metrics, such as energy consumption and resource utilization, is of great importance. A complete introduction and review of the main features of lightweight stream and block ciphers, along with comprehensive comparisons for selecting the appropriate lightweight cipher according to different design constraints, have been reported in (Zinabu u. a., 2025; Naser und Naif, 2022; Manifavas u. a., 2016; Soto-Cruz u. a., 2024; Paar u. a., 2024; Psomiadis, 2020). Based on the evaluations conducted and the results reported in the eSTREAM project and NIST evaluation processes, and considering the structures and characteristics of various lightweight stream ciphers, Trivium (De Cannière und Preneel, 2008) and Grain-128a (Ågren u. a., 2011) are selected to be described in detail in the following subsections. 2.1.1

Trivium

Trivium is one of the selected lightweight stream ciphers of the eSTREAM project and is also a standardized cipher (ISO/IEC 29192-3:2012) (Naser und Naif, 2022; Manifavas u. a., 2016). As shown in Figure 1, the structure of this cipher consists of three shift registers, a 93-bit register (A), an 84-bit register (B), and a 111-bit register (C), combined with nonlinear logic gates (Paar u. a., 2024). By calculating the modulo-2 addition of these three shift registers’ outputs, the keystream (si ) is produced (Paar u. a., 2024). The logical equations defining the structure of this lightweight stream cipher are expressed in (1)-(4). The 80-bit key and Initialization Vector (IV) are placed in the 80 Least Significant Bits (LSBs) of registers A and B, respectively, and all other bits of the registers are set to logical ‘0’, except for bits 109-111 of register C, which are set to a logical ‘1’ (Paar u. a., 2024). In the initialization phase, the cipher’s structure is clocked 1152 times, and the keystream output from the 1153rd cycle onward is used to produce the ciphertext (Paar u. a., 2024). ai ≡ ci−66 + ci−111 + ci−110 · ci−109 + ai−69

(1)

bi ≡ ai−66 + ai−93 + ai−92 · ai−91 + bi−78

(2)

ci ≡ bi−69 + bi−84 + bi−83 · bi−82 + ci−87

(3)

si ≡ ai−66 + ai−93 + bi−69 + bi−84 + ci−66 + ci−111

(4)

Figure 1: Architecture of the stream cipher Trivium (Paar u. a., 2024; De Cannière und Preneel, 2008).

4

2.1.2

Grain-128a

Another selected cipher of the eSTREAM project, which has been standardized in ISO/IEC 29167-13:2015 as one of the applicable ciphers for Radio Frequency Identification (RFID) systems, is Grain-128a (Hell u. a., 2019). This cipher can perform authentication in addition to producing a keystream for data encryption/decryption. Here, only the Grain-128a output generation process is discussed. The structure of this cipher is illustrated in Figure 2 for the pre-initialization and output generation stages. The main components in the structure of this cipher are the shift registers, each of which is 128 bits. (5) and (6) describe the update functions of these two shift registers, respectively (Ågren u. a., 2011). The Boolean functions h(x) and y (the output generation function), are also computed according to (7) and (8), respectively. The 128-bit key and the 96-bit IV are placed in the NFSR and LFSR, respectively (Ågren u. a., 2011). The other 32 bits of the LFSR are initialized from bits 96 to 126 with logical ‘1’, and the last bit (the 127th bit) with logical ‘0’ (Ågren u. a., 2011). The pre-initialization phase occupies 256 cycles, and from the 257th cycle onward, keystream generation begins. In each cycle after the 256th one, the modulo-2 addition (XOR) of the keystream and the input data produces the ciphertext (Ågren u. a., 2011). si+128 = si + si+7 + si+38 + si+70 + si+81 + si+96

(5)

bi+128 = si + bi + bi+26 + bi+56 + bi+91 + bi+96 + bi+3 · bi+67 + bi+11 · bi+13 + bi+17 · bi+18 + bi+27 · bi+59 + bi+40 · bi+48 + bi+61 · bi+65 + bi+68 · bi+84 + bi+22 · bi+24 · bi+25 + bi+70 · bi+78 + bi+82 · bi+88 + bi+92 · bi+93 · bi+95 (6) h(x) = bi+12 · si+8 + si+13 · si+20 + bi+95 · si+42 + si+60 · si+79 + bi+12 · bi+95 · si+94 yi = h(x) + si+93 +

X

A ∈ {2, 15, 36, 45, 64, 73, 89}

bi+j ,

(7) (8)

j∈A

2.2

An introduction to memristor

Capacitors, inductors, and resistors are three basic electrical elements that define the relationship between electric charge and voltage, magnetic flux and current, and voltage and current, respectively. Professor L. Chua of the University of California at Berkeley introduced the memristor in the 1970s as the fourth basic electrical element, defining the relationship between electric charge and magnetic flux (Alhaj Ali, 2020). Four decades later, a team at HP Laboratories led by Professor S. Williams implemented the first physical memristor (Alhaj Ali, 2020). The memristor, whose symbol is illustrated in Figure 3(a), can act as a non-volatile memory cell, and its resistance determines its logic state. If the resistance of the memristor is at its minimum value (RON ), it is generally equivalent to Boolean logic ‘1’, and if the resistance is at its maximum value (ROF F ), it represents Boolean logic ‘0’ (Fatemieh und Reshadinezhad, 2026; Alhaj Ali, 2020). As shown in Figure 3(a), if a voltage/current is applied from the Top Electrode (TE) of the memristor, its resistance decreases, and if applied from the other terminal, its resistance increases (Alhaj Ali, 2020). The resistance of the memristor remains unchanged for years when

(a)

(b)

Figure 2: Architecture of Grain-128a: (a) the pre-output generator, and (b) the key initiaization (Ågren u. a., 2011).

5

(a)

(b)

Figure 3: (a) Memristor symbol, and (b) memristive IMPLY gate (Fatemieh und Reshadinezhad, 2026). no voltage/current is applied to its terminals, making it an emerging non-volatile memory cell. Memristive CIM has another fundamental aspect: the execution of logic and arithmetic operations. Researchers have proposed several methods for implementing logic and arithmetic functions using memristors, which are generally classified into two categories: stateful and non-stateful (Alhaj Ali, 2020). In stateful circuit design methods, the resistance of the memristor determines the logic state at the input, output, and intermediate nodes, whereas in non-stateful methods, the logic state is mainly determined by voltage (Alhaj Ali, 2020). The applicability of stateful design methods in memristive CIM-A has been widely evaluated, and several methods have been developed. Among the most important stateful methods are IMPLY (Kvatinsky u. a., 2013), MAGIC (Kvatinsky u. a., 2014), and fast and energy-efficient logic in memory (FELIX) (Gupta u. a., 2018). Each of these methods has several advantages and disadvantages, and, considering the comparative summary in (Alhaj Ali, 2020; Seiler, 2025; Gupta u. a., 2018) and the reliability factor, the IMPLY method has been selected as the stateful method applied in this paper. The IMPLY gate architecture and its truth table are illustrated in Figure 3(b) and Table 1, respectively. The IMPLY gate is a universal gate; this means that any logic function can be computed by applying it multiple times in addition to the FALSE (zero) function. For example, the OR logic function can be implemented using two input memristors and one work memristor in two IMPLY-based computational steps, in addition to executing the FALSE function once, as shown in (9) (Bagheralmoosavi u. a., 2025). Work memristors are used alongside input memristors to implement logic and arithmetic functions in the IMPLY method. Storing temporary variables during the execution of IMPLY-based logic and arithmetic functions is the main role of work memristors, although these memristors can also be used to store the outputs. To perform the IMPLY operation, two voltage, VCON D and VSET , are simultaneously applied to memrsitors ‘p’ and ‘q’ in Figure 3(b), and the result of “pIM P q” is stored in the memristor ‘q’ (Bagheralmoosavi u. a., 2025). For the correct execution of the IMPLY operation, two conditions, 1) VCON D < Vc < VSET (Vc : the threshold voltage of a memristor) and 2) RON << RG << ROF F , must be met (Fatemieh und Reshadinezhad, 2026; Bagheralmoosavi u. a., 2025). Four architectures, serial, parallel, semi-serial, and semi-parallel, have been introduced for the development of logic and arithmetic circuits using the IMPLY method (Fatemieh und Reshadinezhad, 2026; Bagheralmoosavi u. a., 2025). The serial architecture is the most compatible with the structure of memristive crossbar arrays; hence, it is an acceptable structure for use in CIM-A (Fatemieh und Reshadinezhad, 2026; Bagheralmoosavi u. a., 2025). In the serial architecture, the memristors are placed in a row or column of the crossbar array. This architecture is illustrated in Figure 4. The serial architecture has the least hardware complexity among the others, but it can perform only one IMPLY operation per computational step (Fatemieh und Reshadinezhad, 2026). Numerous logic and arithmetic circuits, such as different types of adders and multipliers, have been implemented in the serial architecture using the IMPLY design method (Fatemieh u. a., 2025b; Fatemieh und Reshadinezhad, 2026; Bagheralmoosavi u. a., 2025; Kvatinsky u. a., 2013; Rohani und TaheriNejad, 2017; Seiler und TaheriNejad, 2025; Seiler u. a., 2025). Table 2 lists the implementation algorithms of serial IMPLY-based basic logic gates and their characteristics. p OR q ≡ (p IM P 0) IM P q ≡ (p → 0) → q

6

(9)

Table 1: The truth table of an universal IMPLY logic gate (Fatemieh und Reshadinezhad, 2026). p 0 0 1 1

q 0 1 0 1

p IM P LY q ≡ p IM P q ≡ p → q 1 1 0 1

Table 2: Implementation of basic Boolean logic gates using IMPLY method’s primitive functions (IMPLY and FALSE) (Bagheralmoosavi u. a., 2025) Boolean logic gate NOT p p OR q p NOR q p NAND q p AND q

Equivalent IMPLY logic p→0 (p → 0) → q ((p → 0) → q) → 0 p → (q → 0) (p → (q → 0)) → 0

Figure 4: The serial architecture of IMPLY-based logical and arithmetic circuits (Fatemieh u. a., 2025b; Fatemieh und Reshadinezhad, 2026; Bagheralmoosavi u. a., 2025).

2.3

Memristive cryptography

Memristors can be applied in cryptography due to their electrical properties. This subsection briefly discusses a few of these applications. In (Yang u. a., 2021), an architecture is introduced that generates random numbers based on variations in the threshold slope of Metal Oxide Semiconductor FETs (MOSFETs) in a 1T-1R crossbar array, driven by manufacturing process variations. According to the architecture designed in (Yang u. a., 2021), the threshold slope value is computed by dedicated extraction circuits and stored in a register. Then, this data is read and converted into a specific voltage (Yang u. a., 2021). By applying a non-stateful method introduced in (Yang u. a., 2021), an XOR operation is performed between the data stored in the memristor and the aforementioned random voltage value, and the output ciphertext is generated. Another application of memristors, widely used in cryptography, is the generation of random numbers. In (Rai u. a., 2018, 2021), different Ring Oscillators (ROs) are designed using varying numbers of transistors and memristors to generate random numbers. In the RO structure introduced in (Rai u. a., 2018), a memristor is placed between the inverters of each row of the RO implemented based on Complementary MOS (CMOS) technology. If the propagation delays of two different rows are not the same, the generated output is logic ‘1’, and otherwise, the output is logic ‘0’ (Rai u. a., 2018). The role of memristors in the RO structure proposed in (Hashim u. a., 2016) is to replace the PMOS transistors of the inverter’s pull-up network. The use of memristors in the RO structure increases the randomness of this circuit compared to an RO designed entirely with transistors (Hashim u. a., 2016). In (Rai u. a., 2021), two rows of single-bit ROs have been used to generate random numbers, where each row consists of multiple CMOS inverters, each connected to a single memristor. The random number generator proposed in (Rai u. a., 2021) can be implemented in serial and parallel architectures, so that, if throughput needs to be increased, more data can be processed in the post-processing stage using a parallel-to-serial converter (Rai u. a., 2021). In the post-processing stage, the Trivium stream cipher is used, and after 1152 cycles, a random

7

output bit is generated in each cycle (Rai u. a., 2021). The chaotic behavior of the memristor in (Kaya, 2020) is used to generate noise. The output of the first stage of the proposed structure in (Kaya, 2020) is a floating-point number. In the second stage, first, the output of the previous stage is multiplied by a fixed number (e.g., xi ), then the result yi = xi mod 64 is calculated and converted to binary format (Kaya, 2020). In the next phase of the second stage, five XOR gates are applied to compute the number of ones in yi . Then, the raw bits of the second stage are applied to the Trivium unit in the third stage, and random outputs are generated bit-by-bit (Kaya, 2020). In the final stage of the proposed structure in (Kaya, 2020), the random outputs are evaluated using the standard NIST statistical tests. Lightweight encryption based on the GIFT block cipher has been applied within the structure of a memristive crossbar array for healthcare applications in (Siddiqi u. a., 2023). The main idea in (Siddiqi u. a., 2023) is to integrate the compact units of the GIFT cipher into the structure of the crossbar array. The assumption in (Siddiqi u. a., 2023) is that the generation and exchange of the encryption key between the transmitter and the receiver are handled by methods implemented in previous works. The XOR gate is one of the basic units of the GIFT cipher, which is implemented in (Siddiqi u. a., 2023) using Scouting Logic and Dual Sense Amplifier crossbar logic (DSA). This enhances parallel computation and is compatible with other parts of the GIFT cipher (Siddiqi u. a., 2023). These two approaches are among the methods integrated into the CIM-P architecture. The LFSR is one of the basic structures for generating random numbers. A 4-bit LFSR implementation algorithm using the IMPLY method within a serial architecture has been developed in (Teimoory u. a., 2015). In the cited paper, the authors designed an algorithm to implement a D-type flip-flop using the IMPLY method, which appears redundant because a memristor can inherently store data like a D flip-flop. However, in (Teimoory u. a., 2015), an algorithm for implementing a 4-bit LFSR using four D flip-flops and an XOR gate is proposed, employing eight memristors. The IMPLY-based LFSR computes the results in 55 computational steps (Teimoory u. a., 2015).

3

Implementation of IMPLY-based Trivium and Grain-128a lightweight stream ciphers

In this section, we first examine how to redesign and implement the standard stream ciphers (Trivium and Grain128a) by applying the serial IMPLY design method for the CIM-A architecture. Then, by applying an innovative method and redesigning the shift register structure, the number of computational steps for both selected stream ciphers is reduced compared to conventional designs, thereby directly reducing energy consumption.

3.1

IMPLY-based Logic Gates Required for the Implementation of Trivium and Grain-128a Lightweight Stream Ciphers

In Figures 1 and 2 and (1)-(8), the architecture and logic equations required for the implementation of Trivium and Grain-128a ciphers can be seen. In the Trivium’s structure, there are nine two-input XOR gates, one three-input XOR gate, and three two-input AND gates. Forty seven two-input XOR gates, 18 two-input AND gates, five three-input AND gates, and two four-input AND gates are the logic blocks that are employed in the Grain-128a cipher. To implement logic gates using the IMPLY method, it should be noted that in some cases it is possible that the input data of the logic gate is overwritten and the original value is deleted, so, in some cases, it is necessary to design and use an algorithm that preserves the value of the input data in the input memristors in these gates. For example, the algorithm introduced in (Bagheralmoosavi u. a., 2025) for implementing a two-input XOR gate incur a loss of one of the inputs. Therefore, to implement a two-input XOR gate, two algorithms are used, as shown in Tables 3 and 4. The first algorithm computes the output in nine computational steps using two work memristors, and one of the input memristors loses its original value (see (10)). The second implementation algorithm obtains the output using three work memristors in eleven computational steps. In the second algorithm, the logic values of the the input memristors are preserved. p XOR q ≡ (p → q) → ((q → p) → 0)

8

(10)

Table 3: Implementation algorithm of a destructive IMPLY-based serial XOR logic gate introduced in (Bagheralmoosavi u. a., 2025) Step 1 2 3 4 5 6 7 8 9

Operation s1 = 0 s2 = 0 ′ a → s1 = s1 ′ b → s2 = s2 ′ ′ ′′ s1 → s2 = s2 s1 = 0 ′′ ′ s2 → s1 = s1 ′ a→b=b ′ ′ ′′ b → s1 = s1

Equivalent Logic FALSE(s1 ) FALSE(s2 ) N OT (a) N OT (b) N OT (a) → N OT (b) FALSE(s1 ) N OT (N OT (a) → N OT (b)) a→b XOR(a, b)

Table 4: Implementation algorithm of a non-destructive IMPLY-based serial XOR logic gate Step 1 2 3 4 5 6 7 8 9 10 11

Operation s1 = 0 s2 = 0 s3 = 0 ′ a → s1 = s1 ′ b → s2 = s2 ′ ′ s2 → s3 = s3 ′ ′ ′′ s1 → s2 = s2 ′ ′′ a → s3 = s3 s1 = 0 ′′ ′ s2 → s1 = s1 ′′ ′ ′′ s3 → s1 = s1

Equivalent Logic FALSE(s1 ) FALSE(s2 ) FALSE(s3 ) N OT (a) N OT (b) b N OT (a) → N OT (b) a→b FALSE(s1 ) N OT (N OT (a) → N OT (b)) XOR(a, b)

Two methods can be used to implement XOR gates with three inputs and beyond. If the designer’s goal is to minimize the computational steps and input data preservation is not important, it is recommended to use the algorithm introduced in (Bagheralmoosavi u. a., 2025), where the XOR result of two inputs is first calculated in nine steps, then in each of the nine steps after that, the XOR result of the previous inputs is XORed with the next input and the final result is calculated. The second method, however, focuses on preserving the input data. In this method, the first two inputs are first XORed using three work memristors according to the algorithm presented in Table 4, and the result is placed in a work memristor. After that, other input is XORed with the result of the previous step using the algorithm written in Table 3. In addition to XOR logic gates, two-input, three-input, and four-input AND gates are required to design the two proposed lightweight ciphers. The two-input AND gate can be implemented in a serial architecture using the IMPLY method, with two work memristors and in five computational steps, as shown in Table 2. For the three-input AND gate, the third input can be ANDed with the result of the AND of the first two inputs, and the result can be stored in one of the work memristors in the final step (the 10th computational step). By examining (11) and Table 5, it can be concluded that by overlapping the computational steps of two two-input AND gates, the number of computational steps of a three-input AND gate can be reduced to six by executing two consecutive AND algorithms. The four-input AND gate can also be implemented in 11 computational steps by executing the algorithms for the three-input AND gate and the two-input AND gate, according to Table 6. AN D(a, b, c) ≡ (c → [((b → (a →)) → 0) → 0]) → 0 ≡ (c → (b → (a → 0))) → 0

3.2

(11)

Proposed method for implementing shift registers required by lightweight stream ciphers

In conventional architectures, it is common to use D flip-flops for implementing shift registers. In the CIM-A architecture, as designers implement different logical and arithmetic circuits with memristors, there is no need to design or use memristive D flip-flops. The purpose of using memristors in the CIM-A architecture is the same as that of a D flip-flop, as memristors are inherently capable of storing data within the crossbar array. Memristors 9

can also be used for data processing. In a shift register, the stored data is shifted bit-by-bit in each cycle from the mth cell to the m + 1th or m − 1th cell, depending on the shift direction. In the LSB or MSB of each register, a new bit is entered in each cycle, computed using a special or random calculation method, or assigned a fixed value. Therefore, in each cycle, a new bit is placed in the register, and one is removed. In a D flip-flop, data is placed on the data input line before the rising edge of the clock pulse, and it is stored after the rising edge. This process is completely different in the IMPLY-based serial architecture considered here. In this architecture, a buffer should be applied to shift each bit from one cell to another. The buffer implementation algorithm (two consecutive inverters, see Table 2) shifts data in four computational steps using two work memristors. Using a buffer to shift data in the shift registers of both the Trivium and Grain-128a ciphers is a reliable method. The disadvantage of this method is that each bit of data is shifted in four computational steps. From Figures 1 and 2, a key point for efficiently designing IMPLY-based shift registers can be identified. In each cycle of processing data in these two ciphers, a few cells in each register are directly involved in the computation, and a large number of bits are only shifted from one cell to its neighbor at the end of each cycle. According to Figure 1, only five bits (bits 66, 69, and 91-93) of Trivium’s shift register A are used in the computations of the two-input XOR, two-input AND, and three-input XOR logic gates. The other bits of this register (in the Trivium cipher structure) are only shifted one position to the right (the more significant bit) at the end of each cycle. In this process, it is important that the aforementioned bits have the correct value in each processing cycle to maintain the security of the cipher. The same principle applies to the other two registers of the Trivium cipher and the two registers of the Grain-128a cipher. Considering the above points, the number of computational steps can be significantly reduced by replacing some buffers with inverters, so that each buffer is replaced by an inverter, reducing the number of computational steps by two. The use of the aforementioned method is accompanied by certain design constraints. Specifically, the correctness of the bits used in the computations must be ensured in every processing cycle. Therefore, an algorithm must be proposed and applied that maintains the integrity of the data stored in the memristors required for logical and arithmetic processing in all cycles. Hence, in addition to improving circuit evaluation criteria such as the number of computational steps, the cipher operates correctly without any functional disruption. The proposed algorithm is as follows: The data located in consecutive memristors and used directly in logical processing (such as two-input XOR gates or two-input AND gates) must be transferred from the main memristor (the first memristor, e.g., k) to its ′ neighbor (the second memristor, e.g., k ) in each cycle using only a buffer. In other words, a work memristor and ′ a target memristor (the second memristor, k ), where the data is stored, are used to transfer the bit of the first memristor to the target memristor in four computational steps using the IMPLY-based buffer implementation algorithm. This process is performed in all processing cycles. For example, consider Figure 1. In shift register ′ A, bits 91 (k) and 92 (k ) are both involved in the computations. In each of the 1152 initialization cycles (in the warm-up phase) or from cycle 1153 onwards, only the buffer implementation algorithm must be used to transfer data from the 91st bit of shift register A to its 92nd bit. The main problem, however, concerns the first memristor (k). If the distance of the input data from the first memristor (k) is an odd number of memristors (for example, a distance of one, three, five, etc.), the desired value is shifted with the buffer in the cycle preceding the data shift. Here, the input data is either the logical value stored in the memristor used in logical and arithmetic computations during the previous cycle, or the input bit of the register. For example, in the Trivium’s first shift register (A), ′′ the input memristor for the 91st cell (k) in the proposed algorithm is the output of the 66th memristor (k ), ′′′ and for the 65th cell (k ), the input of the register is considered as the input. The distance of the data from the target memristor is of great importance. If the distance is equal to “2 × n − 1” bits and ‘n’ buffers and ‘n’ inverters are used, then only an inverter is used in the data shifts of cycle “2 × n + 1” onwards. For example, if Table 5: Implementation algorithm of a three-input IMPLY-based serial AND gate Step 1 2 3 4 5 6

Operation s1 = 0 s2 = 0 ′ a → s1 = s1 ′ ′′ b → s1 = s1 ′′ ′′′ c → s1 = s1 ′′′ ′ s1 → s2 = s2

10

Equivalent Logic FALSE(s1 ) FALSE(s2 ) N OT (a) b → N OT (a) c → (b → N OT (a)) AN D(a, b, c)

Table 6: Implementation algorithm of a four-input IMPLY-based serial AND gate Step 1 2 3 4 5 6 7 8 9 10 11

Operation s1 = 0 s2 = 0 ′ a → s1 = s1 ′ ′′ b → s1 = s1 ′′ ′′′ c → s1 = s1 ′′′ ′ s1 → s2 = s2 s1 = 0 ′ d → s1 = s1 ′′ ′ ′ s2 → s1 = s1 s2 = 0 ′′ ′ s1 → s2 = s2

Equivalent Logic FALSE(s1 ) FALSE(s2 ) N OT (a) b → N OT (a) c → (b → N OT (a)) AN D(a, b, c) FALSE(s1 ) N OT (d) AN D(a, b, c) → N OT (d) FALSE(s2 ) AN D(a, b, c, d)

Table 7: A toy example of an 8-bit conventional (buffer-based) and an 8-bit proposed (inverter-based) shift register with a distance of 5 cells Conventional buffer-based 8-bit shift register Cycle Input 0 1 2 3 4 5 6 7 0 1 1 0 1 1 1 0 1 0 1 0 1 1 0 1 1 1 0 1 2 1 0 1 1 0 1 1 1 0 3 0 1 0 1 1 0 1 1 1 4 0 0 1 0 1 1 0 1 1 5 1 0 0 1 0 1 1 0 1 6 1 1 0 0 1 0 1 1 0 7 0 1 1 0 0 1 0 1 1 8 1 0 1 1 0 0 1 0 1 9 0 1 0 1 1 0 0 1 0 10 1 0 1 0 1 1 0 0 1 11 1 1 0 1 0 1 1 0 0 12 1 1 0 1 0 1 1 0

Proposed inverter-based 8-bit shift register Cycle Input 0 1 2 3 4 5 6 7 0 1 1 0 1 1 1 0 1 0 1 0 0 0 1 0 0 1 0 0 2 1 1 1 1 0 1 1 1 1 3 0 0 0 0 0 1 1 1 0 4 0 1 1 1 1 1 0 1 0 5 1 1 0 0 0 0 1 0 0 6 1 0 0 1 1 1 1 1 1 7 0 0 1 1 0 0 0 1 0 8 1 1 1 0 0 1 1 0 0 9 0 0 0 0 1 1 0 1 1 10 1 1 1 1 1 0 0 0 0 11 1 0 0 0 0 0 1 0 1 12 0 1 1 1 1 1 1 1

the distance between the input data and the target memristor is nine memristors, ‘n’ is equal to five, meaning that the combination of buffers and inverters (respectively) is used for 10 consecutive cycles, and from the 11th cycle onwards, only the inverter is used. If the number of memristors placed between the input data and target memristor is “2×n”, the buffers and inverters are used ‘n’ times, and from the “2×n+1th ” cycle onwards, only the buffer is applied to shift data. If the distance between the input data and the target memristor is 10 memristors, the buffers and inverters are used five times, and only the buffer is used to shift data from the 11th cycle onwards. ′ The other bits stored in the memristors before the first memristor (k) and after the second memristor (k ) are shifted only by the inverters (see the inverter implementation algorithm in Table 2), and the shifted values are stored in the adjacent memristors. If the input data is directly stored in the first memristor, only the buffer should be used. For a better understanding of the proposed algorithm, please refer to Tables 7 and 8. In these toy examples, the conventional and proposed shift register implementation algorithms are shown over 12 cycles. In Tables 7 and 8, ‘n’ equals 5 and 6, respectively. In these tables, the columns marked in green refer ′ to the memristors involved in the computations (k and k ).

3.3

Step-by-step implementation of the Trivium and Grain-128a lightweight stream ciphers applying the IMPLY method in a serial architecture

In this subsection, after examining the implementation details of the basic logical and arithmetic blocks from the previous subsections, the implementation of each IMPLY-based lightweight stream cipher, the Trivium and Grain-128a, is analyzed.

11

Table 8: A toy example of an 8-bit conventional (buffer-based) and an 8-bit proposed (inverter-based) shift register with a distance of 6 cells Conventional buffer-based 8-bit shift register Cycle Input 0 1 2 3 4 5 6 7 0 1 1 0 1 1 1 0 1 0 1 0 1 1 0 1 1 1 0 1 2 1 0 1 1 0 1 1 1 0 3 0 1 0 1 1 0 1 1 1 4 0 0 1 0 1 1 0 1 1 5 1 0 0 1 0 1 1 0 1 6 1 1 0 0 1 0 1 1 0 7 0 1 1 0 0 1 0 1 1 8 1 0 1 1 0 0 1 0 1 9 0 1 0 1 1 0 0 1 0 10 1 0 1 0 1 1 0 0 1 11 1 1 0 1 0 1 1 0 0 12 1 1 0 1 0 1 1 0

3.3.1

Proposed inverter-based 8-bit shift register Cycle Input 0 1 2 3 4 5 6 7 0 1 1 0 1 1 1 0 1 0 1 0 0 0 1 0 0 0 0 1 2 1 1 1 1 0 1 1 1 0 3 0 0 0 0 0 1 0 1 1 4 0 1 1 1 1 1 0 1 1 5 1 1 0 0 0 0 0 0 1 6 1 0 0 1 1 1 1 1 0 7 0 0 1 1 0 0 0 1 1 8 1 1 1 0 0 1 1 0 1 9 0 0 0 0 1 1 0 1 0 10 1 1 1 1 1 0 0 0 1 11 1 0 0 0 0 0 1 0 0 12 0 1 1 1 1 1 1 0

Step-by-step implementation algorithm of the IMPLY-based Trivium cipher

The implementation of this cipher is divided into three parts. The first part deals with the initialization of the shift registers using random numbers, which, as in (Siddiqi u. a., 2023), is assumed by default in this research. The second part is the Trivium initialization phase, where data is calculated and processed 1152 times, and the third part starts at cycle 1153, where one bit of the keystream is generated in each cycle. To store data and implement this cipher using IMPLY logic, 288 memristors are required for implementing the registers, along with five work memristors s0 -s4 and one output memristor. In the computations related to register A in the structure shown in Figure 1, two two-input XOR gates and one two-input AND gate are required, yielding a total of 23 computational steps. Three work memristors are used in the computations of this part, and the results are stored in s0 and s2 . At the end of this stage, the logical values of s1 and A93 are no longer required, so these memristors can be used to store intermediate values in the next stages. The computations for registers B and C, similar to the computations for register A, are performed using two two-input XOR gates and one two-input AND gate, and the number of computational steps required for each of these registers is also 23 steps. The data required for subsequent processing steps related to registers B and C are stored in memristors A93 and s3 (for register B) and in memristors B84 and s4 (for register C). After each processing cycle of the Trivium cipher, a two-input XOR gate is used to calculate the input bit for each of the registers (A1 , B1 , and C1 ). Hence, three two-input XOR gates are required to compute the values to be stored in the first memristor of each register, requiring a total of 27 computational steps for this stage without the need for any new work memristors. After computing the input bits of the registers, the implementation algorithm for the two-input XOR gate must be executed twice to compute the output si , and this is also completed in 18 computational steps. The remainder of this subsection explains the details of how data is shifted bit-by-bit in the shift registers of the Trivium lightweight stream cipher. In the previous subsection, the proposed method for implementing fast, energy-efficient shift registers using IMPLY-based inverters and buffers was described in detail. This method is applied in all three shift registers of the Trivium cipher’s structure, as shown in Figure 1. The functionality of the shift register using the proposed method can be examined and observed in detail in Tables 9-11. As an example, the details of how the stored bits in the memristors of register A are shifted are explained here. Based on the contents of Table 9, it can be concluded that two buffers and 88 inverters must be used to shift data directly in each processing cycle. Special bits are those stored in the memristors before they are involved in logical and arithmetic computations; shifting data in these memristors must be performed carefully. Accordingly, three buffers and three inverters in the first and second cycles, 40 buffers and 20 inverters in the third to 22nd cycles, and 66 inverters and 66 buffers in the 23rd to 66th cycles are required to implement shift register A based on the proposed method, correctly shifting bits to specific memristors. From the 67th cycle onwards, two inverters and one buffer are required in each cycle to transfer data in shift register A. The proposed method of shifting data 12

Table 9: Details of the bit-by-bit data shifting process in Trivium’s shift register A based on the proposed method Memristors involved in data shifting A92 → A93 A91 → A92 A90 → A91 A70 → A90 A69 → A70 A68 → A69 A67 → A68 A66 → A67 A65 → A66 A1 → A65 Input → A1

Applied circuits 1 buffer 1 buffer Buffers and inverters 20 inverters 1 inverter Buffers and inverters 1 inverter 1 inverter Buffers and inverters 64 inverters 1 inverter

Additional Details

21 intervening memristors; 11 buffers/inverters, then inverters only.

2 intervening memristors; 1 buffers/inverters, then buffers only.

65 intervening memristors; 33 buffers/inverters, then inverters only.

Table 10: Details of the bit-by-bit data shifting process in Trivium’s shift register B based on the proposed method Memristors involved in data shifting B83 → B84 B82 → B83 B81 → B82 B80 → B81 B79 → B80 B78 → B79 B77 → B78 B70 → B77 A69 → A70 B68 → B69 B1 → B68 Input → B1

Applied circuits 1 buffer 1 buffer Buffers and inverters 1 inverter 1 inverter 1 inverter Buffers and inverters 7 inverters 1 inverter Buffers and inverters 67 inverters 1 inverter

Additional Details

3 intervening memristors; 2 buffers/inverters, then inverters only.

8 intervening memristors; 4 buffers/inverters, then buffers only. 2 intervening memristors; 1 buffers/inverters, then buffers only. 68 intervening memristors; 34 buffers/inverters, then buffers only.

in shift register A was examined in detail. A similar procedure is employed in the other two shift registers (B and C). Table 12 summarizes the number of inverters and buffers for different cycles of each register used in the Trivium lightweight stream cipher. 3.3.2

Step-by-step implementation algorithm of the IMPLY-based Grain-128a cipher

The lightweight Grain-128a stream cipher also performs computations in three stages, like the Trivium cipher. The first stage initializes the shift registers, which will not be discussed in this research. The second stage, called the “pre-initialization” stage, is performed in 256 processing cycles, and the keystream is generated from the 257th round, which is the third stage of implementing this lightweight stream cipher. In this subsection of the article, each part of the second and third stages of the Grain-128a cipher will be examined and implemented based on stateful IMPLY logic according to Figure 2. Two hundred and sixty two memristors are needed to implement this cipher, including two sets of 128 memristors for implementing 128-bit LFSR and 128-bit NFSR, along with six work memristors. If the designer needs to store the output (y) separately in a specific memristor, the number of memristors increases to 263 (see Figure 2). To calculate the output (y), two values, h(x) and a set of b − terms, must be calculated according to (7) and (8). The final result is obtained by performing the two-input XOR algorithm of the values stored in the 93rd bit of the LFSR with the two-input XOR product of the calculated values h(x) and b − terms. The logical computations of the function h(x) include four two-input AND gates, four two-input XOR gates, and one three-input AND gate, which are performed in 62 computational steps, and the final result is stored in the first work memristor (s0 ). The implementation of the computations related to b − terms is executed using six two-input XOR logic gates, one of which is implemented by applying the algorithm of Table 4, XOR(b2 , b15 ), and the other gates are calculated using the algorithm of Table 3. The important point in these computations is that the values stored in the main memristors remain completely intact. The output of these calculations is obtained after 56 computational 13

Table 11: Details of the bit-by-bit data shifting process in Trivium’s shift register C based on the proposed method Memristors involved in data shifting C110 → C111 C109 → C110 C108 → C109 C88 → C108 C87 → C88 C86 → C87 C67 → C86 C66 → C67 C65 → C66 C1 → C65 Input → C1

Applied circuits 1 buffer 1 buffer Buffers and inverters 20 inverters 1 inverter Buffers and inverters 19 inverter 1 inverter Buffers and inverters 64 inverters 1 inverter

Additional Details

21 intervening memristors; 11 buffers/inverters, then inverters only.

20 intervening memristors; 10 buffers/inverters, then buffers only.

65 intervening memristors; 33 buffers/inverters, then inverters only.

Table 12: Inverter and buffer counts for Trivium’s shift registers A, B, and C per cycle Cycles

No. of No. of buffers inverters Shift register A 1-2 7 179 3-22 80 1780 23-66 154 3938 67 3 90 68 and onwards 3 90

Cycles

No. of No. of buffers inverters Shift register B 1-2 7 161 3-4 7 179 5-8 12 324 9-68 210 4830 69 and onwards 4 80

Cycles

No. of No. of buffers inverters Shift register C 1-20 70 2150 21-22 8 214 23-66 154 4730 67 3 108 68 and onwards 3 108

steps. The output of this step is also stored in a work memristor (s2 ). The final output is also calculated using two two-input XOR gates in 18 computational steps, as shown in (8). The computations related to the LFSR feedback, performed in the “pre-initialization” and keystream stages according to (5), require five two-input XOR logic gates and a total of 47 computational steps. The first XOR is calculated using the algorithm of Table 4, and the other XORs are calculated using the algorithm of Table 3 to preserve the input values. In the second and third stages of the Grain-128a cipher implementation, the NFSR feedback calculation is divided into two parts. It has a higher computational complexity than the other logical parts of this cipher. The first part computes the subsections mentioned in (6), which involve several XOR and AND gates. The first part specifically includes the computation of 15 two-input XOR gates: 14 are computed using the algorithm in Table 3, and one is computed using the algorithm in Table 4. Eight two-input AND gates (five computational steps) and three three-input AND gates (six computational steps) are also used to calculate (6). It should be noted that all the main memristors remain intact after these calculations, with 195 computational steps. The final result of this part is also stored in one of the work memristors (s2 ). In the “pre-initialization” stage of the Grain-128a lightweight stream cipher, which is the first 256 rounds of its execution, the output of each round (the same value stored in the first work memristor (s0 )) is XORed as one of the inputs of two two-input XOR gates with the feedbacks of the registers in this cipher, and the results of each XOR are considered as the input bits of each register. If the data shift in the Grain-128a’s LFSR and NFSR, which consists of 256 bits in total, is performed using IMPLY-based buffers, it takes 1024 computational steps per cycle. By using the proposed method and replacing inverters with buffers in these registers, the computational steps per cycle can be significantly reduced. The implementation of Grain-128a’s shift registers is much more complex than the registers in the Trivium cipher. A similar approach to Tables 9-11 is also applied to determine the details of performing data shifts in the Grain-128a cipher shift registers. According to the proposed approach, 114 inverters and two buffers are required to transfer data from each memristor to its neighbor in the LFSR per cycle. The problem’s complexity, however, stems from the other 12 shifts directly involved in the computations. The number of inverters and buffers applied to shift bits that contribute to computations each cycle is tabulated in Table 13. A similar approach should be applied to examine the bit-by-bit shift process in an NFSR to reduce its hardware complexity significantly. In all processing cycles, 92 inverters and 12 buffers are required to shift data between memristors. The shift of the bits stored

14

Table 13: Number of buffers and inverters required per cycle for the bits involved in the logical computations of the Grain-128a’s LFSR Cycles 1-2 7-8 13-18

No. of buffers 12 12 33

No. of inverters 12 12 39

Cycles 3-4 9-10 19-32

No. of buffers 10 13 63

No. of inverters 14 11 105

Cycles 5-6 11-12 33 and onwards

No. of buffers 10 12 4

No. of inverters 14 12 8

Table 14: Number of buffers and inverters required per cycle for the bits involved in the logical computations of the Grain-128a’s NFSR Cycles 1-2 7-8 13-32

No. of buffers 24 16 15

No. of inverters 24 32 33

Cycles 3-4 9-10 33-34

No. of buffers 18 15 16

No. of inverters 30 33 32

Cycles 5-6 11-12 35 and onwards

No. of buffers 16 15 8

No. of inverters 32 33 16

in the memristors that participate in the logical computations follows the specific characteristics of the proposed algorithm. A summary of the number of buffers and inverters required per cycle for the 24 bits involved in the logical computations of the NFSR is tabulated in Table 14.

4

Simulation results, discussion, and evaluation

In the previous section, the implementation of IMPLY-based Trivium and Grain-128a, two lightweight stream ciphers, in the serial architecture was investigated and reported. In subsection 4.1, first, the memristor model applied in circuit-level simulation is introduced; second, the method of applying this model to perform the simulation is explained; and in the last part of this subsection, the simulation results are presented and expounded in detail. In the second subsection, the behavioral simulation results for the proposed encryption/decryption structures (based on Trivium and Grain-128a) and their application to image steganography will be reported.

4.1

Circuit-level simulation, method, and results

For simulating memristor-based logical and arithmetic circuits, various behavioral and physical SPICE models have been presented. The logical and arithmetic blocks applied in the structure of Trivium and Grain-128a in this paper have been simulated by applying the modified open source Voltage ThrEshold Adaptive Memristor (VTEAM) model, which has been fitted according to the parameters of a discrete memristor (BS-AF-W model) developed by “Knowm inc.” (Fatemieh u. a., 2025b). This model has been applied in (Fatemieh u. a., 2025b; Fatemieh und Reshadinezhad, 2026; Bagheralmoosavi u. a., 2025; Fatemieh u. a., 2023; Asgari u. a., 2024, 2025). The model parameters and standard coefficients are shown in Table 15. The specific parameters considered for the simulation of the IMPLY-based blocks are listed in Table 15, too. The values reported in Table 15 are contemplated based on the parameters reported in (Fatemieh u. a., 2025b; Fatemieh und Reshadinezhad, 2026; Bagheralmoosavi u. a., 2025; Fatemieh u. a., 2023; Asgari u. a., 2024, 2025). LTspice has been used to simulate the IMPLY-based logical and arithmetic blocks described in this paper. Table 15: Setup values of IMPLY logic and VTEAM model (Fatemieh u. a., 2025b; Fatemieh und Reshadinezhad, 2026; Bagheralmoosavi u. a., 2025). Parameter vof f αon kon won aon vcond

Value 0.7 V 3 -0.5 nm s 3 nm 0 nm 900 mV

Parameter von Rof f kof f wC vset RG

15

Value -10 mV 1 MΩ 1 cm s 107 pm 1V 40 KΩ

Parameter αof f Ron wof f aof f vreset tpulse

Value 3 10 kΩ 0 nm 3 nm 1V 30 µs

Figure 5: The output waveform of a destructive IMPLY-based serial two-input XOR gate introduced in (Bagheralmoosavi u. a., 2025) for input state “00”. To analyze the functionality of the simulated circuits, note that, since all logic and arithmetic blocks are stateful, the input and output logical values are determined by the memristors’ resistance. The maximum and minimum resistance values of the memristor are equal to logic ’0’ and logic ’1’, respectively. It should be noted that the functionality of IMPLY-based implementation algorithms must be evaluated by considering all possible input states. First, the input values must be initialized to the input memristors. Then the implementation algorithm runs for each state, and the correctness of the output is evaluated across all states against the intended resistance value of the output memristor. The number of computational steps, which is proportional to the computational delay, determines the performance. The energy consumption is computed using the method described in (Fatemieh u. a., 2025b; Fatemieh und Reshadinezhad, 2026; Bagheralmoosavi u. a., 2025; Fatemieh u. a., 2023; Asgari u. a., 2024, 2025; Fatemieh u. a., 2025a). For each input state, the energy consumption of each memristor in the row or column of the array is calculated, and the average over all states is reported as the estimated energy consumption. The number of memristors required is another criterion examined. In the following, the energy consumption, number of computational steps, and number of memristors required for each basic logic and arithmetic unit are calculated after verification. Then, the circuit evaluation criteria for the IMPLY-based Trivium and Grain-128a stream ciphers are reported in two scenarios. In the first scenario for each of the two stream ciphers, conventional algorithms are used without the improvements detailed in the previous section, whereas in the second scenario, the proposed methods are used to improve the circuit evaluation criteria. As mentioned, the functionality of the implementation algorithms for all IMPLY-based logic gates and basic arithmetic circuits used in the design of the Trivium and Grain-128a ciphers has been examined by considering all possible input states and applying the model in the LTSPICE simulator. Two-input XOR gates (according to the algorithms of Tables 3 and 4), three-input XOR gate with the ability to retain input values, two-input, three-input, and four-input AND gates, buffers, and inverters are the circuits that are evaluated. The output waveforms of three gates, as examples of simulation results, are shown in Figures 5-7. In Figure 5, the output waveform of a two-input XOR logic gate in which the logic value of one of the input memristors is changed is plotted. The inputs are assumed to be equal to “00”, and the output is stored in the output memristor (the first work memristor, s1 ) after nine computational cycles. According to Table 15 and the reported value of tstep , the output is computed and stored in 240-270 µs. The analysis of the functionality of the other two-input XOR gate mentioned in Table 3, which computes the output in the last step of the implementation algorithm (300-330 µs), is shown in Figure 6. The inputs of this gate equal “00”, and the input values stored in the memristors remain unchanged after the execution of the implementation algorithm. The output is stored in the first work memristor (s1 ) in the final step. In Figure 7, the simulation result of the three-input AND gate (Ain Bin Cin =“111”) is calculated in six computational steps according to (11) in the time interval 0 to 180 µs. The output is obtained in the last computational step (150-180 µs) and stored in the second work memristor (s2 ). After introducing the applied model, familiarizing with the simulation method and circuit evaluation criteria,

16

Figure 6: The output waveform of a non-destructive IMPLY-based serial two-input XOR gate for input state “00”.

Figure 7: The output waveform of a IMPLY-based serial three-input AND gate for input state “111”. and presenting a number of simulation outputs, the results of the circuit evaluation criteria for the IMPLY-based basic logic gates required in the implementation of the Trivium and Grain-128a, lightweight stream ciphers, are summarized in Table 16. The numbers reported in Table 16 are widely used in estimating the circuit evaluation criteria. Subsection 3.3.1 explains the implementation details of the Trivium cipher using the IMPLY method. The previous paragraphs of this subsection also mentioned the details of the circuit-level implementation of the basic building blocks of this cipher. Tables 17 and 18 specify the number of required computational units, the number of computational steps, and the estimated energy consumption for the initialization and keystream generation stages of the Trivium cipher, respectively. The initialization phase of the Trivium cipher requires 293 memristors, including 288 input/output memristors and five work memristors, and the number of memristors is not increased in all cycles. Overall, the initialization phase of the IMPLY-based Trivium cipher takes 797266 computational steps and consumes 53.4731 µJ. Each round of generating a keystream bit also requires the same number of memristors, takes 710 computational steps, and consumes 47.8983 nJ. So, if the number of keystream bits required by the user is considered to be ‘n’, the total number of computational steps and energy consumption estimation in the two phases are “710 × n + 797266” steps and “0.0478 × n + 53.4731” µJ, respectively. The circuit evaluation criteria for the Grain-128a cipher can be examined in the same way as for the Trivium cipher. The “pre-initialization” phase of this cipher lasts 256 cycles, and the keystream is generated from the 257th 17

Table 16: Summary of the circuit evaluation criteria for the IMPLY-based basic logic gates required in the implementation of the Trivium and Grain-128a. Gate Inverter Buffer 2-input AND 3-input AND 4-input AND 2-input XOR (destructive) (Bagheralmoosavi u. a., 2025) 2-input XOR (non-destructive) 3-input XOR

No. of Memristors 2 3 4 5 6 4

No. of computational steps 2 4 5 6 11 9

Estimated energy consumption (nJ) 0.1291 0.269 0.3833 0.5025 0.9131 0.7426

5

11

0.9146

6

20

1.711

Table 17: Details of the number of computational blocks, computational steps, and estimated energy consumption of the initialization stage of the serial IMPLY-based Trivium cipher. Block 2-input XOR (Bagheralmoosavi u. a., 2025) 2-input AND Inverter in the proposed algorithm (shift register A) Buffer in the proposed algorithm (shift register A) Inverter in the proposed algorithm (shift register B) Buffer in the proposed algorithm (shift register B) Inverter in the proposed algorithm (shift register C) Buffer in the proposed algorithm (shift register C) Buffer in the conventional algorithm (shift register A) Buffer in the conventional algorithm (shift register B) Buffer in the conventional algorithm (shift register C)

No. of blocks 3456 10386 103637

No. of computational steps 17280 93312 207274

Estimated energy consumption (µJ) 1.3246 7.699 13.3795

3499

13996

0.9412

92196

184392

11.9025

4572

18288

1.2298

124382

248764

16.0577

3490

13960

0.9388

107136

428544

28.8195

96768

387072

26.0305

127872

511488

34.3975

cycle. The number of memristors and logical units, the computational steps, and the estimated energy consumption of the units applied in the “pre-initialization” and keystream generation phases are reported in Tables 19 and 20, respectively. Two hundred and sixty two memristors participate in the “pre-initialization” phase of the Grain-128a cipher. This phase requires 245830 computational steps and estimates an energy consumption of 17.6811 µJ. Each round of keystream generation requires 942 computational steps and 263 memristors. The energy consumption estimation of the second stage of the Grain-128a is 0.0666 µJ. The computed values can be generalized for an n-bit keystream. To produce an n-bit keystream of this cipher, “942×n+245830” computational steps and 263 memristors are required. The energy consumption estimate for this cipher in generating an n-bit keystream is also “0.0666 × n + 17.6811” µJ. A comprehensive comparison of the circuit evaluation criteria for the Trivium and Grain-128a ciphers in the mentioned scenarios is presented in Table 21. The number of memristors required by both ciphers to generate an n-bit keystream is 293 and 263, respectively. The number of computational steps and energy consumption of the Trivium cipher applying the proposed shift register’s implementation algorithm, considering the generation of 10,000 bits, are 18% and 22% less than that of the Grain-128a cipher, respectively, while by increasing the number of output bits by a factor of 10, the number of computational steps and energy consumption estimation

18

Table 18: Details of the number of computational blocks, computational steps, and estimated energy consumption of the keystream generation stage of the serial IMPLY-based Trivium cipher. Block 2-input XOR (Bagheralmoosavi u. a., 2025) 2-input AND Inverter in the proposed algorithm Buffer in the proposed algorithm Buffer in the conventional algorithm

No. of blocks 11 3 278

No. of computational steps 99 15 556

Estimated energy consumption (nJ) 8.1686 1.1499 35.8898

10

40

2.69

288

1152

77.472

Table 19: Details of the number of computational blocks, computational steps, and estimated energy consumption of the “pre-initialization” stage of the serial IMPLY-based Grain-128a cipher. Block 2-input XOR (destructive) (Bagheralmoosavi u. a., 2025) 2-input XOR (non-destructive) 2-input AND 3-input AND 4-input AND Inverter in the proposed algorithm (LFSR) Buffer in the proposed algorithm (LFSR) Buffer in the conventional algorithm (LFSR) Inverter in the proposed algorithm (NFSR) Buffer in the proposed algorithm (NFSR) Buffer in the conventional algorithm (NFSR)

No. of blocks 7936

No. of computational steps 71424

Estimated energy consumption (µJ) 5.8932

768

8448

0.7024

2816 768 256 31195

14080 4608 2816 62390

1.0793 0.3859 0.2237 4.0272

1573

6292

0.4231

32768

131072

8.8145

27650

55300

3.5696

5118

20472

1.3767

32768

131072

8.8145

are improved by 24% and 27%. According to Table 21, it can be concluded that the computational delay and energy consumption of the Trivium cipher are lower than those of Grain-128a, and it is preferred over Grain-128a if computational delay and energy consumption are of higher priority for the designer. Grain-128a has higher priority than Trivium if the designers’ main goal is to use a cipher with a 128-bit key and the circuit-level evaluation metrics are second-tier. As the number of bits tends to infinity, the number of computational steps and the energy consumption of the improved IMPLY-based Trivium and Grain-128a lightweight stream ciphers based on the proposed shift register algorithm, are improved by 38%, 42%, 45%, and 32% compared to the conventional design, respectively.

4.2

Application-level simulation

GNU Octave has been used to simulate the behavior of the serial IMPLY-based Trivium and Grain-128a stream ciphers in cryptography and steganography. All the logic gates and logical/arithmetic units have been simulated behaviorally using the algorithms mentioned in the previous sections, and their functionality has been validated. The proposed IMPLY-based shift register algorithm has also been considered in the behavioral simulation of both ciphers’ implementations. The core structures of each cipher were implemented using basic logical circuits designed with the IMPLY method. After verifying the correctness of the cores’ functionality, both ciphers were implemented in their two main stages of operation: initialization and keystream generation. In the final code, the desired keystream was generated, and encryption and decryption operations were performed based on the

19

Table 20: Details of the number of computational blocks, computational steps, and estimated energy consumption of the keystream generation stage of the serial IMPLY-based Grain-128a cipher. Block 2-input XOR (destructive) (Bagheralmoosavi u. a., 2025) 2-input XOR (non-destructive) 2-input AND 3-input AND 4-input AND Inverter in the proposed algorithm (LFSR) Buffer in the proposed algorithm (LFSR) Buffer in the conventional algorithm (LFSR) Inverter in the proposed algorithm (NFSR) Buffer in the proposed algorithm (NFSR) Buffer in the conventional algorithm (NFSR)

No. of blocks 29

No. of computational steps 261

Estimated energy consumption (nJ) 21.5354

3

33

2.7438

11 3 1 122

55 18 11 244

4.2163 1.5075 0.9131 15.7502

6

24

1.614

128

512

34.432

108

216

13.9428

20

80

5.38

128

512

34.432

Table 21: Comprehensive comparison of circuit evaluation metrics of IMPLY-based serial Trivium and Grain128a ciphers. Lightweight stream cipher Trivium (conventional) Trivium (proposed) Grain-128a (conventional) Grain-128a (proposed)

No. of computational steps 1152 × n + 1437696 710 × n + 797266 1646 × n + 363520 942 × n + 245830

n = 10000

n = 100000

12957696 7897266 16823520 9665830

116637696 71797266 164963520 94445830

Estimated energy consumption (µJ) 0.0867 × n + 98.2711 0.0478 × n + 53.4731 0.09878 × n + 25.9135 0.0666 × n + 17.6811

n = 10000

n = 100000

965.2711 531.4731 1013.7135 683.6811

8768.2711 4833.4731 9903.9135 6677.6811

structures of both ciphers. One goal of implementing lightweight ciphers based on the IMPLY design method is to evaluate their functionality in steganography applications. Hiding data bits in the LSB of each pixel in an image is one of the most common and low-cost methods of steganography (Debnath u. a., 2018). Therefore, to examine the functionality of the implemented circuits in this application, random messages were encrypted using the designed IMPLY-based Trivium and Grain-128a, and the output bits were inserted one bit at a time into the LSBs of the carrier images. The output messages from the previous step were also recovered from the images, and after comparing the recovered data with the original data, it can be deduced that the entire process was performed correctly. Assessing image quality metrics such as Peak Signal to Noise Ratio (PSNR) and comparing the histograms of the steganographic output with the original image, ensures the reliability of the process. The human eye cannot distinguish the original image from the steganographic output if the PSNR is above 30 dB (Mittal, 2016). Hence, if the PSNR is higher than 30 dB, it can be concluded that the steganographic output is acceptable, as it is not possible to easily distinguish the two images from each other, and even human recognition is impossible. By drawing histograms of the original image and the output of the steganography process, and analyzing the outputs, the user can determine how well this technique works for hiding data. According to the method applied to the behavioral simulation discussed in the previous paragraphs, five random messages of varying lengths were generated and, after processing, placed in the LSBs of the pixels of five standard 256×256 grayscale images. The output of the steganography process consists of two images, along with the standard grayscale image. The first and second output images are generated from the outputs of the lightweight stream ciphers Trivium and Grain-128a, respectively. Two sets of outputs from the steganography application are shown in Figures 8 and 9. The PSNR values of this application is tabulated in Table 22. According to the results in Table 22 and Figures 8 and 9, it can be concluded that hiding the outputs of

20

(a)

(b)

(c)

Figure 8: Functionality analysis of IMPLY-based lightweight stream ciphers in steganography:(a) Cover image (“boat”), (b) Trivium-based stego image, and (c) Grain-128a stego image.

(a)

(b)

(c)

Figure 9: Functionality analysis of IMPLY-based lightweight stream ciphers in steganography:(a) Cover image (“circuit”), (b) Trivium-based stego image, and (c) Grain-128a stego image. Table 22: Evaluation of the PSNR values of the stego images generated using the IMPLY-based Trivium and Grain-128a ciphers. Cover image Boat Cameraman Circuit Rice Walkbridge

Trivium-based stego image PSNR value (dB) 71.642 69.822 69.715 73.199 65.139

Grain-128a-based stego image PSNR value (dB) 71.966 69.668 69.783 73.178 65.258

random messages of different lengths encrypted using the Trivium and Grain-128a ciphers in the LSBs of standard grayscale images is acceptable. As the number of encrypted bits increases, image quality decreases, but the results show a large gap between the output PSNR values and the acceptable image quality margin (30 dB). By increasing the number of bits, larger images can be used to improve image quality. Histograms corresponding to Figures 8 and 9 are also shown in Figures 10 and 11; thus, the pixel distributions of the outputs of the steganography application can be compared and evaluated with the reference images.

5

Conclusion

In this paper, two lightweight stream ciphers, Trivium and Grain-128a, are implemented using the serial IMPLYbased design method with a minimum number of memristors harmonized with the structure of the memristive crossbar array. In the redesigned ciphers, an attempt has been made to minimize the number of computational steps. Accordingly, in addition to investigating the possibility of overlapping steps in the implementation of basic logical and arithmetic blocks, by replacing the buffers required in the structure of shift registers of the architectures of the Trivium and Grain-128a ciphers with a combination of inverters and buffers, this study attempts to significantly improve the circuit evaluation criteria, such as the number of computational steps and energy consumption. Applying the proposed method to redesign the shift register structures in these two lightweight stream ciphers has 21

(a)

(b)

(c)

Figure 10: Histograms of (a) Figure 8(a), (b) Figure 8(b), and (c) Figure 8(c).

(a)

(b)

(c)

Figure 11: Histograms of (a) Figure 9(a), (b) Figure 9(b), and (c) Figure 9(c). resulted in improvements of 38%-42% in the number of computational steps and 32%-44% in energy consumption compared to the conventional design. The improved structure of the IMPLY-based Trivium cipher generates the keystream faster and consumes less energy than Grain-128a, such that by generating 10,000 valid bits employing the improved IMPLY-based Trivium cipher, the number of computational steps and energy consumption are 18% and 22% less than circuit evaluation criteria of the IMPLY-based Grain-128a architecture, respectively. By ex22

amining the PSNR criterion and the histogram diagrams of the different output images from the steganography application, it can be concluded that the functionality of the designed circuits in this application was acceptable.

Acknowledgment During the preparation of this work, the authors utilized the “Gemini 3.1 Pro” to improve the readability, grammar, and formal academic tone of the manuscript. After using this tool, the authors meticulously reviewed and edited the content. The authors take full responsibility for the scientific accuracy, originality, and overall integrity of the final publication.

Author Contributions Seyed Erfan Fatemieh: Conceptualization, Data Curation, Formal Analysis, Investigation, Methodology, Software, Validation, Visualization, Writing-original draft, Writing-review & editing. Reza Shahdi Alizadeh: Project Administration, Resources, Supervision, Validation, Writing-review & editing. Esmail Zarezadeh: Project Administration, Supervision, Validation, Writing-review & editing.

Data Availability Statement Data is contained within the article.

Statements & Declarations Competing interests The authors declare no conflict of interest.

Funding No funds, grants, or other support was received.

References [Ågren u. a. 2011] Ågren, Martin ; Hell, Martin ; Johansson, Thomas ; Meier, Willi: Grain-128a: a new version of Grain-128 with optional authentication. In: International Journal of Wireless and Mobile Computing 5 (2011), Nr. 1, S. 48–59 [Alhaj Ali 2020] Alhaj Ali, Khaled: New design approaches for flexible architectures and in-memory computing based on memristor technologies, Ecole nationale superieure Mines-Telecom Atlantique, Dissertation, 2020. – Available at https://theses.hal.science/tel-03134905/file/2020IMTA0197_AlhajAli-Khaled.pdf [Asgari u. a. 2024] Asgari, Samane ; Reshadinezhad, Mohammad R. ; Fatemieh, Seyed E.: Energy-efficient and fast IMPLY-based approximate full adder applying NAND gates for image processing. In: Computers and Electrical Engineering 113 (2024), S. 109053 [Asgari u. a. 2025] Asgari, Samane ; Reshadinezhad, Mohammad R. ; Fatemieh, Seyed E.: Energy-efficient and Fast IMPLY-Based Approximate 4: 2 Compressor Applying NAND Gates for Image Processing. In: Available at SSRN 5375944 (2025) [Bagheralmoosavi u. a. 2025] Bagheralmoosavi, Bahareh ; Fatemieh, Seyed E. ; Reshadinezhad, Mohammad R. ; Rubio, Antonio: Power-area efficient serial IMPLY-based 4: 2 compressor applied in data-intensive applications. In: Cluster Computing 28 (2025), Nr. 8, S. 512 [Carvalho 2002] Carvalho, Carlos: The gap between processor and memory speeds. In: Proc. of IEEE International Conference on Control and Automation, 2002, S. 27–34

23

[De Cannière und Preneel 2008] De Cannière, Christophe ; Preneel, Bart: Trivium. In: New Stream Cipher Designs: The eSTREAM Finalists. Springer, 2008, S. 244–266 [Debnath u. a. 2018] Debnath, Bikash ; Das, Jadav C. ; De, Debashis: Design of image steganographic architecture using quantum-dot cellular automata for secure nanocommunication networks. In: Nano communication networks 15 (2018), S. 41–58 [Farahani u. a. 2024] Farahani, Samira S. ; Reshadinezhad, Mohammad R. ; Fatemieh, Seyed E.: New design for error-resilient approximate multipliers used in image processing in CNTFET technology. In: The Journal of Supercomputing 80 (2024), Nr. 3, S. 3694–3712 [Fatemieh u. a. 2025a] Fatemieh, Seyed E. ; Asgari, Samane ; Reshadinezhad, Mohammad R.: Fast and low energy approximate full adder based on FELIX logic. In: arXiv preprint arXiv:2505.06888 (2025) [Fatemieh u. a. 2025b] Fatemieh, Seyed E. ; Bagheralmoosavi, Bahareh ; Reshadinezhad, Mohammad R.: Energy-efficient and fast memristor-based serial multipliers applicable in image processing. In: Results in Engineering 25 (2025), S. 104013 [Fatemieh u. a. 2025c] Fatemieh, Seyed E. ; Mashayekhi, Negin ; Reshadinezhad, Mohammad R.: A highspeed and low-cost approximate full adder in QCA technology. In: The European Physical Journal Plus 140 (2025), Nr. 8, S. 726 [Fatemieh und Reshadinezhad 2026] Fatemieh, Seyed E. ; Reshadinezhad, Mohammad R.: Energy-Efficient Approximate Full Adders Applying Memristive Serial IMPLY Logic for Image Processing. In: Journal of Electrical and Computer Engineering 2026 (2026), Nr. 1, S. 7627964 [Fatemieh u. a. 2023] Fatemieh, Seyed E. ; Reshadinezhad, Mohammad R. ; TaheriNejad, Nima: Fast and Compact Serial IMPLY-Based Approximate Full Adders Applied in Image Processing. In: IEEE Journal on Emerging and Selected Topics in Circuits and Systems 13 (2023), Nr. 1, S. 175–188 [Gupta u. a. 2018] Gupta, Saransh ; Imani, Mohsen ; Rosing, Tajana: Felix: Fast and energy-efficient logic in memory. In: 2018 IEEE/ACM International Conference on Computer-Aided Design (ICCAD) IEEE (Veranst.), 2018, S. 1–7 [Hashim u. a. 2016] Hashim, Noor Alia Binti N. ; Hamid, Fazrena Azlee B. ; Teo, Julius ; Hamid, Muhammad Saiful A.: Analysis of memristor based ring oscillators for hardware security. In: 2016 IEEE International Conference on Semiconductor Electronics (ICSE) IEEE (Veranst.), 2016, S. 181–184 [Hell u. a. 2019] Hell, Martin ; Johansson, Thomas ; Meier, Willi ; Sönnerup, Jonathan ; Yoshida, Hirotaka: An AEAD variant of the grain stream cipher. In: International Conference on Codes, Cryptology, and Information Security Springer (Veranst.), 2019, S. 55–71 [Hennessy und Patterson 2017] Hennessy, John L. ; Patterson, David A.: Computer architecture: a quantitative approach. Morgan kaufmann, 2017 [Kaya 2020] Kaya, Turgay: Memristor and Trivium-based true random number generator. In: Physica A: Statistical Mechanics and its Applications 542 (2020), S. 124071 [Kvatinsky u. a. 2014] Kvatinsky, Shahar ; Belousov, Dmitry ; Liman, Slavik ; Satat, Guy ; Wald, Nimrod ; Friedman, Eby G. ; Kolodny, Avinoam ; Weiser, Uri C.: MAGIC—Memristor-aided logic. In: IEEE Transactions on Circuits and Systems II: Express Briefs 61 (2014), Nr. 11, S. 895–899 [Kvatinsky u. a. 2013] Kvatinsky, Shahar ; Satat, Guy ; Wald, Nimrod ; Friedman, Eby G. ; Kolodny, Avinoam ; Weiser, Uri C.: Memristor-based material implication (IMPLY) logic: Design principles and methodologies. In: IEEE Transactions on Very Large Scale Integration (VLSI) Systems 22 (2013), Nr. 10, S. 2054–2066

24

[Manifavas u. a. 2016] Manifavas, Charalampos ; Hatzivasilis, George ; Fysarakis, Konstantinos ; Papaefstathiou, Yannis: A survey of lightweight stream ciphers for embedded systems. In: Security and Communication Networks 9 (2016), Nr. 10, S. 1226–1246 [Mittal 2016] Mittal, Sparsh: A survey of techniques for approximate computing. In: ACM Computing Surveys (CSUR) 48 (2016), Nr. 4, S. 1–33 [Naser und Naif 2022] Naser, Noor M. ; Naif, Jolan R.: A systematic review of ultra-lightweight encryption algorithms. In: International Journal of Nonlinear Analysis and Applications 13 (2022), Nr. 1, S. 3825–3851 [Nguyen u. a. 2020] Nguyen, Hoang Anh D. ; Yu, Jintao ; Lebdeh, Muath A. ; Taouil, Mottaqiallah ; Hamdioui, Said ; Catthoor, Francky: A classification of memory-centric computing. In: ACM Journal on Emerging Technologies in Computing Systems (JETC) 16 (2020), Nr. 2, S. 1–26 [Oved u. a. 2022] Oved, Batel ; Leitersdorf, Orian ; Ronen, Ronny ; Kvatinsky, Shahar: HashPIM: High-throughput SHA-3 via memristive digital processing-in-memory. In: 2022 11th International Conference on Modern Circuits and Systems Technologies (MOCAST) IEEE (Veranst.), 2022, S. 1–4 [Paar u. a. 2024] Paar, Christof ; Pelzl, Jan ; Güneysu, Tim: Understanding cryptography: From established symmetric and asymmetric ciphers to post-quantum algorithms. Springer Nature, 2024 [Psomiadis 2020] Psomiadis, Paul A.: Security of lightweight cryptographic algorithms, National And Kapadistrian University of Athens, Dissertation, 2020 [Rai u. a. 2018] Rai, Vikash K. ; Tripathy, Somanath ; Mathew, Jimson: Memristor based random number generator: Architectures and evaluation. In: Procedia Computer Science 125 (2018), S. 576–583 [Rai u. a. 2021] Rai, VK ; Tripathy, S ; Mathew, J: Multi-bit True Random Number Generator for IoT Devices using Memristor. In: Security of Internet of Things Nodes: Challenges, Attacks, and Countermeasures (2021), S. 35 [Rohani und TaheriNejad 2017] Rohani, Shokat G. ; TaheriNejad, Nima: An improved algorithm for IMPLY logic based memristive full-adder. In: 2017 IEEE 30th Canadian Conference on Electrical and Computer Engineering (CCECE) IEEE (Veranst.), 2017, S. 1–4 [Seiler 2025] Seiler, Fabian: Heterogeneous Approximate Adders for Memristive Processing in Array, Technische Universität Wien, Dissertation, 2025 [Seiler und TaheriNejad 2025] Seiler, Fabian ; TaheriNejad, Nima: An improved serial imply adder algorithm for efficient neural network applications. In: 2025 IEEE 16th Latin America Symposium on Circuits and Systems (LASCAS) Bd. 1 IEEE (Veranst.), 2025, S. 1–5 [Seiler u. a. 2025] Seiler, Fabian ; TaheriNejad, Nima u. a.: An Efficient Robust Serial IMPLY-based InMemristor Adder. In: 2025 Cross-Disciplinary Conference on Memory-Centric Computing (CCMCC) IEEE (Veranst.), 2025, S. 1–7 [Siddiqi u. a. 2023] Siddiqi, Muhammad A. ; Hernández, Jan Andrés G. ; Gebreziorgis, Anteneh ; Bishnoi, Rajendra ; Strydis, Christos ; Hamdioui, Said ; Taouil, Mottaqiallah: Memristor-based lightweight encryption. In: 2023 26th Euromicro Conference on Digital System Design (DSD) IEEE (Veranst.), 2023, S. 634–641 [Soto-Cruz u. a. 2024] Soto-Cruz, Jesús ; Ruiz-Ibarra, Erica ; Vázquez-Castillo, Javier ; EspinozaRuiz, Adolfo ; Castillo-Atoche, Alejandro ; Mass-Sanchez, Joaquin: A survey of efficient lightweight cryptography for power-constrained microcontrollers. In: Technologies 13 (2024), Nr. 1, S. 3 [TaheriNejad 2024] TaheriNejad, Nima: In-Memory Computing: Global Energy Consumption, Carbon Footprint, Technology, and Products Status Quo. In: Proc. 24th IEEE Int. Conf. Nanotechnol.(IEEE-NANO), 2024, S. 1–6 25

[Teimoory u. a. 2015] Teimoory, Mehri ; Amirsoleimani, Amirali ; Ahmadi, Arash ; Alirezaee, Shahpour ; Salimpour, Saeideh ; Ahmadi, Majid: Memristor-based linear feedback shift register based on material implication logic. In: 2015 European Conference on Circuit Theory and Design (ECCTD) IEEE (Veranst.), 2015, S. 1–4 [Yang u. a. 2021] Yang, Ling ; Cheng, Long ; Li, Yi ; Li, Haoyang ; Li, Jiancong ; Chang, Ting-Chang ; Miao, Xiangshui: Cryptographic key generation and in situ encryption in one-transistor-one-resistor memristors for hardware security. In: Advanced Electronic Materials 7 (2021), Nr. 5, S. 2001182 [Zinabu u. a. 2025] Zinabu, Nahom G. ; Marye, Yihenew W. ; Tune, Kula K. ; Demilew, Samuel A.: Comprehensive analysis of lightweight cryptographic algorithms for battery-limited internet of things devices. In: International Journal of Distributed Sensor Networks 2025 (2025), Nr. 1, S. 9639728

26

Record · ID 157299 · SHA-256 dca16248fc2bfc66
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.