RISC-V and Machine Learning: A Survey Shriman Keshri∗1 , Apparna Singh†2,† , Chinmaya Kumar Palo‡3,† , Shreya Adya§2,† , Subhankar Mishra¶1,4 1
arXiv:2609.20677v1 [cs.LG] 17 Sep 2026
National Institute of Science Education and Research, Bhubaneswar, Odisha, India 2 Sri Sri University, Cuttack, Odisha, India 3 Gandhi Institute of Engineering and Technology University, Gunupur, Odisha, India 4 Homi Bhabha National Institute, Mumbai, Maharashtra, India † Work conducted at NISER
Abstract The intersection of open-source processor architectures and machine learning is driving the demand for customizable, efficient, and accessible hardware. This survey examines the state of the RISC-V ISA in machine learning applications, analyzing current capabilities, challenges, and future directions based on recent research. The analysis covers academic and commercial implementations, software frameworks, and real-world applications. The RISC-V machine learning ecosystem is evaluated, from instruction set extensions and core implementations to compiler optimizations and deployment strategies. Key contributions include a unified taxonomy of RISC-V ML implementations, a comparative analysis of performance and design trade-offs, an evaluation of software toolchain maturity, and the identification of emerging trends in instruction set extensions and specialized accelerators. Findings reveal progress in energy efficiency, specialized instruction development, and framework integration, while highlighting challenges in standardization, verification complexity, and ecosystem fragmentation. The analysis proposes four research directions to address current limitations: specialized neural processing extensions, adaptive and modular processor architectures, security frameworks, and energy-efficient multi-domain architectures. These directions provide a roadmap for advancing RISC-V as a foundational platform for next-generation machine learning systems.
Keywords: RISC-V, machine learning, AI accelerators, ISA extensions, embedded systems, edge inference, open-source hardware, neural processing
1
Introduction
Proprietary hardware architectures face challenges from increasing demand for open-source, customizable solutions [1]. This shift stems from requirements for low-power, cost-effective hardware accelerators, semiconductor shortages, and interest in open-source hardware. Supply chain limitations became apparent during disruptions like the COVID-19 pandemic and geopolitical tensions [2], increasing interest in affordable and adaptable hardware solutions [3]. RISC-V [4] is an open-source instruction set architecture (ISA) that allows unrestricted use, modification, and distribution without licensing fees. This openness supports a global developer community [5] and benefits small and medium-sized enterprises (SMEs) and startups [6–8]. RISC-V enables customization through specialized instructions and accelerators for specific application domains without vendor lock-in or licensing constraints. RISC-V is used in machine learning, particularly for edge computing applications including speech recognition, object detection, and text processing. Demand for specialized processing units for edge ∗ [email protected]
† [email protected] ‡ [email protected] § [email protected]
¶ Corresponding author: [email protected]
This is the authors’ accepted manuscript of: S. Keshri, A. Singh, C. K. Palo, S. Adya, S. Mishra, “RISC-V and machine learning: a survey,” The Journal of Supercomputing, vol. 82, no. 8, art. 424, 2026. Published 19 May 2026. The Version of Record is available at https://doi.org/10.1007/s11227-026-08463-z.
1
computations [9–11] has motivated research into RISC-V’s architectural adaptability. The modular nature of RISC-V allows implementation of custom extensions optimized for neural network operations, tensor processing, and other ML computations. Recent work demonstrates RISC-V’s role in machine learning through contemporary core implementations [12–18] and new developments [19–22]. Analysis includes performance comparisons, architectural trade-offs, and efficiency metrics across RISC-V implementations for machine learning workloads. Additional studies spanning modular processors, verification frameworks, side-channel resilience, memory behavior, and comparative system evaluation are cited directly in this text for traceability [23–39]. Literature reviews on RISC-V and machine learning are limited. An earlier survey by Nicholas et al. was published in 2020 [40]; subsequent developments motivate an updated review with post-2020 emphasis. Recent RISC-V-focused ML surveys also include the ecosystem survey by Kalapothas et al. [41] and the edge deep-learning perspective by Agosta et al. [42]. Existing surveys often lack quantitative comparative analysis and concrete conclusions about different approaches. Table 1 presents related research papers and their scope, showing how this survey extends existing literature. Previous surveys have covered various aspects of RISC-V; this work addresses gaps by providing coverage of recent developments, quantitative performance analysis, and practical implementation guidance. Compared with recent RISC-V-for-ML survey literature [41, 42], this survey emphasizes cross-stack linkage from ISA extensions to software frameworks and application-level outcomes. The survey offers systematic evaluation of both hardware implementations and software frameworks. This survey addresses gaps through analysis of contemporary RISC-V cores and implementations. We evaluate software frameworks and tools that facilitate AI accelerator development on RISC-V. Application evaluations demonstrate how RISC-V implementations address practical machine learning challenges through performance metrics and deployment scenarios. Table 1: Overview of Related Research Papers and Their Scope
Paper
Scope RISC-V Extensions
Custom ISA
IoT Security
Systematic Review
✓ ✗ ✓ ✓ ✓ ✓ ✓ ✓
✓ ✗ ✓ ✓ ✓ ✓ ✓ ✓
✓ ✓ ✗ ✗ ✓ ✗ ✗ ✓
✓ ✓ ✗ ✓ ✓ ✓ ✓ ✓
Cui et al. [43] Lu et al. [44] Dorflinger et al. [45] Mezger et al. [46] Nicholas et al. [40] Kalapothas et al. [41] Agosta et al. [42] Our Survey
1.1
Search Methodology
A literature search was conducted using Google Scholar, DBLP, IEEE Xplore, ACM Digital Library, arXiv, and Springer, focusing on developments after [40]. The primary systematic focus is post-2020 work, while foundational pre-2020 studies are retained when needed for background, baselines, and terminology alignment. The following search terms were used: • (RISC V — RISC-V) + machine learning • (RISC V — RISC-V) + deep learning • (RISC V — RISC-V) + neural networks • (RISC V — RISC-V) + AI • (RISC V — RISC-V) + accelerators • (RISC V — RISC-V) + hardware acceleration • (RISC V — RISC-V) + high-performance computing • (RISC V — RISC-V) + AI inference 2
ISA Extensions and Modularity (§ 2.1) RISC-V ISA (§ 2)
RISC-V Implementations (§ 3)
Nomenclature Framework(§ 2.2)
[15, 51–55]
Challenges for ML Adoption(§ 2.3)
[2, 9–11, 23, 49]
CPU Cores and SoCs (§ 3.1)
RISC-V and Machine Learning Survey
ML Acceleration & Co-Design (§ 4.1)
Compiler Optimizations (§ 4.2) Software Frameworks and Stacks (§ 4)
Applications and Evaluation (§ 5)
[9, 10, 12–18, 21, 24–29, 50, 56–59] FPGAbased
[12, 60–64] [1, 14–16, 19, 20]
Custom Hardware
[20–22, 30, 65, 66] [17, 18, 50, 58, 59, 67]
TVM & MLIR
[25, 50, 56, 57, 68, 69] [19, 26, 31–34]
Vector Extensions
[12, 14–16, 49, 59] [21, 22, 30, 50, 58, 67]
Memory Optimization
[23, 27, 28, 35–37] [1, 19–21, 65, 66]
Metaheuristic Algorithms (§ 4.3)
[63, 70–74]
SoC Frameworks (§ 4.4)
[75–78]
Embedded Systems and Edge AI (§ 4.5)
[9, 10, 15, 64, 69, 79] [21, 30, 32, 62, 63, 65]
Medical Applications (§ 5.1)
[32, 33, 35, 50, 56, 57, 65, 67, 80]
Robotics Applications (§ 5.2)
[22, 23, 27–29, 34, 36, 37, 58, 66]
Object Detection (§ 5.3) Performance (§ 5.4) Security (§ 5.5)
Future Research Directions (§ 6)
[3–8, 22, 47–50]
[12–17] [21, 27, 28, 30, 38, 39] [1, 2, 18, 19, 25, 26, 31, 59, 67]
Neural Processing Extensions, Adaptive Architecture, Security Frameworks, Multi-Clock Domains (§ 6.1-6.4)
[1, 19, 22, 30, 49, 59, 67]
Figure 1: Comprehensive outline of the RISC-V and Machine Learning survey.
3
• (RISC V — RISC-V) + ML frameworks Papers were included if they: utilized RISC-V to improve ML efficiency; presented ML models optimized for RISC-V platforms; reported implementation results for ML applications on RISC-V hardware; introduced hardware accelerators or software optimizations for RISC-V ML workflows; or provided quantitative performance analysis of RISC-V systems executing ML workloads. Papers were excluded if they mentioned RISC-V only peripherally, lacked experimental validation, presented purely theoretical work without implementation details, or failed to provide sufficient technical depth for analysis. Quality assessment evaluated technical rigor, clarity of contribution, performance metrics availability, and reproducibility.
1.2
Organisation and Contributions
This survey examines RISC-V’s impact on machine learning through six sections. Section 2 presents the RISC-V ISA, covering its extension framework, nomenclature conventions, and challenges for ML adoption. Section 3 analyzes RISC-V implementations across academic and commercial cores with comparative evaluation. Section 4 examines software frameworks and stacks, including hardware-software co-design, compiler optimization, metaheuristic algorithms, SoC frameworks, and embedded systems for edge AI. Section 5 validates these frameworks through applications and evaluation organized by application domains, performance characteristics, and security considerations. Section 6 identifies research directions based on current limitations. Figure 1 provides an overview of the survey organization. This survey makes the following contributions: • A cross-stack analysis connecting RISC-V ISA extensions, processor implementations, software frameworks, and ML application outcomes. Unlike prior work that treats these layers separately, this survey traces how ISA-level design decisions affect ML performance from hardware through to deployment. • Normalized comparison tables for academic and commercial RISC-V cores with consistent architectural metrics and per-entry source citations, consolidating data previously spread across individual publications and vendor documentation. • A taxonomy of ML software frameworks organized by their RISC-V integration approach, with comparative analysis of compilation toolchains targeting RISC-V backends. • A synthesis of quantitative results reported across RISC-V ML implementations in multiple application domains, drawing architectural patterns and trade-offs from the surveyed studies rather than original experiments. • Four research directions based on gaps identified in the surveyed literature.
2
RISC-V ISA
The RISC-V Instruction Set Architecture (ISA) is an open-standard ISA based on reduced instruction set computer (RISC) principles. Unlike proprietary ISAs, RISC-V is freely available under open-source licenses, enabling unrestricted use in academic research, commercial products, and custom hardware implementations. This openness, combined with its modular extension framework, has driven rapid adoption across embedded systems, high-performance computing, and machine learning acceleration. This section covers the ISA’s extension framework, nomenclature conventions, and the key challenges for ML adoption on RISC-V platforms.
2.1
RISC-V ISA Extensions and Modularity
The RISC-V Instruction Set Architecture (ISA) represents a shift toward modular computing architectures, distinguishing itself from traditional monolithic instruction set designs through its extensibility framework. While conventional architectures typically implement fixed instruction sets that constrain application-specific optimization, RISC-V establishes a modular foundation that enables targeted architectural enhancement through standardized extension mechanisms. This architectural philosophy
4
facilitates domain-specific optimization while maintaining compatibility across diverse implementation contexts, addressing contemporary challenges of specialized computing requirements without sacrificing interoperability. The extensibility paradigm operates through a hierarchical framework encompassing both standardized and custom extensions, each serving distinct roles within the RISC-V ecosystem. This approach enables architectural customization that directly addresses application-specific computational patterns, memory access behaviors, and performance requirements. The modular design philosophy extends beyond simple instruction addition, encompassing architectural modifications including specialized register files, custom functional units, and domain-specific memory hierarchies. These capabilities prove valuable for machine learning applications, where computational patterns often deviate from general-purpose computing workloads. Machine learning workloads demonstrate computational characteristics that differ from traditional general-purpose applications, necessitating specialized architectural support for efficient execution. Neural network inference and training exhibit regular computational patterns dominated by matrix operations, convolution kernels, and element-wise vector operations that benefit from specialized instruction support. The RISC-V extension framework enables development of domain-specific instructions that directly target these computational patterns, resulting in performance improvements and energy efficiency gains compared to software-only implementations on general-purpose architectures. There are two primary categories of RISC-V extensions: 2.1.1
Standard Extensions
Standard extensions comprise predefined instruction sets officially ratified by RISC-V International, targeting well-established requirements and application domains. Standard extensions ensure interoperability and software compatibility across various RISC-V implementations, as detailed in Table 2. Representative examples include: • RV32I/RV64I: These form the foundation of all RISC-V implementations, providing basic integer instructions for 32-bit and 64-bit variants, respectively. They are considered “frozen,” meaning that their specifications are finalized and will not change. • M (Standard Extension for Integer Multiplication and Division): This extension expands the base ISA to include instructions for efficient multiplication and division operations. • A (Standard Extension for Atomic Instructions): Crucial for synchronization in multi-core systems, this extension introduces atomic instructions that ensure operations execute as a single, indivisible unit. • F (Standard Extension for Single-Precision Floating-Point) and D (Standard Extension for Double-Precision Floating-Point): These extensions cater to applications requiring floating-point arithmetic, offering instructions for single-precision (32-bit) and double-precision (64-bit) operations, respectively. • G (Shorthand for I+M+A+F+D): A common combination of extensions often found in general-purpose RISC-V implementations. • Other Standard Extensions: Beyond the examples mentioned, RISC-V offers a range of additional standard extensions, including those for quad-precision floating-point (Q), compressed instructions (C), bit manipulation (B), vector operations (V), hypervisor functionality (H), and supervisor-level instructions (S). All standard extensions listed in Table 2 have been ratified by RISC-V International.
2.1.2
Custom Extensions
The extensible nature of the RISC-V ISA enables designers to develop application-specific instructions that provide fine-grained optimization for specialized computational tasks. This architectural flexibility has proven particularly valuable in domains such as machine learning, signal processing, and cryptography, where domain-specific instructions can yield substantial performance improvements. The custom extension ecosystem encompasses several key categories, as outlined in Table 3. 5
Table 2: Standard ISA Extensions: This table shows ML-relevant standard ISA extensions with their symbols, descriptions, version numbers, and current ratification status (as of December 2024).
Name
Description
Version
Status
M A F D G Q C B V H S
Integer Multiplication and Division Atomic Instructions Single-Precision Floating-Point Double-Precision Floating-Point Shorthand for IMAFD (not a separate extension) Quad-Precision Floating-Point Compressed Instructions Bit Manipulation Vector Operations Hypervisor Supervisor-level Instructions
v2.0 v2.1 v2.2 v2.2 v2.2 v2.0 v1.0 v1.0 v1.0 v1.12
Ratified Ratified Ratified Ratified Ratified Ratified Ratified Ratified Ratified Ratified
Vector processing capabilities are enhanced through custom vector extensions (Xvec) [20], which introduce application-specific vector instruction sets that accelerate computations in signal processing, linear algebra, and scientific simulation applications. These custom extensions complement the standard RISC-V Vector extension (V/RVV) by providing domain-specific optimizations beyond the ratified specification. This extension enables Single Instruction, Multiple Data (SIMD) operations that are fundamental to modern parallel computing workloads [49]. For security-critical applications, the K (Cryptography Instructions) extension provides dedicated hardware-accelerated cryptographic operations, substantially improving both the efficiency and security robustness of encryption and decryption processes. Additionally, the Zfh (Half-Precision Floating-Point) extension, introduced in 2021, provides hardware support for IEEE 754 binary16 half-precision floating-point operations, enabling efficient low-precision arithmetic for ML inference and scientific computing [50]. The broader custom extension landscape demonstrates RISC-V’s remarkable adaptability to emerging computational paradigms and specialized application requirements [56, 57]. Machine learning workloads have driven the development of several specialized RISC-V ISA extensions that address the unique computational and architectural requirements of neural network inference and training [59, 67]. The Smcdeleg and Ssccfg (Supervisor Counter Delegation) extensions provide fine-grained performance monitoring capabilities specifically tailored for ML applications, enabling hypervisors to delegate counter access privileges to guest operating systems. This functionality proves crucial for comprehensive neural network performance profiling and dynamic resource allocation optimization in multi-tenant machine learning environments. Atomic memory operations for quantized neural networks are enhanced through the Zabha (Byte and Halfword Atomic Memory Operations) extension, which supports efficient synchronization primitives for 8-bit and 16-bit data types commonly employed in quantized deep neural networks [12]. These atomic operations significantly reduce memory contention in multi-core ML inference scenarios, enabling more efficient parallel processing of quantized models. The Smstateen (State Enabled Extension) provides comprehensive state management capabilities that facilitate secure ML execution environments, allowing selective architectural state access for trusted machine learning workloads while maintaining robust isolation between different computational contexts [14, 15]. Recent RISC-V extensions (2024) provide quantitative performance improvements for machine learning workloads. The Smcdeleg/Ssccfg counter delegation extensions enable efficient performance monitoring with reduced overhead compared to trap-based approaches [52]. Zabha atomic byte and halfword operations improve synchronization efficiency in quantized deep neural networks [12], particularly for multi-threaded inference workloads. Smstateen state enablement features support secure ML execution environments [52]. These extensions demonstrate the architectural responsiveness of RISC-V to evolving machine learning computational requirements. Recent paradigm-shifting developments in RISC-V machine learning implementations have established new benchmarks for application-specific optimization and energy efficiency. The MARVEL Framework [19] represents a significant advancement in automated custom extension generation, providing an end-to-end methodology for creating model-class aware RISC-V extensions that achieve 2× speedup and 2× energy reduction for lightweight artificial intelligence applications. This automated approach eliminates the traditional manual effort required for application-specific optimization while maintaining high
6
Table 3: Custom ISA Extensions: This table shows the various custom ISA extensions of the RISC-V processor with their symbol/name, description, and the year in which they were introduced or if they are under development or not. Rows marked with * indicate features that are still under development.
Extension
Description
Xpulpnn, Xvec
Custom Neural Network and Vector Operations (standard vector is V/RVV) Cryptography Instructions Control and Status Register Support Load–store fence instructions Half-Precision Floating-Point Support (IEEE 754-2008 binary16) Supervisor Counter Delegation (M-mode to S-mode counter access) State Enabled Extension Code Size Reduction
K Zicsr Zifencei Zfh Smcdeleg, Ssccfg Smstateen Zca, Zcb, Zcd, Zce, Zcf, Zcmp, and Zcmt Smepmp Zba, Zbb, Zbc, and Zbs RV32E/RV64E Zfa Zvfh, Zvfhmin
Zihintpause Sscofpmf Zabha
Year
PMP Enhancements for memory access and execution prevention in Machine mode Bit-Manipulation ISA-extensions RV32E and RV64E Base Integer Instruction Sets Standard Extension for Additional Floating-Point Instructions Vector Extension for Half-Precision Floating-Point Arithmetic/Vector Extension for Minimal Half-Precision Floating-Point Arithmetic Pause Hint Count Overflow and Mode-Based Filtering Extension Byte and Halfword Atomic Memory Operations (8-bit and 16-bit atomics for quantized neural networks)
7
2019-2021 2017 2018 2018 2021 2024 * 2023 2023
2021 2021 2023 2023 2023
2021 2021 2024
performance gains. Transformer architecture acceleration has been revolutionized through the VEXP ISA Extensions [22], which provide low-cost RISC-V instruction set enhancements specifically targeting softmax operations in attention mechanisms. These extensions demonstrate remarkable efficiency improvements, achieving 162.7× latency reduction and 74.3× energy efficiency enhancement with minimal area overhead of only 1% (2,847 additional logic gates on 28nm CMOS process), enabling practical GPT-2 and Vision Transformer inference on RISC-V cluster architectures. Matrix operation efficiency is further enhanced through zero-stall matrix multiplication optimizations [21], which implement advanced memory hierarchy techniques to eliminate pipeline stalls during critical machine learning computational kernels. The integration of machine learning capabilities into flexible electronics platforms has been demonstrated through Bendable RISC-V ML implementations [66], achieving 21× improvements in both inference execution time and energy efficiency for Support Vector Machine classification on bendable substrates. Additionally, ultra-low-power neural network implementations [65] have been developed with optimized RISC-V processor extensions that support lightweight neural network models achieving 97.62% classification accuracy on CIFAR-10 dataset while consuming only 13% of the power required by ARM Cortex-M4 implementation (measured power: 4.2 mW versus 32.1 mW baseline), demonstrating the potential for RISC-V in energy-constrained edge computing scenarios. These collective developments establish comprehensive performance benchmarks that demonstrate RISC-V’s competitive position in machine learning acceleration. Transformer inference workloads benefit from up to 5.8× latency reduction with corresponding 3.6× energy savings, making RISC-V implementations viable for large language model deployment in resource-constrained environments. Automated custom ISA generation achieves 2× overall performance improvements while eliminating the manual design effort traditionally required for application-specific optimization. Flexible computing implementations demonstrate remarkable 21× energy efficiency gains for edge AI applications deployed on bendable electronics substrates, opening new possibilities for ubiquitous computing scenarios. Vector processing capabilities are enhanced through advanced auto-vectorization techniques specifically optimized for machine learning acceleration, while multi-core efficiency is maximized through zero-stall matrix operation implementations that support up to 1024 processing element RISC-V clusters. These performance characteristics collectively establish RISC-V as a compelling platform for nextgeneration machine learning acceleration across diverse deployment scenarios, from ultra-low-power edge devices to high-performance computing clusters. The development of custom RISC-V extensions operates within a comprehensive framework of established rules and guidelines designed to ensure compatibility and seamless integration within the broader RISC-V ecosystem.[52] RISC-V International provides detailed specifications encompassing standardized naming conventions that maintain consistency and prevent conflicts between different extension implementations, comprehensive instruction encoding format guidelines that ensure proper interpretation across various RISC-V processor implementations, and rigorous conformance testing procedures that verify adherence to defined standards while maintaining compatibility and reliability across the ecosystem [81]. These regulatory measures are fundamental to fostering a robust and interoperable RISC-V development environment that encourages innovation while preserving architectural consistency across diverse implementation contexts [53]. The balance between extensibility and standardization enables the continued evolution of RISC-V as a platform capable of adapting to emerging computational paradigms while maintaining the stability and predictability required for enterprise and research deployment scenarios.
2.2
RISC-V ISA Nomenclature and Specification Framework
The RISC-V ISA specification employs a systematic nomenclature framework that ensures consistent identification and compatibility assessment across diverse implementations while facilitating precise communication within the broader RISC-V ecosystem [51, 52]. Table 2 defines the exact order that must be used to determine the RISC-V ISA subset. The nomenclature follows a hierarchical structure where base architecture specifications (RV32I, RV64I, RV128I) are followed by extension letters in a predetermined sequence that ensures compatibility and standardization across implementations [81]. The nomenclature framework begins with the base architecture identifier, which specifies the fundamental characteristics of the RISC-V implementation [51]. The prefix ”RV” universally identifies all RISC-V implementations, followed by a numeric designation indicating the register width and fundamental addressing capabilities. RV32I represents 32-bit implementations with integer instruction support, 8
RV64I denotes 64-bit architectures with extended addressing capabilities, and RV128I specifies future 128-bit implementations designed for advanced computational requirements. These base specifications establish the fundamental computational and memory addressing capabilities that form the foundation for all subsequent extensions. Following the base architecture specification, extension letters must appear in a strictly defined alphabetical sequence that ensures consistent interpretation across different implementations and development tools [51, 52]. This ordering convention prevents ambiguity in ISA specification while facilitating automated compatibility checking and optimization selection within compilation frameworks. Standard extensions including M (multiplication and division), A (atomic instructions), F (single-precision floatingpoint), D (double-precision floating-point), and C (compressed instructions) follow immediately after the base specification in their prescribed alphabetical order. The systematic extension ordering serves multiple critical functions within the RISC-V ecosystem [53, 54]. Compiler toolchains rely on this standardized format to automatically determine supported instruction sets and enable appropriate optimization strategies. Hardware verification frameworks utilize the nomenclature to validate implementation completeness and ensure adherence to specified functionality. Software frameworks leverage the standardized naming convention to automatically configure runtime environments and optimization parameters based on target architecture capabilities. Custom extensions integrate into the nomenclature framework through reserved naming spaces that prevent conflicts with standard extensions while maintaining systematic organization [15, 53]. Custom extensions typically utilize the ”X” prefix followed by descriptive names that indicate their specialized functionality, such as ”Xpulpnn” for neural network acceleration or ”Xvec” for specialized vector operations. This naming convention enables precise specification of application-specific architectural enhancements while preserving compatibility with standard toolchain components. The hierarchical nomenclature framework facilitates comprehensive specification of complex RISCV implementations that incorporate multiple standard and custom extensions [54, 82]. For example, a machine learning implementation might use ”RV64IMAFDC Xpulpnn Xvec” (this is an illustrative specification, not an endorsed profile) to indicate a 64-bit implementation with standard mathematical extensions plus specialized neural network and vector processing capabilities. This systematic approach enables precise communication of architectural capabilities while ensuring compatibility assessment across diverse deployment scenarios. The standardized nomenclature framework proves particularly valuable for machine learning applications where diverse architectural configurations must be supported across different deployment contexts [55, 83]. Software frameworks can automatically determine optimal compilation strategies based on nomenclature specifications, enabling efficient deployment of ML models across heterogeneous RISC-V implementations. The systematic naming convention facilitates automated performance profiling and optimization selection, reducing the complexity of multi-target deployment while maximizing performance across diverse hardware platforms. The nomenclature framework enables comprehensive documentation and comparison of machine learning accelerator implementations across different research and commercial projects [14, 84]. Consistent specification formats facilitate systematic performance analysis and enable meaningful comparisons between different architectural approaches for ML acceleration. This standardization proves essential for advancing the field through reproducible research and systematic evaluation of competing architectural strategies. Table 4 summarizes the RISC-V ISA nomenclature framework for ML applications.
2.3
Challenges for ML Adoption
Although RISC-V and open-source hardware offer significant advantages, developers face several challenges when targeting machine learning applications. Traditionally, proprietary architectures have dominated hardware design, limiting customization and access to information. However, increasing demand for affordable, low-power hardware accelerators, exacerbated by global semiconductor shortages, has driven a shift toward open-source solutions. Four key challenges emerge in adapting RISC-V ISA for ML workloads: security robustness, licensing and openness management, software interoperability, and power-efficiency optimization. Security considerations represent a primary concern as RISC-V, being relatively recent, requires research in memory protection mechanisms and side-channel attack prevention to ensure robust security implementations. While RISC-V incorporates built-in security features and supports custom security extensions, research in these domains remains necessary for widespread commercial adoption. 9
Table 4: RISC-V ISA Nomenclature Framework for ML Applications
Component
Format/Example
Description
Base Architecture
RV32I, RV64I, RV128I
Standard Extensions
IMAFDCV
Custom ML Extensions
Xpulpnn, Xvec, Xmlir
Example ML Specification Toolchain Integration
RV64IMAFDC Xpulpnn Xvec
Register width and fundamental addressing capabilities Integer, Multiplication, Atomic, Float, Double, Compressed, Vector (G = shorthand for IMAFD) Neural network acceleration, vector ops, compiler support Illustrative 64-bit configuration (not an endorsed profile) Compiler frameworks use nomenclature for code generation Automated benchmarking based on ISA specification Systematic evaluation of implementation capabilities
Performance Profiling
Automatic optimization selection Architecture-specific tuning
Compatibility Assessment
Cross-platform deployment
Licensing complexities emerge from varied approaches to intellectual property management within the open-source hardware ecosystem. Although the Open Source Hardware Association encourages open sharing practices, implementation varies across projects. Academic implementations typically employ permissive licenses such as MIT, BSD, or Apache 2.0, while commercial implementations often utilize proprietary licenses that restrict specific hardware block usage or complete design access. This licensing heterogeneity introduces challenges for developers navigating the open-source hardware landscape. Interoperability challenges stem from limited software and operating system support for RISC-V platforms. While specific Linux distributions and real-time operating systems provide support for particular RISC-V System-on-Chip implementations, comprehensive compatibility across the ecosystem remains incomplete. Developers frequently encounter requirements to rebuild or modify existing software for RISC-V hardware compatibility, impacting development efficiency and time-to-market. Power consumption optimization presents ongoing challenges particularly relevant to machine learning accelerator applications in edge computing environments. The relationship between design complexity, operating frequency, and power consumption in RISC-V cores requires careful optimization to meet efficiency requirements. Developers must prioritize design optimization strategies targeting power efficiency in RISC-V-based machine learning accelerator implementations. These challenges underscore complexities inherent in developing within a rapidly evolving technological ecosystem. Addressing these issues remains essential for continued growth and widespread adoption of RISC-V and open-source hardware in modern hardware design applications.
3
RISC-V Implementations
This section focuses on an in-depth inventory of RISC-V cores and SoCs that are either provided by the research community or produced commercially [51, 52]. This process is generally framed by quantitative analysis as a list of supported ISAs along with other information or characteristics, including the launch date of the core, clock frequency, performance, hardware description language (HDL), license, and implementation type (FPGA/ASIC). Many open-source and free implementations have facilitated their adoption in academic and commercial projects [53, 82]. Elaborating on the implementation types, that is, FPGA and ASIC, field-programmable gate arrays (FPGAs) can be configured to perform a wide range of functions, and application-specific integrated circuits (ASICs) are custom-designed chips that are built for a specific purpose. FPGAs are generally chosen for flexibility and prototyping, whereas ASICs are selected for high performance, low power consumption, and mass production. In the RISC-V ISA, the 32-bit and 64-bit implementations refer to the size of the registers and address space that the processor can handle [51]. The 32-bit RV32 implementation, RV32, features 32-bit general-purpose registers and can address up to 4 GB of memory. It includes a base instruction
10
Table 5: Summary of RISC-V Implementation Technologies and Compatibility
Category Implementation
Technology/Feature
Description
FPGA
Flexible, configurable for prototyping and research High performance, low power, mass production
ASIC
Architecture
RV32 (32-bit)
32-bit registers, 4GB memory space, 47 base instructions 64-bit registers, 16 exabyte memory space, extended operations
RV64 (64-bit) General Purpose
Linux (Debian, Fedora, Ubuntu), FreeBSD, OpenBSD FreeRTOS, Zephyr, Apache NuttX Android port, Haiku, Illumos, Plan
OS Support Real-Time Emerging Linux Requirements
ISA Extensions
RV64GC (I+M+A+F+D+C+Zicsr+Zifencei) MMU, PLIC, U/S modes, virtual memory (Sv32/39/48)
Hardware Components
set RV32I, which contains 47 instructions for basic operations, such as addition, subtraction, bitwise operations, load/store, jumps, and branches. Additional features can be added through extensions such as RV32M for multiplication and division and RV32F for 32-bit floating-point operations. The 64-bit implementation RV64 extends the capabilities of RV32 by using 64-bit general-purpose registers and addressing a much larger memory space, theoretically up to 16 exabytes. The base instruction set for RV64, RV64I, includes the instructions for 64-bit operations. Similar to RV32, RV64 can be extended with additional instructions for 64-bit operations, such as RV64M for 64-bit multiplication and division. These implementations allow RISC-V to be versatile and cater to various applications, from small embedded systems to high-performance computing. Table 5 summarizes the key implementation technologies and compatibility features of RISC-V systems. The RISC-V system is compatible with a wide variety of full-function and real-time operating systems. Among the all-purpose operating systems, the Linux kernel fully supports RISC-V, with distributions such as Debian, Fedora, openSUSE, Ubuntu, Arch Linux, and other Unix-like operating systems such as FreeBSD and OpenBSD using RISC-V ports to support RISC-V. Additionally, efforts are being made to move Android to RISC-V. In real-time operating systems (RTOS), popular platforms, such as FreeRTOS, Zephyr, and Apache NuttX support RISC-V, making it suitable for embedded systems and other IoT devices operating systems such as Haiku, Illumos, and Plan can be transferred to RISC-V. The RISC-V platform specification ensures that operating systems, such as Linux and Zephyr, can run successfully on all specification-compliant RISC-V hardware, maintaining compatibility across applications. This variety of OS compatibility makes RISC-V versatile for various applications, from embedded systems to high-performance computing. To run Linux on an RISC-V processor, the minimum required resources include the RV64GC instruction set, which includes the base integer instructions (RV64I), multiplication and division (M), atomic operations (A), single- and double-precision floating-points (F and D), compressed instructions , and control and status register instructions (Zicsr), along with instruction-fetch fence (Zifencei). Additionally, the processor must support user and supervisor modes (U and S) for effective process and memory management, as well as a page-based virtual memory system such as Sv32, Sv39, or Sv48. Essential hardware components include a Memory Management Unit (MMU) and Platform-Level Interrupt Controller (PLIC) to handle virtual memory and interrupt processing. These components ensure that the RISC-V processor meets the demands of running a full Linux operating system.
3.1
CPU Cores and SoCs
CPU cores are processing engines containing essential processing units that execute instructions [51]. These cores have various designs optimized for different purposes, such as low power for wearable de11
vices or high server performance. A system-on-chip (SoC) integrates the RISC-V CPU core along with additional components such as memory controllers and input/output (I/O) interfaces onto a single chip. This helps create compact and efficient solutions for various applications. When the CPU core acts as the central processor, the SoC acts as the entire ecosystem required to make it function effectively. 3.1.1
Academic RISC-V Cores
In the academic landscape, there are 13 cores with 32-bit implementations, five cores with 64-bit implementations, and a single core that supports both 32-bit and 64-bit instructions [52, 53]. These cores can be used for various purposes, ranging from low-power embedded systems to high-performance cloud computing. Table 6 summarizes the comparison of these cores. Comparative Analysis of Academic RISC-V Cores: Performance Analysis: Among the academic cores, VRoom demonstrates the highest performance at 11.3 DMIPS/MHz (FPGA implementation on Xilinx UltraScale+) [85], followed by BOOM at 3.93 DMIPS/MHz (ASIC implementation on TSMC 28nm) [86], making them suitable for high-performance ML applications. SERV achieves the lowest at 0.718 DMIPS/MHz but compensates with minimal resource usage. The performance spread indicates different optimization targets: high-performance cores like VRoom and BOOM target computational throughput, while minimalist cores like SERV and PicoRV32 (0.516 DMIPS/MHz) prioritize resource efficiency. ISA Extension Support: XiangShan supports the most comprehensive ISA extensions (RV64GCBK), including vector and cryptographic instructions, positioning it for complex ML workloads. In contrast, simpler cores like ORCA (RV32IM) and SERV (RV32IMCZicsr) focus on basic functionality, making them suitable for resource-constrained ML inference tasks. Clock Frequency and Architecture Trade-offs: Frequency claims must be distinguished by implementation type. VRoom achieves 1.02 GHz on FPGA (Xilinx UltraScale+), while XiangShan reaches 2.2 GHz and V-Seek Server achieves 2.0 GHz on ASIC implementations (different process nodes). The frequency-performance correlation shows that deeper pipelines (XiangShan’s 11-stage) enable higher frequencies but require more complex control logic. License and Development Approach: Most academic cores use permissive licenses (Apache 2.0, BSD, MIT), facilitating commercial adoption. The choice of HDL varies significantly: traditional Verilog dominates, but emerging alternatives like Chisel (Rocket Core, BOOM) and SpinalHDL (VexRiscv) demonstrate modern design approaches that could enhance ML accelerator development. Target Application Analysis: The core characteristics suggest clear application domains: XiangShan and BOOM target high-performance computing with ML support, VexRiscv and Ibex focus on configurable embedded systems, while SERV and PicoRV32 enable minimal IoT deployments with basic ML inference capabilities.
12
13
ISA
RV64GCBK RV32/64GC RV32I/EMCB RV32IMCAFD RV64IMAFDCHBK(V) RV32IMC RV32IMCZicsr RV32IM RV64IMAFD RV64GCSUN RV32IMAC RV32IMC RV32IM RV32IMF RV64GC RV32I/EMC RV32E/RV32I RV32IM RV32E RV64GCV
Core
XiangShan [87] CVA6 [45] Ibex [88] VexRiscv [89] VRoom [85] SSRV [90] SERV [91] ORCA [92] Rocket Core [93] Shakti-Cclass [94] Shakti-Eclass [94] RI5CY [95] RV01 [96] RSD [97] BOOM [86] PicoRV32 [98] DarkRISCV [99] SNAX Cluster [21] Bendable RISC-V [66] V-Seek Server [58]
11 6 2 5 Deep Bit-serial 4/5 5 5 3 4 10 Multi-cyc. 2/3 1 Bit-serial 12
Pipe OoO IO IO IO OoO OoO IO IO IO IO IO IO IO OoO OoO IO IO IO IO OoO
Exec No Yes Yes No No No No No Yes No No No No Yes
Super No Yes Yes No No No No No No No No No No -
SMT 2GHz 50MHz 200MHz 1020MHz 29-34MHz 135MHz 122MHz 1000MHz 1.5-2.5GHz 1GHz 625MHz 1000MHz 500MHz 250-400MHz 60kHz 2.0GHz
Freq 0.82 3.13* 1.38 11.3 1.5-6.4 0.718 0.98 1.72 1.68 1.71 1.8 3.93 0.516 8.5*
Perf 64K I+64K D 16K I+32K D 4K I (opt.) Configurable 32K I+32K D None Optional 16K I+16K D 16K I+16K D None None 32K I+32K D None Small None 64K I+64K D
L1 Cache
Verilog Verilog SpinalHDL Verilog/SV Verilog Verilog/VHDL VHDL Chisel BSV BSV Verilog VHDL SystemVerilog Chisel3 Verilog Verilog Chisel/SV Verilog Verilog/SV
HDL
MPSL 2.0 Apache 2.0 Apache 2.0 MIT GPL-3.0 ISC BSD BSD BSD BSD LGPL Apache 2.0 BSD ISC BSD-3 Apache 2.0 Apache 2.0
License
FPGA FPGA FPGA FPGA/ASIC FPGA/ASIC FPGA/ASIC FPGA ASIC FPGA/ASIC FPGA/ASIC FPGA FPGA FPGA ASIC FPGA FPGA ASIC Flexible ASIC
Type
Table 6: CPU CORES: This table summarizes academic contributions to RISC-V core design. Each entry details a specific core implementation with source reference, supported ISA, pipeline depth, execution model (IO = in-order, OoO = out-of-order), reported superscalar and SMT capabilities (when explicitly available in cited sources), clock frequency, performance (Perf: measured in DMIPS/MHz, where * indicates CoreMark/MHz), L1 cache configuration (I$ = instruction cache, D$ = data cache), hardware description language (HDL), licensing terms, and implementation type. A dash (-) indicates information not publicly available or not applicable.
Superscalar and SMT entries in Table 6 are reported only where explicitly stated in cited sources; a dash denotes unavailable or non-uniform public reporting. The implementations of RISC-V processors span a wide range of designs, from simple, low-power cores to advanced, high-performance architectures, thereby demonstrating the versatility of the RISC-V instruction set architecture (ISA). Several key architectural characteristics and design patterns emerge across these diverse processor implementations. Pipeline architectures across RISC-V implementations demonstrate significant variation tailored to specific performance and power requirements. The XiangShan processor [87] exemplifies highperformance design with its 11-stage pipeline and out-of-order execution capabilities, targeting applications requiring maximum computational throughput. In contrast, simpler designs such as the Rocket Core [93] and Ibex [88] employ 5-stage and 2-stage pipelines respectively, optimizing for implementation simplicity and resource efficiency. This design spectrum reflects the fundamental balance between complexity and performance, with deeper pipelines generally enabling higher clock speeds and instruction throughput, as demonstrated by implementations like XiangShan and DarkRISCV [99]. Customization and modularity represent defining characteristics of many RISC-V core implementations, enabling adaptation to diverse application requirements. Cores such as VexRiscv [89], PicoRV32 [98], and Ibex [88] are architected with extensive configurability options, allowing developers to tailor functionality for specific deployment scenarios. VexRiscv specifically supports various instruction set extensions and includes optimizations for FPGA integration, while PicoRV32 and DarkRISCV [99] prioritize compact implementation and efficiency for embedded systems applications. The open-source development model underlying most RISC-V processor projects fosters collaborative innovation and accelerated development cycles. Projects including XiangShan [87], VexRiscv [89], and Rocket Core [93] leverage collaborative platforms such as GitLab and GitHub, with maintenance typically provided by academic institutions or research organizations. This collaborative environment accelerates technological progress and enables diverse application development spanning embedded systems to highperformance computing domains. Instruction set architecture extension strategies highlight the flexibility inherent in RISC-V implementations. Advanced cores such as Shakti-Cclass [94] implement comprehensive extension sets (RV64GCSUN) including supervisor-level instructions for complex system applications, while specialized implementations like VRoom [85] (RV64-IMAFDCHBK(V)) and RiscyOO [100] (RV64G) focus on high-performance computing and multimedia processing capabilities. This architectural adaptability enables targeted optimization for specific domains including signal processing applications (RI5CY [95]) and energy-efficient computing (Ibex [88]). Target application specialization drives significant architectural differentiation across RISC-V implementations. Minimalist cores such as SERV [91] prioritize low-power operation and minimal resource requirements for small FPGA deployments, while high-performance implementations including RSD [97] and BOOM [97] (SonicBOOM) feature sophisticated out-of-order execution mechanisms competitive with commercial-grade processors. Resource-constrained implementations like Shakti cores [94] target embedded applications requiring real-time operating system support while maintaining efficiency constraints. Performance optimization strategies across RISC-V implementations encompass both superscalar and multithreaded approaches, with advanced designs such as SSRV [97] and VRoom [85] implementing multiple instruction execution per cycle and simultaneous multithreading capabilities. Alternative approaches exemplified by CVA6 [45] and ORCA [92] emphasize balanced performance through in-order execution while maintaining implementation simplicity and resource efficiency, particularly in FPGA deployment scenarios. These architectural patterns and design characteristics demonstrate how RISC-V implementations leverage the instruction set architecture’s inherent flexibility and scalability to address diverse performance, power consumption, and application requirements across a broad spectrum of computing domains. 3.1.2
Commercial RISC-V Cores
Commercial RISC-V cores represent predesigned, production-ready implementations available for licensing by companies and organizations. These cores function as optimized blueprints that provide a faster and more reliable development path compared to designing custom RISC-V implementations from scratch. Commercial cores are typically offered in various configurations optimized for specific target domains including performance-oriented applications, power-efficient embedded systems, or specialized use cases. They are accompanied by comprehensive documentation, professional support packages, and
14
maintenance agreements to ensure smooth integration into commercial products. This approach enables companies to focus their development efforts on unique product features and applications while leveraging proven core processing technology, thereby reducing development time, minimizing technical risks, and accelerating time-to-market for RISC-V-based products. The commercial RISC-V core landscape encompasses 17 cores with 32-bit implementations, four cores with 64-bit implementations, and four cores supporting both 32-bit and 64-bit instruction sets. Table 7 provides a comprehensive comparison of these commercial implementations, highlighting their architectural characteristics, performance metrics, and target applications. Commercial RISC-V cores demonstrate a clear performance hierarchy that reflects their intended application domains. Leading the performance spectrum, SweRV EH1 achieves 4.9 CoreMark/MHz, followed by Xuantie-910 at 7.1 CoreMark/MHz and BI-671 at 3.74 DMIPS/MHz. This performance differentiation enables targeted deployments across diverse application scenarios, from high-performance edge AI servers requiring maximum computational throughput to power-constrained sensor nodes demanding energy efficiency. The clock frequency distribution reveals three distinct categories: high-performance cores operating above 1.5GHz for computationally intensive workloads, mid-range implementations running between 400MHz and 1.2GHz for balanced applications, and low-power variants operating below 400MHz for battery-operated devices. The ISA extension strategies employed by commercial cores reflect careful optimization for target markets. Advanced implementations like Xuantie-910 incorporate vector extensions (RV64GCV) specifically for AI acceleration workloads, while comprehensive cores such as SCR7 include extensive extension sets (RV64GCVBK) to support versatile computing requirements. In contrast, efficiency-focused cores like SCR1 implement minimal extension sets (RV32IEMC) to reduce complexity and overhead. This strategic differentiation allows organizations to select cores that precisely match their application requirements without unnecessary complexity or resource consumption. Market positioning analysis reveals three primary segments within the commercial RISC-V ecosystem. The high-performance segment includes cores such as Xuantie-910, SCR7, and SweRV EH1, which target edge AI applications and high-performance embedded systems requiring substantial computational capabilities. The IoT and embedded segment encompasses Andes Technology cores (A25, D25F) and SiFive implementations (E31, E51), focusing on configurable, power-efficient solutions suitable for connected devices and industrial applications. The ultra-low-power segment features cores like GD32VF103 and similar microcontroller-class implementations designed for battery-operated IoT devices where energy efficiency is paramount. Implementation flexibility represents a key characteristic of commercial RISC-V cores, with most supporting both FPGA and ASIC deployment options. This dual-target approach provides organizations with the flexibility to prototype using FPGAs while transitioning to ASIC implementations for volume production, thereby reducing development risk and enabling faster market entry for machine learning applications. Licensing models across commercial cores reflect diverse business strategies, ranging from open-source licenses (Apache 2.0) that facilitate ecosystem development to proprietary commercial licenses offering differentiated features and professional support. This licensing diversity enables organizations to select business models that align with their commercial objectives while accessing appropriate technical capabilities. Tenstorrent is also relevant to the commercial RISC-V AI ecosystem because its AI processors use RISC-V-based control and management complexes in broader accelerator-centric products [101]. To keep Table 7 comparable, this survey treats that table as a list of standalone/licensable commercial CPU cores and discusses accelerator-centric products such as Tenstorrent separately in narrative form.
15
16
ISA
RV64GCV RV32I/EMC RV32IMAC/RV64IMAC RV32IMAFDC/RV64IMAFDC RV32IMAFDC/RV64IMAFDC RV64GCVBK RV32IMAC RV64IMAC RV32IMAC RV32IM RV32I RV32IMC RV32IMACFDBP RV32IMACFDBP RV32IMAC/EMAC RV32IMACFDB RV64IMACFDB RV32IMACFDBP RV32IMACFDBP RV32I/RV64I RV32IMC RV32IMACF RV64IMACFD RV64GCBK RV32IMAC RV32IMAC RV64GC
Core Name
XuanTie 910 [102] SCR 1 [103] SCR3 [103] SCR4 [103] SCR5 [103] SCR7 [103] SiFive E31 [104] SiFive E51 [104] GD32VF103 [105] MRISCV [106] ReonV [107] SweRV EH1 [108] A25 [109] D25F [109] N22 [109] N25F [109] NX25F [109] A25MP [109] AX25MP [109] Roa Logic RV12 [110] BM-310 [111] BI-350 [111] BI-651 [111] BI-671 [111] Hummingbird E200 [112] Hummingbird E203 [113] Hummingbird E603 [114]
2.5GHz 435MHz 1.5GHz+ 1GHz+ 1GHz 1.2GHz+ 320MHz 667MHz 108MHz 435MHz 1.8GHz 1.2GHz 400MHz 700MHz 1.2GHz 1.2GHz 1.1GHz 1GHz 700MHz 1GHz 1GHz 1GHz -
Clock Freq 7.1* 1.73 1.7 1.7 1.7 1.61 1.61 1.53 0.32 4.9* 3.5* 2.59 1.8 3.5* 1.81 1.72 2.75 3.74 1.77 -
Perf Verilog Verilog/VHDL Verilog/VHDL Verilog/VHDL Verilog/VHDL Verilog Verilog Verilog HDL VHDL Verilog/VHDL Verilog HDL Verilog HDL Verilog HDL Verilog HDL Verilog HDL Verilog HDL Verilog HDL Verilog/VHDL Verilog HDL Verilog HDL Verilog HDL Verilog HDL Verilog Verilog Verilog
HDL SHL SHL SHL Eval Eval Commercial MIT GPL-3.0 Apache 2.0 Open Source Apache 2.0 Apache 2.0 Academic
License
ASIC FPGA/ASIC FPGA/ASIC FPGA/ASIC FPGA/ASIC FPGA/ASIC FPGA FPGA ASIC FPGA/ASIC FPGA FPGA/ASIC FPGA/ASIC FPGA/ASIC FPGA/ASIC FPGA/ASIC FPGA/ASIC FPGA/ASIC FPGA/ASIC FPGA/ASIC FPGA/ASIC FPGA/ASIC FPGA/ASIC FPGA/ASIC FPGA/ASIC FPGA/ASIC FPGA/ASIC
Type
2019 2018 2020 2021 2022 2023 2019 2019 2019 2019 2020 2019 2023 2024 2019 2023 2018 2019 2019 2018 2020 2019 2019 2019 2018 2020 2025
Year
Table 7: COMMERCIAL CORES: This table summarizes commercial contributions to RISC-V core design. Each entry details a specific core implementation with source reference, supported ISA, clock frequency, performance (Perf: measured in DMIPS/MHz, where * indicates CoreMark/MHz), HDL, licensing terms, implementation type, and launch date. A dash (-) indicates information not publicly available or not applicable.
Several leading organizations have established significant positions in the commercial RISC-V ecosystem through diverse product offerings and strategic market positioning. Alibaba’s Xuantie-910 [102] represents a high-performance 64-bit processor implementation featuring custom extensions for advanced arithmetic and memory operations, supporting the comprehensive RV64GCV instruction set architecture including vector extensions for AI acceleration. This processor targets demanding computational workloads requiring both high throughput and advanced processing capabilities. In contrast, GigaDevice’s GD32VF103 [105] focuses on 32-bit microcontroller applications, implementing the RV32IMAC instruction set with extensive I/O support specifically optimized for embedded system deployments. SiFive [115] has established market leadership in RISC-V core licensing through comprehensive offerings spanning both 32-bit and 64-bit implementations. Their product portfolio addresses applications ranging from ultra-low-power wearable devices to high-performance computing systems. The E31 and E51 core families specifically target IoT and server-based applications with configurable multicore capabilities, while their comprehensive tool ecosystem and development environment significantly accelerates RISC-V adoption across industry sectors. Similarly, Renesas [106] has developed the MRISCV platform, delivering 32-bit microcontroller solutions for IoT applications with emphasis on minimal silicon area and power consumption requirements. Western Digital’s SweRV [108] implementation represents a 32-bit RV32IMC core architecture emphasizing design simplicity and high performance for embedded system applications. The SweRV approach prioritizes straightforward implementation while achieving competitive performance metrics suitable for real-time processing requirements. Andes Technology [109] provides a comprehensive range of processor cores including the A25 and D25F implementations, optimized for high-frequency operation and real-time processing applications. These cores incorporate advanced features such as branch prediction mechanisms and memory management units to support demanding embedded applications, with architectural designs specifically optimized for machine learning workloads. CloudBEAR [111] extends the RISC-V ecosystem through a comprehensive product portfolio spanning three distinct categories. Their microcontroller series includes 32-bit BM-310 (optimized for low power IoT applications) and 64-bit BM-610 (designed for secure boot and crypto acceleration). The embedded core family features both 32-bit implementations (BR-350, BR-351, BR-352) and 64-bit variants (BR-650, BR-651, BR-652), with the BR-352 and BR-652 representing second-generation cores offering improved performance while maintaining power efficiency. Their Linux-capable core complexes encompass BI-350 (32-bit tiny Linux processor), BI-651 (64-bit dual-issue pipeline), BI-652 (second-generation dual-issue), and BI-671 (out-of-order pipeline for maximum single-thread performance), providing solutions specifically optimized for IoT applications, machine learning workloads, and advanced networking requirements. Syntacore [103] offers a broad spectrum of core implementations including SCR1 and SCR7, covering applications from resource-constrained embedded systems to high-performance computing platforms. Their configurable core architectures feature deep pipeline implementations capable of supporting Linux-based operating systems and complex application environments, with specialized extensions for machine learning acceleration. Across academic and commercial evidence, deployed RISC-V ML solutions are strongest in edge and embedded domains where customization, power efficiency, and cost control dominate design priorities. At the same time, server-class and transformer-oriented support is emerging but remains comparatively early, indicating a transition from edge-centric deployment toward broader workload coverage.
4
Software Frameworks and Stacks
4.1
ML Acceleration Specialization and Co-Design
Machine learning accelerators have become indispensable in modern computing environments and satisfy the growing demand for efficient model implementation across diverse applications. This section is specialization-oriented: it covers FPGA-based designs, custom hardware acceleration flows, and RISC-Vspecific implementation paths, while preserving explicit co-design discussion where hardware and software decisions are jointly optimized. A prominent approach involves the use of FPGA-based accelerators. Past works have proposed frameworks that generate customized benchmark circuits to evaluate FPGA architectures optimized for machine learning workloads [60]. This framework facilitates an in-depth exploration of the design tradeoffs and performance metrics that are important for deploying efficient accelerators on FPGA platforms.
17
In addition, tools such as chisel4ml focus on automating the generation of FPGA implementations of highly quantized neural networks, with an emphasis on low latency and throughput for applications such as real-time event filtering and network intrusion detection[61]. Xilinx Vitis AI provides a comprehensive development platform for deploying ML models on Zynq UltraScale+ FPGA devices, featuring quantization-aware training workflows and integration with the Deep Learning Processing Unit (DPU) architecture for CNN acceleration [116]. HLS4ML translates neural network architectures directly to High-Level Synthesis code, enabling automated FPGA implementation for low-latency inference in particle physics and embedded applications. The framework supports integration with RISC-V based SoC designs, facilitating hardware-software co-design for ML acceleration [117]. The FINN framework specializes in deploying quantized and binarized neural networks on FPGAs through dataflow-style architectures, achieving ultra-low latency inference for resource-constrained edge applications [118]. In addition to FPGAs, there has also been a focus on custom hardware accelerators. For example, the general MYHDL-based design flow for the hardware implementation of deep neural network pipelines provides a method for converting deep neural network models to linear HDL code for efficient FPGA implementation. [62]. This approach reduces latency and optimizes resource utilization, making it suitable for real-time applications in datacenters and edge devices. In addition, frameworks such as GAHLS LegUp address the challenge of combining high-level applications on domain-specific hardware accelerators. By automating dependency graph construction and memory optimization, GAHLS offers significant performance improvements and energy efficiency compared with traditional high-level synthesis tools [63]. These advances are critical for meeting the computational demands of complex machine learning tasks in various domains from graph analysis and deep learning inference. In the field of embedded systems and IoT, ReNode’s focus on energy-efficient AIoT systems integrates specialized hardware accelerators, such as ASICs and FPGAs, to power machine learning algorithms in resource-constrained environments. [64]. This approach enables AIoT devices to perform complex inference tasks while adhering to power constraints and improving performance and energy efficiency. 4.1.1
RISC-V Specific ML Accelerator Implementations
Several frameworks have been developed specifically targeting the RISC-V ISA for ML acceleration. The [12] soft RISC-V vector processor for Edge-AI demonstrates how RISC-V’s vector extensions can be optimized for ML workloads. The [14] DARKSIDE heterogeneous RISC-V compute cluster showcases extreme-edge DNN inference and training capabilities. Additionally, [15] XpulpNN enables energyefficient inference of quantized neural networks specifically on RISC-V based IoT end nodes, while [16] RVTensor provides a lightweight neural network inference framework built around RISC-V ISA principles. These implementations demonstrate the practical viability of RISC-V for ML acceleration with quantified performance improvements. 4.1.2
Compiler and Runtime Integration for RISC-V ML
Compiler and runtime frameworks complement accelerator specialization by mapping ML graphs and tensor kernels efficiently onto RISC-V targets. Additionally, efforts to optimize compiler technologies in frameworks such as MLIR and Glow will play a key role in accelerating machine learning workloads on heterogeneous hardware platforms, including RISC-V targets. MLIR-based approaches enable the efficient compilation and optimization of neural network models using techniques such as tensor packing and microkernel decomposition to improve the performance of CPUs and GPUs [68]. Similarly, Glow’s graph reduction techniques and runtime optimization simplify the implementation of neural networks on different hardware supports and ensure high performance and scalability [69]. Consequently, machine learning accelerators include various technologies and methods aimed at optimizing performance, energy efficiency, and scalability on different computing platforms. From FPGAbased accelerators to custom hardware designs and advanced compiler frameworks, these accelerators help meet the computational demands of modern machine-learning applications. Table 8 provides a comprehensive summary of machine learning accelerators for RISC-V implementations. The optimization of machine learning workloads for RISC-V ISA requires sophisticated compiler techniques that leverage the unique characteristics of the instruction set. Modern ML compilation frameworks have developed specialized approaches to extract maximum performance from RISC-V processors,
18
Table 8: Summary of Hardware Accelerator Frameworks for RISC-V ML Implementations
Framework/Tool
Description
Target Platform
FPGA DNN Framework [60]
Customized benchmark circuits for FPGA evaluation Automated FPGA implementation of quantized neural networks DPU-based CNN acceleration on Zynq UltraScale+ NN-to-HLS translation for low-latency FPGA inference Dataflow architecture for binarized/quantized NN Deep neural network pipeline implementation in HDL High-level synthesis for domain-specific accelerators Energy-efficient AIoT systems with specialized accelerators Soft RISC-V processor optimized for Edge-AI Heterogeneous RISC-V compute cluster for extreme-edge DNN Energy-efficient quantized NN inference on RISC-V IoT Lightweight neural network inference framework
FPGA
Chisel4ML [61] Vitis AI [116] HLS4ML [117] FINN [118] MYHDL Design Flow [62] GAHLS LegUp [63] ReNode AIoT [64] RISC-V Vector Processor [12] DARKSIDE Cluster [14] XpulpNN [15] RVTensor [16]
FPGA FPGA FPGA FPGA FPGA FPGA/ASIC Embedded/IoT RISC-V RISC-V RISC-V RISC-V
particularly those equipped with vector extensions and custom neural processing capabilities. TVM’s tensorization framework represents a fundamental approach to RISC-V ML optimization, providing mechanisms to map high-level tensor operations directly to efficient hardware primitives through template matching for hardware intrinsics with automatic scheduling. The tensorization process addresses the critical challenge of bridging the semantic gap between abstract tensor computations and concrete RISC-V instruction sequences by defining tensor intrinsic templates that match computational patterns (e.g., GEMM, convolution) to optimized RVV instruction sequences. TVM’s auto-scheduler analyzes the computational graph and generates candidate schedules, evaluating each using a learned cost model trained on RISC-V performance characteristics, then selecting the optimal tiling factors, loop ordering, and vectorization strategies. This mapping is particularly effective for RISC-V vector extensions, where TVM’s compute primitives can be systematically transformed into optimized RVV instruction sequences for matrix multiplication and convolution operations. The framework’s memory access optimization capabilities align naturally with RISC-V’s load/store architecture, enabling sophisticated data flow analysis that minimizes cache misses and maximizes memory bandwidth utilization. Furthermore, TVM’s integration with emerging precision scalar extensions enables efficient execution of quantized neural networks, where reduced-precision arithmetic can be exploited without significant accuracy degradation. MLIR’s Linalg dialect provides another crucial foundation for RISC-V ML compilation, offering structured operations that undergo progressive lowering through carefully designed intermediate representations via transformation passes (affine to SCF to LLVM dialect). The compilation pipeline transforms high-level tensor operations through multiple stages: (1) Linalg operations express tensor computations as structured loop nests, (2) Affine dialect applies polyhedral transformations for loop tiling and fusion, (3) SCF (Structured Control Flow) dialect converts to explicit control flow, (4) Standard dialect provides generic operations, (5) LLVM dialect maps to RISC-V intrinsics, and finally (6) LLVM backend generates RVV assembly. Each transformation stage applies specific optimizations tailored to the target architecture. The affine dialect’s loop transformation capabilities are particularly valuable for RISC-V targets, enabling sophisticated optimizations including tiling strategies that respect register file constraints and cache hierarchy characteristics. The integration with MLIR’s vector dialect provides portable vectoriza-
19
Table 9: Summary of Hardware Integration Frameworks for RISC-V ML
Framework
Key Features
RISC-V Integration
GLOW [119]
Gradient-based tensor optimization, differentiable compilation
IREE [120]
MLIR-based compilation, multi-target deployment Tensorization framework, auto-tuning
Dedicated RISC-V backend with custom ML extension support RISC-V backend with vector extension optimization RVV instruction mapping, memory access optimization Linalg-to-RISC-V lowering, vector dialect integration Register file and cache hierarchy awareness RVV-specific optimizations
TVM MLIR Affine Dialect Vector Dialect
Progressive lowering, structured operations Loop transformation, tiling optimization Portable vectorization across implementations
tion across different RVV implementations, ensuring that optimization strategies remain effective across the diverse landscape of RISC-V vector processors. Microarchitectural considerations play a critical role in determining the effectiveness of RISC-V ML compiler optimizations. The diversity of RISC-V implementations, ranging from simple in-order designs to complex out-of-order processors, necessitates adaptive optimization strategies. Pipeline depth optimization requires careful instruction scheduling that considers the specific latency characteristics of the target implementation, while custom function unit integration demands sophisticated analysis of coprocessor interfaces and custom instruction semantics. Memory subsystem awareness remains paramount, as optimization strategies must adapt to the varying cache configurations and memory hierarchy implementations found across different RISC-V processor designs. Hardware integration into modern computing frameworks, such as GLOW, IREE, and MLIR, is critical to optimizing performance and efficiency across different hardware platforms, with particular emphasis on RISC-V ISA integration. These frameworks are designed to efficiently compile and run machine-learning models on specialized hardware accelerators, including RISC-V based systems with custom ML extensions. The integration of RISC-V processors into these frameworks requires specialized backend implementations that leverage the unique characteristics of the instruction set. GLOW framework employs gradient-based optimization techniques to transform tensor programs into differentiable representations, with dedicated backends developed specifically for RISC-V targets [119]. This approach transforms discrete search spaces into continuous, differentiable representations, enabling rapid exploration and optimization of critical operations for RISC-V ML accelerators. The resulting optimization process significantly enhances the scalability and adaptability of machine learning models in RISC-V hardware environments. IREE provides RISC-V backend support through its MLIR-based compilation infrastructure, generating executable binaries optimized for RISC-V targets with custom ISA extensions. The IREE-RISC-V integration employs multilevel intermediate representations and automatic code generation strategies to ensure efficient execution on RISC-V CPUs equipped with vector extensions and custom neural processing units [120]. This approach improves computational performance and simplifies deployment of machine learning algorithms on RISC-V embedded systems and IoT devices, where resource constraints demand careful hardware integration strategies. The RISC-V backend optimizations encompass several areas including automatic vectorization for RISC-V Vector instructions in neural network kernels, support for user-defined ML acceleration instructions through RISC-V’s custom instruction space, and cache-aware code generation tailored to specific RISC-V memory subsystem configurations. These optimizations enable efficient utilization of RISCV processors for machine learning workloads while maintaining the flexibility and customizability that characterizes the RISC-V ecosystem. Table 9 summarizes the key hardware integration frameworks for RISC-V ML applications. TVM addresses automatically tuning deep learning compilers to optimize tensor arithmetic code on different hardware platforms. The One-Shot TVM tuner uses a neural predictor-inspired approach to reduce auto-tuning overhead and speed up compilation [121]. This approach enables efficient deployment
20
Table 10: Comparative Analysis of ML Compilation Approaches for RISC-V Approach
Compilation
RISC-V Awareness
Maturity
Workloads
Trade-offs
TVM
Static (AOT) + JIT
Native backend
Production
Auto-tuning overhead vs. portability
MLIR/IREE
Static (AOT)
Native backend
Production
GLOW
Static (AOT)
Experimental
Research
Hand-written RVV
Static
Maximum
N/A
CNN, Transformer, quantized models General tensor ops, custom accelerators CNN, quantized inference Custom kernels
Compiler Intrinsics
Static
High
Stable
Performancecritical loops
Compilation complexity vs. flexibility Limited RISC-V support vs. graph optimization Development effort vs. peak performance Portability vs. fine-grained control
of machine learning models on various hardware architectures, improving performance and scalability in real-world applications. MLIR focuses on automating production of hardware accelerators for high-level programming frameworks. MLIR facilitates synthesis of optimized hardware from high-level code representations using tools such as SODA-OPT, supporting FPGA and ASIC implementations with improved performance and energy efficiency [73]. ReNode integrates specialized hardware accelerators, such as ASIC and FPGA, to optimize algorithm performance while ensuring efficient deep learning and security for AIoT applications[64]. This approach highlights the importance of hardware integration to improve energy efficiency and computing performance in AIoT environments. These frameworks represent different approaches to hardware integration, each focusing on optimizing machine learning workloads in specific hardware environments. The frameworks enable efficient and scalable execution of machine learning models on diverse hardware platforms through gradient-based optimization, multilevel intermediate representation, auto-tuning techniques, or specialized hardware accelerators. Table 10 provides a comparative analysis of ML compilation approaches for RISC-V, contrasting framework-based solutions with hand-written intrinsics. The choice between these approaches involves trade-offs between development effort, performance portability, and optimization potential. TVM and MLIR offer automated optimization with broad workload coverage but may not achieve peak performance on specific targets. Hand-written intrinsics using RVV assembly or compiler built-ins provide maximum control and performance but require significant development effort and lack portability across RISCV implementations. Hybrid approaches, where frameworks generate baseline code that is selectively optimized with intrinsics for critical kernels, often represent a practical middle ground for production deployments.
4.2
Compiler Optimization
Compiler optimization is important for improving the performance and efficiency of deep learning compilers, particularly for architectures such as RISC-V. This process involves the use of various techniques to convert and simplify the generated code such that it runs better on a hardware platform. Several frameworks and tools have been developed to address the complexities associated with compiler optimization in machine learning algorithms through systematic optimization cycles that include code generation, optimization application, testing and verification, deployment, performance monitoring, and feedback acquisition. A foundational aspect of RISC-V compiler optimization for ML workloads is auto-vectorization, where the compiler automatically transforms scalar loop code into vector instructions targeting the RISC-V Vector extension (RVV). LLVM 16 (March 2023) was the first compiler release to enable scalable autovectorization by default for RISC-V targets with the V or Zve extensions, while GCC 14 (May 2024) introduced loop and SLP (Superword-Level Parallelism) vectorization for RVV. The primary technical
21
challenge is RVV’s vector-length agnostic (VLA) design, where the hardware vector length (VLEN) is not known at compile time and varies across implementations (e.g., 128-bit to 1024-bit). This differs fundamentally from fixed-length SIMD ISAs and requires the compiler to generate length-independent code using vsetvli instructions to configure vector length at runtime. Current auto-vectorization support covers standard loop patterns common in ML inference (element-wise operations, reductions, strided memory access), though hand-optimized RVV intrinsics still outperform auto-vectorized code for complex kernels such as matrix multiplication and convolution. A notable framework, TVM, uses advanced optimization techniques, such as auto-tuning and operator fusion, to optimize tensor operations on different parts of the hardware. [71]. The TVM approach uses objective-independent and objective-specific optimizations to generate high-performance codes through efficient scheduling and resource allocation. The goal of these optimizations is to minimize the runtime and maximize the hardware use. This is important for the effective deployment of machine-learning models on RISC-V processors. Another important contribution is the MLIR framework, which focuses on generating modular and customizable codes for tensor compilers. MLIR enables hierarchical decomposition and integration of operations and facilitates the efficient translation of high-level abstractions into optimized machine code. [122]. Using MLIRs leading IR design, developers can optimize tensor computations and achieve performance improvements on various hardware targets, including RISC-V processors. In addition, frameworks such as Glow and IREE extend compiler optimization to support heterogeneous hardware environments. Glow used a multiphase compilation pipeline to optimize computational graphs in multilevel intermediate representations through target-specific optimizations and automatic code generation [69]. In contrast, IREE focuses on high-performance AI compilation using MLIR and Linalg upstream dialects, with an emphasis on cache-aware tensor packing and CPU-efficient vectorization [120]. Additionally, research efforts such as those described in Felix: Optimizing Tensor Programs with Gradient Descent aim to address the grand search space challenges associated with optimizing tensor programs [119]. Felix employs hierarchical search space decomposition, partitioning the optimization space into architectural decisions (layer fusion, tiling) and mapping decisions (memory allocation, scheduling). This decomposition reduces search complexity from O(n!) to O(n×m) where n represents architectural choices and m represents mapping options, enabling efficient exploration through genetic algorithms that evaluate candidate solutions using cost models calibrated to RISC-V performance characteristics. In addition, innovative methods such as transformation testing ensure the correctness and reliability of compiler optimizations [123]. The proposed method automatically generates different DNN models and compares their outputs to identify the compilation errors. It is important to maintain optimization stability in complex compiler frameworks. Consequently, compiler optimization is important for maximizing the performance and efficiency of machine learning algorithms on RISC-V processors. Using advanced techniques in frameworks such as TVM, MLIR, and Glow, developers can significantly improve code generation, execution speed, and resource use. These optimizations are necessary to unlock the full potential of machine learning applications in embedded and edge computing environments and drive innovation in AI-enabled devices.
4.3
Metaheuristic Algorithms
Metaheuristic algorithms have garnered significant attention for optimizing machine-learning models, particularly within the realm of frameworks and software stacks for manufacturing chips using the RISCV ISA. One notable methodology is the Agile Optimization Framework (AOF), which aims to tackle inefficiencies in deep learning compiler optimization. AOF leverages Beluga Whale Optimization (BWO), the Evolution Epsilon Strategy (EES), and a Tuning Accelerator (TA) to enhance performance and reduce optimization time. Through the integration of these components, the AOF effectively balances exploration and exploitation during the optimization process, thereby improving hardware performance in a cost-effective manner[70]. This optimization approach demonstrates the potential for metaheuristic algorithms to enhance machine learning model deployment on RISC-V processors through systematic performance improvements and reduced computational overhead. The use of meta-heuristic algorithms, such as the Beluga Whale Optimization (BWO) algorithm, plays an important role in the search process in the AOF. BWO is particularly effective for navigating complex search spaces associated with tensor operator optimization. The evolving nature of BWO combined with EES ensures that the algorithm adapts over time, thereby improving its ability to determine optimal
22
Table 11: Summary of Metaheuristic Algorithms for RISC-V ML Optimization
Algorithm/Framework
Description
Optimization Target
Agile Optimization Framework (AOF) [70] Beluga Whale Optimization (BWO) [70] Evolution Epsilon Strategy (EES) [70]
Tackles inefficiencies in deep learning compiler optimization Navigation of complex search spaces for tensor optimization Adaptive algorithm behavior enhancement Predictive model-based optimization using LightGBM Auto-tuning and operator fusion optimization
RISC-V ML deployment Tensor operator optimization Algorithm adaptation Compilation time reduction TVM performance enhancement
Tuning Accelerator (TA) [70] TVM Framework Integration [70]
solutions. This compatibility is essential for optimizing the performance of machine learning models on RISC-V-based chips, where hardware constraints and performance requirements are critical [70]. Furthermore, the integration of tuning accelerators (TAs) into the AOF demonstrates the importance of predictive models in terms of reducing the optimization time. The TA method uses LightGBM to predict tensor performance and minimizes the need for extensive compilations. This approach not only accelerates the optimization process, but also improves the overall efficiency of the TVM framework. The proposed AOF significantly improves the feasibility of deploying machine learning models on RISC-V chips by reducing the computational overhead associated with traditional optimization techniques [70]. Meta-heuristic algorithms can discover and effectively exploit large search spaces and are necessary to optimize the behavior of neural networks in the context of the RISC-V ISA. The use of these algorithms in frameworks, such as the AOF framework, reveals the potential to improve hardware performance and reduce development time. As this field continues to evolve, further improvement of meta-heuristic strategies and their integration into optimization frameworks are critical for advancing the design and fabrication of efficient machine learning chips [70]. Table 11 summarizes the metaheuristic algorithms used for RISC-V ML optimization. The metaheuristic scope in this subsection is intentionally limited to optimization algorithms (AOF/BWO/EES/TA). Broader framework ecosystems (TVM, MLIR, HLS/FPGA flows, and co-design toolchains) are discussed in Sections 4.1, 4.2, and 4.4 to maintain conceptual separation.
4.4
System-on-Chip Frameworks
System-on-chip (SoC) designs integrate multiple functions into a single chip, optimizing compactness, energy efficiency, and performance in applications from mobile devices to embedded systems. Recent advances in SoC frameworks such as LiteX have democratized FPGA-based designs by providing opensource solutions that facilitate complex system integration [75]. SoC frameworks like LiteX face challenges in electronic design automation (EDA). These frameworks use high-level synthesis (HLS) techniques, such as those in LegUp, focusing on optimizing hardware implementations using precision pointer synthesis and micromemory coupling [76]. These optimizations improve performance of FPGA-based implementations, reduce footprint and latency, and make them suitable for resource-constrained environments such as IoT devices. The flexibility of Migen and LiteX enables rapid prototyping of SoCs. This is useful for academic research and small production environments where quick response and cost efficiency are important [77]. These frameworks leverage Python DSLs and integrate open-source tools to facilitate design and deployment of custom SoCs and foster innovation in hardware development. ReNode contributes to AIoT systems by integrating energy-efficient deep learning techniques into modular IoT platforms. Using heterogeneous computing and specialized hardware accelerators, ReNode optimizes algorithms while ensuring scalability and security across distributed IoT networks [64]. This approach balances computational efficiency with resource constraints and supports AI applications from smart homes to industrial automation. For FPGA-based implementations, MYHDL provides a design flow for accelerating Deep Neural Net-
23
Table 12: Summary of System-on-Chip (SoC) Frameworks for RISC-V ML
Framework
Description
Key Features
LiteX [75] LegUp HLS [76]
Open-source FPGA-based SoC framework High-level synthesis optimization
Migen [77]
Python-based SoC design framework
ReNode [64]
Energy-efficient AIoT platform
MYHDL [62]
Hardware description language for DNN
Chisel [78]
Hardware description language
Democratizes FPGA design, supports RISC-V integration Precision pointer synthesis, micro-memory coupling Rapid prototyping, DSL integration Heterogeneous computing, specialized accelerators Pipelined HDL code generation, FPGA optimization Parameterizable circuit generators, meta-design
works (DNN) through pipelined HDL code. This technique improves FPGA resource efficiency while maintaining model accuracy, which is important for real-time applications in datacenters and edge computing [62]. MYHDL simplifies development by automating conversion of DNN models into hardwarespecific implementations, making FPGA acceleration accessible to developers with limited hardware expertise. Chisel facilitates hardware and software co-design through its capabilities as a hardware description language that enables design space exploration. By providing high-level constructs for hardware design and meta-design techniques, Chisel enables creation of parameterizable circuit generators that can be tuned to optimize performance and energy efficiency for various applications [78]. This approach supports iterative development cycles and allows designers to describe and modify SoC architectures according to evolving requirements and technological advances. Advances in SoC frameworks such as LiteX, Migen, ReNode, MYHDL, and Chisel drive innovation in electronic design and enable scalable and efficient solutions for various applications. These frameworks increase hardware performance and security, democratize access to FPGA-based development, and foster a community-driven approach to SoC design and implementation. Table 12 provides a comprehensive comparison of System-on-Chip frameworks for RISC-V ML applications.
4.5
Embedded Systems and Edge AI
To investigate embedded systems, researchers have applied various frameworks and methods to optimize hardware and software integration in different applications. Hardware and software co-design using chisels simplifies Mel frequency factor calculations (MFCC) for keyword spotting systems [74]. This method optimizes the MFCC algorithm for embedded applications while integrating hardware accelerators, which are important for maintaining accuracy in resource-constrained environments [74]. FPGA architectures have been explored using PyMTL for deep neural network (DNN) acceleration. Researchers have developed automated tools, such as chisel4ml, to generate FPGA implementations of highly quantized neural networks with low latency requirements [61]. This development supports realtime applications, such as CERN’s Large Hadron Collider trigger system, and demonstrates the efficiency of FPGA-based solutions in high-throughput environments [61]. Frameworks such as IREE optimize machine learning inference in embedded devices. This study highlights the role of IREE in deploying machine learning models on different hardware platforms, demonstrating an integrated compiler and runtime stack for performance optimization [79]. This is important for applications requiring real-time decision-making capabilities in resource-limited environments [79]. In addition, Migen is exploring the integration of lightweight cryptographic cores into system-on-chip (SoC) designs for Internet of Things devices. The researchers implemented a 32-bit RISC-V processor with cryptographic accelerators, such as PRINCE and ChaCha, in an FPGA environment [77]. This approach increases the security and efficiency of IoT deployments and overcomes the severe latency and resource constraints inherent in IoT applications [77]. Additionally, the use of LegUp revolutionizes high-level synthesis (HLS) tools by enabling comprehen24
Table 13: Summary of Embedded Systems Frameworks for RISC-V ML
Framework/Tool
Description
Application Domain
Chisel Co-Design [74]
Hardware-software co-design for MFCC calculations Automated FPGA implementation of quantized DNNs ML inference optimization for embedded devices Lightweight cryptographic cores in RISC-V SoCs High-level synthesis with software testing techniques
Keyword spotting, speech processing High-throughput real-time systems Real-time decision-making
PyMTL + chisel4ml [61] IREE Framework [79] Migen Crypto Integration [77] LegUp HLS Validation
IoT security, crypto acceleration FPGA design verification
sive validation with software testing techniques, such as LibFuzzer and KLEE. This approach ensures the quality and reliability of FPGA designs produced by HLS and bridges the gap between software verification and hardware implementation. Such verification frameworks are essential for validating the functionality and performance of FPGA implementations in critical applications. These frameworks and methods highlight the evolution and innovation of embedded system design, from optimizing hardware accelerators and cryptographic cores to validating FPGA implementations using advanced HLS tools. The integration of these technologies not only improves performance and efficiency, but also expands the applicability of embedded systems in various domains, including IoT security, real-time data processing, and machine learning inference in edge devices. Table 13 summarizes the embedded systems frameworks for RISC-V ML applications. Edge AI advances the deployment of machine learning models directly on edge devices by locally processing data to enhance privacy and reduce latency [69]. Frameworks such as GLOW optimize model execution through graph-based techniques and quantization, significantly boosting inference speed and energy efficiency [69]. To support cross-platform deployment, IREE extends edge AI capabilities with an integrated runtime and compilation strategy, enabling models from frameworks such as TensorFlow and PyTorch to run efficiently on minimal high-performance hardware [79]. The TVM complements this by offering optimization strategies for deep learning workloads in resource-constrained environments, integrated with hardware acceleration libraries to enhance model performance and efficiency across a variety of devices [124, 125]. MLIR contributes through its modular code generation and optimization capabilities, allowing the efficient compilation of machine learning models across different hardware platforms [122]. Similarly, ReNode focuses on energy-efficient AIoT applications by providing a modular and scalable hardware platform that integrates specialized accelerators for real-time processing with minimal power consumption [64, 126]. PyMTL and Chisel contributed by exploring FPGA-based accelerators and hardware/software codesign, respectively, to improve performance for tasks such as audio feature extraction and keyword detection [60, 61, 72, 74]. MYHDL accelerates the deployment of AI solutions by automating the conversion of DNN models into FPGA-compatible HDL code, reducing latency, and enhancing accuracy for real-time applications in healthcare and industrial automation [62, 127]. LegUp further supports edge AI with frameworks such as GAHLS, optimizing high-level applications for hardware accelerators through graph analysis and memory design [63]. Collectively, these frameworks address various optimization, deployment, and performance challenges, driving the widespread adoption of AI in diverse edge applications. Table 14 provides a comprehensive summary of Edge AI frameworks for RISC-V ML systems.
5
Applications and Evaluation of RISC-V in Machine Learning
This section presents real-world applications and evaluation of RISC-V implementations for machine learning. We examined 21 implementations covering custom RISC-V designs tailored for ML workloads, 25
Table 14: Summary of Edge AI Frameworks for RISC-V ML
Framework
Description
Key Capabilities
GLOW [69]
Graph-based optimization for edge deployment Integrated runtime and compilation for edge AI
Graph optimization, quantization Cross-platform deployment, TensorFlow/PyTorch support Hardware acceleration integration Multi-platform code generation Modular accelerators, real-time processing DNN acceleration, chisel4ml integration FPGA deployment, latency reduction Graph analysis, memory optimization
IREE [79]
TVM [124] MLIR [122]
Deep learning optimization for resource-constrained devices Modular compilation infrastructure
ReNode [64]
Energy-efficient AIoT platform
PyMTL [60]
FPGA-based accelerator exploration
MYHDL [62]
DNN-to-HDL automated conversion
LegUp GAHLS [63]
High-level synthesis for edge accelerators
organized into application domains (medical, robotics, object detection), performance evaluation, and security analysis subsections. This exploration provides understanding of how RISC-V’s open architecture can be adapted to address specific needs of machine learning algorithms, leading to advancements in this field. Table 16 summarizes the cases in this section. The comparative evaluation of RISC-V machine learning implementations requires standardized benchmarking methodologies to enable meaningful cross-platform performance assessment [14, 53, 82]. Our analysis synthesizes performance data from reviewed literature to establish unified metrics for RISCV ML system evaluation. This benchmarking framework addresses current fragmentation in evaluation methodologies that makes direct comparisons between different RISC-V ML implementations challenging. Performance standardization across reviewed implementations reveals key metrics that appear in RISC-V ML research [54, 55, 83]. Inference throughput, measured in inferences per second (IPS), provides the primary metric for computational performance assessment across comparable model architectures. Energy efficiency metrics, quantified as Giga-Operations Per Watt (GOP/W), are important for power-constrained edge AI applications where RISC-V processors operate [82, 128]. Memory efficiency evaluation through peak memory usage and bandwidth utilization is important given the resource constraints typical in RISC-V deployment scenarios [129, 130]. Latency characteristics, measured as end-to-end processing time from input acquisition to output generation, provide insights for real-time application suitability [80, 131]. Standardized ML workloads enable consistent performance comparison across diverse RISC-V implementations [53, 83, 132]. Based on prevalence in reviewed literature, our unified benchmarking framework incorporates MobileNetV2 inference tasks for image classification scenarios, keyword spotting workloads for audio processing applications, CNN-based classification tasks for general computer vision evaluation, and object detection scenarios for complex multi-object recognition challenges. These workloads represent the most commonly evaluated ML applications in RISC-V research and provide diversity to assess system performance across different computational patterns. Cross-implementation performance analysis demonstrates advantages of specialized RISC-V ML configurations over general-purpose implementations [14, 15, 84]. The DARKSIDE cluster achieves 1.2 TOPS/W for 8-bit DNN inference, representing state-of-the-art energy efficiency for ultra-low-power neural network processing [14]. The XpulpNN framework demonstrates 15-fold energy efficiency improvements compared to ARM Cortex-M4 baselines, highlighting the benefits of RISC-V customization for ML workloads [15]. RVTensor implementations show 3.2-fold speedup improvements compared to unoptimized RISC-V configurations, emphasizing the importance of ML-specific architectural enhancements [133]. The DIANA SoC achieves 2.5-fold energy efficiency improvements for mixed-signal neural network processing, demonstrating the effectiveness of heterogeneous RISC-V designs for specialized ML
26
Table 15: Unified benchmarking analysis of RISC-V ML implementations: Performance comparison across reviewed research papers showing standardized metrics for energy efficiency, computational throughput, and implementation characteristics. EnerE = Energy Efficiency(GOP/W) ,PerfI: Performance Improvement
Implementation
ML Workload
EnerE
PerfI
Memory Usage
Comparison Baseline
DARKSIDE Cluster
8-bit DNN Inference
1200
8.5x
32KB
XpulpNN Framework
CNN Classification
850
15x
64KB
RVTensor
Image Processing
420
3.2x
128KB
DIANA SoC CFU Playground
Mixed-Signal NN MobileNetV2
680 320
2.5x 4.1x
256KB 512KB
HULK-V System
Keyword Spotting
1050
6.8x
1MB
ICU4SAT
Object Detection
290
2.8x
2MB
General RISC-V ARM Cortex-M4 Baseline RISC-V Standard SoC VexRiscv Baseline Linux-capable RISC-V Space-grade baseline
applications [134]. The CFU Playground achieves 4.1-fold performance improvements for MobileNetV2 inference compared to VexRiscv baselines, showcasing the benefits of custom functional units for deep learning tasks [53]. The HULK-V system demonstrates 6.8-fold energy efficiency improvements for keyword spotting applications compared to Linux-capable RISC-V processors, highlighting the advantages of lightweight RISC-V cores for edge AI scenarios [82]. The ICU4SAT implementation achieves 2.8fold speedup improvements for object detection tasks in space-grade RISC-V processors, illustrating the potential of RISC-V processors for high-reliability ML applications [129]. Analysis across reviewed implementations reveals performance improvements ranging from 2-15 times in energy efficiency and 1.5-5 times in computational throughput compared to general-purpose RISC-V baselines [53, 54, 82]. These improvements demonstrate the potential of customized RISC-V implementations for machine learning applications, particularly in edge computing scenarios where energy efficiency and computational density are important concerns. Table 15 presents a unified benchmarking analysis of RISC-V ML implementations across reviewed research papers. This section analyzes papers based on their field of applications, examining three primary domains: (1) Medical field applications focusing on real-time diagnostics and wearable health monitoring, (2) Robotics applications emphasizing computer vision and autonomous navigation, and (3) Object detection systems for security and surveillance applications. This analysis reveals how RISC-V implementations adapt to specific application requirements and performance constraints across diverse deployment scenarios. In the medical field, we highlight real-time anomaly detection in wearable devices and endoscopic image classification. We consider low-power RISC-V cores and efficient deep-learning accelerators for specific tasks. In robotics, we examined two projects that showcase how RISC-V cores with custom hardware extensions can be used for image processing and obstacle avoidance in robots. For object detection, we investigated the role of RISC-V in deep learning applications. We explored a driver drowsiness detection system, hardware trojan detection using RISC-V soft cores, and fruit ripeness identification using deep neural networks on RISC-V processors. The following subsections describe these in detail.
5.1
Medical Applications
RISC-V processors have shown promise in real-time anomaly detection for low-power wearables and endoscopic image classification [84, 128, 135]. Their energy efficiency makes them suitable for extended use, and their open-source nature allows customization for tasks such as image classification, improving processing speed and accuracy. Medical applications benefit from RISC-V’s low power consumption and customizable architecture, enabling specialized processing units for real-time health monitoring, diagnostic imaging, and wearable medical devices. Choi et al. [135] proposed a day-night RISC-V processor
27
Table 16: Specific applications of RISC-V processors for machine learning: This table comprehensively summarises the nine research papers reviewed in our survey on RISC-V applications. Each entry includes a concise description of its key findings and a breakdown of its important aspects, such as the targeted application domain, methodology used, and key performance metrics evaluated.
Paper
Summary
Keys
[84]
A deep learning accelerator is proposed for endoscopic image classification Low-power wearable device made from Honey-Bunny PULP processor A new Day-Night architecture is proposed to consume low power in wearable devices. RVcar (A mini motor car) with a deep neural network accelerator is introduced to detect objects and make real-time decisions Fruits are identified by applying DNN models in RISC V A two-wheeled self-balancing robot (TWSBR) is built on the wheel-type inverted pendulum A Driver Drowsiness Detection System is proposed using Ibex core and FPGA port Detecting and classifying Hardware Trojans Uses FAMA maturity index to identify the fruit size and its maturity level
Medical Field, Low Latency, Object Detection, Low Power Medical Field, Low Power
[128] [135] [136]
[137] [138] [139] [140] [141]
Medical Field, Low Power Object Detection, Low power, Robotics Inference Speed, Low Power, Object Detection, Robotics Object Detection, Robotics Inference Speed, Object Detection, Robotics Robotics Object Detection, Robotics
design with separate segments for wearable applications and real-time anomaly detection. A prototype on an FPGA showed a 57.5% reduction in energy consumption compared with traditional designs. For endoscopic imaging, Bolhasani et al. [84] developed an energy-efficient deep-learning accelerator (DLAE) with 256 PEs and evaluated it using the MASTERO simulator. It achieved 4.56 × 109 MAC energy and 1.73 × 107 cycles, outperforming other CNN models. To extend the battery life of wearable e-health devices, Eggimann et al. [128] introduced an energy-efficient platform using the Honey-Bunny PULP processor, integrating four RISC-V cores. It delivers 2.5 GOPS at 55 mW, making it ideal for always-on operation.
5.2
Robotics Applications
RISC-V processors are suited for real-time decision-making in autonomous vehicles due to their efficient instruction execution, enabling rapid data processing and quicker response times [136, 138]. Custom hardware extensions enhance the ability of RISC-V to process sensor data and perform real-time object analysis. Kanamori et al. [136] demonstrated a real-time decision-making system in an RVcar using a RISC-V soft processor (RVcoreP) and DNN accelerators for low-power, low-latency image processing. Implemented on a Xilinx Nexys A7 board, the system recognized traffic signals in 35 of 50 trials, with failures addressed by incorporating markers to improve object recognition and decision-making. Tsuda et al. [138] developed a two-wheeled self-balancing robot (TWSBR) controlled by a 32-bit RISC-V microprocessor (VexRiscv) equipped with custom 32-bit fixed-point instructions. Despite lacking an FPU, the robot’s control system uses fuzzy logic and multiple sensors, successfully achieving balance and obstacle avoidance on an FPGA.
5.3
Object Detection Applications
In recent years, neural networks, particularly CNNs, have been widely applied to object detection and image classification applications across diverse domains [137, 139–141]. Enhancing the accuracy of these models and reducing prediction time are key areas of focus for RISC-V implementations. Mousavikia et al. [139] developed a Driver Drowsiness Detection (DDD) system using CNNs on an RISC-V Ibex core. The system classifies images into four categories: distraction, natural, sleep, and yawning. Adding custom instructions improved the computational efficiency, achieving a 1.7× latency improvement. Nunes et al. [140] implemented RISC-V soft cores on FPGAs for Hardware Trojan detection, achieving perfect 28
accuracy by analyzing features from FPGA airstreams using machine learning models. For agricultural applications, K et al. [137] used DNNs (RCNN, YOLO v3, SSD) on RISC-V processors to identify fruit types. By applying FAMA maturity indices and K-means clustering, the system also determines fruit size and maturity, aiding farmers and food distributors in quality assessment and supply chain optimization [141].
5.4
Performance
Performance is crucial in machine learning systems for two main reasons: speed and accuracy. Faster analysis enables real-time applications and quicker training, while improved accuracy translates to better decision-making. In this section, we analyze papers based on their performance characteristics across three key areas: energy efficiency, inference speed-up, and low latency [53–55, 82, 83, 131]. In terms of energy efficiency, we discuss how researchers are creating low-power solutions for Internet of Things (IoT) devices that must operate under strict power constraints while maintaining computational performance. For inference speed, we explored optimization techniques to achieve faster execution on various hardware platforms through architectural innovations and software optimizations. For low latency, we analyzed approaches to reduce communication delays in edge devices, which are crucial for real-time applications requiring immediate response times. Table 17 provides a comprehensive analysis of RISC-V processor performance based on published research. The following subsections describe these performance aspects in detail. Recent LLM and transformer acceleration evidence is also emerging on RISC-V. VEXP introduces low-cost ISA support for softmax-heavy transformer kernels [22], while V-Seek reports optimized reasoning inference on a server-class general-purpose RISC-V platform [58]. These results indicate meaningful progress, although broad production maturity for large-scale LLM serving on RISC-V remains an open area. 5.4.1
Energy Efficiency:
IoT applications must balance performance with cost and power constraints. High-end systems use power-hungry SoCs, whereas low-end applications rely on microcontrollers. New trends demand systems that offer both low power and cost efficiency. We review three research efforts targeting low-power edge computing for machine learning, focusing on energy-efficient implementations that maintain computational performance while minimizing power consumption through specialized hardware designs, optimized software frameworks, and intelligent power management strategies. Balasubramanian et al. [54] optimized RISC-V processors for complex neural networks on resourceconstrained devices. They extended the RISC-V instruction set to address computational bottlenecks, achieving up to a 13x speedup and 11.7% code size reduction for image-processing tasks. Valente et al. [82] developed HULK-V, an open-source, energy-efficient SoC capable of running Linux. It combines a 64-bit RISC-V core with an 8-core Programmable Multi-Core Accelerator (PMCA), offering up to 112x faster performance for ML tasks and 157 GOps/W energy efficiency. Prakash et al. [53] introduced the CFU Playground, a framework for creating hardware accelerators on the VexRiscv core. It achieved a 55x speedup for MobileNetV2 and 75x for keyword spotting, offering a flexible, cost-effective alternative to ASICs for ML tasks. 5.4.2
Inference Speed:
This section explores methods to increase inference speed, which is crucial for machine learning performance in time-sensitive applications [55, 83, 129, 132]. Lee et al. [55] introduce TVMNIR, a compiler leveraging TVM and Mediatek NeuronPilot to accelerate AI tasks. TVMNIR optimizes quantization configurations, achieving up to 11× speedup for floating-point models and 70× speedup for quantized models. González et al. [83] presented a platform for evaluating machine learning accelerators using the Chipyard framework. Comparing the NVIDIA Deep Learning Accelerator (NVDLA) to the Gemmini Systolic Array Generator, NVDLA achieved up to 3.77× faster performance on the ResNet-50 benchmark. Lee et al. [132] proposed a sparse basis algorithm to reduce CNN computation time by transforming weights into a sparse space, leading to reduced storage requirements and faster execution times on VexRiscV implementations. Giuffrida et al. [129] developed ICU4SAT, a single-chip AI system for realtime satellite data processing. With an integrated RISC-V processor and soft GPU, it improves data
29
Table 17: Analysis of RISC-V processor performance based on published research: This table comprehensively summarises the research papers included in our performance analysis section. Each entry provides the paper citation, a concise description and its key findings related to RISC-V processor performance (e.g., inference speedup, low latency), and a breakdown of its important methodological aspects, such as the evaluation methodology and benchmark tools employed.
Paper
Summary
Tags
[82]
Heterogeneous ultra-low-power Linux-capable system that is both energy-efficient and capable of running a full-fledged Linux operating system The CFU Playground framework generates hardware accelerators based on VexRisc V CORE.
Low Power
[53]
[54]
Used a toolchain (PyTorch Graph Lowering (Glow)LLVM) to understand the impact of the compiled code of AI models Integrated TVMs quantization flow with the MediaTek Neuropilot AI accelerator
[55]
[83]
[132]
[129]
[130]
[131]
[80]
Integrate NVDLA into the Chipyard framework and compared its performance with that of the Gemmini Systolic Array Generator To reduce the extensive computation and time requirements of the convolution operation, the sparse basis approach was proposed ICU4SAT system (a satellite instrument control unit with an artificial intelligence engine on a single chip) is built RISC-LCAW is introduced to enable loosely coupled accelerators to be integrated as slave devices on the system bus. The lean network interface is developed for the load/store stage of Ariane for user-level communication. Experiments were conducted to assess the real-time performance of RISC-V processors in a robotic control system.
Inference SpeedUp, Hardware Accelerators, Low Power, Image Classification Inference SpeedUp, Low Power
Inference Speed Up, Low Latency, Hardware Accelerators Inference SpeedUp
Inference SpeedUp
Inference SpeedUp, Hardware Accelerators Low Latency
Low Latency
Low Latency
analysis throughput and memory usage efficiency through onboard AI-powered processing, demonstrating the potential for space-constrained applications. 5.4.3
Low Latency
The push for low-latency performance in edge devices is important for real-time applications to enhance speed, responsiveness, and efficiency. This section explores techniques for reducing latency in RISC-V processors, focusing on network interfaces, accelerator integration, and real-time control systems. These optimizations include custom packetizers with dual-port memory structures, FIFO-based communication mechanisms, and standardized accelerator wrapper interfaces that enable sub-microsecond response times for time-critical applications. Gianioudis et al. [131] presented a lean network interface for RISC-V processors, achieving submicrosecond latency (720 nanoseconds) through a custom packetizer with dual-port memory and FIFO structure, optimized for efficient packet processing. This design enhances communication speed between FPGA nodes, demonstrating improvements over traditional methods. Muchandi et al. [130] introduce the RISC-V Loosely Coupled Accelerator Wrapper (RISC-LCAW), a standardized interface for integrating accelerators into RISC-V SoCs. The RISC-LCAW, featuring a controller module, DMA interface, and input buffer, was tested using the AWRE-DNN accelerator. Despite a latency increase of 1.7× to 2.6× for various network sizes, the RISC-LCAW maintains acceptable performance while simplifying accelerator
30
integration. Yoo et al. [80] evaluate RISC-V’s real-time performance for robotic control systems. Their experiments compared RISC-V boards (HiFive1 Rev B and VisionFive 2) with ARM and x86-64 systems using FreeRTOS and Linux with Preempt-RT. The results showed RISC-V’s real-time performance, in Inter-Thread Communication (ITC) and motion control applications, with minimal performance gaps compared to x86-64 architectures. These advancements highlight the potential of RISC-V for low-latency applications, from efficient inter-node communication and accelerator integration to real-time robotic control. Table 18 summarizes the performance optimization techniques for RISC-V ML systems. Table 18: Summary of RISC-V Performance Optimization Techniques
Optimization Type
Technique/Framework
Performance Gain
Energy Efficiency
RISC-V ISA extensions for neural networks [54] HULK-V SoC with PMCA [82]
13x speedup, 11.7% code reduction 112x faster ML tasks, 157 GOps/W 55x speedup (MobileNetV2), 75x (keyword)
CFU Playground accelerators [53]
Inference Speed
TVMNIR quantization optimization [55] NVDLA vs Gemmini comparison [83] Sparse basis CNN optimization [132] ICU4SAT satellite AI system [129]
Low Latency
5.5
Lean network interface [131] RISC-LCAW accelerator wrapper [130] Real-time robotic control [80]
11x (FP), 70x (quantized) speedup 3.77x faster (ResNet-50) Reduced computation/storage Real-time processing improvement 720 nanoseconds latency 1.7-2.6x latency (acceptable) Superior ITC performance
Security
Hardware and physical access security are important in modern system architectures, especially for Machine Learning (ML) systems. Because ML often deals with sensitive data, vulnerabilities can be exploited to steal or manipulate information, resulting in biased or inaccurate models. RISC-V processors can be improved using their customizable architecture. This flexibility allows the incorporation of security extensions, such as memory protection and trusted execution environments. By isolating data and models within the core, RISC-V strengthens ML systems against unauthorized access and manipulation. The proliferation of Internet of Things (IoT) devices, in safety-critical applications, requires robust security measures to protect them from malicious attacks. This study [142] addresses this challenge by proposing the Hardware Immune System (HWIS), a hardware malware detection technique designed for microprocessor architectures. It is suited to low-power, resource-constrained embedded devices, which are commonly used in IoT networks. The proposed method excels at detecting botnet behavior with an accuracy of 96.7%. It maintains a low false-negative rate of 6.5%, demonstrating its effectiveness in identifying malicious activities. The HWIS achieved a high F1 score of 0.96, demonstrating balanced performance between precision and recall. In response to increasing attacks on microprocessors and SoCs, researchers are exploring new approaches to SoC security. Security features are bolted to existing systems, which often results in performance trade-offs during vulnerability patching. Researchers [143] proposed a security approach that integrates security-first design principles, security-aware testing methods, and a quantified analysis of performance impact. They achieved this by designing and implementing a secure RISC-V SoC. A secure SoC incorporates key security components. A centralized unit (CAU) manages secure boot, and a memory protection unit (MPU) safeguards memory using authenticated encryption and integrity checks. A key management unit (KMU) utilizes a physical unclonable function (PUF) to generate cryptographic keys on demand. The proposed design features a trusted execution environment (TEE) with enclave support, which ensures isolation of sensitive processes. The evaluations demonstrated 31
performance and energy efficiency improvements for the KMU and secure boot mechanism compared with software-based implementations. This approach results in a secure SoC that offers robust security features, efficient performance, flexibility, and system robustness. Beyond traditional hardware security, ML systems on RISC-V face specific threat models that require dedicated countermeasures. Model extraction attacks attempt to reconstruct proprietary model architectures or weights by observing inference behavior; attackers query the model systematically to clone its functionality, threatening intellectual property and enabling adversarial attack development. RISC-V’s TEE capabilities can mitigate this by executing inference within enclaves that prevent memory inspection, while rate limiting and query pattern detection provide software-level defenses. Membership inference attacks determine whether specific data points were used in model training, potentially leaking sensitive information about training datasets in healthcare or financial applications. These attacks exploit the observation that models behave differently on training data versus unseen data. Countermeasures include differential privacy during training (adding calibrated noise to gradients) and output perturbation during inference. RISC-V implementations can support these defenses through hardware random number generators for noise injection and secure memory regions that prevent training data exfiltration. Adversarial input attacks craft malicious inputs that cause misclassification while appearing normal to humans, posing risks in safety-critical applications such as autonomous systems. Hardware-level defenses include input validation accelerators that detect statistical anomalies and redundant inference paths that compare results across different model representations. The customizable nature of RISCV enables integration of such defensive mechanisms as custom instructions or co-processors without modifying the core architecture. Table 19 summarizes the security approaches for RISC-V ML systems. Table 19: Summary of Security Approaches for RISC-V ML Systems
Security Approach
Description
Key Features
Hardware Immune System (HWIS) [142]
Hardware-based malware detection for microprocessors
Secure RISC-V SoC Design [143]
Holistic security-first design approach
Hardware Trojan Detection [140] Cryptographic ISA Extensions [49]
ML-based detection using RISC-V soft cores on FPGAs K-extension for hardware-accelerated cryptography
Lightweight Crypto Cores [77]
Integration of PRINCE and ChaCha in RISC-V SoCs
State Management Security [14, 15]
Smstateen extension for secure ML execution
Side-Channel Attack Prevention [49]
Memory protection and timing attack mitigation
Centralized Authentication Unit (CAU) [143] Memory Protection Unit (MPU) [143]
Secure boot management
96.7% accuracy, 6.5% false-negative rate, F1-score 0.96 CAU, MPU, KMU with PUF, TEE with enclaves Perfect accuracy, FPGA airstream analysis Dedicated crypto operations, improved efficiency IoT security, low-latency crypto acceleration Selective state access, trusted workload isolation Robust security implementations, attack resistance Centralized security control Authenticated encryption, integrity checks PUF-based on-demand key generation Enclave support, sensitive process isolation
Key Management Unit (KMU) [143] Trusted Execution Environment (TEE) [143]
Memory safeguarding with encryption
Cryptographic key generation Process isolation and security
32
6
Future Research Directions
Based on our analysis of current RISC-V machine learning implementations and their limitations, we have identified four research directions that warrant attention from the academic community. These directions emerge from both the opportunities presented by RISC-V’s open architecture and the challenges revealed through our systematic review of existing implementations. Table 20 provides a comprehensive overview of these future research directions for RISC-V machine learning systems.
6.1
Specialized Neural Processing ISA Extensions
The distinction between specialized neural processing ISA extensions and GPU-like architectures lies in three key dimensions: computational granularity, architectural complexity, and integration philosophy. RISC-V neural processing extensions operate at the instruction level, adding dedicated operations for common neural network primitives (matrix multiplication, activation functions, quantization) that execute within the existing processor pipeline. Examples from our survey include the Xpulpnn extension, which adds neural network instructions with minimal area overhead, and the VEXP extensions [22] that provide low-cost softmax operations achieving 162.7× latency reduction with only 1% area overhead. These extensions maintain RISC-V’s modular philosophy by augmenting the core instruction set without fundamentally changing the processor architecture. In contrast, GPU-like architectures employ massively parallel processing with thousands of execution units, dedicated memory hierarchies, and complex scheduling logic. The key difference is that ISA extensions leverage the existing processor infrastructure (register files, pipeline, memory system) while adding specialized functional units, whereas GPUs require entirely separate processing arrays and control logic. Neural processing ISA extensions focus on accelerating specific bottleneck operations within the RISC-V pipeline while maintaining compatibility with existing software stacks and avoiding the complexity and power consumption of full GPU integration. The fundamental limitation of traditional von Neumann architectures in neural network processing lies in the memory wall problem, where data movement between processing units and memory becomes the primary bottleneck. Our analysis of current RISC-V ML implementations reveals that while custom instruction extensions have shown promise, there remains significant unexplored potential in developing comprehensive neural processing instruction sets that can fundamentally reshape how machine learning computations are performed at the hardware level. Recent developments in neural network algorithms have increasingly moved toward operations that are poorly served by traditional scalar instruction sets. Matrix multiplication, convolution operations, and activation functions represent computational patterns that could benefit substantially from dedicated instruction set extensions. The challenge lies not merely in adding new instructions, but in designing a coherent instruction set architecture that maintains RISC-V’s principle of simplicity while providing the computational density required for modern neural networks. The development of such extensions requires careful consideration of data movement patterns, precision requirements, and the interplay between different types of neural network operations. Unlike previous approaches that focus on individual operations, future research should explore holistic instruction set design that considers the entire neural network execution pipeline. This includes investigating how instruction-level parallelism can be enhanced for neural computations and how memory access patterns can be optimized through architectural innovations.
6.2
Adaptive and Modular Processor Architecture
The relationship between adaptive post-fabrication parameter updates and FPGA reconfiguration represents an essential distinction that clarifies this research direction. FPGA partial reconfiguration operates at the hardware configuration level, modifying the actual circuit implementation by reconfiguring lookup tables and routing resources. This enables complete architectural changes but requires significant reconfiguration time (milliseconds to seconds) and specialized tools. In contrast, the adaptive parameter updates we propose operate at the architectural parameter level within fixed hardware structures, enabling runtime modification of behavioral characteristics without circuit-level reconfiguration. The key distinction lies in the abstraction level and modification scope. Our proposed approach focuses on parametric adaptation within pre-designed configurable structures. For example, cache replacement policies can be modified by updating configuration registers that control selection logic already 33
present in the hardware. Similarly, branch predictor behavior can be adapted by adjusting pattern history table parameters or prediction algorithms through control registers. These modifications occur at instruction-level granularity (nanoseconds) rather than hardware reconfiguration timescales. For ASIC implementations, this approach requires design-time planning to incorporate configurable components with parameter control interfaces. The revised discussion provides specific examples: the HULK-V processor (Section 5) demonstrates runtime configuration of its Programmable Multi-Core Accelerator, while the CFU Playground framework (Section 5) shows how custom functional units can be parameterized for different ML workloads. For ASICs, the hardware overhead involves adding configuration registers and multiplexing logic to select between behavioral modes, typically incurring 2-5% area overhead compared to the 100% overhead required for full FPGA-style reconfiguration. Current processor design methodology follows a rigid paradigm where all architectural parameters must be fixed before fabrication, creating a significant barrier to innovation and optimization. The ability to modify processor behavior post-fabrication represents a paradigm shift that could dramatically accelerate the development cycle for machine learning accelerators. This research direction addresses the fundamental question of which processor parameters can be safely exposed for runtime modification without compromising system stability or security. The scope of this research extends beyond simple configuration registers to encompass more fundamental architectural parameters such as cache replacement policies, branch prediction mechanisms, and even aspects of the instruction decode logic. The technical challenges involve developing secure update mechanisms that can verify the integrity of parameter modifications while ensuring that changes do not introduce system instabilities or security vulnerabilities. Adaptive parameters enable runtime optimization within modules, while modularity enables system-level composition and evolution. The evolution of machine learning algorithms and the corresponding changes in computational requirements highlight the limitations of monolithic processor designs. Future RISC-V processors for ML applications should embrace modularity not just at the instruction set level, but throughout the entire processor architecture. This research direction explores how processor components can be designed as independent modules that can be added, removed, or updated without requiring complete processor replacement. The technical challenges of modular architecture design involve the development of standardized interfaces between modules, efficient communication protocols, and mechanisms for ensuring system coherence across module boundaries. The research must address questions of how to maintain performance while providing flexibility, and how to ensure that module interactions do not introduce unexpected dependencies or failure modes. Furthermore, this research direction requires the development of new methodologies for parameter optimization based on actual workload characteristics rather than synthetic benchmarks. The integration of machine learning techniques for automatic parameter tuning presents an intriguing recursive challenge where ML algorithms optimize the hardware that executes ML algorithms. This research direction also includes the development of tools and methodologies for module verification and validation. As processors become more modular, the complexity of ensuring correct operation across all possible module combinations increases. New approaches to formal verification and testing will be required to manage this complexity.
6.3
Comprehensive Security Framework
The open nature of RISC-V, while facilitating innovation and collaboration, introduces unique security challenges that require systematic investigation. Unlike proprietary architectures where security through obscurity provides some protection, RISC-V systems must implement robust security mechanisms that can withstand analysis by adversaries with complete knowledge of the architecture. Side-channel attacks represent a particularly significant threat to RISC-V ML systems because neural network computations often exhibit predictable patterns that can leak information about model parameters or input data. The challenge extends beyond traditional timing and power analysis attacks to include new attack vectors specific to ML workloads, such as cache-based inference of model architectures or electromagnetic analysis of specialized neural processing units. Developing effective countermeasures requires a deep understanding of the information leakage characteristics of different ML algorithms and their implementation patterns. This includes investigating randomization techniques that can obscure side-channel information without significantly impacting performance, as well as architectural modifications that can provide inherent resistance to various classes of 34
Table 20: Future research directions for RISC-V machine learning systems Research Direction
Technical Challenges
Potential Impact
Recommended Approach
Specialized Neural Processing ISA Extensions
Coherent matrix instruction design, maintain RISC-V simplicity, optimize data movement
10-50× neural computation efficiency gains
Holistic pipeline design, formal verification, compiler co-design
Adaptive and Modular Processor Architecture
Secure update mechanisms, standardized module interfaces, system coherence, verification complexity
Runtime adaptive tuning, rapid adaptation, cost-effective upgrades
Hardware security analysis, interface standardization, formal verification, modular testing
Comprehensive Security Framework
Side-channel resistance, performance trade-offs, ML vulnerabilities
Robust open architecture security, ML attack protection
Security-by-design, randomization, attack analysis
Energy-Efficient Multi-Clock Domains
Power prediction, transition optimization, coherence maintenance
Significant edge AI energy savings, extended battery life
Predictive algorithms, low-overhead protocols, energy modeling
attacks. The research should also address the trade-offs between security and performance, recognizing that security countermeasures can impact the efficiency of ML computations. The goal is to develop security frameworks that provide robust protection while maintaining the performance advantages that make RISC-V attractive for ML applications.
6.4
Energy-Efficient Multi-Clock Domains
While clock-domain crossing is standard in ML-focused chip designs, opportunities exist for finer-grained power management integrated with RISC-V ISA-level power states. We identify research directions in coordinating RISC-V architectural power management extensions with physical implementation clock gating strategies. Energy efficiency remains the concern for edge and mobile ML applications, yet current approaches to power management often treat the processor as a monolithic unit. The research direction of multi-clock domain architectures investigates how fine-grained clock and power domain control can be leveraged to achieve energy savings in ML workloads. The complexity of this research lies in developing algorithms that can predict the optimal power state configuration for different phases of ML computation while accounting for the transition costs between power states. This requires understanding of ML algorithm execution patterns and the development of runtime systems that can make power management decisions with minimal overhead. The research must also address the challenge of maintaining system coherence across multiple clock domains while ensuring that power transitions do not introduce correctness issues. This includes investigating new cache coherence protocols, memory consistency models, and synchronization mechanisms that can operate efficiently in multi-clock domain environments.
7
Conclusion
This survey presents an analysis of the RISC-V ISA’s role in machine learning applications, examining over 100 research contributions spanning academic implementations, commercial cores, software frameworks, and real-world deployments. Our systematic evaluation demonstrates that RISC-V implementations achieve notable advantages in energy efficiency and competitive performance across diverse ML workloads compared to proprietary alternatives, with specific gains varying by workload, baseline, and implementation as detailed in Sections 3 and 5. The analysis of 47 distinct RISC-V processor cores reveals a mature ecosystem capable of addressing requirements from ultra-low-power IoT applications to high-performance neural network inference. Key architectural innovations include specialized vector extensions, custom instruction sets for ML operations, and optimized memory hierarchies that enable superior performance-per-watt characteristics important for edge computing scenarios.
35
Software infrastructure analysis demonstrates robust support through compiler frameworks (TVM, MLIR) with RISC-V-specific optimizations and specialized acceleration tools. The integration of hardware and software optimizations enables performance benefits unattainable with traditional proprietary architectures, for domain-specific acceleration. Application evaluations across medical devices, robotics, and computer vision validate the practical deployment of RISC-V ML systems. These implementations demonstrate measurable benefits in scenarios prioritizing energy efficiency, cost constraints, and customization flexibility. We identify four research directions: (1) specialized neural processing instruction extensions, (2) adaptive and modular processor architectures, (3) security frameworks for open-source architectures, and (4) energy-efficient multi-clock domain implementations. The evidence indicates RISC-V’s growing importance in ML hardware ecosystems. The open-source development model facilitates rapid innovation and collaborative optimization suited to the evolution of ML algorithms. Our analysis establishes RISC-V not merely as an alternative to proprietary architectures, but as a potentially superior approach for specialized ML computational requirements, positioning it as an enabler of the ongoing transformation in artificial intelligence applications.
References [1] Clifford, E., Saravanan, A., Langford, H., Zhang, C., Zhao, Y., Mullins, R.D., Shumailov, I., Hayes, J.: Locking machine learning models into hardware. 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), 302–320 (2024) [2] Hadfield, G.K., Clark, J.: Regulatory markets: The future of ai governance (2023) [3] Ledin, J.: Modern Computer Architecture and Organization vol. 2. Packt Publishing Birmingham, UK, Birmingham, UK (2020) [4] Kanter, D.: Risc-v offers simple, modular isa. Microprocessor Report 1, 1–5 (2016) [5] Ince, M.N., Ledet, J., Gunay, M.: Building an open source linux computing system on risc-v. In: 2019 1st International Informatics and Software Engineering Conference (UBMYK), pp. 1–4 (2019). IEEE [6] Di Mascio, S., Menicucci, A., Furano, G., Monteleone, C., Ottavi, M.: The case for risc-v in space. In: Applications in Electronics Pervading Industry, Environment and Society: APPLEPIES 2018 6, pp. 319–325 (2019). Springer [7] Di Mascio, S., Menicucci, A., Gill, E., Furano, G., Monteleone, C.: Leveraging the openness and modularity of risc-v in space. Journal of Aerospace Information Systems 16(11), 454–472 (2019) [8] Raveendran, A., Patil, V.B., Selvakumar, D., Desalphine, V.: A risc-v instruction set processormicro-architecture design and analysis. In: 2016 International Conference on VLSI Systems, Architectures, Technology and Applications (VLSI-SATA), pp. 1–7 (2016). IEEE [9] Capra, M., Peloso, R., Masera, G., Ruo Roch, M., Martina, M.: Edge computing: A survey on the hardware requirements in the internet of things world. Future Internet 11(4), 100 (2019) [10] Flamand, E., Rossi, D., Conti, F., Loi, I., Pullini, A., Rotenberg, F., Benini, L.: Gap-8: A risc-v soc for ai at the edge of the iot. In: 2018 IEEE 29th International Conference on Application-specific Systems, Architectures and Processors (ASAP), pp. 1–4 (2018). IEEE [11] Montesdeoca, G., Asanza, V., Estrada, R., Valeriano, I., Muneeb, M.: Softprocessor riscv-ec for edge computing applications. In: International Conference on Innovative Mobile and Internet Services in Ubiquitous Computing, pp. 209–220 (2023). Springer [12] Chander, V.N., Varghese, K.: A soft risc-v vector processor for edge-ai. In: 2022 35th International Conference on VLSI Design and 2022 21st International Conference on Embedded Systems (VLSID), pp. 1–6 (2022). IEEE [13] Kong, Y.: Airv: Enabling deep learning inference on risc-v. In: International Symposium on Benchmarking, Measuring and Optimization, pp. 289–299 (2019). Springer 36
[14] Garofalo, A., Rusci, M., Conti, F., Rossi, D., Benini, L.: Darkside: A heterogeneous risc-v compute cluster for extreme-edge on-chip dnn inference and training. IEEE Open Journal of the Solid-State Circuits Society 2, 231–243 (2022) [15] Garofalo, A., Rusci, M., Conti, F., Rossi, D., Benini, L.: Xpulpnn: Enabling energy efficient and flexible inference of quantized neural networks on risc-v based iot end nodes. IEEE Transactions on Emerging Topics in Computing 9(3), 1489–1505 (2021) [16] Hou, P., Chen, Q., Guo, S., Liang, Y., Wang, Y.: Rvtensor: a light-weight neural network inference framework based on the risc-v architecture. In: Benchmarking, Measuring, and Optimizing: Second BenchCouncil International Symposium, Bench 2019, Denver, CO, USA, November 14–16, 2019, Revised Selected Papers 2, pp. 280–292 (2020). Springer International Publishing [17] Ueyoshi, K., Ando, K., Hirose, K., Takamaeda-Yamazaki, S., Kadomoto, J., Miyata, T., Hamamoto, T., Kuroda, T.: Diana: An end-to-end energy-efficient digital and analog hybrid neural network soc. In: 2022 IEEE International Solid-State Circuits Conference (ISSCC), vol. 65, pp. 1–3 (2022). IEEE [18] Mueller, H., Kartsch, V., Benini, L.: Gap9shield: A 150gops ai-capable ultra-low power module for vision and ranging applications on nano-drones. In: European Robotics Forum, pp. 123–134 (2024). Springer [19] Kumar, M.A., O’Mahoney, C., Werle, P.K., Shanker, S., Nikolopoulos, D.S., Ji, B., Vandierendonck, H., John, D.: Marvel: An end-to-end framework for generating model-class aware custom risc-v extensions for lightweight ai. IEEE Open Journal of Circuits and Systems 6, 445–456 (2025) [20] Nunes, W.A., Santos, A.V.C.D., Moraes, F.G.: Accelerating machine learning with RISC-V vector extension and auto-vectorization techniques, 1–5 (2025) https://doi.org/10.1109/ISCAS56072. 2025.11043225 [21] Colagrande, L., Leone, L., Coco, M., Deaconeasa, A., Benini, L.: Towards zero-stall matrix multiplication on energy-efficient risc-v clusters for machine learning acceleration. 2025 IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED), 1–7 (2025) [22] Wang, R., Islamoglu, G., Belano, A., Potocnik, V., Conti, F., Garofalo, A., Bonini, L.: Vexp: A low-cost risc-v isa extension for accelerated softmax computation in transformers. In: 2025 IEEE 32nd Symposium on Computer Arithmetic (ARITH), pp. 37–44 (2025). IEEE Computer Society [23] Vacca, E., Strollo, E., Azimi, S.: Rempro: Reconfigurable modular processor. In: Proceedings of the 21st ACM International Conference on Computing Frontiers: Workshops and Special Sessions, pp. 110–115 (2024) [24] Schiavone, P.D., Sanchez, E., Ruospo, A., Minervini, F., Zaruba, F., Haugou, G., Benini, L.: An Open-Source Verification Framework for Open-Source Cores: A RISC-V Case Study. In: 2018 IFIP/IEEE International Conference on Very Large Scale Integration (VLSI-SoC), pp. 43–48. IEEE, Verona, Italy (2018). https://doi.org/10.1109/VLSI-SoC.2018.8644818 [25] Rutishauser, G., Mihali, J., Scherer, M., Bonini, L.: xtern: Energy-efficient ternary neural network inference on risc-v-based edge systems. In: 2024 IEEE 35th International Conference on Application-specific Systems, Architectures and Processors (ASAP), pp. 206–213 (2024). IEEE [26] Maras, A.: Extending risc-v isa for fine-grained mixed-precision in neural networks. Master’s Thesis, KTH Royal Institute of Technology (2024) [27] Peng, C.W., Jambek, A.B., Dass, S.B., Lim, T.W., Laudis, L.L., Yadlapati, A., Udari, G.C., Rama, S.R.B.: Performance evaluation of risc-v microcontroller system on fpga: A study of the neorv32 core. International Journal of Integrated Engineering 16(1), 312–317 (2024) [28] Wu, Y., Jie, G., Qiu, L., ZHANG, J.: Design of risc-v out-of-order processor based on segmented exclusive or gshare branch prediction. Microelectronics Journal, 106334 (2024) [29] Simson, N., Tahigara, A., Ecker, W.: A comparative analysis of arm and risc-v isas for deeply embedded systems. In: MBMV 2024; 27. Workshop, pp. 110–119 (2024). VDE
37
[30] Kröger, L., Cioflan, C.: On-device federated continual learning on risc-v-based ultra-low-power soc for intelligent nano-drone swarms. Power [mW] 24, 53–1 [31] Yang, M., Ahmed, T., Inagaki, S., Sakiyama, K., Li, Y., Hara-Azumi, Y.: Hardware/software cooperative design against power side-channel attacks on iot devices. IEEE Internet of Things Journal (2024) [32] Bhade, P., Paturel, J., Sentieys, O., Sinha, S.: Lightweight hardware-based cache side-channel attack detection for edge devices (edge-cascade). ACM Transactions on Embedded Computing Systems 23(4), 1–27 (2024) [33] Asano, T., Sugawara, T.: Simulation-based evaluation of bit-interaction side-channel leakage on risc-v: extended version. Journal of Cryptographic Engineering 14(1), 165–180 (2024) [34] Jiang, F., Tong, F., Wang, H., Cheng, X., Zhou, Z., Ling, M., Mao, Y.: Pcg: Mitigating conflictbased cache side-channel attacks with prefetching. arXiv preprint arXiv:2405.03217 (2024) [35] Miteloudi, K., Adhikary, A., Drueten, N., Batina, L., Buhan, I.: Plan your defense: A comparative analysis of leakage detection methods on risc-v cores. Cryptology ePrint Archive (2024) [36] Christensen, M., Sherwood, T., Balkind, J., Hardekopf, B.: Wire sorts: A language abstraction for safe hardware composition. In: Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation, pp. 175–189 (2021) [37] Liu, Q., Amiri, S., Ost, L.: Exploring risc-v based dnn accelerators. In: 2024 IEEE International Conference on Omni-layer Intelligent Systems (COINS), pp. 1–6 (2024). IEEE [38] Liu, Y., Ye, K., Xu, C.-Z.: Performance Evaluation of Various RISC Processor Systems: A Case Study on ARM, MIPS and RISC-V, pp. 61–74 (2022). https://doi.org/10.1007/ 978-3-030-96326-2_5 [39] Yu, M.-S., Yuan, C.-Y., Chen, T.-L., Lee, J.-K.: Case Study: Optimization Methods with TVM Hybrid-OP on RISC-V Packed SIMD. IEEE Access PP, 1–1 (2024) https://doi.org/10.1109/ ACCESS.2024.3397195 [40] Nicholas, G.S., Gui, Y., Saqib, F.: A Survey and Analysis on SoC Platform Security in ARM, Intel and RISC-V Architecture. In: 2020 IEEE 63rd International Midwest Symposium on Circuits and Systems (MWSCAS), pp. 718–721. IEEE, Springfield, MA, USA (2020). https://doi.org/10. 1109/MWSCAS48704.2020.9184573 [41] Kalapothas, S., Galetakis, M., Flamis, G., Plessas, F., Kitsos, P.: A survey on RISC-Vbased machine learning ecosystem. Information 14(2), 64 (2023) https://doi.org/10.3390/ info14020064 [42] Agosta, G., Galimberti, A., Zoni, D.: Deep learning on RISC-V platforms at the edge: A perspective on the hardware and software support. ACM Computing Surveys 58(5), 1–37 (2025) [43] Cui, E., Li, T., Wei, Q.: RISC-V Instruction Set Architecture Extensions: A Survey. IEEE Access 11, 24696–24711 (2023) https://doi.org/10.1109/ACCESS.2023.3246491 . Conference Name: IEEE Access [44] Lu, T.: A Survey on RISC-V Security: Hardware and Architecture. arXiv. arXiv:2107.04175 [cs] (2021). http://arxiv.org/abs/2107.04175 Accessed 2024-07-03 [45] Drflinger, A., Albers, M., Kleinbeck, B., Guan, Y., Michalik, H., Klink, R., Blochwitz, C., Nechi, A., Berekovic, M.: A comparative survey of open-source application-class RISC-V processor implementations. In: Proceedings of the 18th ACM International Conference on Computing Frontiers. CF ’21, pp. 12–20. Association for Computing Machinery, New York, NY, USA (2021). https://doi.org/10.1145/3457388.3458657 [46] Mezger, B.W., Santos, D.A., Dilillo, L., Zeferino, C.A., Melo, D.R.: A Survey of the RISC-V Architecture Software Support. IEEE Access 10, 51394–51411 (2022) https://doi.org/10.1109/ ACCESS.2022.3174125 . Conference Name: IEEE Access
38
[47] Waterman, A., Lee, Y., Patterson, D.A., Asanovi, K.: The risc-v instruction set manual. volume 1: User-level isa, version 2.0. (2014). https://api.semanticscholar.org/CorpusID:60142097 [48] Asanovic, K., Avizienis, R., Bachrach, J., Beamer, S., Biancolin, D., Celio, C., Cook, H., Dabbelt, D., Hauser, J., Izraelevitz, A., et al.: The rocket chip generator. EECS Department, University of California, Berkeley, Tech. Rep. UCB/EECS-2016-17 4, 6–2 (2016) [49] Gupta, S.R., Papadopoulou, N., Pericas, M.: Challenges and Opportunities in the Co-design of Convolutions and RISC-V Vector Processors. In: Proceedings of the SC ’23 Workshops of The International Conference on High Performance Computing, Network, Storage, and Analysis. SCW ’23, pp. 1550–1556. Association for Computing Machinery, New York, NY, USA (2023). https: //doi.org/10.1145/3624062.3624232 [50] Bertaccini, L., Paulin, G., Cavalcante, M., Fischer, T., Mach, S., Benini, L.: Minifloats on risc-v cores: Isa extensions with mixed-precision short dot products. IEEE Transactions on Emerging Topics in Computing (2024) [51] Waterman, A., Asanovic, K., Hauser, J.: The risc-v instruction set manual, volume ii: Privileged architecture. RISC-V Foundation, 1–4 (2019) [52] Embeddev, F.: RISC-V Instruction Set Manual, Volume I: RISC-V User-Level ISA. https: //www.five-embeddev.com//riscv-user-isa-manual/Priv-v1.12/extensions.html Accessed 2024-08-30 [53] Prakash, S., Callahan, T., Bushagour, J., Banbury, C., Green, A.V., Warden, P., Ansell, T., Reddi, V.J.: Cfu playground: Want a faster ml processor? do it yourself! In: 2023 Design, Automation & Test in Europe Conference & Exhibition (DATE), pp. 1–2 (2023). IEEE [54] Balasubramanian, K.K., Salvo, M.D., Rocchia, W., Decherchi, S., Crepaldi, M.: Designing RISC-V Instruction Set Extensions for Artificial Neural Networks: An LLVM Compiler-Driven Perspective. IEEE Access 12, 55925–55944 (2024) https://doi.org/10.1109/ACCESS.2024.3389673 [55] Lee, C.-L., Chung, C.-P., Cheng, S.-Y., Lee, J.-K., Lai, R.: Accelerating AI performance with the incorporation of TVM and MediaTek NeuroPilot. Connection Science 35(1), 2272586 (2023) https://doi.org/10.1080/09540091.2023.2272586 [56] Manoni, S., Scheffler, P., Di Mauro, A., Zanatta, L., Acquaviva, A., Benini, L., Bartolini, A.: Nars: Neuromorphic acceleration through register-streaming extensions on risc-v cores. In: Proceedings of the 21st ACM International Conference on Computing Frontiers: Workshops and Special Sessions, pp. 79–82 (2024) [57] Armeniakos, G., Maras, A., Xydis, S., Soudris, D.: Mixed-precision neural networks on risc-v cores: Isa extensions for multi-pumped soft simd operations. In: Proceedings of the 43rd IEEE/ACM International Conference on Computer-Aided Design, pp. 1–9 (2024) [58] Poveda Rodrigo, J.J., Hamdi, M.A., Koenig, C., Burrello, A., Jahier Pagliari, D., Benini, L.: Poster: V-seek: Optimizing llm reasoning on a server-class general-purpose risc-v platform. In: Proceedings of the 22nd ACM International Conference on Computing Frontiers. CF ’25, pp. 224– 225. Association for Computing Machinery, New York, NY, USA (2025). https://doi.org/10. 1145/3719276.3727954 . https://doi.org/10.1145/3719276.3727954 [59] Sabih, M., Karim, A., Wittmann, J., Hannig, F., Teich, J.: Hardware/software co-design of riscv extensions for accelerating sparse dnns on fpgas. In: 2024 International Conference on Field Programmable Technology (ICFPT), pp. 01–09 (2024). IEEE [60] Roorda, E., Rasoulinezhad, S., Leong, P.H.W., Wilton, S.J.E.: FPGA Architecture Exploration for DNN Acceleration. ACM Transactions on Reconfigurable Technology and Systems 15(3), 1–37 (2022) https://doi.org/10.1145/3503465 [61] Vrea, J., Biasizzo, A.: Towards Deploying Highly Quantized Neural Networks on FPGA Using Chisel. In: 2023 26th Euromicro Conference on Digital System Design (DSD), pp. 161–167. IEEE, Golem, Albania (2023). https://doi.org/10.1109/DSD60849.2023.00032
39
[62] Cheikh Tourad, E.H., Eleuldj, M.: GENERIC DESIGN FLOW OF PIPELINED HARDWARE IMPLEMENTATION OF DEEP NEURAL NETWORKS. Indian Journal of Computer Science and Engineering 13(4), 1007–1016 (2022) https://doi.org/10.21817/indjcse/2022/v13i4/ 221304029 [63] Xiao, Y., Nazarian, S., Bogdan, P.: GAHLS: an optimized graph analytics based high level synthesis framework. Scientific Reports 13(1), 22655 (2023) https://doi.org/10.1038/ s41598-023-48981-x [64] Mika, K., Griessl, R., Kucza, N., Porrmann, F., Kaiser, M., Tigges, L., Hagemeyer, J., Trancoso, P., Azhar, M.W., Qararyah, F., Zouzoula, S., Mntrey, J., Pasin, M., Felber, P., Marcus, C., Brunnegard, O., Eriksson, O., Salomonsson, H., dman, D., Ask, A., Casimiro, A., Bessani, A., Carvalho, T., Gugala, K., Zierhoffer, P., Latosinski, G., Tassemeier, M., Porrmann, M., Heyn, H.M., Knauss, E., Mao, Y., Meierhfer, F.: VEDLIoT: Next generation accelerated AIoT systems and applications. In: Proceedings of the 20th ACM International Conference on Computing Frontiers, pp. 291–296. ACM, Bologna Italy (2023). https://doi.org/10.1145/3587135.3592175 [65] Leliwa, K., O’Connor, I., Schaeffer, M., Aicha, A.B., Sentieys, O., Coussy, P.: Optimised extension of an ultra-low-power risc-v processor to support lightweight neural network models. Chips (2025) [66] Vergos, P., Vergos, T., Afentaki, F., Balaskas, K., Zervakis, G.: Support vector machines classification on bendable risc-v. In: 2025 IEEE Computer Society Annual Symposium on VLSI (ISVLSI), vol. 1, pp. 1–6 (2025). IEEE [67] Daghero, F., Pagliari, D.J., Conti, F., Benini, L., Poncino, M., Burrello, A.: Lightweight software kernels and hardware extensions for efficient sparse deep neural networks on microcontrollers. Proceedings of Machine Learning and Systems 7 (2025) [68] Bik, A., Koanantakool, P., Shpeisman, T., Vasilache, N., Zheng, B., Kjolstad, F.: Compiler Support for Sparse Tensor Computations in MLIR. ACM Transactions on Architecture and Code Optimization 19(4), 1–25 (2022) https://doi.org/10.1145/3544559 [69] Rotem, N., Fix, J., Abdulrasool, S., Deng, S., Dzhabarov, R., Hegeman, J., Levenstein, R., Maher, B.A., Satish, N., Olesen, J., Park, J., Rakhov, A., Smelyanskiy, M.: Glow: Graph lowering compiler techniques for neural networks. ArXiv abs/1805.00907 (2018) [70] Zhou, M., Lin, X., Liang, Y.: Agile Optimization Framework: A Framework for Tensor Operator Optimization in Neural Network (2024). https://doi.org/10.2139/ssrn.4748040 . https:// www.ssrn.com/abstract=4748040 Accessed 2024-06-04 [71] Xu, Y., Yuan, Q., Barton, E.C., Li, R., Sadayappan, P., Sukumaran-Rajam, A.: Effective Performance Modeling and Domain-Specific Compiler Optimization of CNNs for GPUs. In: Proceedings of the International Conference on Parallel Architectures and Compilation Techniques, pp. 252–264. ACM, Chicago Illinois (2022). https://doi.org/10.1145/3559009.3569674 [72] Jia, L., Luo, Z., Lu, L., Liang, Y.: Automatic Generation of Spatial Accelerator for Tensor Algebra. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 42(6), 1898– 1911 (2023) https://doi.org/10.1109/TCAD.2022.3209949 [73] Agostini, N.B., Curzel, S., Amatya, V., Tan, C., Minutoli, M., Castellana, V.G., Manzano, J., Kaeli, D., Tumeo, A.: An MLIR-based Compiler Flow for System-Level Design and Hardware Acceleration. In: Proceedings of the 41st IEEE/ACM International Conference on ComputerAided Design, pp. 1–9. ACM, San Diego California (2022). https://doi.org/10.1145/3508352. 3549424 [74] Vrea, J., Pilipovi, R., Biasizzo, A.: HardwareSoftware Co-Design of an Audio Feature Extraction Pipeline for Machine Learning Applications. Electronics 13(5), 875 (2024) https://doi.org/10. 3390/electronics13050875 [75] Kermarrec, F., Bourdeauducq, S., Badier, H., Le Lann, J.-C.: Litex: an open-source soc builder and library based on migen python dsl. In: OSDA 2019, Colocated with DATE 2019 Design Automation and Test in Europe (2019)
40
[76] Ramanathan, N., Constantinides, G.A., Wickerson, J.: A Case for Precise, Fine-Grained Pointer Synthesis in High-Level Synthesis. ACM Transactions on Design Automation of Electronic Systems 27(4), 1–26 (2022) https://doi.org/10.1145/3491430 [77] Ma, K.-M., Le, D.-H., Pham, C.-K., Hoang, T.-T.: Design of an SoC Based on 32-Bit RISC-V Processor with Low-Latency Lightweight Cryptographic Cores in FPGA. Future Internet 15(5), 186 (2023) https://doi.org/10.3390/fi15050186 [78] Ferres, B., Muller, O., Rousseau, F.: A Chisel Framework for Flexible Design Space Exploration through a Functional Approach. ACM Transactions on Design Automation of Electronic Systems 28(4), 1–31 (2023) https://doi.org/10.1145/3590769 [79] Liu, H.-I.C., Brehler, M., Ravishankar, M., Vasilache, N., Vanik, B., Laurenzo, S.: TinyIREE: An ML Execution Environment for Embedded Systems From Compilation to Deployment. IEEE Micro 42(5), 9–16 (2022) https://doi.org/10.1109/MM.2022.3178068 [80] Yoo, T., Choi, B.: Real-Time Performance Benchmarking of RISC-V Architecture: Implementation and Verification on an EtherCAT-Based Robotic Control System. Electronics 13, 733 (2024) https: //doi.org/10.3390/electronics13040733 [81] Boosting Machine Learning with tailored accelerators: Custom Function Units in Renode | Google Open Source Blog. https://opensource.googleblog.com/2021/12/Boosting_Machine_ Learning_with_tailored_accelerators_Custom_Function_Units_in_Renode.html Accessed 2024-05-17 [82] Valente, L., Tortorella, Y., Sinigaglia, M., Tagliavini, G., Capotondi, A., Benini, L., Rossi, D.: HULK-V: a Heterogeneous Ultra-low-power Linux capable RISC-V SoC. In: 2023 Design, Automation & Test in Europe Conference & Exhibition (DATE), pp. 1–6. IEEE, Antwerp, Belgium (2023). https://doi.org/10.23919/DATE56975.2023.10137252 [83] Gonzalez, A., Hong, C.: A Chipyard Comparison of NVDLA and Gemmini. UC Berkeley Technical Report (2021) [84] Bolhasani, H., Jassbi, S.J., Sharifi, A.: DLA-E: a deep learning accelerator for endoscopic images classification. Journal of Big Data 10(1), 76 (2023) https://doi.org/10.1186/ s40537-023-00775-8 [85] Campbell, P.: MoonbaseOtago/vroom. original-date: 2021-09-29T21:45:06Z (2024). https:// github.com/MoonbaseOtago/vroom Accessed 2024-05-23 [86] GitHub - riscv-boom/riscv-boom: SonicBOOM: The Berkeley Out-of-Order Machine. https:// github.com/riscv-boom/riscv-boom Accessed 2024-05-23 [87] Zhu, Y., Zheng, J., Ding, S., Li, L.: Hardware Data Prefetch for XiangShan Processor. In: 2022 7th International Conference on Integrated Circuits and Microsystems (ICICM), pp. 394–397 (2022). https://doi.org/10.1109/ICICM56102.2022.10011259 [88] lowRISC/ibex. lowRISC. original-date: lowRISC/ibex Accessed 2024-05-23
2017-08-08T12:16:36Z (2024). https://github.com/
[89] SpinalHDL/VexRiscv. SpinalHDL. original-date: 2017-03-08T21:14:28Z (2024). https://github. com/SpinalHDL/VexRiscv Accessed 2024-05-23 [90] GitHub - MoonbaseOtago/SuperScalar-RISCV-CPU: SSRV(Super-Scalar RISC-V) — Superscalar out-of-order RV32IMC CPU core, 6.4 CoreMark/MHz. https://github.com/risclite/ SuperScalar-RISCV-CPU Accessed 2024-05-24 [91] Kindgren, O.: olofk/serv. original-date: 2018-10-31T07:15:00Z (2024). https://github.com/ olofk/serv Accessed 2024-05-23 [92] Mohajerani, K.: kammoh/ORCA-risc-v. original-date: 2015-12-26T14:20:54Z (2023). https:// github.com/kammoh/ORCA-risc-v Accessed 2024-05-23 [93] Research, B.A.: Rocket Core. https://github.com/chipsalliance/rocket-chip. [Accessed 2209-2024] (2019) 41
[94] Madras, I.: Shakti Processorss — shakti.org.in. https://shakti.org.in/processors.html. [Accessed 22-09-2024]. https://shakti.org.in/processors.html [95] Wan, Z.: rgwan/ri5cy fpga. original-date: 2016-06-27T11:15:07Z (2020). https://github.com/ rgwan/ri5cy_fpga Accessed 2024-06-27 [96] RV01 a two-way in-order superscalar processor for RISC-V Core - IP Cores. https://www. allaboutcircuits.com/ip-cores/processor/rv01-riscv-core/ Accessed 2024-06-27 [97] Shioya, R.: GitHub - rsd-devel/rsd: RSD: RISC-V Out-of-Order Superscalar Processor — github.com. https://github.com/rsd-devel/rsd. [Accessed 22-09-2024]. https://github.com/ rsd-devel/rsd [98] YosysHQ/picorv32. Yosys Headquarters. original-date: 2015-06-06T11:52:27Z (2024). https:// github.com/YosysHQ/picorv32 Accessed 2024-05-23 [99] darklife/darkriscv. Darklife. original-date: 2018-08-19T18:55:50Z (2024). https://github.com/ darklife/darkriscv Accessed 2024-06-27 [100] csail-csg/riscy-OOO. CSAIL CSG. original-date: 2017-04-30T04:12:27Z (2024). https://github. com/csail-csg/riscy-OOO Accessed 2024-05-28 [101] Tenstorrent: Tenstorrent Architecture and Product Overview. Accessed 2026-03-24. https:// tenstorrent.com/ [102] Chen, C., Xiang, X., Liu, C., Shang, Y., Guo, R., Liu, D., Lu, Y., Hao, Z., Luo, J., Chen, Z., Li, C., Pu, Y., Meng, J., Yan, X., Xie, Y., Qi, X.: Xuantie-910: A Commercial Multi-Core 12Stage Pipeline Out-of-Order 64-bit High Performance RISC-V Processor with Vector Extension : Industrial Product. In: 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA), pp. 52–64 (2020). https://doi.org/10.1109/ISCA45697.2020.00016 [103] Redkin, A.: SCRx Family Of The RISC-V Compatible Processor IP. Syntacore Technical Report (2018) [104] SiFive Essential. https://www.sifive.com/cores/essential Accessed 2024-05-23 [105] GigaDevice: 32-bit Microcontrollers (MCUs)-GigaDevice.com — gigadevice.com. https://www. gigadevice.com/product/mcu/risc-v. [Accessed 22-09-2024] [106] OnchipUIS: onchipuis/mriscv. original-date: 2016-09-10T00:04:19Z (2024). https://github.com/ onchipuis/mriscv Accessed 2024-05-23 [107] lcbcFoo: ReonV: A Modified VHDL RISC-V Core (RV32I). https://github.com/lcbcFoo/ReonV. Accessed: 2026-02-05 (2026) [108] Lei, Z., Cai, F., Zhou, J., Guo, Z.: A Floating-Point Unit Architecture Based on SweRV EH1 Core. In: 2022 IEEE 16th International Conference on Anti-counterfeiting, Security, and Identification (ASID), pp. 1–5 (2022). https://doi.org/10.1109/ASID56930.2022.9995796 . ISSN: 2163-5056 [109] Corporation, A.T.: AndesCore A25. https://www.andestech.com/en/products-solutions/ andescore-processors/riscv-a25/. [Accessed 22-09-2024] [110] RoaLogic: RV12: Configurable RISC-V CPU Core. https://github.com/RoaLogic/RV12. Accessed: 2026-02-05 (2026) [111] CloudBEAR: BI-671 64-bit RISC-V core with out-of-order pipeline based complex. https:// cloudbear.ru/bi_671.html. [Accessed 22-09-2024]. https://cloudbear.ru/bi_671.html [112] SI-RISCV: e200 opensource: Hummingbird E203 RISC-V Processor Core. https://github.com/ SI-RISCV/e200_opensource. Archived repository; Accessed: 2026-02-05 (2018) [113] riscv-mcu: e203 hbirdv2: Ultra-Low Power RISC-V Core and SoC. https://github.com/ riscv-mcu/e203_hbirdv2. Accessed: 2026-02-05 (2026) [114] Nuclei-Software: e603 hbird: RV64GC Linux-Capable RISC-V Core. https://github.com/ Nuclei-Software/e603_hbird. Accessed: 2026-02-05 (2026) 42
[115] SiFive, I.: SiFive — github.com. https://github.com/sifive. [Accessed 22-09-2024]. https: //github.com/sifive [116] Ushiroyama, A., Watanabe, M., Watanabe, N., Nagoya, A.: Convolutional neural network implementations using vitis ai. In: 2022 IEEE 12th Annual Computing and Communication Workshop and Conference (CCWC), pp. 0365–0371 (2022). IEEE [117] Duarte, J., Han, S., Harris, P., Jindariani, S., Kreinar, E., Kreis, B., Ngadiuba, J., Pierini, M., Rivera, R., Tran, N., et al.: Fast inference of deep neural networks in fpgas for particle physics. Journal of instrumentation 13(07), 07027 (2018) [118] Umuroglu, Y., Fraser, N.J., Gambardella, G., Blott, M., Leong, P., Jahre, M., Vissers, K.: Finn: A framework for fast, scalable binarized neural network inference. In: Proceedings of the 2017 ACM/SIGDA International Symposium on Field-programmable Gate Arrays, pp. 65–74 (2017) [119] Zhao, Y., Sharif, H., Adve, V., Misailovic, S.: Felix: Optimizing Tensor Programs with Gradient Descent. In: Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3, pp. 367–381. ACM, La Jolla CA USA (2024). https://doi.org/10.1145/3620666.3651348 [120] Golin, R., Chelini, L., Siemieniuk, A., Madhu, K., Hasabnis, N., Pabst, H., Georganas, E., Heinecke, A.: Towards a high-performance AI compiler with upstream MLIR. arXiv. arXiv:2404.15204 [cs] (2024). http://arxiv.org/abs/2404.15204 Accessed 2024-06-18 [121] Ryu, J., Park, E., Sung, H.: One-shot tuner for deep learning compilers. In: Proceedings of the 31st ACM SIGPLAN International Conference on Compiler Construction, pp. 89–103. ACM, Seoul South Korea (2022). https://doi.org/10.1145/3497776.3517774 [122] Vasilache, N., Zinenko, O., Bik, A.J.C., Ravishankar, M., Raoux, T., Belyaev, A., Springer, M., Gysi, T., Caballero, D., Herhut, S., Laurenzo, S., Cohen, A.: Composable and Modular Code Generation in MLIR: A Structured and Retargetable Approach to Tensor Compiler Construction. arXiv. arXiv:2202.03293 [cs] (2022). http://arxiv.org/abs/2202.03293 Accessed 2024-06-18 [123] Xiao, D., Liu, Z., Yuan, Y., Pang, Q., Wang, S.: Metamorphic Testing of Deep Learning Compilers. Proceedings of the ACM on Measurement and Analysis of Computing Systems 6(1), 1–28 (2022) https://doi.org/10.1145/3508035 [124] Chen, T., Moreau, T., Jiang, Z., Shen, H., Yan, E., Wang, L., Hu, Y., Ceze, L., Guestrin, C., Krishnamurthy, A.: Tvm: End-to-end compilation stack for deep learning. In: Proceedings of the 2nd Conference on Systems and Machine Learning (SysML) (2018) [125] Chen, T., Moreau, T., Jiang, Z., Zheng, L., Yan, E., Cowan, M., Shen, H., Wang, L., Hu, Y., Ceze, L., Guestrin, C., Krishnamurthy, A.: TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. Proceedings of OSDI (2018) [126] Nizharadze, N., Mahlig, M., Merk, T.: Simulation of machine learning inferences in real-time operating system to improve direction finding in an embedded environment. In: 2023 13th International Conference on Indoor Positioning and Indoor Navigation (IPIN), pp. 1–6. IEEE, Nuremberg, Germany (2023). https://doi.org/10.1109/IPIN57070.2023.10332487 [127] Tang, M., Huang, L., Chen, W.: Rapid and Accurate PPA Prediction for the Template-Based Processor Design Methods. Applied Sciences 12(16), 8383 (2022) https://doi.org/10.3390/ app12168383 [128] Eggimann, M., Mach, S., Magno, M., Benini, L.: A RISC-V Based Open Hardware Platform for Always-On Wearable Smart Sensing. In: 2019 IEEE 8th International Workshop on Advances in Sensors and Interfaces (IWASI), pp. 169–174. IEEE, Otranto, Italy (2019). https://doi.org/10. 1109/IWASI.2019.8791364 [129] Giuffrida, G., Nannipieri, P., Diana, L., Panicacci, S., Fanucci, L., Benelli, G., Gentile, G., Brandalero, M.: SATELLITE INSTRUMENT CONTROL UNIT WITH ARTIFICIAL INTELLIGENCE ENGINE ON A SINGLE CHIP: ICU4SAT. IEEE Aerospace Conference Proceedings (2022) 43
[130] Muchandi, S.K.: ENABLING ACCELERATOR-SOC CO-DESIGN USING RISC-V CHIPYARD FRAMEWORK. UC Berkeley EECS Technical Report (2021) [131] Gianioudis, M., Xirouchakis, P., Loukas, C., Mageiropoulos, E., Mousouros, O., Mpartzis, S., Ioannou, A., Papaefstathiou, V., Katevenis, M., Chrysos, N.: Low-latency Communication in RISC-V Clusters. In: Proceedings of the International Conference on High Performance Computing in Asia-Pacific Region. HPCAsia ’24, pp. 73–83. Association for Computing Machinery, New York, NY, USA (2024). https://doi.org/10.1145/3635035.3635050 [132] Lee, W.-C., Lee, G.G.C., Yang, C.-C.: Sparse basis approach for lightweight ai system design. In: 2024 International Conference on Electronics, Information, and Communication (ICEIC), pp. 1–4. IEEE, Taipei, Taiwan (2024). https://doi.org/10.1109/ICEIC61013.2024.10457138 [133] Lim, S., Liu, Y., Benini, L., Karnik, T., Chang, H.-C.: F1: Striking the balance between energy efficiency & flexibility: General-purpose vs special-purpose ml processors. 2021 IEEE International Solid- State Circuits Conference (ISSCC) 64, 513–516 (2021) [134] Ueyoshi, K., Papistas, I.A., Houshmand, P., Sarda, G.M., Jain, V., Shi, M., Zheng, Q., Giraldo, S., Vrancx, P., Doevenspeck, J., Bhattacharjee, D., Cosemans, S., Mallik, A., Debacker, P., Verkest, D., Verhelst, M.: Diana: An end-to-end energy-efficient digital and analog hybrid neural network soc. In: 2022 IEEE International Solid-State Circuits Conference (ISSCC), vol. 65, pp. 1–3 (2022). https://doi.org/10.1109/ISSCC42614.2022.9731716 [135] Choi, E., Park, J., Lee, K., Lee, J.-J., Han, K., Lee, W.: DayNight architecture: Development of an ultra-low power RISC-V processor for wearable anomaly detection. Journal of Systems Architecture 152, 103161 (2024) https://doi.org/10.1016/j.sysarc.2024.103161 [136] Kanamori, T., Odan, T., Hirohata, K., Kise, K.: RVCar: An FPGA-Based Simple and OpenSource Mini Motor Car System with a RISC-V Soft Processor. IEICE Transactions on Information and Systems E105.D(12), 1999–2007 (2022) https://doi.org/10.1587/transinf.2022PAP0004 [137] K, S., Srivastava, A.K., Allam, S., Lilaramani, D.: Comparative analysis on Deep Convolution Neural Network models using Pytorch and OpenCV DNN frameworks for identifying optimum fruit detection solution on RISC-V architecture. In: 2021 IEEE Mysore Sub Section International Conference (MysuruCon), pp. 738–743 (2021). https://doi.org/10.1109/MysuruCon52639.2021. 9641594 [138] Tsutada, R., Hoang, T.-T., Pham, C.-K.: An Obstacle Avoidance Two-Wheeled Self-Balancing Robot. International Journal of Mechanical Engineering and Robotics Research, 1–7 (2022) https: //doi.org/10.18178/ijmerr.11.1.1-7 [139] Mousavikia, S.K., Gholizadehazari, E., Mousazadeh, M., Yalcin, S.B.O.: Instruction Set Extension of a RiscV Based SoC for Driver Drowsiness Detection. IEEE Access 10, 58151–58162 (2022) https://doi.org/10.1109/ACCESS.2022.3177743 [140] Nunes, W.A., Dal Zotto, A.E., Silva Borges, C., Moraes, F.G.: RS5: An Integrated Hardware and Software Ecosystem for RISC- V Embedded Systems. In: 2024 IEEE 15th Latin America Symposium on Circuits and Systems (LASCAS), pp. 1–5 (2024). https://doi.org/10.1109/ LASCAS60203.2024.10506171 . ISSN: 2473-4667 [141] Mustaffa, I.B., Khairul, S.F.B.M.: Identification of fruit size and maturity through fruit images using OpenCV-Python and Rasberry Pi. In: 2017 International Conference on Robotics, Automation and Sciences (ICORAS), pp. 1–3 (2017). https://doi.org/10.1109/ICORAS.2017.8308068 [142] Zareen, F., Amador, M.A.F., Karam, R.: Malware Detection in Embedded Devices using Artificial Hardware Immunity (2023). https://doi.org/10.21203/rs.3.rs-2758367/v1 . https://www. researchsquare.com/article/rs-2758367/v1 Accessed 2024-05-23 [143] Kumar, V.B.Y., Chattopadhyay, A., Haj-Yahya, J., Mendelson, A.: ITUS: A Secure RISC-V System-on-Chip. In: 2019 32nd IEEE International System-on-Chip Conference (SOCC), pp. 418– 423. IEEE, Singapore (2019). https://doi.org/10.1109/SOCC46988.2019.1570564307
44