1
Versatile On-device Adaptation at the Edge by Unifying Few-shot, Zero-shot, Continual, and In-context Learning
Abstract—With the ever-increasing pervasiveness of smart edge devices, the demand is growing for applications that can be tailored to users (e.g., custom keyword spotting) or patients (e.g., adaptive health monitoring). Yet, most edge devices rely on fixed inference algorithms and thus cannot learn on-device to personalize predictions. When they can, devices typically support only a specific learning scenario, such as few-shot learning (FSL): going beyond this requires resorting either to another specialized device or to cloud-based retraining, which implies significant energy and latency overheads, a lack of real-time capabilities, and privacy concerns. In this work, we introduce embedder-centric learning (ECL), a framework that unifies four different online learning scenarios: FSL for on-the-fly customization, continual learning (CL) for knowledge accumulation, zero-shot learning (ZSL) for leveraging semantic data, and in-context learning (ICL) for adapting beyond classification. We demonstrate in silicon that ECL can be deployed on resource-constrained devices across four real-world use cases representative of the aforementioned learning scenarios. Our approach establishes a new state-ofthe-art performance for FSL character recognition (Omniglot: 96.8% for 5-way 1-shot, 83.3% for 32-way 1-shot), and the first hardware baseline for CL in keyword spotting (NeuroBench keyword FSCIL: 71.8% for 200-way 5-shot). Moreover, we present the first hardware demonstrations of ZSL with semantic data (60.6% for 5-way spoken sentence classification) and ICL (46.2% at the 500th token of RegBench) operating at micro-to-milliwatt power budgets. Therefore, by unifying multiple learning scenarios, we pave the way for smart and versatile devices that can adapt right at the edge, without reliance on the cloud.
I. I NTRODUCTION
I
N the last decade, the integration of neural networks (NNs) on resource-constrained devices at the edge has become increasingly commonplace. Example use cases include navigation for drones [1], keyword spotting (KWS) on smart speakers [2], [3], [4], [5], and health monitoring on wearables [6], [7], [8], [9], [10]. However, most of these devices rely on fixed, pre-trained NNs that cannot be adapted post-deployment. Consequently, these devices cannot locally handle shifting data distributions, emerging features, specific or additional users, or evolving task requirements [11], such as the sim-to-real gap for drones, new keywords for KWS, and user- or patient-specific tailoring for wearables. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible. This publication was funded by the Dutch Research Council (NWO) as part of the projects AdaptEdge (file number 20267 in the NWO Talent Programme – Veni) and Transforming the Adaptability of Decentralized AI (file number NGF.1609.242.038 in the NGF - AiNed AiNed XS Europe programme). Douwe den Blanken, Martin Lefebvre and Charlotte Frenkel are with the Microelectronics Department (EEMCS Faculty), Delft University of Technology, 2628 CD Delft, Netherlands (e-mail: [email protected]; [email protected]; [email protected]; Corresponding author: Douwe den Blanken).
Approach Strength Limitations
Off-device adaptation
Energy-efficient High accuracy No on-device learning
High energy, latency and privacy costs
☹
High latency, power, mem. overhead Adapt on-chip
Memory-efficient Limited accuracy & scalability Constrained inference perf.
☹
Lack of flexibility across multiple learning scenarios
Learning scenarios
FSL
Sensory modalities
Images
CL
ZSL
Audio
ICL
Tokens
Breadth = Fundamental challenge
Support samples Query sample
Low dim. Temporal embedder
Wide range of modalities
☺☺
Support Query embeddings
Scenariospecific head
arXiv:2607.29353v1 [cs.LG] 31 Jul 2026
Douwe den Blanken, Graduate Student Member, IEEE, Martin Lefebvre, Member, IEEE and Charlotte Frenkel, Member, IEEE
FSL CL ZSL ICL
Flexible learning scenario selection
Fig. 1. (a) Overview of current approaches to learning at the edge and their limitations. (b) Learning at the edge requires supporting different learning scenarios, such as few-shot learning (FSL), continual learning (CL), zeroshot learning (ZSL), and in-context learning (ICL), while accommodating for different sensory modalities, forming a fundamental challenge. (c) Outline of the proposed embedder-centric learning (ECL) framework, which splits each learning scenario into a shared embedder and a scenario-specific head. This embedder-centric framing lets ECL support a wide range of sensory modalities while unifying four learning scenarios.
Addressing these needs is challenging, due to the various limitations of existing strategies to adapt (Fig. 1(a)). We group these strategies into three categories. The first category, comprising inference accelerators, relies on the cloud for off-device adaptation, which comes at the expense of the energy cost for the cloud link [12], [13], latency penalties that preclude online learning [13], [14], and risks of exposing user private data [13], [14]. The second trains a model from scratch with backpropagation (BP) directly on the device [15], [16], [17] but requires storing all intermediate activations, which is especially prohibitive for long temporal signals [11]. The third aims to alleviate this overhead of BP by implementing specialized learning algorithms in hardware [18], [19], [20], [21], or by exploring algorithms based on local learning rules [11], [22]. However, while multi-layer variants start emerging [23], [24],
2
their hardware implementation remains an open challenge to the best of our knowledge. Furthermore, custom learningCorresp. learning scenarios Different learning use cases optimized hardware typically trades inference efficiency for Customize on the fly learning accuracy [25]. Therefore, enabling on-chip learning “Morning” “Hey” at minimal hardware cost remains a challenge. “Wake up!” Beyond efficiently supporting a single learning scenario Accumulate knowledge over time on-chip, real-world deployment must cope with diverse data dynamics, label availability, and task structures. This demands support for additional learning scenarios, also across sensory Adapt using semantic information only modalities, from images to long-timescale audio (Fig. 1(b)). beak This breadth forms a fundamental challenge: current accelerafin shell tors cannot accommodate more than one learning scenario, yet designing a specialized chip for every combination of scenario Adapt beyond classification and sensory modality is impractical. Therefore, we aim to Cat → Chat Apple → Pomme Car → Voiture Book → ???? answer the following question: How can we optimally support different learning scenarios for versatility and efficiency across Fig. 2. (a) Comparison of four different learning use cases, from customizing to new tasks on the fly to adapting beyond classification tasks. (b) To enable use cases at the edge? of these use cases, a learning scenario can be adopted: FSL, CL, ZSL, In this work, we propose the embedder-centric learning each or ICL, each corresponding to a use case. (ECL) framework, which unifies few-shot learning (FSL), continual learning (CL), zero-shot learning (ZSL), and incontext learning (ICL) with support for a wide range of sensory Finally, we offer concluding remarks in Section V. To promote modalities (Fig. 1(c)). Our core idea is to frame each learning reproducibility, reuse, and improvement, all code for this paper scenario in an embedder-centric way, splitting it into two is open-source, including the training frameworks, accelerator components: a strong embedder NN, and a scenario-specific source code, and the test/simulation setup.1 head of one or more fully-connected (FC) layers that processes the resulting embeddings into a prediction. This embedder- II. BACKGROUND : O N - DEVICE L EARNING U SE C ASES AT centric design is effective for two reasons. First, embeddings THE E DGE represent data in a reduced-dimensional space, keeping the Figs. 2(a) and (b) respectively depict four different use knowledge memory small enough to maintain and reuse on- cases for learning at the edge and the learning scenarios that device across FSL, CL, ZSL, and ICL. Second, an embedder are used to tackle them. The first highlighted use case is that excels on temporal data, especially across long temporal to customize to new data or tasks on the fly. Examples of dependencies, lets ECL support different sensory modalities this include learning with only a few examples to detect new efficiently: audio (provided raw or sequentially as Mel- keywords [30], to classify novel image classes [31], and to frequency cepstral coefficients (MFCCs) [26] frames), images detect custom gestures [32]. Few-shot learning (FSL) addresses (provided sequentially as pixels), or, more generally, tokens. this problem of recognizing new classes that were not seen We validate the performance and low cost of the ECL frame- during pre-training using only a few labeled samples per class. work for each of the learning scenarios both in software and The new classes are referred to as ways while the samples hardware using our recently introduced Chameleon system- per class are referred to as shots [28], [33], [34], [35], [36]. on-chip (SoC) [25]. We surpass state-of-the-art (SotA) FSL As FSL only requires limited data, it is a particularly good accuracies on the Omniglot dataset [27] (96.8% for 5-way 1- fit for the edge, where the online appearance of new data is shot, 83.3% for 32-way 1-shot), and we demonstrate for the first often scarce (e.g., examples given by users or rare events not time on-device CL on the NeuroBench keyword few-shot class- accommodated for during training). FSL, however, requires incremental learning (FSCIL) dataset to classify 200 classes at every new class to be available at once. a 9.5 µW real-time power. We also perform ZSL using semantic When classes instead emerge incrementally, the model data for the first time on-device, using only 3.1 µJ to learn five must accumulate knowledge over time rather than acquire new classes. In addition, we present the first demonstration of it in a single step. Examples include incrementally learning on-device ICL on a formal language, requiring only 16.8 µJ additional image classes [37], such as a growing alphabet of per token. These results establish new baselines across learning handwritten characters, or learning to detect an expanding scenarios and sensory modalities, and underline our embedder- set of keywords [38]. To handle this second set of use cases, centric approach as an enabler for learning at the edge. continual learning (CL) can be applied. CL allows an NN to The remainder of this article is structured as follows. First, gain knowledge of new tasks or data distributions over time, Section II introduces the four learning use cases at the edge without forgetting the previously acquired knowledge [39]. and the corresponding learning scenarios considered in this However, both FSL and CL still depend on labeled data, which work. Then, Section III presents our proposed ECL framework the edge cannot always provide. and details its implementations of the FSL, CL, ZSL, and A third category of use cases instead offers only semantic ICL learning scenarios. Then, Section IV outlines our test information from different sensory modalities, such as highsetup and presents hardware and software results for each 1 https://github.com/cogsys-tudelft/ecl learning scenario, followed by a synopsis of the overall results.
3
Query sample
Embeddings Support
Prediction Scenariospecific head
Query Compressed accumulated knowledge
FC
Avg. embeds.
Support samples 2 ways 2 shots
Support
Embedder
Query Query sample
T=1 T=2
Semantic properties
Diff. sensory modalities Support samples α
β
γ
δ
T=1 T=2
L2 distance calc. =
Supp.
Semantic embedder
Query
Embedder
Support
Query sample
Query
FC layer
FC layer
“Relation head”
3 ways
Embedder
Query
Query sample
“Prototype”
Logit
T=0
Supp.
Embedder
Concat
T=0
FC
L2 distance calc. =
Concat
Classes
Temporal embedder
Convert to w+b
Support samples
MLP layer
MLP layer
γ
δ
ε
Fig. 3. Overview of the ECL framework. (a) ECL unifies four online learning scenarios through an embedder-centric formulation: support and query samples are embedded before a scenario-specific head, consisting of one or more FC layers, produces a prediction on the query sample. (b) FSL in ECL uses prototypical networks (PNs) [28]: support embeddings are averaged class-wise into prototypes, and the query sample is assigned to its nearest prototype. Prototypes are transformed into an equivalent representation in terms of weights and biases (w + b) of an FC layer, to compute the distance between the query sample and the prototypes efficiently. (c) CL in ECL reuses the same PN-based approach, but appends the parameters corresponding to a new class (learned at time T ≥ 1) to existing FC layers’ parameters (learned during an initial training phase at T = 0), instead of overwriting them. (d) ZSL in ECL uses a relation network (RN) [29] with two embedders, one for the semantic support data and one for the query sample; their embeddings are concatenated and passed to a relation head, here an multi-layer perceptron (MLP), for the final prediction. (e) ICL in ECL concatenates the query and support samples before embedding, and the scenario-specific head feeds the joint embedding to an MLP to perform the prediction.
level attributes, audio, or textual descriptions. Examples of this include animal classification from high-level properties only [40] or spoken-sentence detection from transcriptions alone. Zero-shot learning (ZSL) enables learning from such semantic information alone, removing the need for labeled test-domain data for each new task [28], [40]. In effect, ZSL learns new tasks using a sensory modality different from that of the test data. Each scenario so far stops at classification. Yet many edge use cases call for more, for example, forecasting temporal data or few-shot language modeling, forming the fourth and final group of use cases. Using in-context learning (ICL), use cases such as regression or next-token prediction tasks [41], [42], [43], where the next symbol should be predicted given a vocabulary and labeled input-output pairs, can also be dealt with. ICL relies purely on the forward pass of a model: learning happens by building memory associations between a sequence of labeled examples combined with a test input inside a context window, without modifying any model parameters [43]. ICL was first observed as an emergent property of the transformer architecture [44] in large-scale language modeling [43]. However, it has recently been shown
that it is not exclusive to these models [45]. Supporting these four learning scenarios at the edge conventionally requires a separate specialized system for each. However, this is impractical under the tight compute and memory budgets of edge devices. Hence, how to efficiently support all four within one device remains an open problem. III. T HE ECL F RAMEWORK FOR V ERSATILITY AT THE E DGE In this work, we propose a framework called ECL, that unifies the above four learning scenarios by framing each of them in an embedder-centric way. ECL thus allows splitting each learning scenario into two key components illustrated in Fig. 3(a): (i) an embedder NN and (ii) a scenario-specific head. First, the NN maps support samples, the labeled data for learning, and a query sample, the input to process after learning, to an embedding, i.e., a feature vector of substantially lower dimensionality than the input. Together, the embedded support samples represent a compressed form of accumulated knowledge. Second, the scenario-specific head processes embedded knowledge together with the query embedding using one or several fully-connected (FC) layers, in a way that
4
is dependent on the learning scenarios. Through this shared present at adaptation time. In CL, on the other hand, we assume structure, ECL efficiently supports versatile adaptation. that new classes appear over time, denoted as T ≥ 1, after an While our framework removes the need for designing a initial adaptation phase, denoted as T = 0 (Fig. 3(c)). Therefore, specialized implementation per scenario, supporting adaptation instead of calculating the FC parameters once and overwriting across sensory modalities, however, would still require a the previous ones, we append new entries to the PN’s FC different NN type for each. To avoid this, ECL requires parameters over time for each additional class while keeping an NN that can embed samples containing long temporal the previous FC parameters unchanged, thereby alleviating dependencies, such as audio consisting of MFCC frames or catastrophic forgetting. 3) ZSL Implementation in ECL: To perform ZSL in ECL, we even raw samples, images converted into pixel streams, or effectively any data type provided as tokens. We therefore use a relation network (RN) [29]. Although PNs also support propose to use temporal convolutional networks (TCNs) [46] as ZSL [28], we choose an RN since it learns a distance function the temporal embedder for ECL. TCNs are NNs that use stacked rather than assuming one (for example, L2 in PNs), which causal 1D convolutions combined with residuals to capture can improve performance for ZSL [29]. Since ZSL uses two long-range relationships in sequential data [46]. These NNs can data modalities, we require two embedder NNs for the RN extract high-quality embeddings from long sequences while (Fig. 3(d)). The first NN produces support embeddings on a their memory cost scales only logarithmically with sequence per-class basis to learn from the provided semantic information. length [25], allowing them to efficiently support a wide range This NN can either be a TCN or a multi-layer perceptron (MLP), of sensory modalities. While a new embedder still has to be depending on whether the semantic information is temporal or pre-trained off-chip for each new sensory modality, the target not. The second NN then produces the embedding for the todeployment platform for ECL only needs to support one NN be-classified sample in the target domain. All embeddings are then concatenated and fed into an MLP, the so-called relation type, simplifying the hardware requirements. Independent of the chosen embedder, by construction, the head, for prediction. These last two operations constitute the embedder-centric stance of ECL itself yields three further scenario-specific head for ZSL. In this way, the two embedder advantages for hardware implementation. First, the target NNs specialize in producing high-quality embeddings in their hardware only needs to support regular NN inference, since input domain while the relation head learns to combine this generating an embedding is equivalent to a forward pass. information into a prediction. 4) ICL Implementation in ECL: For ICL, the dataflow is Second, the knowledge memory that stores embeddings can slightly different from the previous learning scenarios. Instead stay small (a few kB), because embeddings have far lower of concatenating support and query samples after the embeddimensionality than the input samples. Third, formulating ding step like in ZSL, we concatenate all samples before prothe scenario-specific head as one or more FC layers lets cessing with the embedder network (Fig. 3(e)). For ZSL, this is ECL execute within the same NN inference pipeline, requiring not possible as the scenario incorporates two different data types only minor control logic to construct the layer. Overall, and hence requires two different embedders. For ICL, however, ECL avoids learning-specific memories and introduces only this order reversal has two key implications. First, it gives the minor control overhead while relying on standard inference embedder network the ability to directly model relationships hardware, bringing efficient learning across FSL, CL, ZSL, and between all inputs, instead of the relation head in ZSL. Second, ICL to extreme-edge devices. it allows the TCN to produce one output prediction for each inFurthermore, let us now explain how the ECL framework put, which is necessary for tasks such as regression or language implements each learning scenario (Fig. 3), by detailing the modeling. After this joint-embedding step, a single next-output operation of the embedders and scenario-specific heads. prediction can be calculated using one or more FC layers in 1) FSL Implementation in ECL: For FSL deployment, we an MLP, which forms the scenario-specific head for ICL. use a technique called prototypical networks (PNs) [28]. In PNs, an embedder NN is used to embed the available support samples. IV. R ESULTS A prototype for a new class is then formed by averaging the This section presents the experimental results from using the embeddings from the support samples for that class (Fig. 3(b)). ECL framework. First, Section IV-A introduces our hardware After embedding, an unknown query sample is then classified test platform. Then, for each learning scenario (Sections IV-B as the class of the prototype with the lowest L2 distance to to IV-E), we describe the representative benchmark or dataset its embedding [28]. To efficiently support PNs in ECL, we used in this work, followed by software results before employ the equivalent transformation of PNs into weights presenting ECL’s performance when deployed in hardware. and biases that parametrize a single FC layer [28]: exact L2Section IV-F finally outlines a synopsis of the obtained results. based nearest neighbor classification can then be performed in a single matrix-vector multiplication instead of requiring additional hardware for distance computation, a technique also A. Test Platform used in Chameleon [25]. This equivalent FC layer constitutes To validate ECL in silicon, we use the Chameleon SoC [25] the scenario-specific head for FSL. as our test platform. Chameleon is an NN accelerator optimized 2) CL Implementation in ECL: For CL, we largely follow for generating and processing dense embeddings from temporal the same procedure as for FSL. In FSL, however, the equivalent data using TCNs. Fig. 4(a) outlines Chameleon’s architecture. FC layer can be computed instantaneously, as all support data is The SoC’s matrix-vector processing element (PE) array uses
5
Mat.-vec. PE array
Mul-free PE
Learn. controller
4-bit log2
FSL
shots
ways
Weight mem. (65 kB)
CL
shots
ways
FC
x << w
Infer. controller
W
Bias Proto. to b mem. FC params (3.5 kB)
4-bit
Act. mem. (2 kB) Inner loop step
# inner # outer Extract Append Load steps steps W + b? to mems? ... ? FC
ZSL
ways
1
Train embds.
ICL
1
1
-
Outer loop step Learning (1 way)
Learning controller
Learning (1 way)
1 shot
1 shot
Inference
ECL component Embedder Controller used
Embedder
Embedder
Embedder
Embedder FC
PE array
PE array
PE array
PE array
PE array
Active module
Time Store embeddings
Prot.
2 ways learned
Learning (2 ways; T = 0)
FC
Use PE array for Load embeddings averaging embeddings
1 shot
Learning (1 way; T = 1)
Inference
Inference
1 shot
3 ways learned
ECL component Embedder Controller used
Embedder
Embedder
FC Embedder
Embedder
PE array
PE array
PE array
PE array
PE array
Active module
Time
Replace embedder Inference
Learning (5 ways)
ECL component Sem. emb. Controller used
Sem. emb.
Sem. emb.
Sem. emb.
Sem. emb.
Embedder
PE array
PE array
PE array
PE array
PE array
PE array
Active module
Time
ECL component
MLP
Load train embeds.
Learn + infer Embedder
FC
Learn + infer
Learn + infer
Learn + infer
MLP Embedder MLP Embedder MLP Embedder MLP
Controller used Active module
PE array
PE array
PE array
PE array
Time
Fig. 4. (a) Simplified architecture overview of Chameleon. The SoC includes a matrix-vector processing element (PE) array which uses 4-bit log2 weights for multiplier-free PEs and 4-bit activations. Together with the inference control logic, it forms the SoC’s NN inference data path. (b) Comparison of Chameleon’s configurations per learning scenario, which are implemented by the SoC’s learning controller. (c)-(f) Qualitative illustration of ECL’s execution schedules on Chameleon for FSL, CL, ZSL and ICL respectively. We show per scenario which main module and controller are used over time for the two ECL components. Overall, ECL’s mapping to the SoC can be partitioned in three phases: (i) embedding all support samples for a single way (inner loop), (ii) optionally converting them to prototypical parameters (outer loop), and (iii) using the support embeddings, FC or MLP with the embedded query sample to predict. Dashed brackets indicate inner loop steps, while regular black brackets indicate outer loop steps. Blue brackets indicate the inference phase. Colors in (c)-(f) match the blocks in (a) and components in Fig. 3.
inference-optimized quantization with 4-bit activations and 4bit log2 weights, to shrink the memory footprint and reduce the multipliers to bit shifters. The NN weights and biases reside in 65 kB and 3.5 kB memories respectively while activations and embeddings share a 2 kB memory. An inference controller manages inference and embedding generation while a separate learning controller handles embedding processing for learning. Fig. 4(b) compares how the latter controller is configured to support each of ECL’s four learning scenarios. These configurations result in distinct execution schedules. Fig. 4(c)-(f) show how ECL’s two key components map onto Chameleon over time for FSL, CL, ZSL, and ICL respectively. Additionally, they indicate which main module and controller are used over time. All schedules execute ECL in three phases, as annotated in Fig. 4(c) for FSL. First, the TCN embeds a support sample, and the SoC stores the resulting embedding
in its activation memory. We call this an inner-loop step, which repeats a number of times dependent on the scenario (see Fig. 4(b)). Second, if the scenario requires it, the stored support embeddings are converted into equivalent prototypical parameters and written to the same weight and bias memories that hold the embedder’s parameters. Together, the first two phases form one outer-loop step, which also repeats a a number of times dependent on the scenario. Third, the query sample is embedded and passed to the scenario-specific head, i.e., one or a few FC layers, to make a prediction. Fig. 4(d) shows these three phases for CL under ECL in a 1-shot use case. At T = 0, a single-step inner loop and a two-step outer loop learn two ways, enabling inference over those two ways. A single inner and outer loop then learn an additional class, whose weights and bias are appended to the weight and bias memories, enabling inference over three ways. In contrast, for ZSL (Fig. 4(e)), no parameter conversion is needed, reducing the outer loop to a single step but changing the inner loop step count to the number of ways. After embedding the semantic samples, the semantic embedder is replaced by the querydata embedder. Then, the SoC is ready for inference. For ICL (Fig. 4(f)), all samples are concatenated before processing by the embedder, collapsing both the inner and outer loops to a single step. These schedules verify a key advantage of our embedder-centric formulation: extensive reuse of the inference data path for both TCN embedding generation and FC-layer computation. As a result, ECL’s hardware efficiency hinges primarily on how efficiently inference is performed. Since Chameleon is optimized for processing TCNs, it underpins the learning efficiency results that follow. B. ECL Supports On-the-fly Customization with FSL We first demonstrate FSL under ECL, where an NN can learn to recognize new classes from only a few labeled examples. 1) Representative Benchmark: To evaluate the performance of our framework on FSL for classification tasks, we use the popular Omniglot dataset [27]. The dataset consists of a total of 1623 different handwritten characters across a large variety of alphabets, containing 20 sample images per character of 28 × 28 pixels. Since our framework is designed for processing sequential data, we flatten each input image to shape it into a 1D sequence. We test on Omniglot across the standard evaluation settings of 1, 5-shot and 5, 20-way, as well as 32-way 1-shot. We choose this dataset as it is widely used across both software and hardware works to measure FSL performance [18], [20], [28], [33]. Hence, Omniglot enables a direct comparison to these previous works. 2) Software Results: Table I shows the FSL accuracies across the standard FSL evaluation settings for both the TCN embedder that ECL employs and a commonly used convolutional neural network (CNN) alternative [28], [33]. Compared to the similarly-sized CNN, the TCN incurs a drop of 0.2 to 1.6 accuracy points across all settings. However, this is acceptable and expected, as neighboring pixels from the original image can be tens of timesteps apart when flattened for the TCN. Hence, by using a sequential embedder in ECL, we maintain support for learning from image data while we extend support
6
TABLE I. Accuracy comparison between similarly sized FP32 TCN and CNN across standard evaluation settings for FSL on Omniglot using PNs.
CNN [28] TCN
Accuracy (%)
100 +3.4
5-way
5-way 1-shot 5-shot
20-way 1-shot 5-shot
98.8% 97.9%
96.0% 94.4%
+0.5
100
99.7% 99.5%
20-way +16.6
98.9% 98.5%
100
90
90
90
80
80
80
70
70 5-shot FSL-HDnn [18]
70
1-shot 5-shot Kim et al.* [21]
1-shot SAPIENS* [20]
32-way
Power Latency Frequency Model size All weights on-chip FSL logic overhead Total core area
FSL-HDnn [18] 27 mW 53 ms 100 MHz 5.5 MB ✗ 25% 11.3 mm2
+11.3
90
1-shot This work
Fig. 5. FSL test accuracy comparison on the Omniglot dataset between FSL accelerators that have reported Omniglot results in silicon. We set new accuracy records across all shot-way scenarios, with improvement ranging from 0.5 to 16.6 accuracy points. Missing bars originate from non-reported accuracies. Error bars indicate 95%-confidence intervals for our work. Green annotations indicate accuracy point improvements over the SotA. ∗ indicates the use of an off-chip FP32 embedder.
This work 11.6 mW 12.9 µW 0.59 ms 0.54 s 100 MHz 100 kHz 59 kB ✓ 0.5% 0.83 mm2
(a)
100
Accuracy (%)
Model
TABLE II. Comparison to FSL-HDnn [18] on power, latency, model size and FSL area overhead. At iso-frequency, ECL is significantly more efficient, while relaxing latency constraints yields micro-watt level power. Bold indicates best.
+4.0
80 70 60 50
-8.4
M5 CNN [38] Rec. SNN [38] ECL (SW, FP32) ECL (HW, 4-bit log2 )
40 100 110 120 130 140 150 160 170 180 190 200 for learning from variable-length sequences across modalities, Number of classes which we will demonstrate in the remainder of this work. (b) 3) Hardware Demonstration: Fig. 5 compares the FSL M5 CNN [38] Rec. SNN [38] ECL (FP32) ECL (4-bit log2 ) accuracies of the ECL framework on Chameleon to those Parameters 1.40M 2.12M 115k 115k of other silicon designs that perform FSL on Omniglot. Bit width 32 32 32 4 These results for FSL were previously reported in [25] and Model size 5.61 MB 8.50 MB 459 kB 58.2 kB are included here to show the performance of ECL on a representative use case for FSL. It can be seen that through 6. (a) CL accuracy comparison between ECL in software (SW) and in our framework, we achieve new accuracy records across all Fig. hardware (HW) on Chameleon and the NeuroBench baselines on the keyword shot-way scenarios versus prior works. When comparing in FSCIL dataset [38]. While ECL in FP outperforms both baselines after learning detail with the FSL-HDnn chip [18] (Table II), which is, to 20 new classes, our framework incurs a 4.4 accuracy-point drop compared to the 96× larger M5 baseline after continually learning the full dataset under tight the best of our knowledge, the only other end-to-end work ∼ quantization. Shaded areas indicate 95%-confidence intervals. (b) Parameter with silicon results on Omniglot, we additionally demonstrate count, bit width, and total model size comparison between the four models. major gains in hardware efficiency. We find that, at iso-clockfrequency, our framework reduces power by 2.3×, latency by 90×, and model size by 93×. Furthermore, while FSL- benchmark from the NeuroBench initiative [38]. The dataset HDnn [18] requires external memory accesses for the NN, we consists of 100 spoken keyword classes for pre-training and eliminate these in our framework by specifically training a 100 extra keyword classes for class-incremental CL, all with quantized NN that can be stored fully in on-chip SRAM. In a duration of 1 s. This high number of classes combined addition, the total silicon area overhead of the hardware required with the temporal sensory modality of the samples makes the for learning with our framework is only 0.5%, compared keyword FSCIL benchmark an ideal fit for measuring the CL to 25% for FSL-HDnn [18]. This stems from FSL-HDnn’s performance of ECL. During CL, we follow the original setting reliance on hyperdimensional computing (HDC) [47], which and learn the 100 keywords over ten 10-way 5-shot sessions. is significantly more complex than our equivalent FC-layer To ensure a fair comparison with the baseline models from forward pass and requires dedicated hardware. By setting new the original NeuroBench work [38], each 48 kHz audio sample accuracy records while significantly reducing the power, latency, is pre-processed off-chip into MFCCs. By requiring a preand area, we validate ECL’s ability to efficiently support on- trained model to learn new classes while retaining knowledge of previously learned ones, this benchmark aims to emulate a the-fly customization at the edge using FSL. realistic use case at the edge. 2) Software Results: Fig. 6 shows how the test accuracy C. ECL Supports Knowledge Accumulation with CL varies with an increasing number of classes learned Second, we demonstrate CL under ECL, where an NN can continually compared to two baseline models published in learn multiple tasks over time without overwriting knowledge NeuroBench [38]: the M5 CNN [48] and a two-layer recurrent of previously learned tasks, unlike with FSL. leaky integrate-and-fire (LIF) spiking neural network (SNN). 1) Representative Benchmark: To validate the CL capabil- Compared to the CNN, ECL’s floating-point (FP) NN has a ities of our framework, we test it with the keyword FSCIL slightly lower accuracy from pre-training but starts to outper-
7
TABLE III. CL performance summary of ECL deployed on Chameleon for the NeuroBench keyword FSCIL dataset. @ 0.73 V, 14.4 kHz Final accuracy (200 classes) Latency per sample Average power Energy per sample Embedder network size Total embeddings size (200 classes)
71.8% 1.0 s (real-time) 9.5 µW 9.5 µJ 58.2 kB 6.4 kB
TABLE IV. ZSL performance of ECL deployed on Chameleon for the FSC dataset. On this dataset, the latency and energy for inference are significantly higher than those for learning. This is due to the large size of the audio sample used for inference compared to the size of the semantic data used for learning. @ 0.73V, 5.8 kHz Uses semantic data? 5-way accuracy Learning power Total learning latency (5 ways) Total learning energy (5 ways) Inference power Inference latency (per sample) Inference energy (per sample)
✓ 60.6% 8.2 µW 383 ms 3.1 µJ 8.0 µW 13 s (real-time)1 104 µJ
form it after learning only twenty classes with a final increase 1 Processing time matches maximum audio length of four accuracy points, even though ECL’s NN is ∼ 12× (13 s); computation averages 16.4 ms/frame, matchsmaller. We postulate that our smaller embedder is forced by its ing one MFCC frame’s duration. capacity limit to learn more generic, reusable features, leading to a higher accuracy after the novel classes are introduced. Furthermore, we find that our framework improves significantly Every sentence concerns a single action, object, and location. over the SNN, even though ECL’s FP NN is ∼ 18.4× smaller. For example: “Bathroom (location) heat (object) down (action)”. Overall, our framework sets a new standard for CL on the We use 75% of these phrases for training, 10% for validation, keyword FSCIL dataset at a significant reduction in model size. and 15% for testing, with no class overlap across splits. To 3) Hardware Demonstration: Fig. 6 also displays the CL enable ZSL on FSC, we use an MLP to create the semantic accuracies of the ECL framework when the model is quantized transcription embeddings, while a TCN embeds the test audio and executed on the Chameleon chip. The ∼ 8× reduced model sample, which is converted to MFCCs beforehand. These size of only 58.2 kB, compared to the FP ECL model, results embeddings are then combined and passed into the relation head in a drop of 8.4 accuracy points after learning 200 classes. to predict which never-heard-before transcription was spoken. Compared to the 96× larger M5 model, our quantized model While FSC was originally proposed as a benchmark for spoken only incurs a loss of 4.4 accuracy points. Table III reports the language understanding through supervised classification, we corresponding hardware performance of ECL for CL on this adopt it under a zero-shot setting. In this setup, the semantic benchmark. To the best of our knowledge, this is the first work zero-shot data are the transcriptions of the to-be-classified to demonstrate end-to-end fully on-chip CL on the keyword sentences. We opt for this benchmark as there are currently no FSCIL dataset from NeuroBench. The only other end-to-end CL datasets based on sequential data suitable for zero-shot learning work with silicon results is Clo-HDnn [19]. However, it only in extreme-edge deployments. supports learning up to 128 classes due to high dimensionality 2) Software Results: Under a 5-way ZSL scenario on of the used embeddings (1024-8192). Additionally, similar to FSC, ECL achieves a test accuracy of 80.8% with FP FSL-HDnn [18], it requires reloading the on-chip weights durprecision. Since ZSL on FSC is introduced in this paper ing inference as it cannot store all weights on-chip, introducing as a small-scale alternative to existing benchmarks, no significant additional latency and energy penalties. ECL addirect comparisons currently exist: Relation Network [29], dresses both limitations. First, ECL’s much lower embedding for instance, demonstrated ZSL using embedding networks dimension for CL (64) enables 4–32× smaller embedding data ranging from approximately 10 to 50 million parameters to sizes, so that 200 classes can be stored on-chip using only achieve 84.5% test accuracy on a 10-way task. By contrast, 6.4 kB. Second, by compressing ECL’s embedding NN, we can ECL uses a semantic embedder with only 120k parameters, a store all model parameters fully on-chip. Due to being fully test embedder with 96.7k parameters, and a relation head with end-to-end on-chip, ECL uses only 9.5 µW on Chameleon 28.6k parameters, making the total network approximately 40 for learning and inference while processing the keywords in to 200× smaller than those in reference works. real time. By demonstrating the first end-to-end fully on-chip 3) Hardware Demonstration: Table IV reports the perfordeployment on the NeuroBench keyword FSCIL dataset, we mance of ECL for ZSL on the FSC dataset when deployed show our framework’s ability to accurately learn 100 new on Chameleon. To the best of our knowledge, this is the first classes with only a few samples and validate how ECL can be demonstration of ZSL using semantic data at a µW power used to accumulate knowledge over time at the edge with CL. budget suitable for the edge. The only other work that performs ZSL in silico with a similar power budget is the work by D. ECL Leverages Semantic Data with ZSL Liu et al. [50]. In that work, EEG data are used to adapt a Third, we demonstrate ZSL under ECL, where an NN learns seizure-prediction model to unseen patients, enabling learning to recognize new classes from semantic information alone, i.e., without requiring seizure recordings from the new patient. without labeled data in the query domain, unlike FSL and CL. However, their approach has three key limitations. First, it 1) Representative Benchmark: To demonstrate ZSL with does not support semantic data, so it cannot generalize to our framework, we use the Fluent Speech Commands (FSC) truly unseen classes: all learned classes were seen during dataset [49]. The FSC dataset consists of 248 unique spoken training. Second, it requires replay of data from previous phrases up to 13 s long, each with a corresponding transcription. patients, incurring energy-expensive off-chip memory accesses.
8
(sharing)
-19.6 (quant + sharing)
Cumulative accuracy (%)
55 50 45
Third, it relies on backpropagation for learning, requiring 16-bit data widths and incurring a 7% area overhead. Our proposed ECL framework addresses all of these shortcomings. By leveraging semantic data, ECL enables genuine zero-shot generalization to five previously unseen classes without any replay requirement. In addition, by using an RN as part of ECL for ZSL, we can use inference-level bit widths during learning. Combined, this yields a learning latency of 383 ms to learn five new classes for a total learning energy of 3.1 µJ. During inference, we process each sample within its duration to match the audio data rate, by setting the clock frequency to 5.8 kHz, which leads to a real-time inference power of 8.0 µW. While advantageous for tight energy or power constraints, our low-bit-width RN does impact ZSL accuracy, going from 80.8% to 60.6% when quantized. Two limitations of Chameleon explain this gap: (i) the lack of per-channel scaling support requires both embedder NNs to share a single embedding quantizer (Fig. 7(a)), which leads to a 4.7-point training accuracy drop at 16-bit quantization, and (ii) while 6-bit quantization only loses 5.8 points compared to FP32, the 4-bit log2 weight format incurs a 19.6-point drop (Fig. 7(b)). By demonstrating end-to-end ZSL under ECL using semantic data at a µW budget for the first time, we enable learning even when labeled target-domain data are unavailable and add a level of versatility beyond what FSL and CL alone can provide. E. ECL Supports Adaptation Beyond Classification with ICL Fourth, we demonstrate ICL under ECL, where an NN can learn to perform regression or predict the next token in a sequence, going beyond the previous learning scenarios focused only on classification.
-5.3
40 35 30
Fig. 7. Effect of weight and activation quantization on ZSL accuracy for FSC. (a) Training accuracy at 2k steps. (b) Test accuracy at the checkpoint with the highest validation accuracy. The FP32 bars show the baseline model that initializes all quantized runs. (a) As Chameleon does not support perchannel scaling, both embedder NNs must share a single embedding quantizer. A unique quantizer per NN recovers the FP32 training accuracy at 16 bit, whereas sharing one costs 4.7 accuracy points, although the confidence intervals overlap. (b) Test accuracy increases slightly from 16 to 6 bit but collapses at 4 bit. The 4-bit log2 weights supported by Chameleon perform better than standard 4-bit weights, but trail FP32 by 19.6 accuracy points, against 5.8 points for 6-bit quantization. 4-bit log2 accuracy here is higher than in Table IV because this ablation uses a larger bias bit width than Chameleon supports. Hyperparameters are identical across runs. Bars show the mean across n runs, points individual runs, and error bars 95% confidence intervals.
-1.0
Chance level = 10%
TF (FP16) ECL (SW, FP32) ECL (HW, 4-bit log2 )
0 50 100 150 200 250 300 350 400 450 500 Sequence length
Fig. 8. Cumulative ICL test accuracy comparison between ECL in software (SW) and in hardware (HW) on Chameleon and a parameter-matched TF on the RegBench dataset. Accuracy is measured over increasing sequence length for 5k query examples. While our FP model nears the transformer (TF) to within one accuracy point, when quantized, the final accuracy drops by 5.3 points.
1) Representative Benchmark: To demonstrate ICL under ECL, we use RegBench [51], a dataset designed to assess in-context language learning ability. The RegBench dataset consists of a collection of problem instances, each represented as a sequence of examples. All examples within a given instance are sampled from the same probabilistic language [51]. The goal in each example is to predict the next token based on the previous input. Since, with increasing context length, more examples have been shown, prediction performance should increase. In this work, we follow the standard setting of the benchmark, with a vocabulary size of twenty tokens and between ten and twenty examples of at most fifty tokens. This benchmark dataset aims to be simple enough for analysis in small models while still capturing the core characteristics of ICL in large language models. This makes RegBench a good fit for this work, as we aim to demonstrate our framework’s capabilities at the edge, constrained by limited memory and compute. 2) Software Results: Fig. 8 shows how the cumulative test accuracy varies with increasing sequence length for both ECL at 133k parameters and a parameter-matched transformer (TF) baseline. Despite ECL’s TCN model not having direct connections from all to all input tokens, unlike a TF [44], the TCN’s FP accuracy approaches that of the TF within one accuracy point at the 500th token. This result shows that, at a small scale, ECL achieves near-TF ICL performance while requiring 31× less activation memory storage (9.2 vs. 288 kB) by exploiting the TCN’s dilation-induced sparsity [25]. 3) Hardware Demonstration: Fig. 8 shows that the quantized model deployed on Chameleon reduces the cumulative accuracy by 5.3 points compared to ECL in FP at the 500th token. The quantized network occupies only 66.4 kB, making it suitable for execution on severely memory-constrained edge devices. We achieve a latency of 2.15 ms per token and an average power consumption of 7.83 mW to yield an energy cost of 16.8 µJ per token at a clock frequency of 100 MHz. To the best of our knowledge, this is the first demonstration of ICL at a milliwatt power budget. By showing how ICL can tackle this type of task, we demonstrate how ECL extends beyond plain classification, increasing the framework’s versatility.
9
F. Discussion
of these choices suit the target scenario well, they prevent the support of additional learning scenarios. Alternatively, Kwon et al. [56] selectively update parts of the full NN with BP, relaxing the fixed-embedder assumption entirely. Yet, in this work too, this greater flexibility was not leveraged beyond FSL. The key insight of ECL is therefore not in any individual scenario implementation, but in recognizing this shared embedder-centric structure and exploiting it for hardware unification across four learning scenarios. With regard to future work, we propose four key avenues. First, one limitation of our evaluation is that we demonstrated ECL on relatively small datasets and NNs, reflecting the constraints of Chameleon as our extreme-edge demonstration platform. However, the compute capabilities of edge devices are still increasing: they can then also deal with larger and more complex datasets. Hence, future research should investigate how our framework scales, not only in terms of model and data size but also in task complexity. Second, we postulate that ECL can be expanded to support simple few-shot reinforcement learning (RL) tasks by repurposing ECL’s ICL mechanism, since RL problems can be viewed as sequence modeling problems that can also be learned incontext [57], [58]. Third, ECL is currently formulated to be completely gradient-free for maximum hardware efficiency through simplicity. Hence, a natural next step would be to explore whether allowing gradient descent steps, restricted to the scenario-specific heads, could further enhance performance while retaining hardware efficiency. Fourth, weight-transportfree [59] or forward-learning [60] methods could also be considered for the weight updates without incurring the cost of full backpropagation. How such mechanisms can be efficiently incorporated into our unified embedder-centric framework is a promising avenue for future work.
Putting together the results in Sections IV-B to IV-E, we identify four common threads. First, while FSL is the only scenario that explicitly requires few labeled samples by definition, we also demonstrate low-sample regimes for CL (5 samples per class), ZSL (0 samples in the target domain), and ICL (20-sample sequences). Thus, across all four learning scenarios, ECL remains suitable for edge deployment, where labeled data are inherently scarce [25]. Second, by choosing the TCN as the embedder rather than modality-specific NN architectures, ECL operates across images (FSL on Omniglot), audio (CL and ZSL), and tokens (ICL on RegBench). To the best of our knowledge, no prior work has demonstrated learning across all three modalities on a single device at a micro-to-milliwatt power budget. Third, since all learning scenarios share the same embedder NN type, and since all scenario-specific embedding processing logic can be incorporated into a single hardware block with only 0.5% area overhead, ECL significantly expands accelerator versatility at minimal added cost. Fourth, by supporting learning scenarios ranging from FSL to ICL, we support a progression of relaxing assumptions about available data at test time. Namely, FSL and CL assume labeled samples in the target modality while ZSL removes that assumption by substituting these with semantic data. ICL then removes the class structure entirely. This property allows ECL to be matched to the data available in a given use case. Although above we presented the four learning scenarios as a progression, they are not necessarily mutually exclusive. For example, combining FSL with CL yields few-shot classincremental learning (FSCIL) [37], where a model incrementally learns new classes with only a few shots per class. Likewise, combining ZSL with CL gives continual ZSL [52], [53], [54] or lifelong ZSL [55]. More generally, ICL learns next-token prediction tasks, so it is compatible with FSL, CL, V. C ONCLUSION and ZSL once they are framed in this format. One way to do this is to predict a class token given a context of the support In this work, we presented ECL, a framework that unifies embeddings and a query embedding. However, each of these the FSL, CL, ZSL, and ICL online learning scenarios across learning scenarios still requires an embedder network to embed sensory modalities. We enabled this unification by framing all samples: adding an ICL mechanism on top increases the every scenario in an embedder-centric way, which allowed operation count, latency, and weight storage requirements, an close compatibility with existing inference hardware while overhead to be considered in light of the target deployment keeping the embeddings to a few kB in memory. To learn scenario specifications. across sensory modalities, we used TCNs to support samples Taking a step back, we find that most prior works already with long temporal dependencies. rely on a fixed NN embedder for on-chip learning, yet To validate our framework’s performance on a resourcenone recognized this commonality explicitly or exploited it constrained edge device, we used the Chameleon SoC. In total, for hardware unification. Instead, prior works differentiated we considered four real-world use cases, one per learning themselves by how the embeddings are used and which scenario. First, when performing FSL on the Omniglot dataset, learning scenario is targeted. For example, SAPIENS [20] and we set new accuracy records across all shot-way scenarios, Kim et al. [21] apply L1 distance on top of embeddings for improving on the SotA by 0.5 to 16.6 accuracy points at 2.3× FSL, while FSL-HDnn [18] and Clo-HDnn [19] use HDC [47] lower power. Second, using CL under ECL, we provided the on embeddings for FSL and CL respectively. HDC expands first end-to-end, fully on-chip hardware results on the keyword embedding dimensionality to encode information, but this FSCIL task from NeuroBench. Compared to NeuroBench’s requires specialized hardware blocks [18], [19]; by contrast, original FP baseline, our model is 96× smaller, incurring a ECL uses dense embeddings that map directly onto standard loss of only 4.4 accuracy points after 200 classes at a real-time matrix-vector PE arrays, removing this overhead. Liu et al. [50] power of 9.5 µW. Third, we demonstrated the first end-totake a different approach and apply BP to update the last two end in-silico implementation of ZSL using semantic data. In FC layers of a fixed CNN embedder for ZSL. While each particular, we showed how ECL learned to classify five spoken
10
sentences from the FSC dataset using only their transcription at an accuracy of 60.6%. Learning consumes 8.2 µW, while realtime inference afterward consumes only 8.0 µW. Fourth, we also demonstrated for the first time end-to-end, fully on-chip ICL at a milliwatt power budget. On the RegBench dataset, ICL under ECL requires 16.8 µJ per token at a clock frequency of 100 MHz, showing ECL’s ability to adapt beyond classification. Together, these results establish ECL as a unified framework for versatile learning across scenarios at the edge, thereby enabling privacy-friendly, energy-efficient, and low-latency ondevice adaptation. ACKNOWLEDGEMENT The authors thank Prof. Makinwa for his feedback, Dr. Marco P. Apolinario for his detailed input and for our fruitful discussions, and Dr. Johannes von Oswald for his RegBench implementation. R EFERENCES [1] P. McEnroe, S. Wang, and M. Liyanage, “A survey on the convergence of edge computing and ai for uavs: Opportunities and challenges,” IEEE Internet Things J., vol. 9, no. 17, pp. 15 435–15 459, 2022. [2] J. S. P. Giraldo, S. Lauwereins, K. Badami, and M. Verhelst, “Vocell: A 65-nm speech-triggered wake-up soc for 10- µ w keyword spotting and speaker verification,” IEEE J. Solid-State Circuits, vol. 55, no. 4, pp. 868–878, 2020. [3] K. Kim, C. Gao, R. Graça, I. Kiselev, H.-J. Yoo, T. Delbruck, and S.-C. Liu, “A 23-uw keyword spotting ic with ring-oscillator-based time-domain feature extraction,” IEEE J. Solid-State Circuits, vol. 57, no. 11, pp. 3298–3311, 2022. [4] F. Tan, W.-H. Yu, J. Lin, K.-F. Un, R. P. Martins, and P.-I. Mak, “A 1.8% far, 2 ms decision latency, 1.73 nj/decision keywords-spotting (kws) chip incorporating transfer-computing speaker verification, hybrid-if-domain computing and scalable 5t-sram,” IEEE J. Solid-State Circuits, vol. 60, no. 3, pp. 1103–1112, 2025. [5] S. Park, K. Shin, D. Lee, M. Kang, S. Lee, Y. Park, M. Seok, and D. Jeon, “A 5.6µw 10-keyword end-to-end keyword spotting system using passive-averaging sar adc and sign-exponent-only layer fusion with 92.7% accuracy,” in 2024 IEEE Symp. VLSI Technol. Circuits (VLSI Technol. Circuits), 2024, pp. 1–2. [6] J. Liu, J. Fan, Z. Zhong, H. Qiu, J. Xiao, Y. Zhou, Z. Zhu, G. Dai, N. Wang, Q. Liu, et al., “An ultra-low power reconfigurable biomedical ai processor with adaptive learning for versatile wearable intelligent health monitoring,” IEEE Trans. Biomed. circuits Syst., vol. 17, no. 5, pp. 952–967, 2023. [7] L. Yan, J. Bae, S. Lee, T. Roh, K. Song, and H.-J. Yoo, “A 3.9 mw 25electrode reconfigured sensor for wearable cardiac monitoring system,” IEEE J. Solid-State Circuits, vol. 46, no. 1, pp. 353–364, 2011. [8] S. Yin, M. Kim, D. Kadetotad, Y. Liu, C. Bae, S. J. Kim, Y. Cao, and J.-S. Seo, “A 1.06- µ w smart ecg processor in 65-nm cmos for real-time biometric authentication and personal cardiac monitoring,” IEEE J. Solid-State Circuits, vol. 54, no. 8, pp. 2316–2326, 2019. [9] X. Zhang, Z. Zhang, Y. Li, C. Liu, Y. X. Guo, and Y. Lian, “A 2.89 µ w dry-electrode enabled clockless wireless ecg soc for wearable applications,” IEEE J. Solid-State Circuits, vol. 51, no. 10, pp. 2287– 2298, 2016. [10] J. Liu, J. Fan, Z. Zhong, H. Qiu, J. Xiao, Y. Zhou, Z. Zhu, G. Dai, N. Wang, Q. Liu, Y. Xie, H. Liu, L. Chang, and J. Zhou, “An UltraLow Power Reconfigurable Biomedical AI Processor With Adaptive Learning for Versatile Wearable Intelligent Health Monitoring,” IEEE Trans. Biomed. Circuits Syst., vol. 17, no. 5, pp. 952–967, Oct. 2023, Conference Name: IEEE Transactions on Biomedical Circuits and Systems, ISSN: 1940-9990. Accessed: Nov. 1, 2024. [11] C. Frenkel and G. Indiveri, “Reckon: A 28nm sub-mm2 task-agnostic spiking recurrent neural network processor enabling on-chip learning over second-long timescales,” in 2022 IEEE Int. Solid-State Circuits Conf. (ISSCC), vol. 65, 2022, pp. 1–3. [12] V. Jain, S. Giraldo, J. De Roose, L. Mei, B. Boons, and M. Verhelst, “Tinyvers: A tiny versatile system-on-chip with state-retentive emram for ml inference at the extreme edge,” IEEE J. Solid-State Circuits, vol. 58, no. 8, pp. 2360–2371, 2023.
[13] M. Verhelst and B. Moons, “Embedded deep neural network processing: Algorithmic and processor techniques bring deep learning to iot and edge devices,” IEEE Solid-State Circuits Mag., vol. 9, no. 4, pp. 55–65, 2017. [14] V. Sze, Y.-H. Chen, T.-J. Yang, and J. S. Emer, “Efficient processing of deep neural networks: A tutorial and survey,” Proc. IEEE, vol. 105, no. 12, pp. 2295–2329, 2017. [15] J. K. Kim, P. Knag, T. Chen, and Z. Zhang, “A 640m pixel/s 3.65 mw sparse event-driven neuromorphic object recognition processor with on-chip learning,” in 2015 Symp. VLSI Circuits (VLSI Circuits), IEEE, 2015, pp. C50–C51. [16] S. K. Gonugondla, M. Kang, and N. Shanbhag, “A 42pj/decision 3.12 tops/w robust in-memory machine learning classifier with on-chip training,” in 2018 IEEE Int. Solid-State Circuits Conference-(ISSCC), IEEE, 2018, pp. 490–492. [17] A. Amravati, S. B. Nasir, S. Thangadurai, I. Yoon, and A. Raychowdhury, “A 55nm time-domain mixed-signal neuromorphic accelerator with stochastic synapses and embedded reinforcement learning for autonomous micro-robots,” in 2018 IEEE Int. Solid-State Circuits Conference-(ISSCC), IEEE, 2018, pp. 124–126. [18] H. Yang, C. E. Song, W. Xu, B. Khaleghi, U. Mallappa, M. Shah, K. Fan, M. Kang, and T. Rosing, “Fsl-hdnn: A 5.7 tops/w end-toend few-shot learning classifier accelerator with feature extraction and hyperdimensional computing,” in 2024 IEEE Eur. Solid-State Electron. Res. Conf. (ESSERC), 2024, pp. 33–36. [19] C. E. Song, W. Xu, K. Fan, S. Jain, G. Hota, H. Yang, L. Liu, K. Akarvardar, M.-F. Chang, C. H. Diaz, G. Cauwenberghs, T. Rosing, and M. Kang, “Clo-hdnn: A 4.66 tflops/w and 3.78 tops/w continual on-device learning accelerator with energy-efficient hyperdimensional computing via progressive search,” in 2025 Symp. VLSI Technol. Circuits (VLSI Technol. Circuits), 2025, pp. 1–3. [20] H. Li, W.-C. Chen, A. Levy, C.-H. Wang, H. Wang, P.-H. Chen, W. Wan, H.-S. P. Wong, and P. Raina, “One-shot learning with memoryaugmented neural networks using a 64-kbit, 118 gops/w rram-based non-volatile associative memory,” in 2021 Symp. VLSI Technol., 2021, pp. 1–2. [21] S. Kim, W. Lee, S. Kim, S. Park, and D. Jeon, “An in-memory computing sram macro for memory-augmented neural network,” TCAS-II, vol. 69, no. 3, 2022. [22] G. K. Chen, R. Kumar, H. E. Sumbul, P. C. Knag, and R. K. Krishnamurthy, “A 4096-neuron 1m-synapse 3.8-pj/sop spiking neural network with on-chip stdp learning and sparse weights in 10-nm finfet cmos,” IEEE J. Solid-State Circuits, vol. 54, no. 4, pp. 992–1002, 2018. [23] M. P. E. Apolinario, K. Roy, and C. Frenkel, “Tess: A scalable temporally and spatially local learning rule for spiking neural networks,” in 2025 Int. Joint Conf. Neural Netw. (IJCNN), 2025, pp. 1–9. [24] T. Bohnstingl, S. Woźniak, A. Pantazi, and E. Eleftheriou, “Online spatio-temporal learning in deep neural networks,” IEEE Trans. Neural Netw. Learn. Syst., vol. 34, no. 11, pp. 8894–8908, 2023. [25] D. d. Blanken and C. Frenkel, “Chameleon: A multiplier-free temporal convolutional network accelerator for end-to-end few-shot and continual learning from sequential data,” IEEE J. Solid-State Circuits, pp. 1–16, 2026. [26] S. Davis and P. Mermelstein, “Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,” IEEE Trans. Acoust., speech, signal Process., vol. 28, no. 4, pp. 357–366, 1980. [27] B. M. Lake, R. Salakhutdinov, and J. B. Tenenbaum, “Human-level concept learning through probabilistic program induction,” Science, vol. 350, no. 6266, 2015. [28] J. Snell, K. Swersky, and R. Zemel, “Prototypical networks for few-shot learning,” in Advances Neural Inf. Process. Syst., vol. 30, 2017. [29] F. Sung, Y. Yang, L. Zhang, T. Xiang, P. H. Torr, and T. M. Hospedales, “Learning to compare: Relation network for few-shot learning,” in Proc. IEEE Conf. Comput. Vis. pattern Recognit., 2018, pp. 1199–1208. [30] M. Mazumder, C. Banbury, J. Meyer, P. Warden, and V. J. Reddi, “Few-shot keyword spotting in any language,” arXiv preprint arXiv:2104.01454, 2021. [31] O. Vinyals, C. Blundell, T. Lillicrap, D. Wierstra, et al., “Matching networks for one shot learning,” Advances neural Inf. Process. Syst., vol. 29, 2016. [32] T. Pfister, J. Charles, and A. Zisserman, “Domain-adaptive discriminative one-shot learning of gestures,” in Eur. Conf. Comput. Vis., Springer, 2014, pp. 814–829. [33] C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in Int. Conf. Mach. Learn., PMLR, 2017, pp. 1126–1135.
11
[34] B. Lake, R. Salakhutdinov, J. Gross, and J. Tenenbaum, “One shot learning of simple visual concepts,” in Proc. Annu. meeting Cogn. Sci. Soc., vol. 33, 2011. [35] G. Koch, R. Zemel, R. Salakhutdinov, et al., “Siamese neural networks for one-shot image recognition,” in ICML deep Learn. workshop, Lille, vol. 2, 2015, pp. 1–30. [36] L. Fei-Fei, R. Fergus, and P. Perona, “One-shot learning of object categories,” IEEE Trans. pattern analysis Mach. Intell., vol. 28, no. 4, pp. 594–611, 2006. [37] X. Tao, X. Hong, X. Chang, S. Dong, X. Wei, and Y. Gong, “Fewshot class-incremental learning,” in Proc. IEEE/CVF Conf. Comput. Vis. pattern Recognit., 2020, pp. 12 183–12 192. [38] J. Yik, K. Van den Berghe, D. den Blanken, Y. Bouhadjar, M. Fabre, P. Hueber, W. Ke, M. A. Khoei, D. Kleyko, N. Pacik-Nelson, et al., “The neurobench framework for benchmarking neuromorphic computing algorithms and systems,” Nature Commun., vol. 16, no. 1, p. 1545, 2025. [39] J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, et al., “Overcoming catastrophic forgetting in neural networks,” Proc. Nat. Acad. Sci., vol. 114, no. 13, pp. 3521–3526, 2017. [40] C. H. Lampert, H. Nickisch, and S. Harmeling, “Attribute-based classification for zero-shot visual object categorization,” IEEE Trans. pattern analysis Mach. Intell., vol. 36, no. 3, pp. 453–465, 2013. [41] R. Zhang, S. Frei, and P. L. Bartlett, “Trained transformers learn linear models in-context,” J. Mach. Learn. Res., vol. 25, no. 49, pp. 1–55, 2024. [42] S. Garg, D. Tsipras, P. S. Liang, and G. Valiant, “What can transformers learn in-context? a case study of simple function classes,” Advances neural Inf. Process. Syst., vol. 35, pp. 30 583–30 598, 2022. [43] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al., “Language models are few-shot learners,” Advances neural Inf. Process. Syst., vol. 33, pp. 1877–1901, 2020. [44] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances neural Inf. Process. Syst., vol. 30, 2017. [45] W. L. Tong and C. Pehlevan, “Mlps learn in-context on regression and classification tasks,” arXiv preprint arXiv:2405.15618, 2024. [46] S. Bai, J. Z. Kolter, and V. Koltun, “An empirical evaluation of generic convolutional and recurrent networks for sequence modeling”, 2018. arXiv:1803.01271. [47] P. Kanerva, “Hyperdimensional computing: An introduction to computing in distributed representation with high-dimensional random vectors,” Cogn. computation, vol. 1, no. 2, pp. 139–159, 2009. [48] W. Dai, C. Dai, S. Qu, J. Li, and S. Das, “Very deep convolutional neural networks for raw waveforms,” in 2017 IEEE Int. Conf. Acoust., speech signal Process. (ICASSP), IEEE, 2017, pp. 421–425. [49] L. Lugosch, M. Ravanelli, P. Ignoto, V. S. Tomar, and Y. Bengio, “Speech Model Pre-training for End-to-End Spoken Language Understanding”, arXiv:1904.03670 [eess], Jul. 2019. Accessed: Nov. 10, 2025. [50] J. Liu, X. Liu, X. Wang, Z. Xie, C. Guo, Z. Zhong, J. Fan, H. Qiu, Y. Xu, H. Qin, et al., “A high accuracy and ultra-energy-efficient zeroshot-retraining seizure detection processor,” IEEE J. Solid-State Circuits, vol. 59, no. 11, pp. 3549–3565, 2024. [51] E. Akyürek, B. Wang, Y. Kim, and J. Andreas, “In-context language learning: Architectures and algorithms,” arXiv preprint arXiv:2401.12973, 2024. [52] A. Chaudhry, M. Ranzato, M. Rohrbach, and M. Elhoseiny, “Efficient lifelong learning with a-gem,” arXiv preprint arXiv:1812.00420, 2018. [53] I. Skorokhodov and M. Elhoseiny, “Normalization matters in zero-shot learning,” 2020. [54] C. Gautam, S. Parameswaran, A. Mishra, and S. Sundaram, “Tfgczsl: Task-free generalized continual zero-shot learning,” Neural Netw., vol. 155, pp. 487–497, 2022. [55] K. Wei, C. Deng, X. Yang, et al., “Lifelong zero-shot learning.,” in IJCAI, 2020, pp. 551–557. [56] Y. D. Kwon, R. Li, S. I. Venieris, J. Chauhan, N. D. Lane, and C. Mascolo, “Tinytrain: Resource-aware task-adaptive sparse training of dnns at the data-scarce edge,” arXiv preprint arXiv:2307.09988, 2023. [57] L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch, “Decision transformer: Reinforcement learning via sequence modeling,” Advances neural Inf. Process. Syst., vol. 34, pp. 15 084–15 097, 2021. [58] M. Janner, Q. Li, and S. Levine, “Offline reinforcement learning as one big sequence modeling problem,” Advances neural Inf. Process. Syst., vol. 34, pp. 1273–1286, 2021.
[59] C. Frenkel, M. Lefebvre, and D. Bol, “Learning without feedback: Fixed random learning signals allow for feedforward training of deep neural networks,” Frontiers Neuroscience, vol. Volume 15 - 2021, 2021, ISSN: 1662-453X. [60] G. Hinton, “The forward-forward algorithm: Some preliminary investigations”, 2022. arXiv:2212.13345.
Douwe den Blanken (Graduate Student Member, IEEE) received the M.Sc. degree (with honors) in embedded systems from Delft University of Technology (TU Delft), Delft, The Netherlands, in 2023, where he is currently pursuing the Ph.D. degree, under the supervision of Prof. C. Frenkel. His current research interests include efficient learning algorithms and their implementation in silicon, as well as the quantization and acceleration of modern DNNs.
Martin Lefebvre (Member, IEEE) received the M.Sc. and Ph.D. degrees in engineering sciences from the Université catholique de Louvain (UCLouvain), Belgium, in 2017 and 2024. His research interests include hardware-aware machine learning algorithms, mixed-signal vision chips for embedded image processing, and low-power current reference architectures. He currently is a postdoctoral researcher in the cognitive sensor nodes and systems (CogSys) laboratory led by Prof. Frenkel at TU Delft, The Netherlands, working on neuromorphic hardware/software co-design for efficient on-chip learning. Dr. Lefebvre serves as a reviewer for various IEEE journals and conferences including IEEE Journal of Solid-State Circuits and IEEE Transactions on Circuits and Systems I and II.
Charlotte Frenkel (Member, IEEE) received the M.Sc. degree (summa cum laude) in Electromechanical Engineering and the Ph.D. degree in Engineering Science from Université catholique de Louvain (UCLouvain), Louvain-la-Neuve, Belgium in 2015 and 2020, respectively. In February 2020, she joined the Institute of Neuroinformatics, UZH and ETH Zurich, Switzerland, as a postdoctoral researcher. She is an Assistant Professor at Delft University of Technology, Delft, The Netherlands, since July 2022, and a Research Scientist at Google since February 2026. Her research aims at bridging the bottom-up (bio-inspired) and top-down (engineering-driven) design approaches toward neuromorphic intelligence, with a focus on hardware-algorithm co-design for (Neuro)AI, digital hardware accelerators, and brain-inspired on-device learning. Dr. Frenkel received a best paper award at the IEEE International Symposium on Circuits and Systems (ISCAS) 2020 conference in the Neural Networks track, and her Ph.D. thesis was awarded the FNRS-FWO / Nokia Bell Scientific Award 2021 and the FNRS-FWO / IBM Innovation Award 2021. In 2023, she was awarded prestigious Veni and AiNed Fellowship grants from the Dutch Research Council (NWO). She presented several invited talks, including keynotes at the tinyML EMEA technical forum 2021 and at the Neuro-Inspired Computational Elements (NICE) neuromorphic conference 2021. She serves or has served as a program co-chair of NICE 2023-2024 and of the tinyML Research Symposium 2024, as a TPC member of IEEE ISSCC for 2027 and IEEE ESSERC for 2022-2024, and as an associate editor for the IEEE Transactions on Biomedical Circuits and Systems for 2022-2025.