Conceptio › Archive › arXiv CS
arXiv CSopen access

Neural Cellular Automata Learn General Features in their Hidden Channels

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Neural Cellular Automata Learn General Features in their Hidden Channels

arXiv:2609.21870v1 [cs.LG] 18 Sep 2026

Etienne Guichard1 [0009-0005-5300-8182] and Stefano Nichele1 [0000-0003-4696-9872] Østfold University of Applied Sciences, BRA Veien 4, 1757 Halden, Norway {etienne.guichard, stefano.nichele}@hiof.no

Abstract. Modern deep learning models achieve impressive generalization through over-parameterization, but this paradigm often struggles with overfitting and memorization in few-shot regimes. Neural Cellular Automata (NCAs) offer a highly parameter-efficient alternative, yet research has focused primarily on their output, leaving the role of their internal hidden channels largely unexplored. In this paper, we investigate the internal dynamics of NCA hidden channels and introduce a novel transfer-learning mechanism that injects a pretrained teacher’s hidden states into a student model to guide early optimization. Evaluated on few-shot and scale-variant MNIST benchmarks, NCAs outperform comparable recurrent and feed-forward architectures, demonstrating superior generalization with a minimal parameter budget (≈ 9,800 parameters). Mechanistic analysis reveals that the hidden channels decouple feature extraction from uniform classification consensus by absorbing morphological complexity and converging to mutually orthogonal states. Furthermore, we demonstrate that these hidden channels capture general, scale-invariant topological primitives rather than class-specific templates. This allows a student model to achieve strong few-shot performance on unseen classes using features transferred from a teacher trained only on a subset of digits (0–5). Our results highlight the potential of utilizing hidden-state dynamics as a robust, decentralized computational substrate for parameter-efficient transfer learning. Keywords: Neural Cellular Automata · Few-Shot Learning · Transfer Learning.

1

Introduction

Modern deep learning relies heavily on over-parameterization to achieve state-ofthe-art performance across a myriad of domains [14, 1]. Empirical scaling laws suggest a power-law relationship in which continuous increases in parameter counts necessitate commensurate increases in both training data and computational resources [11, 10]. While this paradigm has yielded exceptional performance in complex tasks where massive, high-quality datasets are readily available [2], it presents a fundamental challenge for few-shot learning [24]. In strictly

2

E. Guichard, S. Nichele

data-limited environments, massively over-parameterized models are highly susceptible to overfitting, often tending to memorize the sparse training dataset rather than learning generalizable abstractions [3, 6]. Consequently, the community has increasingly sought parameter-efficient architectures and strong inductive biases capable of robust generalization without requiring exhaustive data or computational budgets [7, 19]. Neural Cellular Automata (NCAs) represent a promising alternative, a class of decentralized machine learning models that may offer a compelling solution to the generalization bottlenecks of few-shot learning. Inspired by the biological processes of cell morphogenesis and self-organization [21, 12, 15], NCAs operate by repeatedly applying highly localized, shared update rules across a spatial grid. Initially popularized by Mordvintsev et al. [13] for image generation and regeneration tasks, NCAs have recently demonstrated remarkable computational capabilities across various domains [20, 5, 25, 9]. Crucially, despite possessing drastically lower parameter counts—often magnitudes smaller than traditional convolutional or feed-forward networks—NCAs exhibit emergent, complex behaviors and strong robustness [22, 8]. Their strict adherence to local messagepassing acts as a structural inductive bias, effectively preventing the memorization of global spatial coordinates and forcing the network to learn robust, coordinate-free structural rules [18]. Despite these advantages, the majority of existing NCA research has predominantly focused on their external capabilities, evaluating models almost exclusively by the values of their output channels [17, 16]. Studies frequently assess an NCA’s ability to generate specific textures, segment images, or classify inputs based solely on the final convergence of these visible states [23]. In contrast, relatively few papers have rigorously investigated the internal dynamics of the NCA’s hidden channels [4]. These hidden channels—analogous to the internal chemical gradients of biological cells—act as a decentralized, dynamic computational substrate. Understanding how information is processed, spatially distributed, and geometrically represented within this hidden space remains an open challenge critical for interpreting the model’s emergent capabilities. Traditional transfer learning relies on the extraction and fine-tuning of pretrained synaptic weights [26, 27]. In this paper, we explore a different method exclusive to NCAs: transfer learning via hidden channels rather than weights. Because an NCA’s hidden channels store distributed, spatial representations of features across developmental steps, and are incorporated into its computational graph, we hypothesize that injecting a “teacher” NCA’s mature hidden state into a “student” NCA can serve as a form of bootstrapping. By bypassing the chaotic initial phases of state formation, a student network can immediately leverage well-structured topological primitives—such as edges and intersections—without requiring identical synaptic parameters or architecture, paving the way for highly efficient transfer.

Neural Cellular Automata Learn General Features in their Hidden Channels

2

Experimental Setup

2.1

Generalisation Experiments

3

The experimental setup for this experiment is as follows: we test an NCA against other models (mainly a Globally perceptive recurrent model, and two variants of a feed-forward model), all models have the same parameter count and are trained on k examples per class (on the MNIST training dataset) to see how well they generalize with few examples. Each model is trained for 3,000 iterations using the AdamW optimizer with a learning rate of 1e − 3 and is finally evaluated on the entire MNIST test set. The NCA model is evaluated with MSE loss, while the other models (producing logits) are evaluated using cross-entropy loss. To further test generalization, we also evaluate the model on scaled-down versions of the MNIST dataset, where MNIST digits are linearly scaled down (using a nearestneighbor downscaler) and padded back up to 28 × 28 images with background pixels. All experiments were conducted on 100 random seeds. We also set an 80% threshold (a reasonably good score) as the metric for the model’s ability to generalize, since with more data points, the models converge faster, with the idea being that at this lower score, we should be able to distinguish between models better. 2.2

Inductive Bias

We repeat the same experiment as in 2.1, with one caveat: during training, for the first 1500 steps of backpropagation, we inject the non-solution hidden state of an NCA pretrained on the entire MNIST dataset and has seen the k examples we are training on. By ’non-solution,’ we mean distinguishing the NCA’s different channels. Channel 0 is used to inject the MNIST digit, channels 1-10 are used for one-hot classification, and channels 11-15 are used for computation; these are the non-solution channels. From steps 1500-3000 of backpropagation, we remove the hidden-state injection; the NCA has to learn to reconstruct this hidden state on its own. We ensure that no gradient information from the "teacher" NCA is available to the "student" NCA by detaching the injected tensor from the computational graph before training the student. The injected hidden state is only injected at the student’s NCA developmental step T = 0, see Figure 1 for more details. Everything else is as before. During testing, the NCA has to reconstruct the hidden channels itself, without aid from the injection model. 2.3

Transfer learning

Here, we repeat the same experiment as in 2.2, but the injected state from the pretrained model is always present throughout training and testing, injected only at the student’s NCA developmental step T = 0. This effectively turns the NCA into a conditioned readout model, where it learns to identify the digit and the hidden state to produce the one-hot encoding.

4

E. Guichard, S. Nichele

Fig. 1: Diagram showing how the hidden channels are extracted from a teacher NCA and used in a Student NCA. The teacher NCA sees a digit, produces a classification, and some hidden channels are used in computing. These hidden channels are extracted and detached and passed to the student NCA alongside the digit before the forward pass in training.

3

Experimental Results

3.1

NCA Generalisation

Fig. 2: All 4 models, tested on k ∈ {1, 5, 10, 25, 50, 100} example per class, all have ≈ 9800 parameters. Since there are 10 classes, 101 examples equals one example per class. Scales range from 100 percent MNIST to 25 percent MNIST, with 25 percent intervals.

As shown in Figure 2, the NCA outperforms all other models in K-shot generalization, except for the mapping classic in 1-shot experiments. The model also crosses the 80% threshold considerably earlier than all other models at about 4 examples per class, as opposed to 10 (for mapping classic and Global Recurrent) and 11 (for Mapping). When evaluated on different scales, the NCA outperforms all models across all k-shot tests at 75% and 50%; at 25%, it is not statistically distinguishable from other models.

Neural Cellular Automata Learn General Features in their Hidden Channels

5

Fig. 3: All 4 models, tested on k ∈ {1, 5, 10, 25, 50, 100} example per class, all have ≈ 9800 parameters. Since there are 10 classes, 101 examples equals one example per class. Scales range from 100 percent MNIST to 25 percent MNIST, with 25 percent intervals. NCA receives a non-solution hidden state across a portion of training.

3.2

Inductive Bias

As shown in Figure 3, when allowed to see useful hidden channel representations across a portion of the K-shot learning regime, the NCA outperforms all models on all K-shot learning problems, crossing the 80% threshold somewhere between 3 and 4 examples. The NCA also outperforms all models in scaled-down evaluations, and surprisingly achieves a statistically significant (yet modest) result on even the 25% scale, something it was not capable of without this mechanism.

Fig. 4: Training loss (blue) and validation loss (orange) of one run of the NCA in a 1-shot environment trained with the hidden injections.

Figure 4 shows the training loss (blue) and validation loss (orange) of a single run of a student NCA in a one-shot environment. The dotted red line represents when the injections of the hidden state are no longer provided. As shown in

6

E. Guichard, S. Nichele

the figure, training loss quickly reaches a plateau (of near 0) while validation loss also stabilizes as the injection is provided. Once the injection is no longer provided, training loss briefly spikes, and validation loss drops considerably. This has mostly to do with the difference in training and testing before the injections are removed. During training, the NCA is already given well-constructed hidden channels, while during testing, it has to construct them itself, a task it has not been trained on yet. 3.3

Transfer Learning

Fig. 5: All 4 models, tested on k ∈ {1, 5, 10, 25, 50, 100} example per class, all have ≈ 9800 parameters. Since there are 10 classes, 101 examples equals one example per class. Scales range from 100 percent MNIST to 25 percent MNIST, with 25 percent intervals. NCA receives the non-solution hidden channels across all of training and validation.

As shown in Figure 5, when the student NCA is allowed to see the hidden channels throughout both training and testing, the performance on K-shot learning vastly outperforms all other models. 3.4

Reduced Knowledge Transfer Learning

Here we perform the same experiment as before, but the pretrained NCA only sees digits 0–5, while the Decoder NCA is trained on K-shot generalization on all MNIST digits. As shown in Figure 6, even when the teacher NCA is trained on a 5-way MNIST dataset, the features extracted are sufficient for the student model to achieve good generalization in 1-100-shot learning. The drop-off in evaluationscale performance is also similar in magnitude to that of a teacher NCA trained on the entire MNIST dataset.

4

Hidden Channel Analysis

While the experimental results in Section 3 demonstrate the NCA’s capability in few-shot generalization and transfer learning, they do not explain the underlying

Neural Cellular Automata Learn General Features in their Hidden Channels

7

Fig. 6: All 4 models, tested on k ∈ {1, 5, 10, 25, 50, 100} example per class, all have ≈ 9800 parameters. Since there are 10 classes, 101 examples equals one example per class. Scales range from 100 percent MNIST to 25 percent MNIST, with 25 percent intervals. NCA receives the non-solution hidden channels across all of training and validation.

mechanisms driving this performance. Unlike standard feed-forward networks, where intermediate activations represent static feature hierarchies, the hidden channels of an NCA (channels 11–15) act as a decentralized, dynamic computational substrate, capable of storing representations as well. In this section, we aim to probe the internal dynamics of these hidden channels to understand how NCAs leverage the development of their hidden states for classification. The models represented in this analysis are a fully trained teacher NCA (akin to an NCA from Section 3.1 trained on all of MNIST) and a student NCA (such as that in Section 3.3) trained using injected inputs from the teacher NCA on a 1-shot MNIST problem. 4.1

Scale Free Representation

As shown in Figures 2, 3, 5, and 6, the NCA model generally outperformed all other models on varying scales of the MNIST data set. Figure 7 hints at the underlying mechanism. As can be seen, the broad features the NCA builds in its hidden channel, such as topological intersections (channel 14) and edges (channel 12) are scale invariant. Because the NCA operates exclusively via local neighborhood communication, it is incapable of memorizing global spatial coordinates and features. Consequently, the hidden channels are forced to learn scale-invariant structures rather than pixel-level patterns. 4.2

Temporal Evolution of Spatial Variance

To quantify how the NCA separates feature extraction from classification, we tracked the mean spatial variance of the hidden channels versus the readout channels over time (Figure 8). We compute the mean spatial variance Vt (C) for a specific subset of channels (i) C at developmental step t. Let xt,c,h,w represent the activation value for the i-th batch sample at channel c and spatial coordinates (h, w). The temporal evolution of spatial variance is defined as:

8

E. Guichard, S. Nichele

Fig. 7: Top: Hidden channel representations built by the NCA on a full-size MNIST 4. Bottom: Hidden channel representations built by the same NCA on a 50% scale version of the same 4.

" # N H W 2 1 X X  (i) 1 XX (i) xt,c,h,w − µt,c Vt (C) = N |C| i=1 HW w=1 c∈C

(1)

h=1

where N is the batch size, H and W are the spatial dimensions of the grid, (i) and µt,c is the spatial mean of that specific channel map, given by: H

(i)

µt,c =

W

1 X X (i) xt,c,h,w HW w=1

(2)

h=1

In our analysis, we evaluate this metric separately for the hidden channels (Chidden = {11, . . . , 15}) and the readout channels (Creadout = {1, . . . , 10}). As the NCA processes an input, the hidden channels exhibit a rapid increase in spatial variance, corresponding to the complex morphological formation of features, followed by a plateau, corresponding to the stabilization of the features. Conversely, the classification readout channels maintain a near-zero spatial variance, reflecting a uniform grid-wide consensus. This decoupling demonstrates that the hidden channels actively absorb the task’s structural complexity, providing a stabilized representation that enables the readout channels to converge on a uniform classification. 4.3

Cross-Channel Cosine Similarity Over Time

Given the restricted parametric capacity of our models (≈ 9,800 parameters), we investigated how the NCA maximizes its representational efficiency. Figure 9 illustrates the cross-channel cosine similarity between all pairs of hidden channels over time. Crucially, the global average similarity across all pairs converges to approximately zero. This orthogonality confirms that the NCA eliminates representational redundancy, forcing distinct channels to adopt highly specialized, non-overlapping algorithmic roles (such as synergistic or inhibitory func-

Neural Cellular Automata Learn General Features in their Hidden Channels

(a) Teacher NCA

9

(b) Student NCA

Fig. 8: Temporal evolution of the spatial variance of the hidden channels of (a) the teacher NCA and (b) the student NCA.

(a) Teacher NCA

(b) Student NCA

Fig. 9: Cross-Channel Cosine Similarity between all channel pairs over time for (a) the teacher NCA and (b) the Student NCA.

tions). Furthermore, while the pre-trained Teacher model exhibits smooth, preorchestrated channel relationships from T = 0, the Student model exhibits a brief, high-variance transient phase. This suggests the Student actively translates and reorganizes the injected state to align with its specific parametric topology, rather than passively copying it. 4.4

Latent Trajectory of Hidden States

Finally, we visualize the dynamical attractors of the NCA by projecting the continuous temporal trajectory of a single digit’s hidden state into a 2D UMAP embedding (Figure 10). When evaluating the Student model from a randomstate initialization (red trajectory), the system must traverse a chaotic transient manifold before finally settling into its global attractor basin. In contrast, the state-injected Student (green trajectory) initiates its sequence at the exact terminus of the Teacher’s mature developmental trajectory. Because the Student possesses distinct weights, it smoothly glides down the gradient of its own loss landscape to its own attractor, entirely bypassing the high-variance phase of morphogenesis. This confirms that hidden-channel injection effectively acts as

10

E. Guichard, S. Nichele

Fig. 10: 2D UMAP trajectory of the hidden channels for the Teacher (blue), Student with injection (green), and Student with no injection (red).

developmental scaffolding, placing the model directly into a stable, mature representational basin. The fact that the student, even when starting from the teacher’s terminal state, follows its own trajectory shows that it is not acting merely as a readout. If this were the case, one would expect the student’s hidden channels to have no trajectory at all.

5

Discussion

This study is not intended to prove that NCAs are superior to other machine learning methods; it seeks to shed light on how these models differ and what advantages and disadvantages they entail. Traditionally, the hidden space of NCAs has been seen as a hidden computational storage space, primarily used to pass intermediate messages between local steps until a classification consensus is reached. Results from Section 3.1 and Section 3.3 indicate that NCAs may have strong generalization capabilities, particularly in few-shot regimes. Unlike standard feed-forward networks, which often rely on high parameter capacity to fit complex datasets, the recurrent, local nature of NCAs forces them to capture structural patterns with a highly restricted parameter budget. As shown in Section 4.1, these representations are scale invariant, helping the NCA generalize further on unseeded data. Both the inductive bias and transfer learning experiments show that the hidden channels are not merely useful information storage spaces but also core computational elements that can be exploited. Rather than serving as intermediate information storage channels, they function more as spatially distributed feature maps that the output cells utilize during classification. When we inject a pre-trained teacher’s hidden state, we provide the student model with precomputed feature maps. This bootstrapping allows the student to bypass the chaotic initial phases of state formation and exploit features more readily in

Neural Cellular Automata Learn General Features in their Hidden Channels

11

few-shot regimes, as visualized in the latent trajectory analysis (Section 4.4). We also theorize that the bootstrapping mechanism provides the student with more orthogonal features to learn from, even in few-shot regimes, partially explaining why its performance increases across the entire dataset after seeing only one example. Our analysis of the internal dynamics, however, reveals a more nuanced division of labor. As shown in the spatial variance analysis (Section 4.2), the hidden channels capture the complex morphological structure of features, while the readout channels maintain near-zero spatial variance to preserve grid-wide consensus. Furthermore, the near-zero average cosine similarity across channels suggests that the NCA minimizes representational redundancy, forcing distinct channels to adopt highly specialized, orthogonal algorithmic roles. Perhaps the most notable observation comes from the reduced-knowledgetransfer experiments. The teacher’s ability to extract general features from a subset of classes (digits 0–5) that are immediately useful for a student learning unseen classes (digits 6–9) suggests that the hidden channels are capturing fundamental, scale-invariant topological primitives rather than class-specific templates. Because the NCA operates exclusively via local neighborhood communication, it is unable to memorize global spatial coordinates, forcing the network to rely on these robust geometric primitives instead. Despite these promising qualities, several limitations must be acknowledged: – Dataset Complexity: These experiments were conducted on variants of the MNIST dataset. While MNIST serves as a valuable benchmark for topological and geometric reasoning, it is structurally simple and monochromatic. It remains to be seen whether these scale-invariant features and state-transfer mechanisms scale effectively to high-resolution, complex color datasets (such as CIFAR-10 or ImageNet). – Computational Overhead: Although the NCA achieves high representational efficiency with a low parameter count (≈9,800 parameters), simulating dozens of developmental steps during both training and inference introduces considerable computational and time overhead. Backpropagation through time (BPTT) over many steps remains a scaling bottleneck compared to standard single-pass feed-forward architectures. This also means the comparisons to other models (with the exception of the global recurrent model) do not fully reflect the capability differences, as the recurrent models are afforded a much higher computational budget.

6

Conclusion

In this study, we investigated the internal representation and transferability of hidden channels in Neural Cellular Automata (NCAs) within data-scarce and scale-variant environments. While traditional research has viewed the hidden space of NCAs merely as intermediate message-passing channels, our work demonstrates a highly coordinated division of labor. Through spatial variance and cross-channel cosine similarity analyses, we showed that the hidden channels

12

E. Guichard, S. Nichele

actively absorb morphological complexity, allowing the readout channels to converge uniformly on a classification consensus. Crucially, these hidden channels strictly minimize representational redundancy, dynamically self-organizing into mutually orthogonal algorithmic roles over time. Our experiments on bootstrapping demonstrate that transfer learning can be successfully achieved by injecting pre-computed hidden states rather than solely relying on standard synaptic weight adjustment. This mechanism allows a student model to bypass chaotic initial morphogenesis phases and initialize within a stable, mature attractor basin. Furthermore, our reduced-knowledge transfer experiments confirm that the hidden channels capture scale-invariant topological primitives rather than class-specific templates, enabling robust generalization on unseen classes. Despite these promising results, several limitations must be acknowledged. Our evaluation was confined to structurally simple, monochromatic MNIST variants, and the scaling of backpropagation through time (BPTT) over many developmental steps introduces non-trivial computational and time overhead. Future work will focus on scaling these dynamics to higher-resolution color datasets, investigating the behavior of larger multi-channel hidden spaces, and systematically reverse-engineering the orthogonal roles of individual channels to improve the mechanistic interpretability of decentralized learning architectures.

Acknowledgment This work was supported by MishMash - Research Council of Norway, GrantId: 357438.

References [1]

[2] [3]

[4]

[5]

Mikhail Belkin et al. “Reconciling modern machine-learning practice and the classical bias–variance trade-off”. In: Proceedings of the National Academy of Sciences 116.32 (2019), pp. 15849–15854. Tom Brown et al. “Language models are few-shot learners”. In: Advances in Neural Information Processing Systems. Vol. 33. 2020, pp. 1877–1901. Nicholas Carlini et al. “Extracting training data from large language models”. In: 30th USENIX Security Symposium (USENIX Security 21). 2021, pp. 2633–2650. Hugo Cisneros, Josef Sivic, and Tomas Mikolov. “Visualizing computation in large-scale cellular automata”. In: The 2020 Conference on Artificial Life. ALIFE 2020. MIT Press, 2020, pp. 239–247. doi: 10.1162/isal_a_ 00277. url: http://dx.doi.org/10.1162/isal_a_00277. Sam Earle et al. “Illuminating diverse neural cellular automata for level generation”. In: Proceedings of the Genetic and Evolutionary Computation Conference (GECCO). 2022, pp. 68–76. url: https://arxiv.org/abs/2109. 05489.

Neural Cellular Automata Learn General Features in their Hidden Channels

[6]

[7]

[8] [9]

[10] [11] [12]

[13] [14]

[15]

[16] [17] [18] [19]

[20]

[21]

13

Vitaly Feldman. “Does learning require memorization? a short tale about a long tail”. In: Proceedings of the 52nd Annual ACM SIGACT Symposium on Theory of Computing (STOC). 2020, pp. 954–959. url: https://arxiv. org/abs/1906.05271. Chelsea Finn, Pieter Abbeel, and Sergey Levine. “Model-agnostic metalearning for fast adaptation of deep networks”. In: International conference on machine learning. PMLR. 2017, pp. 1126–1135. William Gilpin. “Cellular automata as convolutional neural networks”. In: Physical Review E 100.3 (2019), p. 032402. Etienne Guichard et al. ARC-NCA: Towards Developmental Solutions to the Abstraction and Reasoning Corpus. 2025. arXiv: 2505.08778 [cs.AI]. url: https://arxiv.org/abs/2505.08778. Jordan Hoffmann et al. “Training compute-optimal large language models”. In: arXiv preprint arXiv:2203.15556 (2022). Jared Kaplan et al. “Scaling laws for neural language models”. In: arXiv preprint arXiv:2001.08361 (2020). Michael Levin. “Morphogenetic fields in embryogenesis, regeneration, and cancer: non-local control of complex patterning”. In: Biosystems 109.3 (2012), pp. 243–261. Alexander Mordvintsev et al. “Growing neural cellular automata”. In: Distill 5.2 (2020), e23. Preetum Nakkiran et al. “Deep Double Descent: Where Bigger Models and More Data Hurt”. In: International Conference on Learning Representations (ICLR). 2020. url: https://arxiv.org/abs/1912.02292. Stefano Nichele and Andreas Molund. “Deep learning with cellular automatonbased reservoir computing”. In: Complex Systems 26.4 (2017), pp. 319–340. url: https://www.complex-systems.com/abstracts/v26_i04_a03/. Eyvind Niklasson et al. “Self-Organising Textures”. In: Distill 6.2 (2021), e00027–003. url: https://distill.pub/2021/selforg-textures/. Rasmus Berg Palm et al. Variational Neural Cellular Automata. 2022. arXiv: 2201.12360 [cs.NE]. url: https://arxiv.org/abs/2201.12360. Ettore Randazzo et al. “Self-classifying MNIST digits with neural cellular automata”. In: Distill 5.8 (2020), e27. Jake Snell, Kevin Swersky, and Richard Zemel. “Prototypical networks for few-shot learning”. In: Advances in neural information processing systems. Vol. 30. 2017. Shyam Sudhakaran et al. “Growing 3D Artefacts and Functional Machines with Neural Cellular Automata”. In: Proceedings of the 2021 Conference on Artificial Life (ALIFE 2021). 2021. url: https://arxiv.org/abs/2103. 08737. Alan Mathison Turing. “The chemical basis of morphogenesis”. In: Philosophical Transactions of the Royal Society of London. Series B, Biological Sciences 237.641 (1952), pp. 37–72.

14

[22]

[23]

[24] [25]

[26] [27]

E. Guichard, S. Nichele

Alexandre Variengien, Elias Najarro, and Sebastian Risi. “Towards robust and generalizable representations in neural cellular automata”. In: arXiv preprint arXiv:2111.00287 (2021). Kathryn Walker et al. “Physical Neural Cellular Automata for 2D Shape Classification”. In: Proceedings of the Artificial Life Conference (ALIFE). 2022. url: https://arxiv.org/abs/2203.07548. Yaqing Wang et al. “Generalizing from a few examples: A survey on fewshot learning”. In: ACM computing surveys (csur) 53.3 (2020), pp. 1–34. Kevin Xu and Risto Miikkulainen. Neural Cellular Automata for ARCAGI. 2025. arXiv: 2506.15746 [cs.NE]. url: https://arxiv.org/abs/2506. 15746. Jason Yosinski et al. “How transferable are features in deep neural networks?” In: NeurIPS 27 (2014). url: https://arxiv.org/abs/1411.1792. Fuzhen Zhuang et al. A Comprehensive Survey on Transfer Learning. 2020. arXiv: 1911.02685 [cs.LG]. url: https://arxiv.org/abs/1911.02685.

Record · ID 1006857 · SHA-256 a07bd37f4d0ab35e
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.