ConceptioArchivearXiv CS
arXiv CSopen access

Intrinsic Motivation in Reinforcement Learning: A Research Agenda for Adaptive Self-Organisation

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

Intrinsic Motivation in Reinforcement Learning: A Research Agenda for Adaptive Self-Organisation Anatoly Belikova

arXiv:2609.17325v1 [cs.AI] 15 Sep 2026

a SingularityNET Foundation,

Abstract Biological cells can be viewed as individual, interacting agents whose collective dynamics give rise to adaptive behaviour at multiple levels of organisation, from individual cells through tissues to whole multicellular organisms. In this perspective and tutorial article we discuss whether intrinsic rewards in artificial neural systems can support adaptation, functional specialisation and higher-level selforganisation without a shared external objective. We review empowerment, curiosity, learning progress, information gain, unsupervised skill discovery, mutual information estimation and the use of world models for intrinsic reward computation. Particular attention is given to failure modes showing when such objectives do not produce sustained exploration or increasingly complex behaviour. We argue that more capable systems may require complementary objectives, communication, memory, learning at multiple temporal scales and environmental constraints. Based on this perspective, we outline three experimental directions. These include a resource-constrained environment in which otherwise stable behavioural attractors become unsustainable, allowing us to test whether environmental constraints can mitigate characteristic failure modes of intrinsic objectives. The network of recurrent agents with per-agent intrinsic rewards, and a hierarchical world-model agent in which exploratory motor competence develops before goal-directed behaviour. These experiments are intended to test whether intrinsic learning can lead to adaptive organisation at progressively higher levels. Keywords: intrinsic motivation, mutual information, empowerment, curiosity, unsupervised skill discovery, world models, hierarchical reinforcement learning

Email address: [email protected] (Anatoly Belikov)

1. Introduction In this article, we consider intrinsic motivation in reinforcement learning. RL methods are well developed both theoretically and at the implementation level. They are highly effective when provided with a dense reward function. Some progress has been made in defining universal intrinsic reward functions that are applicable across different environments. Here, intrinsic means that the reward is computed only from what an agent observes during its lifetime, without using hidden properties of the environment. This is especially interesting given that nature does not define a reward function for biological organisms; instead, evolution must have led to the development of some kind of intrinsic preferences. Existing intrinsic objectives exhibit characteristic failure modes, but some of them are naturally complementary and may mitigate one another’s limitations. We also analyse environmental constraints, such as limited energy, and structural features of biological systems, such as the existence of relatively global communication channels, that could be introduced into artificial environments to support the development and self-organisation of intrinsically motivated agents. 1.1. Article structure and contribution: We propose a research direction focused on adaptability in artificial neural networks, with the goal of developing learning systems that could demonstrate adaptive behaviour similar to that observed in biological networks. We ask whether some of their functional properties—adaptation, memory, communication, functional differentiation and the formation of higher-level organisation—can emerge in non-biological recurrent neural networks. A broader question motivating the research agenda is whether such a learning system can produce progressively more complex adaptive behaviour and organisation at higher levels of organisation, similar to the formation of multicellular organisms. The main contribution of the article is a comparison of what existing objectives optimise, an analysis of their failure modes, and a research agenda for studying how complementary per-agent objectives may produce higher-level adaptation. We also propose a biologically inspired recurrent architecture and a specific environment suitable for evaluating intrinsic motivation methods. We begin with motivation from biology. The next few sections provide a tutorial-style overview and comparison of empowerment, DIAYN (skill discovery), learning progress and information gain.

2

Figure 1: Kidney tubules in the newt are made with a constant size, whereas cell size can vary drastically under polyploidy. The same shape achieved through different molecular mechanisms: cell to cell communication vs cytoskeleton bending. Adapted from [1].

The tutorial part, together with the appendix, contains derivations of the main identities and discusses interesting corner cases, including an analysis of why meaningful intrinsic rewards may fail to produce useful behaviour. We provide an overview of some methods of world modelling and their possible use for intrinsic reward computation, and conclude with a biologically motivated research agenda: experiments intended to improve our understanding of the intrinsic motivation methods discussed. This article is inspired by the work of Jürgen Schmidhuber, who developed the first formal computational methods for intrinsic motivation and artificial curiosity in reinforcement learning, as well as by Michael Levin’s work on adaptive behaviour in biological systems across different scales. 2. Motivation from biology AI as a field often draws inspiration from biology. Recent discoveries point that there is at least some intelligence on different levels of organization. Molecular networks have memory including pavlovian conditioning, associative learning etc. Also there is no large difference between neural cells and somatic cells. Somatic cells exhibit learning, problem solving, they also use electricity for communication. Some notable examples of plasticity and problem solving in image 2

3

Figure 2: Classical conditioning in biological network, for example drug-drug conditioning. Adapted from Figure 1 in [2].

4

Figure 3: Adaptation of neurons in planaria to barium chloride. After exposure to BaCL2 neurons in planaria’s head die off. New head has resistant neurons. It’s unlikely that planaria has ever been exposed to BaCL2 before. Adapted from Figure 5 in [3].

5

But perhaps the most striking and important case is cancer plasticity. Tumors switch between metabolic pathways (e.g., glycolysis to oxidative phosphorylation) to survive fluctuating conditions, a form of tissue-level adaptability. For more detailed overviews, see Michael Levin’s interview Michael Levin: Intelligence Beyond the Brain and his presentation Bioelectricity: A Bridge between Physics and Cognition, by Way of Biology [4, 5]. evolutionary trend hypothesis There appears to be a trend in growing intelligence across biosphere. Both - upper cap and average(measured by total biomass). For insects it’s estimated that eusocial species make now around 50% [6] of insects biomass while paradoxically constituting only about 2% of species. Modern colony size are relatively recent invention [7] and there is growing evidence that colony size is a primary drive of specialisation [8]. This might appear obvious in hindsight given growing specialisation of people in modern economy. This transition provides evidence that evolution may increase not only the upper bound of intelligence, but also the weighted by biomass prevalence of complex adaptive organization. ∑ 𝐵 (𝑡)𝐶 (𝑡) 𝐶𝑏𝑎𝑟 (𝑡) = 𝑖∑ 𝑖 𝐵 (𝑡)𝑖 𝑖 𝑖 here 𝐶 is intelligence measure 𝐵 - biomass i - cells, organisms or colonies. Possible mechanism is random specialization of some species in intelligencebased adaptation with later evolutionary arms race. 2.1. Learning hierarchy We can very roughly sort different learning types from more simple to more advanced forms, that likely appeared later in evolutionary history. 1. Non-associative learning Habituation - decrease of response to non-harmful repeated stimulus Sensitization - increase of response after exposure to strong or harmful stimulus. 2. Associative learning Classical conditioning: learning of stimulus-outcome association . Operant Conditioning: learning of action-outcome association. 6

3. Flexible individual ans socially mediated learning Metacognition. This includes capacity to model agent’s own understanding e.g. being uncertain, seek more information, being able to assess own likelihood of error. Spatial Learning & Navigation Insightful Problem Solving, probably based on some internal modelling as opposed to trial and error. Teaching & Pedagogy Play 4. Cumulative Culture: Cultural learning over generations, learning from purely symbolic input e.g. reading instructions. There is no universally accepted hierarchy covering all forms of learning and cognition. So this classification is rather a heuristic. But there is a good evidence that non-associative mechanisms are evolutionarily ancient, whereas flexible planning, metacognitive control, teaching, and cumulative culture appeared much later. Level 2 and Level 3 learning types can be observed already in insects. Bumblebee for example would play with appropriately sized balls without any external reward from only intrinsic motivation [10]. Another example is Portia spiders which employ a sophisticated hunting strategy involving mimicry, detours, and ambush tactics when hunting on other spiders. Portia often attacks other spiders in their own web, so it has to be clever. Sometimes it will literally lure the prey by mimicking vibrations of an insect being stuck. It has been reported that Portia can learn to hunt a new, never seen before spider by trial and error. Portia is a very good example since it has a very small brain. Bumblebees have approximately 950k - 1 mln neurons vs Portia around 100k. From observations we can conclude that it has: Internal Representation: The ability to take a detour where the prey is out of sight for extended periods implies the spider is not just reacting to the prey's presence. It must be operating from a stored representation of the environment and the prey's location within it. Object Permanence and Spatial Memory: Portia's behavior indicates it has a grasp of object permanence (knowing the prey still exists even when it can't be seen) and strong spatial memory to navigate the planned route. 7

Figure 4: Xenobots - small, self-organizing robots made from frog embryonic cells (Xenopus laevis). This one is made of skin and cardiac cells. Image adapted from [9], Wikimedia Commons, licensed under CC BY 4.0.

8

Systematic Scanning: The process of systematically scanning its surroundings before a hunt is interpreted as the spider building a detailed mental map of the area, which it then uses to execute its plan. Expectancy Violation: When a spider takes a detour and finds wrong number of pray it spent more time inspecting the scene. This implies it had an expectation of what it should find, based on its internal model, and that expectation was violated. Also there is growing evidence about tool use by insects. Image 5 shows the experiment: researchers gave hungry ants containers with sugar water. Researchers then altered the surface tension of the sugar water by adding surfactant. When surfactant concentrations were over 0.05%, representing considerable drowning risk, ants were observed building the sand structures to syphon sugar water out of the container. These structures were never observed when ants foraged in containers of pure sugar water, indicating an adaptable approach to this novel tool use. 1. Biological organisms are adaptive on different levels of organization 2. Learn without backpropagation 3. All cells types can communicate with each other 4. Demonstrate goal-directed behaviour Current neural architectures are more fragile than biological ones. Our motivation is to determine if current intrinsic motivation methods would lead to learning and behaviour types observable in insects. Whether these types of behaviours are reproducible with intrinsic rewards. Could intrinsic motivation in RL lead to the development of level 2 and level 3 learning types: 1) Does it lead to emergent communication and cooperation? 2) Lead to observational learning? 3) Does it lead to play and episodic learning?

9

Figure 5: Examples of learned behaviour, play, complex hunting strategy and tool use in arthropods. The trained bumblebee ball-rolling image is a generated illustration of the experiment reported by Loukola et al. [11]. The tool-selection image is a frame from the supplementary video by Chow et al. [12], licensed under CC BY 4.0. The spontaneous bumblebee ball-rolling image is adapted from Galpayage Dona et al. [10] under CC BY 4.0. The Portia image, “Male Portia waiting for opportunity to predate buttonspider,” is by i_c_riddell under CC BY. The ant tool-use image is a generated illustration of the experiment reported by Zhou et al. [13]. Each image links to its source or related publication.

10

3. Basic mathematical definitions Entropy For a discrete random variable 𝑌 with probability mass function 𝑝(𝑦) we define entropy H: 𝐻(𝑌 ) = − ∑ 𝑝(𝑦) log 𝑝(𝑦).

(1)

𝑦

𝐻(𝑌 |𝑍) = − ∑ 𝑝(𝑦, 𝑧) log 𝑝(𝑦|𝑧).

(2)

𝑦,𝑧

Kullback–Leibler divergence For two probability distributions 𝑃 and 𝑄 over the same sample space, 𝐷KL (𝑃‖𝑄) = ∑ 𝑝(𝑥) log 𝑥

𝑝(𝑥) = 𝔼𝑥∼𝑃 [log 𝑝(𝑥) − log 𝑞(𝑥)] . 𝑞(𝑥)

(3)

Mutual information Mutual information of two random variables Y, Z 𝐼(𝑌 ; 𝑍) ∶= 𝐷𝐾𝐿 (𝑃(𝑌 , 𝑍) ‖ 𝑃(𝑌 )𝑃(𝑍))

(4)

Mutual information defined via entropy 𝐼(𝑌 ; 𝑍) ∶= 𝐻(𝑌 ) − 𝐻(𝑌 ∣ 𝑍) = 𝐻(𝑍) − 𝐻(𝑍 ∣ 𝑌 )

(5)

𝑝(𝑥,𝑦) log 𝑝(𝑥)𝑝(𝑦) , where x and y are

Pointwise mutual information PMI(𝑥, 𝑦) = events, not random variables. It’s possible for PMI to be negative, but it’s expectation is always non-negative. 𝐼(𝑋; 𝑌 ) = 𝔼(𝑥,𝑦)∼𝑝(𝑥,𝑦) [PMI(𝑥, 𝑦)] 3.1. Reinforcement learning Reinforcement learning objective - maximising expected discounted reward 𝐽(𝜋) ∶= 𝔼𝑝𝜋 (𝜏) [∑ 𝛾𝑡 𝑟𝑡 ] 𝑡

𝜋 - policy. 𝛾 - discount factor in range (0, 1] 𝑠𝑡 - state at time t. 𝑝𝜋 (𝜏) - trajectory distribution under policy. For a concise introduction to policy-gradient methods, see [14]. 11

(6)

4. Empowerment in reinforcement learning Mutual information between agent’s action and observable states can be used as reward. In this section, we derive such a reward and discuss its relation to empowerment. Empowerment is defined as channel capacity from agent’s action to subsequent state [15]: ℰ(𝑠) = max 𝐼(𝐴; 𝑆 ′ ∣ 𝑆 = 𝑠). 𝑝(𝑎∣𝑠)

(7)

We can optimise policy to increase 𝐼𝜋 (𝐴; 𝑆 ′ |𝑠); 𝐼𝜋 (𝐴; 𝑆 ′ |𝑠) ≤ ℰ(𝑠) gives us lower-bound estimation of empowerment. We will refer to this quantity as on-policy empowerment. In order to gain some intuition we will consider different formulations of the same quantity: 𝐼(𝐴; 𝑆 ′ ∣ 𝑠) = 𝐷𝐾𝐿 (𝑝(𝑎, 𝑠′ ∣ 𝑠) ‖ 𝑝(𝑎 ∣ 𝑠)𝑝(𝑠′ ∣ 𝑠)).

(8)

𝐼(𝐴; 𝑆 ′ ∣ 𝑠) = 𝐻(𝑆 ′ ∣ 𝑠) − 𝐻(𝑆 ′ ∣ 𝐴, 𝑠).

(9)

𝐼(𝐴; 𝑆 ′ ∣ 𝑠) = 𝐻(𝐴 ∣ 𝑠) − 𝐻(𝐴 ∣ 𝑆 ′ , 𝑠).

(10)

We can rewrite 𝐼(𝐴; 𝑆 ′ ∣ 𝑠) as expectation over events: First expand 𝐻(𝑆 ′ ∣ 𝑠): 𝐻(𝑆 ′ ∣ 𝑠) = − ∑ 𝑝(𝑠′ ∣ 𝑠) log 𝑝(𝑠′ ∣ 𝑠).

(11)

𝑠′

Since 𝑝(𝑠′ ∣ 𝑠) = ∑𝑎 𝑝(𝑎, 𝑠′ ∣ 𝑠), 𝐻(𝑆 ′ ∣ 𝑠) = − ∑ 𝑝(𝑎, 𝑠′ ∣ 𝑠) log 𝑝(𝑠′ ∣ 𝑠) = −𝔼𝑝(𝑎,𝑠′ ∣𝑠) log 𝑝(𝑠′ ∣ 𝑠).

(12)

𝑠′ ,𝑎

Second term 𝐻(𝑆 ′ ∣ 𝐴, 𝑠) 𝐻(𝑆 ′ ∣ 𝐴, 𝑠) = − ∑ 𝑝(𝑎, 𝑠′ ∣ 𝑠) log 𝑝(𝑠′ ∣ 𝑎, 𝑠) = −𝔼𝑝(𝑎,𝑠′ ∣𝑠) log 𝑝(𝑠′ ∣ 𝑎, 𝑠). (13) 𝑎,𝑠′

Substituting Eqs. (12) and (13) into Eq. (9) gives 12

𝐼(𝐴; 𝑆 ′ ∣ 𝑠) = −𝔼𝑝(𝑎,𝑠′ ∣𝑠) log 𝑝(𝑠′ ∣ 𝑠) + 𝔼𝑝(𝑎,𝑠′ ∣𝑠) log 𝑝(𝑠′ ∣ 𝑎, 𝑠) = 𝔼𝑝(𝑎,𝑠′ ∣𝑠) [log 𝑝(𝑠′ ∣ 𝑎, 𝑠) − log 𝑝(𝑠′ ∣ 𝑠)] .

(14)

Here, s is the current state, 𝑆 ′ is the next state and 𝐴 is the action taken in state s. 𝐴 and 𝑆 ′ are random variables but 𝑠 is not. It’s important to note that removing conditioning on the current state 𝑠, as in 𝐼(𝐴; 𝑆 ′ ) or 𝐼(𝐴; 𝑆) will lead to a different, possibly degenerate solution. Here 𝐼(𝐴; 𝑆 ′ ) is action - future state mutual information and 𝐼(𝐴; 𝑆) - action current state. Maximising 𝐻(𝑆 ′ ) − 𝐻(𝑆 ′ |𝐴) means that future state 𝑆 ′ should be predictable from action alone, and 𝐼(𝐴; 𝑆) = 𝐻(𝑆) − 𝐻(𝐴 ∣ 𝑆) means the agent should directly map current state to action. For example always choose turn left in one room and turn right in an another room. In this case 𝐼(𝐴; 𝑆) will be high, but 𝐼(𝐴; 𝑆 ′ |𝑠) close to zero. What is the behaviour that is encouraged by this type of reward? First term that is maximised 𝐻(𝑆 ′ ∣ 𝑠) is conditional entropy of future state 𝑆 ′ . Entropy is low for spiky, concentrated probability distributions and high for more even distributions. A binary random variable with probabilities (0.5,0.5) has an entropy of one bit. A variable with 10 possible outcomes with probability 𝑝(𝑠𝑖 ) = 0.1 entropy is ~3.32 bits. So 𝐻(𝑆 ′ ∣ 𝑠) is high then our agent visits a large and diverse set of states. To be more precise it means for given state 𝑠 policy could achieve diverse set of future states. Global state visitation entropy 𝐻(𝑆 ′ ) is related to conditional as 𝐻(𝑆 ′ ) = 𝐻(𝑆 ′ ∣ 𝑆) + 𝐼(𝑆 ′ ; 𝑆), so 𝐻(𝑆 ′ ) is at least as large as 𝐻(𝑆 ′ ∣ 𝑆). What it means in practice agent could stay in an relatively small area with a lot of achievable states it could switch between. Second term is minimised: 𝐻(𝑆 ′ ∣ 𝐴, 𝑠), this is the entropy of the future state given the current state and action. This term encourages policy to take actions that have predictable outcomes. In a deterministic environment we have 𝐻(𝑆 ′ ∣ 𝐴, 𝑠) = 0, so only the first term is optimised. Consider this grid world, with actions up, down, left, right:

13

3

Being in the central cell will have highest value of 𝐻(𝑆 ′ ∣ 𝑠), − ∑𝑖=0 0.25 𝑙𝑜𝑔2 0.25 = 2 bits. We can formalise empowerment use as intrinsic reward: ∑𝑡 𝛾𝑡 PMI(𝑠𝑡+1 , 𝑎| 𝑠𝑡 ), following the objective in Eq. (6). Since MI is an expected value of PMI it’s close to original definition: 𝔼 [𝑟𝑡PMI |𝑆𝑡 = 𝑠] = 𝐼𝜋 (𝐴; 𝑆 ′ |𝑆 = 𝑠) ≤ ℰ(𝑠) Technical note: Empowerment is a property of a state, environment dynamics and the action space, not of our policy. It’s already defined as maximum capacity and we can’t literally optimise it. A behavioural policy may be trained both to approach this capacity and to visit states in which the capacity is high. We use the term ”empowerment maximisation” for such cases. 4.1. Inverse model trick We can train a model to predict the action that caused a transition from state 𝑠𝑡 to state 𝑠𝑡+1 and then use it as empowerment estimator. Here is how it works: 14

We can then estimate empowerment with equation 10. 𝐼(𝐴; 𝑆 ′ ∣ 𝑠) = 𝐻(𝐴 ∣ 𝑠) − 𝐻(𝐴 ∣ 𝑆 ′ , 𝑠) By definition we have for entropies: 𝐻(𝐴 ∣ 𝑠) = −𝔼𝑎∼𝜋(⋅|𝑠) log 𝜋(𝑎 ∣ 𝑠).

(15)

𝐻(𝐴 ∣ 𝑆 ′ , 𝑠) = −𝔼 𝑎∼𝜋(⋅|𝑠) log 𝑝(𝑎 ∣ 𝑆 ′ , 𝑠).

(16)

𝑠′ ∼𝑝(⋅|𝑎,𝑠)

The distribution 𝑝(𝑎 ∣ 𝑠) is known. It is the action distribution under the current policy 𝜋(𝑎|𝑠). Distribution of actions given previous and future states 𝑝(𝑎 ∣ 𝑠′ , 𝑠) is not known, but we can approximate it with additional distribution 𝑞(𝑎 ∣ 𝑠′ , 𝑠). First we plug it in conditional entropy 𝐻(𝐴 ∣ 𝑆 ′ , 𝑠): 𝐻(𝐴 ∣ 𝑆 ′ , 𝑠) = −𝔼 𝑎∼𝜋(⋅|𝑠) log 𝑠′ ∼𝑝(⋅|𝑎,𝑠)

𝑝(𝑎 ∣ 𝑆 ′ , 𝑠)𝑞(𝑎 ∣ 𝑆 ′ , 𝑠) . 𝑞(𝑎 ∣ 𝑆 ′ , 𝑠)

(17)

Splitting the logarithm gives 𝐻(𝐴 ∣ 𝑆 ′ , 𝑠) = − 𝔼 𝑎∼𝜋(⋅|𝑠) log 𝑞(𝑎 ∣ 𝑆 ′ , 𝑠) 𝑠′ ∼𝑝(⋅|𝑎,𝑠)

− 𝔼 𝑎∼𝜋(⋅|𝑠) log 𝑠′ ∼𝑝(⋅|𝑎,𝑠)

𝑝(𝑎 ∣ 𝑆 ′ , 𝑠) . 𝑞(𝑎 ∣ 𝑆 ′ , 𝑠)

(18)

The second term is, by definition, the KL divergence. Since 𝑆 ′ is itself random, this divergence must subsequently be averaged over 𝑝(𝑠′ |𝑠): 𝐷𝐾𝐿 (𝑝(𝑎 ∣ 𝑆 ′ , 𝑠), ‖𝑞(𝑎 ∣ 𝑆 ′ , 𝑠)) = 𝔼𝑝(𝑠′ |𝑠) 𝐷KL (𝑝(𝑎|𝑠′ , 𝑠)‖𝑞(𝑎|𝑠′ , 𝑠)) .

(19)

Substitute in Eq. 18: 𝐻(𝐴 ∣ 𝑆 ′ , 𝑠) = −𝔼 𝑎∼𝜋(⋅|𝑠) log 𝑞(𝑎 ∣ 𝑆 ′ , 𝑠) − 𝐷𝐾𝐿 (𝑝(𝑎 ∣ 𝑆 ′ , 𝑠), ‖𝑞(𝑎 ∣ 𝑆 ′ , 𝑠)). (20) 𝑠′ ∼𝑝(⋅|𝑎,𝑠)

Substituting Eq. (20) into Eq. (10) gives 𝐼(𝐴; 𝑆 ′ ∣ 𝑠) = 𝔼 𝑎∼𝜋(⋅|𝑠) log 𝑞(𝑎 ∣ 𝑆 ′ , 𝑠) − 𝔼𝑎∼𝜋(⋅|𝑠) log 𝜋(𝑎 ∣ 𝑠) 𝑠′ ∼𝑝(⋅|𝑎,𝑠)

+ 𝐷𝐾𝐿 (𝑝(𝑎 ∣ 𝑆 ′ , 𝑠) ‖ 𝑞(𝑎 ∣ 𝑆 ′ , 𝑠)). 15

(21)

p(a|s) is defined by our policy 𝜋: 𝐼𝜋 (𝐴; 𝑆 ′ |𝑠) = 𝔼 𝑎∼𝜋(⋅|𝑠) [log 𝑞(𝑎|𝑠′ , 𝑠) − log 𝜋(𝑎|𝑠)] 𝑠′ ∼𝑝(⋅|𝑎,𝑠)

+ 𝔼𝑝(𝑠′ |𝑠) 𝐷KL (𝑝(𝑎|𝑠′ , 𝑠)‖𝑞(𝑎|𝑠′ , 𝑠)) .

(22)

This gives us a variational lower bound on on-policy empowerment, and therefore lower bound on empowerment. This construction is similar to ELBO objective in variational autoencoder [16, 17]. 𝐼𝜋 (𝐴; 𝑆 ′ |𝑠) ≥ 𝔼 𝑎∼𝜋(⋅|𝑠) [log 𝑞(𝑎|𝑠′ , 𝑠) − log 𝜋(𝑎|𝑠)]

(23)

𝑠′ ∼𝑝(⋅|𝑎,𝑠)

We can minimise divergence 𝐷𝐾𝐿 term by maximising log-likelihood log 𝑞(𝑎|𝑠, 𝑠′ ) on the on-policy samples (𝑎, 𝑠, 𝑠′ ). We need only positive samples for this case, unlike MINE or InfoNCE. For example if actions are continuous 𝑞𝜃 (𝑎 ∣ 𝑠, 𝑠′ ) = 𝒩(𝑎; 𝜇𝜃 (𝑠, 𝑠′ ), 𝜎𝜃 (𝑠, 𝑠′ ))

(24)

Then our reward function is: 𝑟𝑡emp = log 𝑞𝜃 (𝑎𝑡 |𝑠𝑡 , 𝑠𝑡+1 ) − log 𝜋(𝑎𝑡 |𝑠𝑡 )

(25)

Our new reward function has this relation with empowerment and on-policy empowerment 𝐼𝜋 : 𝔼[𝑟𝑡emp |𝑆𝑡 = 𝑠] ≤ 𝐼𝜋 (𝐴; 𝑆 ′ |𝑠) ≤ ℰ(𝑠)

(26)

Left bound becomes tight when 𝑞(𝑎|𝑠, 𝑠′ ) = 𝑝(𝑎|𝑠, 𝑠′ ), right bound becomes tight when policy achieves maximum capacity for given state. We can obtain similar results for forward model that predicts future state from the equation 𝐼(𝐴; 𝑆 ′ ∣ 𝑠) = 𝐻(𝑆 ′ ∣ 𝑠) − 𝐻(𝑆 ′ ∣ 𝐴, 𝑠) Though in this case we have to learn distribution of future state given current ′ 𝑝(𝑆 |𝑠) which is usually a much harder problem than learning the inverse model. In most environments there are much less options for an action that have caused transition between states (𝑠, 𝑠′ ) than there are possible future states under 𝑝(𝑠′ |𝑠). For example in our grid world it’s always only one possible action that causes transition between two cells, but there are 4 possible future states for the central cell. 16

The difficulty comes from simple models such as 𝑞(𝑠′ , 𝑠) = 𝒩(𝑠′ ; 𝜇𝜃 (𝑠), 𝐼) being inherently unimodal, not suitable for modelling a multimodal distribution. It’s possible to make 𝜇𝜃 (𝑠) output parameter for e.g. mixture of several Gaussians, or use more advanced and difficult techniques.

Figure 6: Maximum-likelihood estimation failure with a single normal distribution.

A unimodal predictor will place mean between modes and give diffuse(high variance) estimation for future state S as illustrated in Fig. 6. 4.2. Continuous(differential) entropy Let’s start with a definition of continuous entropy, it’s usually noted as lower ℎ to distinguish from discrete case: ℎ(𝑋) = − ∫ 𝑝(𝑥) log 𝑝(𝑥) 𝑑𝑥

(27)

Differential entropy of normal distribution: 1 log2 (2𝜋𝑒𝜎2 ). (28) 2 Continuous entropy behaves quite differently from discrete, in fact for a variable 𝑋 with Delta-function density we have ℎ𝑋 = −∞. Mutual information on the other hand stay non-negative, but there is another issue. Consider this transition function: 𝑆 ′ = 𝐴 + 𝜖 ℎ(𝐴) =

17

With action and noise following normal distribution: 𝐴 ∼ 𝒩(0, 𝜎), 𝜖∼ 𝒩(0, 𝑁). If action reconstruction becomes possible to arbitrary precision, in such deterministic environments empowerment may become infinite. For our inverse model we will have: ℎ(𝐴|𝑆 ′ , 𝑠) = 12 log(2𝜋𝑒𝑁 2 ) ⟶ −∞ as 𝑁 → 0 𝐼(𝐴; 𝑆 ′ ∣ 𝑠) = ℎ(𝐴 ∣ 𝑠) − ℎ(𝐴 ∣ 𝑆 ′ , 𝑠) = ℎ(𝐴 ∣ 𝑠) − (−∞) = ∞

(29)

Also for continuous actions it’s possible to maximise empowerment by increasing action magnitude: as 𝜎 → ∞ ℎ(𝑎|𝑠) = 12 log(2𝜋𝑒𝜎2 ) ⟶ ∞ Thus, the policy will learn to maximise ℎ(𝐴 ∣ 𝑠) by increasing the action range. A natural solution is to constrain the range to 𝐴 ∈ [𝑎min , 𝑎max ] or restrict covariance. To prevent entropy explosion in inverse model p(a∣s’,s) we could add noise to observations or actions: 𝑝(𝑎 ∣ 𝑠′ + 𝜖, 𝑠) cannot perfectly recover 𝑎 anymore keeping entropy and mutual information finite. For a more detailed treatment of continuous entropy and mutual information, see [18]. 5. Multistep and single-step empowerment Consider this simple 2 step environment. Two actions are possible a0=0, and a1=1. The environment always returns 0 at t=0 and t=1, and xor of two actions at t=2: T = 0 T = 1 T = 2, return a0 xor a1 observe 0→a0 →observe 0 → a1 → observe 1 observe 0→a0 →observe 0 → a0 → observe 0 observe 0→a1 →observe 0 → a0 → observe 1 observe 0→a1 →observe 0 → a1 → observe 0 The first action produces no visible one-step change: 𝐼(𝐴0 ; 𝑆1 ∣ 𝑆0 ) = 0. 𝑂2 is independent from 𝑂1 and 𝐴1 and 𝐼(𝐴1 ; 𝑆2 | 𝑠1 ) = 0 On the other hand if agent is recurrent it(and inverse model) remembers it’s actions. In such case we have 𝐼(𝐴0 , 𝐴1 ; 𝑆2 ∣ 𝑆0 ) = 1 bit. Note that final state doesn’t identify each action individually: 𝑆2 = 0 implies (𝐴0 , 𝐴1 ) ∈ 00, 11; 𝑆2 = 1 implies (𝐴0 , 𝐴1 ) ∈ 01, 10

18

This example shows that summing local quantities, ∑𝑡 𝐼(𝐴𝑡 ; 𝑆𝑡+1 ∣ 𝑆𝑡 ) can miss information carried jointly by multiple actions. 𝐼(𝐴0 𝐴1 ; 𝑆2 ∣ 𝑠0 ) in our example would be two-step empowerment. We can define k-step empowerment as: 𝐼𝑘 (𝑠0) ≜ max𝑝(𝑎0∶𝑘−1 |𝑠0 ) 𝐼(𝐴0∶𝑘−1 ; 𝑆𝑘 ∣ 𝑠0 ) With intermediate states added it is called trajectory empowerment. 𝐼𝑘 (𝑠0 ) ≜ max𝑝(𝑎0∶𝐻−1 |𝑠0 ) 𝐼(𝐴0∶𝑘−1 ; 𝑆1∶𝑘 ∣ 𝑠0 ) Knowing intermediate states makes it easier to recover actions so 𝐼(𝐴0∶𝑘−1 ; 𝑆𝑘 ∣ 𝑆0 ) ≤ 𝐼(𝐴0∶𝑘−1 ; 𝑆1∶𝑘 ∣ 𝑆0 ). Let’s expand first equation with inverse model: 𝐼(𝐴0∶𝑘−1 ; 𝑆𝑘 ∣ 𝑠0 ) ≥ 𝔼 log 𝑞(𝐴0∶𝑘−1 |𝑆𝑘 , 𝑠0 ) − log 𝑝𝜋 (𝐴0∶𝑘−1 ∣ 𝑠0 )

(30)

5.1. Marginal and Causal objective We can sample all actions at s0 and then execute them one by one. In this case 𝐻−1 actions probability term is easy to compute: 𝑝𝜋 (𝑎0∶𝐻−1 ∣ 𝑠0 ) = ∏𝑡=0 𝜋(𝑎𝑡 ∣ 𝑠0 , 𝑎0∶𝑡−1 ) Or we are sampling actions one by one, a new action at each new state. In this case to compute actions probability we need expectation over all possible trajectories between 𝑠0 and 𝑠𝑇 . 𝑝𝜋 (𝑎0∶𝐻−1 ∣ 𝑠0 ) = ∫ 𝑝𝜋 (𝑎0∶𝐻−1 , 𝑠1∶𝐻−1 ∣ 𝑠0 ) 𝑑𝑠1∶𝐻−1 This is intractable to compute besides most simple cases. If we have a model of the environment(see sec. 9 ) we can compute Monte-Carlo approximation by doing virtual rollouts with fixed initial state and fixed actions. Then approximation is average probability of sequence of actions over all virtual episodes. We can call it marginal MI objective since it marginalises over trajectory. 𝐻−1

𝑤𝑗 = ∏ 𝜋(𝑎𝑡̄ ∣ 𝐶𝑡(𝑗) ) 𝑡=0

1 𝑁 𝑝𝜋 ̂ (𝑎0∶𝐻−1 ̄ ∣ 𝑠0 ) = ∑ 𝑤𝑗 𝑁 𝑗=1

(31)

Here 𝐶𝑡(𝑗) is all relevant inputs to the policy for given episode 𝑗. Another option is to plug in on-policy likelihood: 𝐻−1

𝐼(𝐴0∶𝑘−1 ; 𝑆𝑘 ∣ 𝑠0 ) ≥ 𝔼 log 𝑞(𝐴0∶𝐻−1 ; 𝑆𝐻 |𝑠0 ) − ∑ log 𝜋(𝐴𝑡 ∣ 𝑆0∶𝑡 , 𝐴0∶𝑡−1 ) (32) 𝑡=0

19

This objective can be named causal since it uses causal action entropy ∑𝑡 𝐻𝜋 (𝐴𝑡 |𝑆𝑡 , 𝐴𝑖<𝑡 ) term. 5.2. cards and notebook environment Consider k-step environment: At each time dealer draws a random card 𝐶𝑡 and stacks it on the table. The robot performs action 𝐴𝑡 = 𝐶𝑡 such as: • wave hand left if the card is red and right if card is black • write what card it sees to the notebook After H cards are drawn it’s possible to examine stacked cards and determine which actions the robot took. There are total 2𝑘 possible actions sequences. So we will get up to k bits of information 𝑟 = log 𝑞(𝑎0∶𝑘−1 ̄ ∣ 𝑆𝑘 , 𝑠0 ) − log 𝑝𝜋 ̂ (𝑎0∶𝑘−1 ̄ ∣ 𝑠0 ) −𝐻 ⟶ 0 − log 2 = 𝐻 log 2. On other hand entropy term in causal objective is always zero since each action is determined by the card. So copying environment is not rewarded. Thus marginal objective ≥ causal. It’s important to note that causal objective is not a mutual information, it can be negative. Consider our card environment but this time cards are removed after being shown to the robot, 𝑆2 = 𝑏𝑙𝑎𝑛𝑘. Policy still copies actions so 𝜋(𝐴1 = 𝐶|𝑆1 = 𝐶) = 1. Then causal entropy is 𝐻(𝐴1 ∣ 𝑆0 , 𝑆1 ) = 0. Inverse model can’t reconstruct 𝐴1 from 𝑆0 , 𝑆2 , so it assigns equal probability 1/2 and log2 𝜋(𝐴1 ∣ 𝑆0 , 𝑆1 ) = −1. Thus we have 𝐼𝑐 𝑎𝑢𝑠𝑎𝑙 = −1 − 0 = −1. On other hand for an agent with a notebook maximum is achieved with uniform policy 𝐻(𝐴𝑡 ∣ 𝑆𝑡 ) = log 𝐾 and 𝐽 = 𝑘 log 𝐾. Thus the policy might learn to write down random symbols to notebook. Similar failure modes exist for all informationbased rewards. This issue is discussed in more details in the next section. We can also sample so-called latent plan 𝑈, or skill at the begging and optimise for this objective with objective ℰ𝑈 𝐻 (𝑠0 ) = max𝑝(𝑢|𝑠0 ) 𝐼(𝑈; 𝑆𝐻 |𝑠0 ) with variational lower bound 𝐼(𝑈; 𝑆𝐻 ∣ 𝑠0 ) ≥ 𝔼 [log 𝑞(𝑈 ∣ 𝑆𝐻 , 𝑆0 ) − log 𝑝(𝑈 ∣ 𝑠0 )] This option is discussed in sec. 10 and 15.

20

6. Local Optima Many people enjoy computer games, games can even hook some humans into addictive patterns. Empowerment is suitable for modelling this effect. The reward is high where there are many visited states, but the environment is controllable. An agent might get obsessed with flipping a switch that has a lot of settings because it’s easy to manipulate and offers immediate, predictable feedback. Prediction-error driven agents can get stuck observing a noise source. Similarly, empowermentdriven agent can stuck interacting with a controllable objects with many states. This problem arises for all information-based rewards and it’s likely not solvable on agent’s side.. People like playing video games, but even ones with addiction won’t stick there because of resource constraints. Any real environment has resource constraints and energy and resource sources. We can introduce such constraints into virtual environments. Lets add to the state energy coordinate e∈[0,1] 𝑆 = (𝑠, 𝑒) Every action consumes Δ𝑒 > 0. When 𝑒 < 𝑒𝑐𝑟𝑖𝑡 the motor controller becomes noisy or certain actions are disabled. Formally, p(s’∣s,a) grows broader ⇒ H(S’∣a,s) increases, thus lowering empowerment. There are recharge states 𝑠𝑐ℎ𝑎𝑟𝑔𝑒 (pickups, charging pads, food) that reset e upward. Therefore empowerment maximiser can obtain an intrinsic drive to: 1. stay away from low-energy states 2. periodically reach recharge regions 3. optimise its action sequence for long-term discounted reward I.e. the agent behaves like a living organism: forage → play/act → return to base → repeat. It is easy to introduce to an environment with direct biological analogy, like predator-prey, but harder for other cases. On the other hand in a multi-agent world it might not be necessary. Other agents can provide a natural brake against ”camping on a toy” behaviour. • Disturbance: other agents policy injects variability into transition probability 𝑝(𝑠𝑡+1 | 𝑎𝑡 , 𝑠𝑡 ) leading to increase of entropy of future state 𝐻(𝑆 ′ ∣ 𝐴, 𝑠) conditional on action. • resource competition: agent must keep discovering new high-control niches or defend the old one. 21

• Non-stationary dynamics: camping behaviour might become impossible as time goes due to the non-stationary environment itself. When the multi-agent cure might fail: • Predictable opponents If other agents adopt highly regular or submissive policies 𝐻(𝑆 ′ ∣ 𝐴, 𝑠) can still be low → empowerment can stay high. • Collusion / territorial partition Agents may implicitly agree on “you keep toy #1, I keep toy #2”. Each finds a private controllable niche. This has analogies in biology. Other agents can destabilise simple controllable niches, but they can also create new intrinsically rewarding attractors. Their effect therefore depends on both their behaviour, environment constraints and the intrinsic objective. 7. Communication and collective intrinsic motivation Communication extend an agent’s action channel beyond its own actuators. Other agents may disrupt simple controllable niches, but communication also creates a new controllable subsystem. Empowerment can therefore be defined not only through direct physical actions, but also through causal influence mediated by other agents. As mentioned in seq. 2 communication is ubiquitous in multicellular organisms so it’s important to consider how it affects agents with intrinsic motivation. In the simplest case there are two agents 𝐴 and 𝐵 that can send messages to each other from some alphabet M. They take turns: 𝐴 sends a message and waits for a response from 𝐵. Then 𝐵 sends its message and waits for response. High empowerment reward is achievable for both agents in such a situation. Agent 𝐴 selects message m randomly and sends it to 𝐵. If 𝐵 responds predictably, e.g. 𝑓 (𝑚) = (𝑚+1) mod |ℳ| then we have high accuracy for 𝑃(𝑆 ′ ∣ 𝑚, 𝑠0 ) but also high 𝐻(𝑆 ′ ∣ 𝑠0 ) (since response can’t be predicted without knowing the message). Since next state now is completely predictable we get empowerment equal to 𝐻(𝐴 ∣ 𝑠). In this case both agents will achieve maximum reward with uniform distribution over M. Achievable empowerment is log|A| Thus agents may obtain intrinsic reward by exchanging arbitrary messages without producing useful collective behaviour. 22

Prediction error could reward unexpected messages. Learning progress - learning of changes in other agent’s policy. Information gain - messages exposing other agents internal state diversity objectives - role and protocol specialisation. With a shared channel in which simultaneous messages interfere, agents may develop turn-taking schedules. But with energy constraints agents might refuse sending a response, lowering other agent’s empowerment. We can think of different levels and objectives in such systems: • self, or per-agent empowerment - increasing each own control • transfer empowerment - increasing control of other agents • assistance empowerment - increasing own influence on other agents [19] • joint empowerment - group of communicating agents constitute an organism, or meta-agent and can optimise their joint objective Joint empowerment is not automatically optimised when every agent independently optimises its own empowerment. From biology we can guess that local self-empowerment may lead to competition, domination, no-communication and possibly even to organism level cooperation. Communication predates complex multicellular organisms, but the emergence of integrated multicellular organisms required communication to be combined with adhesion, functional differentiation and mechanisms that limit conflict between cells. The long evolutionary delay, possibly billions of years, between early life and complex multicellular organisation suggests that local communication and adaptive behaviour are not by themselves sufficient to produce a stable higher-level organisation. Synchronisation of organism state across many cells using signals between adjacent cells may face scaling limits due to delays and error accumulation. There are examples of relatively broad low-capacity channels in biology. Quorum sensing(QS) is a widespread mechanism of cell-to-cell communication and coordination using signalling molecules. It is based on release of signal molecules to the outside of cells. The phenomenon has not only been described between cells of the same species (intraspecies), but also between species (interspecies) and between bacteria and higher organisms (inter-kingdom) [20]. Hormones are later development of the same mechanism.

23

There’s also communication based on electricity. Both cell-to-cell using ion pumps embedded in cell membranes [21]. And using relatively broad, macroscopic electric fields. Important example is electric gradient that guides body formation in embryogenesis. Optical signalling may represent another possibility, although its functional role in cell-to-cell communication remains less established [22]. We can conclude that global communication channel might enable more complex behaviour development in agents optimising self-empowerment. Analogously to chemical communication we can introduce global message 𝑔𝑡 that aggregates individual messages: 𝑁 𝑔𝑡 = 𝑁1 ∑𝑖=1 𝑚𝑖,𝑡 with aggregated message broadcasted to all agents: 𝑜𝑖,𝑡+1 = (𝑜local 𝑖,𝑡+1 , 𝑔𝑡 ) 8. MINE and InfoNCE 8.1. Mutual information neural estimation is defined as this objective: Mutual Information Neural Estimation (MINE) [23] uses the Donsker–Varadhan representation to estimate mutual information with a neural network.

𝐼(𝑋; 𝑌 ) = 𝐷KL (𝑃𝑋𝑌 ∥ 𝑃𝑋 ⊗ 𝑃𝑌 ) = sup [𝔼(𝑋,𝑌 )∼𝑃𝑋𝑌 [𝑇 (𝑋, 𝑌 )] − log 𝔼𝑋∼𝑃𝑋 [𝑒𝑇 (𝑋,𝑌 ) ]] . 𝑇 ∶Ω→ℝ

𝑌 ∼𝑃𝑌

(33) Here first expectation is over joint distribution and second is over marginal distributions of X and Y. T is a function returning real number 𝑇 (𝑋, 𝑌 ) → 𝑅. We approximate T by neural network 𝑇𝜃 (𝑋, 𝑌 ) and use gradient ascend to find supremum. At training time we just need Monte-Carlo samples for both terms. 𝑁 𝑁 ̂ Loss is 𝐼MINE = 𝑁1 ∑𝑖=1 𝑇𝜃 (𝑥𝑖 , 𝑦𝑖 ) − log [ 𝑁1 ∑𝑖=1 exp 𝑇𝜃 (𝑥𝑖̃ , 𝑦𝑖̃ )] 𝑥,̃ 𝑦 ̃ - samples from marginal distributions, to obtain them we can draw separate batches for x and y or just shuffle one of them e.g. y from the same batch that is used in 𝑇𝜃 (𝑥, 𝑦). It can be proofed that supremum achieved when 𝑇 ∗ (𝑥, 𝑦) = log

𝑝𝑋𝑌 (𝑥, 𝑦) +𝐶 𝑝𝑋 (𝑥)𝑝𝑌 (𝑦)

24

(34)

𝑃

log 𝑄𝑋,𝑌 = log 𝑃 − 𝑙𝑜𝑔 𝑄 = 𝑙𝑜𝑔 𝑝(𝑋, 𝑌 ) − log 𝑝(𝑋)𝑃(𝑌 ) is just pointwise 𝑋,𝑌 mutual information. For empowerment estimation we have two random variables A and S, but this is conditioned on current state s, so function T will have 3 inputs. Plugging A and S in T gives 𝑇 (𝐴; 𝑆 ′ |𝑠) = 𝑙𝑜𝑔𝑃(𝐴; 𝑆 ′ |𝑠) − 𝑙𝑜𝑔(𝑃(𝐴|𝑠)𝑃(𝑆 ′ |𝑠) + 𝑐 𝑇 (𝐴; 𝑆 ′ |𝑠) = 𝑙𝑜𝑔(𝑝(𝑆 ′ |𝐴, 𝑠)𝑝(𝐴|𝑠)) − 𝑙𝑜𝑔(𝑃(𝐴|𝑠)𝑃(𝑆 ′ |𝑠) + 𝑐 𝑇 (𝐴; 𝑆 ′ |𝑠) = 𝑙𝑜𝑔 𝑝(𝑆 ′ |𝐴, 𝑠) + 𝑙𝑜𝑔 𝑝(𝐴|𝑠) − 𝑙𝑜𝑔 𝑝(𝐴|𝑠) − 𝑙𝑜𝑔 𝑝(𝑆 ′ |𝑠) + 𝑐 𝑇 (𝐴; 𝑆 ′ |𝑠) = 𝑙𝑜𝑔 𝑝(𝑆 ′ |𝐴, 𝑠) − 𝑙𝑜𝑔 𝑝(𝑆 ′ |𝑠) + 𝑐 𝐼(𝐴; 𝑆 ′ ∣ 𝑠) = 𝐻(𝑆 ′ ∣ 𝑠)−𝐻(𝑆 ′ ∣ 𝐴, 𝑠) = − 𝐸 [ 𝑙𝑜𝑔 𝑃(𝑆 ′ ∣ 𝑠)] − (− 𝐸 [𝑙𝑜𝑔 𝑝(𝑆 ′ |𝐴, 𝑠)]) 𝐼𝐴,𝑆′ ∼𝜋 (𝐴; 𝑆 ′ ∣ 𝑠) = 𝐸[ 𝑙𝑜𝑔 𝑝(𝑆 ′ |𝐴, 𝑠) − 𝑙𝑜𝑔 𝑃(𝑆 ′ ∣ 𝑠) ] Thus trained T-function directly gives a point estimate of mutual information up to an additive constant. However it wont give empowerment estimation when trained on shuffled states, actions pairs. Empowerment requires outcomes from distribution conditioned on current state 𝑝(𝑆 ′ |𝑠 = 𝑠𝑡 ). Shuffling batch will give us 𝐼(𝑆 ′ ; (𝑆, 𝐴)) because shuffling will approximate 𝑆 ′ ~𝑝(𝑆 ′ ) - unconditional distribution instead of 𝑃(𝑆 ′ ∣ 𝑠). Compare these equations with explicitly written expectations: 𝑝(𝑆′ ∣𝑆,𝐴) 𝐼((𝑆, 𝐴); 𝑆 ′ ) = 𝐸𝑆∼𝑝(𝑠) 𝐸𝐴∼𝜋(⋅∣𝑆) 𝐸𝑆′ ∼𝑝(⋅∣𝑆,𝐴) 𝑙𝑜𝑔 𝑝(𝑆′ ) 𝑝(𝑆′ ∣𝑆=𝑠,𝐴)

𝐼(𝑆 ′ ; 𝐴|𝑠) = 𝐸𝐴∼𝜋(⋅∣𝑆) 𝐸𝑆′ ∼𝑝(⋅∣𝑆,𝐴) 𝑙𝑜𝑔 𝑝(𝑆′ |𝑆=𝑠) With 𝐼((𝑆, 𝐴); 𝑆 ′ ) = 𝐼(𝑆; 𝑆 ′ ) + 𝐼(𝐴; 𝑆 ′ ∣ 𝑆) Compare 𝐼(𝑆; 𝑆 ′ ) to 𝐼(𝐴; 𝑆 ′ ∣ 𝑠): 𝐼(𝑆; 𝑆 ′ ) = 𝐻(𝑆 ′ ) − 𝐻(𝑆 ′ |𝑆) 𝐼(𝐴; 𝑆 ′ ∣ 𝑠) = 𝐻(𝑆 ′ ∣ 𝑠) − 𝐻(𝑆 ′ ∣ 𝐴, 𝑠) So we have conflicting term 𝐻(𝑆 ′ |𝑆) that cancels out. 𝐼((𝑆, 𝐴); 𝑆 ′ ) = 𝐻(𝑆 ′ ) − 𝐻(𝑆 ′ ∣ 𝐴, 𝑠) Policy-level failure mode. Using this term as a reward is problematic. 𝐻(𝑆 ′ ) is fixed for all states in a minibatch, so our reward function does not distinguish if concrete action leads to better state space exploration. With all episodes being equal in this term the only option to optimise remains inverse model. This might lead to policy visiting just a few states that are easy to predict. This is an example when an intrinsic reward may correlate positively with exploration under the current policy, yet optimising that reward may fail to shift the policy-induced state-visitation distribution towards broader coverage.

25

8.2. InfoNCE Another method for mutual information estimation is Information Noise-Contrastive Estimation. 𝐿InfoNCE = −

exp 𝑓𝜃 (𝑥𝑖 , 𝑦𝑖 ) 1 𝑁 ∑ log 𝑁 . 𝑁 𝑖=1 ∑ exp 𝑓 (𝑥 , 𝑦 ) 𝑗=1

𝜃

𝑖

(35)

𝑗

Here (𝑥𝑖 , 𝑦𝑖 ) — positive pair, 𝑦𝑗 , 𝑗 ≠ 𝑖, — negatives, sampled from marginal distribution. Function 𝑓𝜃 is analogous to 𝑇𝜃 in MINE. It should return high value for samples coming from joint distribution and low values for marginal distribution. It is identical to softmax operator applied to positive and negative samples and can be viewed as a standard multi-class classification problem. This estimator is often more stable than MINE. It defines lower bound on mutual information: 𝐼(𝑋; 𝑌 ) ≥ log 𝑁 − 𝐿InfoNCE 9. WORLD MODELS World models are models that infer world state from observations and predict how this state evolves given actions. This allows to train policy in virtual episodes. Here we describe prominent models Dreamer v3 and TWISTER and their components. Contrastive Predictive Coding [24] DreamerV3 [25] TWISTER [26] Key components of DREAMER/TWISTER ℎ𝑡 = 𝐹𝜙(ℎ𝑡−1 , 𝑧𝑡−1 , 𝑎𝑡−1 ) - context encoder e.g. RNN or transformer. 𝑧𝑡̂ ∼ 𝑑𝜙 (𝑧𝑡̂ | ℎ𝑡 ) - dynamics predictor: models distribution of world state given context. This predicts how world changes given previous state and action 𝑎𝑡−1 which is also encoded in ℎ𝑡 𝑧𝑡 ∼ 𝑞𝜙 (𝑧𝑡 | 𝑜𝑡 , ...) - observation encoder, models distribution of 𝑧𝑡 given current observation 𝑜𝑡 e.g. VAE encoder. In dreamer v3 its 𝑧𝑡 ∼ 𝑞𝜙 (∗ | 𝑜𝑡 , ℎ𝑡 ) . 𝑠𝑡 = {ℎ𝑡 , 𝑧𝑡 } - concatenation of context and world state. 𝑜𝑡̂ ∼ 𝑝𝜙 (𝑜𝑡̂ | 𝑧𝑡 ) - observation decoder, models distribution for observations e.g. VAE decoder, in dreamer it’s 𝑜𝑡̂ ∼ 𝑝𝜙 (𝑜𝑡̂ | 𝑠𝑡 ). 𝑒𝑡̂ = 𝑢(𝑧𝑡 ) - model that projects world states to embedding space. 𝑒𝑡∶𝑡+𝑘 ̂ = 𝑤𝜙 (𝑠𝑡 , 𝑎𝑡∶𝑡+𝑘 ) - predicted embeddings given state and actions 𝑟𝑡̂ ∼ 𝑟𝜙 (𝑟𝑡̂ | 𝑠𝑡 ) - predicted reward for current transition 𝑐𝑡̂ ∼ 𝑐𝜙 (𝑐𝑡̂ | 𝑠𝑡 ) - predicted episode continuation/termination 26

Figure 7: Key components of DREAMER/TWISTER

27

losses: 𝑧𝑡̂ ∼ 𝑑 𝜙 is learnt with KL divergence with observation encoder 𝑞𝜙 (𝑧𝑡 | 𝑜𝑡 , ...). This is very similar to recursive Bayesian estimation or Bayesian filtering, very similar to the inference in the hidden-markov model. 𝑞𝜙 and 𝑝𝜙 are trained with whatever encoder-decoder loss is suitable e.g. ELBO loss. 𝑤𝜙 is trained with constructive loss with projection 𝑢(𝑧𝑡 ) of observed 𝑧𝑡 to predicted 𝑒𝑡∶𝑡+𝑘 ̂ It is also possible to learn other useful signals such as rewards and episode termination. 9.1. Action Conditioned (AC-CPC) This is TWISTER component that is responsible for training 𝑒𝑡̂ = 𝑢(𝑧𝑡 ). This is an architecture that allows compressing high-dimensional observations into useful representations. RSSM stands for Recurrent State-Space Model. This is a part of dreamer that constitutes the world model. In TWISTER it is TSSM - transformer instead of recurrent model. There are two key differences: 1) Dreamer updates distribution 𝑧𝑡̂ ∼ 𝑑𝜙 (∗ | ℎ𝑡 ) to match different encoder: 𝑧𝑡 ∼ 𝑞𝜙 (∗ | 𝑜𝑡 , ℎ𝑡 ). Unlike TWISTER which uses 𝑧𝑡 ∼ 𝑞𝜙 (∗ | 𝑜𝑡 ) 2) AC-CPC/TWISTER contrastive loss for embeddings additionally to reconstructive loss in Dreamer. Another minor difference decoder in dreamer is 𝑜𝑡̂ ∼ 𝑝𝜙 (𝑜𝑡̂ | 𝑠𝑡 ) Having said that it helps to concatenate many observations 𝑜𝑡 or use dreamer style encoder 𝑞𝜙 (𝑧𝑡 | 𝑜𝑡 , ℎ𝑡 ) for AC-CPC. Both methods are reported to work best with discrete distribution of world states 𝑧𝑡 . though can be used with different distributions. 𝑜0 , 𝑎0 , 𝑜1 , 𝑎1 , 𝑜2 , 𝑎2 , 𝑜3 , 𝑎3 ... - states, action sequence ℎ0 = 0, posterior 𝑧0 = 𝑞𝜙 (ℎ0 , 𝑜0 ) 𝑠0 = [ℎ0 , 𝑧0 ] 𝑠0 − > 𝑎0 ℎ1 = 𝐹𝜙 (𝑧0 , ℎ0 , 𝑎0 ) prior 𝑧̂1 = 𝑑𝜙 (ℎ1 ) posterior 𝑧1 = 𝑞𝜙 (ℎ1 , 𝑂1 ) Conceptually there information flow: 𝑧0 = 𝑞(ℎ0 , 𝑜0 ) → ℎ1 = 𝐹(ℎ0 , 𝑧0 , 𝑎0 ) → 𝑎1 28

So 𝑎1 is generated from 𝑜0 , 𝑎0 , 𝑜1 𝑠0 maps to 𝑎0 , 𝑠1 maps to 𝑎1 . and for ac-cpc starting from h1: 𝑒10 = 𝑓 (ℎ1 ) - conditioned by 𝑎0 via ℎ1 𝑒12 = 𝑓 (ℎ1 , 𝑧̂1 , 𝑎1 ) 𝑒13 = 𝑓 (ℎ1 , 𝑧̂1 , 𝑎1 , 𝑎2 ) with positives e10 <-> z1, e12 <-> z2, e13 <-> z3… TWISTER applies InfoNCE objective introduced in sec. 8.2 to action-conditioned predictions of future latent representations: 𝑘

1 𝑁 𝑒𝑠 𝑖𝑖 𝐿𝐼𝑛𝑓 𝑜𝑁𝐶𝐸 = − ∑ 𝑙𝑜𝑔 𝑘 𝑘 𝑁 𝑖=0 𝑒𝑠 𝑖𝑖 + ∑𝑗≠𝑖 𝑒𝑠 𝑖𝑗

(36)

Here s being similarity measure between embedding and latent variable z(z is being projected to the embedding space with a multi-layer perception). N is a batch size. Mutual information between matching embeddings e and latent states z then: 𝐼(𝑒; 𝑧) >= 𝑙𝑜𝑔 𝑁 − 𝐿𝐼𝑛𝑓 𝑜𝑁𝐶𝐸 In twister this is used only for descriptor training, but what if we use this MI estimation for reward? Is it similar to empowerment? If we assume that e and z are good enough representations of it’s inputs we could write: 𝐼(𝑒; 𝑧) ≈ 𝐼((ℎ𝑡 , 𝑎𝑡∶𝑡+𝑘 ), ℎ𝑡+𝑘 ) ≈ 𝐼((𝑠𝑡 , 𝐴), 𝑆 ′ )

(37)

By chain rule for MI(see appendix) 𝐼((𝑠, 𝐴); 𝑆 ′ ) = 𝐼(𝑠 ; 𝑆 ′ ) + 𝐼(𝐴; 𝑆 ′ ∣ 𝑠)

(38)

As you can see this quantity can reward behaviour when future trajectory is predictable even without knowing action. Expand: 𝐼(𝑠 ; 𝑆 ′ ) = 𝐻(𝑆 ′ ) − 𝐻(𝑆 ′ |𝑠) 𝐼(𝐴; 𝑆 ′ ∣ 𝑠) = 𝐻(𝑆 ′ ∣ 𝑠) − 𝐻(𝑆 ′ ∣ 𝐴, 𝑠) - multistep empowerment 𝐻(𝑆 ′ ) − 𝐻(𝑆 ′ |𝑠) + 𝐻(𝑆 ′ ∣ 𝑠) − 𝐻(𝑆 ′ ∣ 𝐴, 𝑠) = 𝐻(𝑆 ′ ) − 𝐻(𝑆 ′ ∣ 𝐴, 𝑠) 𝐼((𝑠, 𝐴); 𝑆 ′ ) = 𝐻(𝑆 ′ ) − 𝐻(𝑆 ′ ∣ 𝐴, 𝑠) This is difference 𝐻(𝑆 ′ ) vs 𝐻(𝑆 ′ ∣ 𝑠) tells us what we are trying to distinguish/predict from s and A. AC-CPC uses for negatives all possible states S in a

29

batch. While 𝐻(𝑆 ′ ∣ 𝑠)(as in empowerment) would require us to construct negatives only from states we could arrive from the current state 𝑠! Despite equation 38 being sound using global pool of states(or their embeddings) has issues. One issue mentioned in sec. 8.1 is that 𝐻(𝑆 ′ ) will be the same for all episodes in the minibatch. Another issue is first term 𝐼(𝑆; 𝑆 ‵ ) doesn’t involve actions. Suppose the positive pair from a RTS-game environment is: current frame: your base, five workers, daytime; future frame: nearly the same scene after a few seconds. A global negative might include completely different map locations, enemy type, time elapsed from start, different army and resource count. It’s easy for discriminator to distinguish even without using actions. At this point 𝐿𝑛𝑐𝑒 is close to zero and learning stops. Empowerment formulation on other hand requires ”hard negatives”. That is states produced from the same initial condition, but under alternative actions. The good thing is that world models allow us to generate realistic virtual episodes! 10. Diversity is all you need The goal of Diversity Is All You Need (DIAYN) [27] is to make an agent learn different skills in an unsupervised manner. In this method we pass to the policy additional parameter z that is sampled from a random distribution(categorical or normal). Z ∼ p(z). Policy conditioned on z is called skill. Training objective has tree parts: 1) Mutual information between states and skills I(S;Z) is maximised 2) MI between actions and skills given the state is minimised I(A;Z | S) 3) Entropy of actions given state is maximised H[A | S] max 𝐹(𝜃) ≜ 𝐼(𝑆; 𝑍) + 𝐻[𝐴|𝑆] − 𝐼(𝐴; 𝑍|𝑆) = (𝐻[𝑍] − 𝐻[𝑍|𝑆]) + 𝐻[𝐴|𝑆] − (𝐻[𝐴|𝑆] − 𝐻[𝐴|𝑆, 𝑍]) = 𝐻[𝑍] − 𝐻[𝑍|𝑆] + 𝐻[𝐴|𝑆, 𝑍]

(39)

𝐻[𝑍] is entropy of skill distribution - constant. Conditional entropy of skill distribution can be rewritten using chain rule for conditional entropy as: 𝐻[𝑍|𝑆] = 𝐻[𝑆, 𝑍] − 𝐻[𝑆] = 𝐻[𝑆|𝑍] − 𝐻[𝑆] + 𝐻[𝑍] This will help us to understand when exactly the reward will be high and when low. 30

Substituting it into original equation: 𝐹(𝜃) = 𝐻[𝑍] − (𝐻[𝑆, 𝑍] − 𝐻[𝑆]) + 𝐻[𝐴|𝑆, 𝑍] = 𝐻[𝑍] − 𝐻[𝑆, 𝑍] + 𝐻[𝑆] + 𝐻[𝐴|𝑆, 𝑍] Expand joint entropy of skills and states: 𝐻[𝑆, 𝑍] = 𝐻[𝑆|𝑍] + 𝐻(𝑍) = 𝐻(𝑍|𝑆) + 𝐻(𝑆) 𝐹(𝜃) = 𝐻[𝑍] − (𝐻(𝑆|𝑍) + 𝐻(𝑍)) + 𝐻(𝑆) + 𝐻[𝐴|𝑆, 𝑍] = 𝐻[𝑍] − 𝐻(𝑆|𝑍) − 𝐻(𝑍) + 𝐻(𝑆) + 𝐻[𝐴|𝑆, 𝑍] = 𝐻(𝑆) + 𝐻[𝐴|𝑆, 𝑍] − 𝐻[𝑆|𝑍] This form helps to understand what behaviour is encouraged by this reward. First term H(S) maximises entropy of states = encourages policy to visit many states. 𝐻[𝐴 | 𝑆, 𝑍] - conditional entropy of action given state and skill - encourages policy to apply different actions given skill and state. Equivalently this means it should be hard to guess skill from action alone. Also maximising this term makes it so knowing state + action doesn’t give more information about skill than knowing just state alone. H(S|Z) is minimised. If it is low we have a small number of states which are visited by given skill. Using definition I(Z;A∣S)=H(A∣S)−H(A∣S,Z) H(A∣S,Z) = H(A∣S) - I(Z;A∣S) So, 𝐹(𝜃) = 𝐻(𝑆) + 𝐻(𝐴 ∣ 𝑆) − 𝐻(𝑆|𝑍) − 𝐼(𝑍; 𝐴 ∣ 𝑆) = 𝐻(𝑍) + 𝐻(𝐴 ∣ 𝑆) − 𝐻(𝑍 | 𝑆) − 𝐼(𝑍; 𝐴 ∣ 𝑆) Max H(S) = maximise the number of visited states. Max H(A∣S) = maximise the number of actions taken in any particular state. Min H(S|Z) = minimise the number of states visited by a particular skill. Min I(Z;A∣S) = all actions are more or less the same for all skills. Min H(S|Z) = skill can be used to predict S This objective can be implemented as: 𝑝(𝑠, 𝑧) 𝐻(𝑍 | 𝑆) = − ∑ 𝑝(𝑠, 𝑧) 𝑙𝑜𝑔 𝑝(𝑠) = 𝐸[−𝑙𝑜𝑔 𝑝(𝑧|𝑠)] 𝐻(𝑍) = − ∑ 𝑝(z) log p(z) = E[-log p(z)] 𝐻(𝐴 ∣ 𝑆, 𝑍) = − ∑ 𝑝(𝑠, 𝑧) ∑ 𝑝(𝑎|𝑠, 𝑧) 𝑙𝑜𝑔 𝑝(𝑎| 𝑠, 𝑧) = 𝐸𝑠,𝑧 [− ∑ 𝑝(𝑎|𝑠, 𝑧)𝑙𝑜𝑔 𝑝(𝑎| 𝑠, 𝑧) ] = −𝐸 𝑠,𝑧,𝑎 [𝑙𝑜𝑔 𝑝(𝑎| 𝑠, 𝑧) ] 𝐹(𝜃) = 𝐸[−𝑙𝑜𝑔 𝑝(𝑧)] − 𝐸[−𝑙𝑜𝑔 𝑝(𝑧|𝑠)] + 𝐸[−𝑙𝑜𝑔 𝑝(𝑎| 𝑠, 𝑧) ] = 𝐸[ 𝑙𝑜𝑔 𝑝(𝑧|𝑠) − 𝑙𝑜𝑔 𝑝(𝑧) − 𝑙𝑜𝑔 𝑝(𝑎| 𝑠, 𝑧) ]

31

Probability of skill given state p(z|s) is approximated by a neural network predictor 𝜙𝜃 (𝑠)− > 𝑧 log 𝑝(𝑎| 𝑠, 𝑧) is just policy entropy log 𝜋𝜃(a| s,z). It can be included in the reward or treated separately, like entropy reguliser as in soft actor-critic. r(s, a| z) = log 𝜙(z | s) − log p(z) In principle it’s possible to use forward density p(s|z), but practically choice p(z|s) is much easier to implement, since we don’t know the ground truth distribution of S given skill. Estimating 𝜙(z∣s) is just a classification/regression problem; the target distribution over z is simple (uniform or Gaussian). Estimating p(s∣z) is often a high-dimensional density problem (very hard for image observations). Side by Side comparison with Empowerment

pushed up variance pushed up Encouraged property Discouraged property

Empowerment p(s‘∣a,s) 𝑝(𝑠′ ∣ 𝑠) Predictable consequences per action Small marginal next-state variability

DIAYN p(z∣s) p(a∣s,z) States discriminate skills; actions stay diverse Peaky action distribution inside a skill

In equation 1 variable Z is sampled ones per episode, S is picked uniformly from the whole trajectory, so the term 𝐻(𝑍 | 𝑆) is computed with respect to the entire trajectory. But there are situations when examining multiple states might be desirable. We might be interested in achieving the same state by different means, for example robot can place a spoon in a mug either by taking a spoon and placing it or by trying to scoop the spoon. In this case intermediate states are clearly distinguish skill vector, but the end state is the same. By default DIAYN will try to avoid visiting the same state from two skills. We could counter it by either by adding external reward, curiosity, or adding structure to the skill vector as discussed below: Structural DIAYN For example let vector z be concatenation of vectors 𝑧𝑖 . If we choose to use just two vectors we will have: • two latents z(1) and z(2) are concatenated 𝑧 = [𝑧(1) , 𝑧(2) ]; • the policy 𝜋(𝑎 ∣ 𝑠, 𝑧(1) , 𝑧(2) ) can use both parts all the time; 32

• the discriminator at an early time slice tries to predict 𝑧(1) only, ignoring 𝑧(2) ; • the discriminator at the final state (or any late slice) predicts 𝑧(2) only. Any two skills that differ only in 𝑧(1) therefore must converge to (almost) the same final state, reached via recognisably distinct trajectories. We can also try to reconstruct all 𝑧(𝑖 < 𝑡) . In this case policy will be encouraged to treat early 𝑧𝑖 as a high-level coarse plan or “style” and later 𝑧𝑖 further nuancing the behaviour. Other options such as using a sliding window are possible. 11. Curiosity and Learning progress Curiosity reward can be a combination of a few things: surprise, for example in the form of prediction error, novelty and learning progress. The theoretical foundations of curiosity and learning progress were developed by Jürgen Schmidhuber; see [28], [29]. 11.1. Prediction error and Novelty Prediction error can be estimated from forward model: 𝑓 (𝑠, 𝑎) − > 𝑠𝑡+1 ̂ Reward then is ||𝑠𝑡+1 − 𝑓 (𝑠, 𝑎)|| Novelty is inversely related to state visitation count: it is high for new states and low for frequently visited states. It encourages large entropy H(S) of the state visitation distribution. Implementation note. In practice, novelty is often computed from embeddings of observed states. As discussed in the Policy-level failure mode paragraph and Section 9, some novelty-based rewards may fail to distinguish episodes that produce broader exploration. A second failure mode can occur when the score lacks a persistent scale across policy updates. If novelty is normalised using statistics of the current batch, small differences between increasingly similar episodes are rescaled. The normalised reward can therefore continue to rank episodes within each batch even as the absolute diversity and coverage decreases.

33

11.2. Learning progress Learning progress is a bit more complex. We have forward model 𝑓𝜃 (𝑠𝑡+1 |𝑠, 𝑎) and distribution over its weights 𝑝(𝜃|𝑂𝑡 ) given the history of transitions; 𝑂𝑡 = ′ )} for 𝜏 < 𝑡. {(𝑠𝜏 , 𝑎𝜏 , 𝑠𝜏 Learning progress is modeled as a change of distribution over parameters of forward model 𝑓𝜃 (𝑠𝑡+1 |𝑠, 𝑎) given a new observation. 𝐿𝑃𝑡 ≜ 𝐾𝐿 [ 𝑝(𝜃 | 𝑂𝑡 ∪ 𝑂𝑡+1 ) ‖ 𝑝(𝜃 | 𝑂𝑡 ) ] This can be approximated simply with improvement in prediction error. That is 𝑟 = || 𝑓𝜃𝑘 (𝑠, 𝑎) − 𝑠𝑡+1 || − ||𝑓𝜃𝑘+1 (𝑠, 𝑎) − 𝑠𝑡+1 || Novelty and prediction error, unlike learning progress are prone to noise-staring behaviour. That is, the reward is high when the agent observes pure noise. But there are easy workarounds. We can train a model that predicts action given previous and current states. 𝑔𝜃 (𝑠𝑡+1 , 𝑠𝑡 ) − > 𝑎. Higher layers of this model can be used as feature extractors, that keep information only about controllable aspects of the environment. These features can be used then instead of raw states in novelty or surprise rewards. Let’s compare curiosity with empowerment. Empowerment I(A;S’∣s)=H(S’∣s)−H(S’∣s,a) Prediction error H(S’∣s,a) State novelty H(S) Given identity 𝐻(𝑆 ′ ) = 𝐼(𝑆 ′ ; 𝑆) + 𝐻(𝑆 ′ ∣ 𝑆) we have H(S’∣s) is less or equal to 𝐻(𝑆 ′ ). That is possible to have large 𝐻(𝑆 ′ ) and but small 𝐻(𝑆 ′ ∣ 𝑆). Reverse is not true: High 𝐻(𝑆 ′ ∣ 𝑆) implies that 𝐻(𝑆 ′ ) at least that large. So maximising 𝐻(𝑆 ′ ) is not identical to maximizing 𝐻(𝑆 ′ ∣ 𝑆). 𝐻(𝑆 ′ ∣ 𝑆) encourages rewards states from which many different states can occur immediately(or in k-steps with k-steps empowerment). But 𝐻(𝑆 ′ ) rewards visiting the whole state space. 12. Information gain Assume our world model estimates world state with z and tracks history in h. Then we can define (point) information gain about world state as 𝐾𝐿(𝑞(𝑧 ∣ 𝑜, ℎ) ∥ 𝑝(𝑧 ∣ ℎ)) = 𝐼𝑝𝑚𝑖 (𝑜; 𝑧 ∣ ℎ) That is how much have we learned about world given a new observation o.

34

Note: term ”information gain” is applicable to different things, including model parameters. In this case we have change in distribution over model parameters which is directly related to learning progress. This quantity could be used as addition to prediction error. We have this relation: 𝐻(𝑂 ∣ ℎ) = 𝐸𝑧∼𝑝(𝑧∣ℎ) [𝐻(𝑂 ∣ 𝑧, ℎ)] + 𝐼(𝑂; 𝑍 ∣ ℎ) The first quantity is unreducible or so called aleatoric uncertainty. This is uncertainty that remains even if we have good state estimation. The second term is the uncertainty caused by not knowing z. If agent encounters a source of noise prediction entropy(and error) will be high, but information gain small hinting at large aleatoric uncertanty. On other case consider agent exploring unknown part of the map, in this case prediction error will be high, but information gain also large. We could use this to reward useful exploration much more than just watching random events. Similar to other information-based rewards IG is prone to ”camping on a toy” issue. For example with fair dice we have 𝐼(𝑂; 𝑍|𝐻) = 𝐻(𝑍|𝐻) − 𝐻(𝑍|𝑂, 𝐻) = 𝑙𝑛(6) − 0 = 𝑙𝑛(6) nats for each throw. 13. SFA Invented by Laurenz Wiskott and Terrence Sejnowski [30], Slow Feature Analysis is an unsupervised learning rule that extracts features whose values change as slowly as possible over time, although they are computed from an input stream that may itself vary quickly. Suppose that an encoder neural network transforms each observation 𝑜𝑡 into a feature vector 𝑦𝑡 : 𝑦𝑡 = 𝑔(𝑜𝑡 ) Slowness is achieved with loss 𝐿slowness =

1 𝑇 ∑ ‖𝑦 − 𝑦𝑡−1 ‖22 . 𝑇 − 1 𝑡=2 𝑡

(40)

In order to avoid degenerate solutions such as encoding each observation as zeros we require decorrelation, zero mean and unit variance for y. Variance loss is defined as: 𝐶𝑌 = 𝑁1 𝑌 𝑇 𝑌 where Y is a concatenation of zero-centred vectors y. 𝐿variance =

1 𝑑 2 ∑ (𝐶𝑌 [𝑗, 𝑗] − 1) . 𝑑 𝑗=1 35

(41)

The correlation loss is the normalised squared Frobenius norm of the off-diagonal part of 𝐶𝑌 : 𝐿correlation =

1 ∑ 𝐶 [𝑖, 𝑗]2 . 𝑑(𝑑 − 1) 𝑖≠𝑗 𝑌

(42)

Then objective is 𝐿𝑠𝑓 𝑎 = 𝐿𝑠𝑙𝑜𝑤𝑛𝑒𝑠𝑠 + 𝐿𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒 + 𝐿𝑐𝑜𝑟𝑟𝑒𝑙𝑎𝑡𝑖𝑜𝑛 This is potentially useful augmentation for action conditioned state embeddings. AC-CPC in TWISTER must store relatively fast details needed for next observation prediction. We could extract slow-changing features from action-conditioned embeddings. Action conditioning is important: it may help distinguish controllable features from features that do not depend on the agents actions. Applying SFA on top of AC-CPC might therefore extract slow, controllable features suitable for computing of intrinsic reward over longer time scales. 14. Predictive information bonus and MDL Imagine an agent that can see part of an image, can move on it and can change pixels. Empowerment alone won’t produce interesting non-random looking images, it does not express a preference for regularity, coherence, or semantic content. It encourages diverse, but predictive states, so the agent might change many pixels randomly producing images visually similar to noise. We can use MI between patches of the image 𝐼(𝑡𝑜𝑝 − 𝑙𝑒𝑓 𝑡, 𝑡𝑜𝑝 − 𝑟𝑖𝑔ℎ𝑡) = 𝐻(𝑡𝑜𝑝 − 𝑙𝑒𝑓 𝑡) − 𝐻(𝑡𝑜𝑝 − 𝑙𝑒𝑓 𝑡|𝑡𝑜𝑝 − 𝑟𝑖𝑔ℎ𝑡) as reward in this case. This distinguishes structured images from two trivial solutions: Blank images: both patches are predictable, but there is no variation so H(top-left) = 0. Independent random noise: the patches vary, but one does not predict the other. So 𝐻(𝑡𝑜𝑝 − 𝑙𝑒𝑓 𝑡|𝑡𝑜𝑝 − 𝑟𝑖𝑔ℎ𝑡) ≈ 𝐻(𝑡𝑜𝑝 − 𝑙𝑒𝑓 𝑡) again giving 0 mutual information. Structured and variable images: patches vary across images but share regularities, giving positive mutual information. The objective can be extended to many random partitions and multiple spatial scales: 𝑅structure (𝑥) = 𝔼(𝑈,𝑉 )∼ℳ [𝐼(𝑋𝑈 ; 𝑋𝑉 )] , For discrete pixels or image tokens, one can train: - a marginal model 𝑝𝜙 (𝑋𝑉 ), and - a conditional model 𝑞𝜓 (𝑋𝑉 ∣ 𝑋𝑈 ). 36

A sample-level structure reward is then 𝑟structure (𝑥) = log 𝑞𝜓 (𝑥𝑉 ∣ 𝑥𝑈 ) − log 𝑝𝜙 (𝑥𝑉 ). The first term rewards predictability from context, while the second prevents the predictor from receiving high reward merely because the patch is constant everywhere. log 𝑝𝜙 (𝑥𝑉 ) term corresponds to minimum description length(see the section below). Predictive information can still collapse to a small family of highly regular images—for example, the same checkerboard in every episode. We therefore need to distinguish within-image structure from across-image diversity. A diversity objective can reward entropy in a global image representation 𝐻(𝑓𝜃 (𝑋)) where f should preferably be a frozen or slowly changing perceptual encoder. Another option is to sample a latent intention 𝑍 at the beginning of an episode and maximize 𝐼(𝐶; 𝑋𝑇 ). It is closely related to unsupervised skill discovery: each latent code corresponds to a different controllable mode of image generation. The two objectives are complementary: predictive information discourages noise generation, global diversity discourages generation of single or a few patterns. Therefore we could combine these terms into one reward: 𝑅 = 𝛼𝑅𝑒 𝑚𝑝𝑜𝑤𝑒𝑟𝑚𝑒𝑛𝑡 + 𝛽𝑅𝑠 𝑡𝑟𝑢𝑐𝑡𝑢𝑟𝑒 + 𝛾𝑅𝑑 𝑖𝑣𝑒𝑟𝑠𝑖𝑡𝑦 − 𝜆𝑅𝑐 𝑜𝑠𝑡 New term 𝑅𝑐 𝑜𝑠𝑡 here represents action or complexity cost. This term is needed to avoid useless image modification e.g. changing one pixel black -> white -> black. This objective does not formally guarantee aesthetically or semantically interesting images. For example, repeated textures, barcodes, or hidden high-frequency signals may score highly despite looking uninteresting to humans. The term used for a processes that could achieve more and more diverse and complex artifacts is ”open-endedness”. Our reward combination does not guarantee an open-ended process. Fixed intrinsic objectives can still be exhausted or exploited. Once the agent discovers a finite family of highly controllable, structured images, it may cycle among them indefinitely without producing genuinely new organization. Open-endedness additionally requires a continually expanding space of challenges, or niches — for example through co-evolving competing agents, procedurally generated environments, or learned objectives that change as previous behaviours become common. Possible formalisation is given in [31].

37

14.1. VAE and MDL VAE could be used as an approximator for description length. According to Shannon’s source coding theorem optimal code length(for a prefix code) that one can assign to a datapoint 𝑥 is its negative log-likelihood −𝑙𝑜𝑔𝑝(𝑥). VAE models empirical distribution 𝑝𝑡𝑟𝑢𝑒 (𝑥) as 𝑝𝜃 (𝑥) = ∫𝑧 𝑝𝜃 (𝑥|𝑧)𝑝(𝑧)𝑑𝑧, approximating the intractable posterior 𝑝𝜃 (𝑧|𝑥) with the variational posterior 𝑞(𝑧|𝑥). For detailed derivation see [16, 17]. ELBO objective ℒELBO (𝑥) = log 𝑝𝜃 (𝑥) − 𝔻𝐾𝐿 (𝑞𝜙 (𝑧|𝑥)||𝑝𝜃 (𝑧|𝑥)) = 𝔼𝑞𝜙 (𝑧|𝑥) [log 𝑝𝜃 (𝑥|𝑧)] − 𝔻𝐾𝐿 (𝑞𝜙 (𝑧|𝑥) ∥ 𝑝(𝑧))

(43)

Thus we have log 𝑝𝜃 (𝑥) ≥ ℒELBO (𝑥) for unknown true distribution this inequality holds in expectation: 𝐸𝑥∼𝑝𝑡𝑟𝑢𝑒 log 𝑝𝜃 (𝑥) ≥ 𝐸𝑥∼𝑝𝑡𝑟𝑢𝑒 ℒELBO (𝑥) 15. Problems with Empowerment and DIAYN Single-step empowerment is short-sighted. It can even be zero for obviously easy environments as shown in the xor environment example. Multistep empowerment is better, but it can’t be applied without modifications to high-frequency long-term environments like RTS games. On the other hand DIAYN objective is global in a sense that we can sample individual states from a trajectory, and it will provide valid lower bound estimation 𝐼(𝑍; 𝑠𝑖𝑛𝑔𝑙𝑒 𝑠𝑡𝑎𝑡𝑒) ≤ 𝐼(𝑍; 𝑤ℎ𝑜𝑙𝑒 𝑡𝑟𝑎𝑗𝑒𝑐𝑡𝑜𝑟𝑦). But empowerment can be zero given one-step or two step estimation. DIAYN’s objective is limited in a sense that it is just for learning skills. Humans or animals use skills to achieve certain goals, skills are combined and used in certain order. DIAYN objective says nothing about how skills should be combined. In more realistic architecture an agent would pursue long-term goals switching between skills as the situation evolves. It’s possible to aggregate multiple steps to enable DIAYN to develop different temporal patters, for example achieving the same goal with different locomotion gaits. However it’s become very easy for policy to develop pathological solution for skill discrimination e.g. just waiving manipulators in different pattern that are easy to classify. Another example of valid, but uninteresting optimum is in Figure 8. Assume out agent can move in 2d circle, starting from the centre and observe it’s position. It’s possible for policy to slice circle along radius and thus recover ”skill” vector z while staying near 38

the start. Other rewards such as novelty, e.g. euclidean distance from all states in minibatch could drive agent to explore much larger state-space. Our example shows that combination of intrinsic rewards can help counter each-others failure modes in some cases. Issues with empowerment can be addressed by introducing different time scales and state compression. In such case empowerment would be used to reward highlevel actions that take many environmental steps to execute. DIAYN skills can be such low-level actions. And skill selection can be trained by n-step empowerment. If we choose to use empowerment for only high-level action selection we have to decide which states to use and which to omit. Consider RTS game with 640x480 video frames as observations. Suppose we have high-level actions pursue, build expand, build unit, attack, flee etc. We can’t use just pixel observations at skill boundary. Both policy and empowerment estimator would require a feature vector that encodes strategic situations in the game + more detailed description of the current situation, that is exact unit position, recent events such as “hero just died”. Multi-layer SFA is a possible front-end to obtain the macro strategic state on which we measure empowerment or DIAYN MI, but exact architecture remains an open research problem.

Figure 8: State-space partition and traversal. Left: DIAYN partitions the state space into skillconditioned regions but does not guarantee that an arbitrary goal 𝑔0 can be reached; increasing the number of skills may only make these regions thinner. Right: a sequence of reachable goalconditioned regions 𝑔0 , 𝑔1 and 𝑔2 can support traversal through the state space.

39

16. Research agenda We conclude with discussion of possible experiments intended to test if existing intrinsic motivation methods can produce biological adaptability on different levels. The proposed direction is related to Schmidhuber’s formal theory of creativity and intrinsic motivation, in which agents actively generate experiments and receive intrinsic reward for discovering novel but learnable regularities that improve prediction or data compression [29]. 16.1. Environment for intrinsic reward evaluation Based on previous discussion of MI-based rewards and their ”camping on a toy” limitations we propose an environment with natural constraints on camping(or other similarly useless) behaviour. There are two properties of biological environments that counter failure modes of information-based rewards: 1. World is adversarial. 2. Energy is scarce. We design an environment where agents have an energy level. Analogously to physiological deficits, a low energy level reduces action amplitude and/or makes actions less predictable. Thus energy level affects the capacity of action to future state channel. We propose a predator–prey environment with the following observed state: 𝑜𝑡 = (𝑠𝑒𝑛𝑠𝑜𝑟𝑡 , 𝑒𝑡 ) where 𝑒𝑡 is the energy level. Agents have a finite storage capacity 𝑒𝑚𝑎𝑥 and each action consumes energy: 𝑒𝑡+1 = 𝑐𝑙𝑖𝑝(𝑒𝑡 − 𝑐(𝑎𝑡 ) + 𝑓 𝑜𝑜𝑑𝑡 , 0, 𝑒𝑚𝑎𝑥 ) Energy level affects actions: 𝑎𝑒𝑓 𝑓 𝑒𝑐𝑡𝑖𝑣𝑒 = 𝑚(𝑒𝑡 )𝑎𝑡 + 𝜎(𝑒𝑡 )𝜖𝑡 Here 𝑚 is a magnitude function, might be as simple as 𝑚(𝑒𝑡 ) = max(𝑒𝑡 , 𝑚min ). The function 𝜎 determines the magnitude of motor noise and increases as 𝑒𝑡 decreases. for example we can define 𝜎(𝑒𝑡 ) = min(𝜎𝑚𝑎𝑥 , − ln(𝑒𝑡 /𝑒𝑚𝑎𝑥 )) 𝜖𝑡 ∼ 𝒩(0, 𝐼) Finding food or capturing prey increases the energy level without providing direct reward signal. Note: as discussed earlier continuous actions requires non-zero observation or action uncertainty to keep mutual information finite.

40

This environment would allow to test directly whether different combinations of intrinsic rewards lead to behaviour that maintains a stable energy level, analogous to energy homeostasis in biological organisms. Possible tests include environments with: 1. controllable toy; 2. ”noisy-tv”; 3. communication channel between predators; 4. adversarial prey; Important ablations include disabling the effects of energy on actions and comparing stationary food with adversarial prey. This comparison is necessary because an intrinsic reward may cause the agent to follow prey for reasons unrelated to energy regulation. Primary metrics to monitor: 1. mean energy level 2. time spent on toys 3. spatial coverage 4. energy level at which the agent starts sustained movement towards prey 5. energy recovery time after reaching the minimum energy level Interesting extension is test conditions under which communication might develop e.g. when it’s hard for an agent to catch the prey on it’s own. 16.2. Recurrent model We propose to study a network of recurrent modules where every module is treated as a local agent. Each agent has its own hidden state, observations, incoming messages, actions and intrinsic reward. There is no reward defined for the network as a whole. The modules can communicate during the forward pass, but messages are detached before being passed between agents. Consequently, gradients from one agent cannot propagate through the internal computations of another agent. For a system of 𝑁 agents communicating with messages 𝑚 the shared recurrent update can be written as ℎ𝑖,𝑡+1 = 𝐹𝜃 (ℎ𝑖,𝑡 , 𝑜𝑖,𝑡 , 𝑚𝑖,𝑡 ),

𝑖 ∈ {1, … , 𝑁},

(44)

where all agents use the same parameters 𝜃, while ℎ𝑖,𝑡 , 𝑜𝑖,𝑡 and the position of an agent in the communication graph are different. The shared parameters are 41

analogous to a common genome, while different hidden states, inputs and network positions provide different local contexts. Functional specialisation may therefore emerge without assigning a permanent identity or a separate set of parameters to every module. Every agent computes an intrinsic reward 𝑟𝑖,𝑡 only from its own interaction history. The primary experiment is to test whether the recurrent modules develop stable and complementary roles and whether their joint dynamics exhibit adaptive behaviour that is not explicitly rewarded at the system level. The basic ablation removes weight sharing. In this condition each module has an independent recurrent function ℎ𝑖,𝑡+1 = 𝐹𝜃𝑖 (ℎ𝑖,𝑡 , 𝑜𝑖,𝑡 , 𝑚𝑖,𝑡 ).

(45)

Independent parameters may make specialisation easier, because different roles can be stored directly in 𝜃𝑖 . Shared weights provide a stronger test: different roles must emerge from local state, experience and network context. As discussed in section 7, a global communication channel might facilitate the formation of collective behaviour. The role of communication can be tested through the following ablations: 1. no communication; 2. local communication; 3. global low-capacity broadcast; 4. local + global communication; This type of agent may be evaluated in a collectively embodied task e.g., sensor + motor policies jointly controlling motion in a maze. In such an environment communication is mandatory. The architecture can also be tested with independently embodied agents. A good example is a predator-prey environment where each recurrent module controls one predator. This type of environment is especially suitable for a communication ablations because the agents can remain independently function when message passing is disabled. This experiment is not intended to prescribe a final architecture. Its purpose is to test whether local intrinsic objectives, recurrent memory, communication and a shared learning rule are sufficient ingredients for the emergence of functional differentiation and higher-level adaptive organisation.

42

16.3. Automatic curriculum in a two-level world model agent The limitations discussed in Section 15 suggest an experiment in which motor control and long-term goal selection are learned by separate agents potentially operating at different time scales. The proposed architecture consists of a low-level agent, a high-level agent and a world model. The world model supplies learned state representations and intrinsic signals for training both agents. Training starts with the low-level agent acting without a valid goal. Its reward is a mixture of learning progress, prediction error and diversity in the learned embedding space. This stage has two purposes: to collect diverse experience for the world model and to train a low-level policy capable of producing non-trivial transitions before it is asked to follow goals. For initial goal conditioning training obvious choices are hindsight training and virtual hindsight training on wm-generated episodes once wm is stable. After this initial stage, the high-level agent generates a target embedding 𝑔𝑡 . The low-level policy receives both the current representation 𝑠𝑡 and the target 𝑔𝑡 , and is rewarded for reducing their distance. A simple progress reward is 𝑟𝑡low = 𝑑(𝑧𝑡 , 𝑔𝑡 ) − 𝑑(𝑧𝑡+1 , 𝑔𝑡 ).

(46)

The high-level action is held for several environment steps(determined by separate switch head), so selecting one target embedding corresponds to a temporally extended action. The two agents can therefore learn different functions: the lowlevel agent learns how to realise changes in the representation space, while the high-level agent learns which changes are informative, reachable and useful for continued exploration. The target space should discard high-frequency details that cannot be controlled over the selected horizon. Slowly varying representations, including representations obtained with SFA-like objectives, are a possible goal space. Long-term empowerment may then be estimated between high-level actions and future states in this compressed representation rather than between individual motor commands and raw observations. Overall process should achieve state-space traversal similar to one in Figure 8. This training process resembles developmental motor learning: initially unstructured self-generated actions establish sensorimotor regularities, after which achieved outcomes can become goals for increasingly directed behaviour. Motor competence and goal selection may then form a coupled curriculum in which each expands the learning opportunities of the other.

43

This experiment tests whether temporal hierarchy addresses two complementary limitations. Skill-discovery objectives such as DIAYN do not specify how independently learned skills should be ordered, while short-horizon empowerment does not represent consequences separated from motor actions by many environment steps. A two-level agent instead treats goal-conditioned behaviour as the low-level action space of a slower decision process. Code availability. Code for the proposed experiments will be published as the implementations are developed at https://github.com/noskill/reinf. 17. Mathematical Appendix 17.0.1. Chain rule for Mutual information 𝐼(𝑋, 𝑌 ; 𝑍) = 𝐼(𝑋; 𝑍) + 𝐼(𝑌 ; 𝑍 ∣ 𝑋)

(47)

By definition we have 𝐼(𝑋, 𝑌 ; 𝑍) = 𝐻(𝑍) − 𝐻(𝑍 ∣ 𝑋, 𝑌 ) Expand 𝐼(𝑋; 𝑍) and 𝐼(𝑌 ; 𝑍 ∣ 𝑋) to entropies: 𝐼(𝑋; 𝑍) = 𝐻(𝑍) − 𝐻(𝑍 ∣ 𝑋) 𝐼(𝑌 ; 𝑍 ∣ 𝑋) = 𝐻(𝑍 ∣ 𝑋) − 𝐻(𝑍 ∣ 𝑋, 𝑌 ) Compute sum: 𝐼(𝑋; 𝑍) + 𝐼(𝑌 ; 𝑍 ∣ 𝑋) = 𝐻(𝑍) − 𝐻(𝑍 ∣ 𝑋) + 𝐻(𝑍 ∣ 𝑋) − 𝐻(𝑍 ∣ 𝑋, 𝑌 ) 𝐼(𝑋; 𝑍) + 𝐼(𝑌 ; 𝑍 ∣ 𝑋) = 𝐻(𝑍) − 𝐻(𝑍 ∣ 𝑋, 𝑌 ) = 𝐼(𝑋, 𝑌 ; 𝑍) 17.0.2. Donsker-Varadhan MI 𝑝(𝑢) We are starting from KL divergence 𝐷𝐾𝐿 (𝑃||𝑄) = ∫ 𝑝(𝑢)𝑙𝑜𝑔 𝑞(𝑢) 𝑑𝑢 Step 1. Define Gibbs density. Let q(u) be a probability density and T(u) be our arbitrary function. We create a new valid probability density, 𝑔(𝑢), by weighting q(u) with 𝑒𝑇 (𝑢) . To ensure g(u) integrates to 1, we must divide by a normalizing constant Z: 𝑔(𝑢) =

𝑒𝑇(𝑢) 𝑞(𝑢) 𝑍

Where the constant Z is just the expected value over the distribution Q: 𝑍 = ∫ 𝑒𝑇 (𝑢) 𝑞(𝑢)𝑑𝑢 = 𝐸𝑄 𝑒𝑇 (𝑢) Step 2: Use the Non-Negativity of KL Divergence 𝑝(𝑢) 𝐷𝐾𝐿 (𝑃||𝐺) = ∫ 𝑝(𝑢)𝑙𝑜𝑔 𝑔(𝑢) 𝑑𝑥 ≥ 0 𝑞(𝑢)

Step 3: Multiply by 𝑞(𝑢) 44

𝑝(𝑢)

𝑝(𝑢) 𝑞(𝑢)

𝑝(𝑢)

𝑞(𝑢)

𝑙𝑜𝑔 𝑔(𝑢) = 𝑙𝑜𝑔 𝑞(𝑢) 𝑔(𝑢) = 𝑙𝑜𝑔 𝑞(𝑢) + 𝑙𝑜𝑔 𝑔(𝑢) 𝑇 (𝑢) 𝑔(𝑢) 𝑒𝑇(𝑢) 𝑞(𝑢) => 𝑞(𝑢) = 𝑒 𝑍 𝑍 𝑞(𝑢) 𝑍 𝑔(𝑢) = 𝑒𝑇(𝑢) 𝑞(𝑢) 𝑙𝑜𝑔 𝑔(𝑢) = 𝑙𝑜𝑔 𝑒𝑇𝑍(𝑢) = 𝑙𝑜𝑔 𝑍 − 𝑇 (𝑢)

𝑔(𝑢) =

Step 4: Substitute and Solve 𝑝(𝑢) 𝑝(𝑢) 𝑞(𝑢) ∫ 𝑝(𝑢) 𝑙𝑜𝑔 𝑞(𝑢) 𝑔(𝑢) 𝑑𝑢 = ∫ 𝑝(𝑢)( log 𝑞(𝑢) + log 𝑍 − 𝑇 (𝑢))𝑑𝑢 ≥ 0 𝑝(𝑢)

First term ∫ 𝑝(𝑢) log 𝑞(𝑢) 𝑑𝑢 is 𝐷𝐾𝐿 (𝑃 || 𝑄) Second term ∫ 𝑝(𝑢) log 𝑍𝑑𝑢 = log 𝑍 ∫ 𝑝(𝑢)𝑑𝑢 = log 𝑍 ∗ 1 = log 𝑍 Third term ∫ 𝑝(𝑢)𝑇 (𝑢) 𝑑𝑢 is 𝐸𝑃 𝑇 (𝑢) 𝐷𝐾𝐿 (𝑃 || 𝑄) + 𝑙𝑜𝑔 𝐸 𝑄 𝑒𝑇 (𝑢) − 𝐸𝑃 𝑇 (𝑢) ≥ 0 𝐷𝐾𝐿 (𝑃 || 𝑄) ≥ 𝐸𝑃 𝑇 (𝑢) − 𝑙𝑜𝑔 𝐸𝑄 𝑒𝑇 (𝑢) By definition: 𝐼(𝑋; 𝑌 ) = 𝐷𝐾𝐿 (𝑃(𝑋, 𝑌 ) ∥ 𝑃(𝑋)𝑃(𝑌 )) We get mutual information(from definition) with P having density p(x, y) - joint density; and Q having density p(x)p(y) - the product of the marginal densities: 𝐼(𝑋; 𝑌 ) ≥ 𝐸𝑃(𝑋,𝑌 ) 𝑇 (𝑥, 𝑦) − 𝑙𝑜𝑔 𝐸𝑃(𝑋)𝑃(𝑌 ) 𝑒𝑇 (𝑥,𝑦) We got lower bound for KL and mutual information. We can optimise this estimator by finding better function 𝑇 e.g. with gradient ascent. References [1] M. Levin, Self-improvising memory: A perspective on memories as agential, dynamically reinterpreting cognitive glue, Entropy 26 (2024) 481. URL: https://doi.org/10.3390/e26060481. doi:10.3390/e26060481. [2] S. Biswas, W. Clawson, M. Levin, Learning in transcriptional network models: Computational discovery of pathway-level memory and effective interventions, International Journal of Molecular Sciences 24 (2023) 285. URL: https://doi.org/10.3390/ijms24010285. doi:10. 3390/ijms24010285. [3] R. Chis-Ciure, M. Levin, Cognition all the way down 2.0: Neuroscience beyond neurons in the diverse intelligence era, Synthese 206 (2025) 257. URL: https://doi.org/10.1007/s11229-025-05319-6. doi:10.1007/ s11229-025-05319-6.

45

[4] M. Levin, Michael levin: Intelligence beyond the brain, YouTube video, Principles of Intelligence, 2022. URL: https://youtu.be/RwEKg5cjkKQ, published September 11, 2022; accessed September 3, 2026. [5] M. Levin, Bioelectricity: A bridge between physics and cognition, by way of biology, YouTube video, Michael Levin’s Academic Content, 2025. URL: https://www.youtube.com/watch?v=GiL6wtg3U0I, published September 23, 2025; accessed September 15, 2026. [6] E. O. Wilson, The Insect Societies, Belknap Press of Harvard University Press, Cambridge, Massachusetts, 1971. [7] T. Vida, Z. T. Calamari, P. Barden, Post K–Pg rise in ant and termite prevalence underlies convergent dietary specialization in mammals, Evolution 79 (2025) 2315–2324. URL: https://pubmed.ncbi.nlm.nih.gov/ 40455576/. doi:10.1093/evolut/qpaf121. [8] L. Bell-Roberts, The Evolution of Division of Labour in Social Insects, Ph.D. thesis, University of Oxford, 2024. URL: https://ora.ox.ac.uk/ objects/uuid:781db1e5-929a-485e-a1bc-5bdf0be99d3a. [9] S. Kriegman, D. Blackiston, M. Levin, J. Bongard, Xenobot—a tall quadruped, Wikimedia Commons, 2020. URL: https://commons. wikimedia.org/wiki/File:Xenobot_-_A_tall_quadruped.jpg, licensed under Creative Commons Attribution 4.0 International (CC BY 4.0); accessed September 15, 2026. [10] H. S. Galpayage Dona, C. Solvi, A. Kowalewska, K. Mäkelä, H. MaBouDi, L. Chittka, Do bumble bees play?, Animal Behaviour 194 (2022) 239–251. URL: https://doi.org/10.1016/j.anbehav.2022.08.013. doi:10.1016/j.anbehav.2022.08.013, creative Commons Attribution 4.0 International License. [11] O. J. Loukola, C. J. Perry, L. Coscos, L. Chittka, Bumblebees show cognitive flexibility by improving on an observed complex behavior, Science 355 (2017) 833–836. URL: https://doi.org/10.1126/science.aag2360. doi:10.1126/science.aag2360. [12] P. K. Y. Chow, T. K. Lehtonen, V. Näreaho, O. J. Loukola, Prior associations affect bumblebees’ generalization performance in a tool-selection task, 46

iScience 25 (2022) 105466. URL: https://doi.org/10.1016/j.isci. 2022.105466. doi:10.1016/j.isci.2022.105466, creative Commons Attribution 4.0 International License. [13] A. Zhou, Y. Du, J. Chen, Ants adjust their tool use strategy in response to foraging risk, Functional Ecology 34 (2020) 2524–2535. URL: https:// doi.org/10.1111/1365-2435.13671. doi:10.1111/1365-2435.13671. [14] J. Peters, Policy gradient methods, Scholarpedia 5 (2010) 3698. URL: http: //www.scholarpedia.org/article/Policy_gradient_methods. doi:10.4249/scholarpedia.3698. [15] A. S. Klyubin, D. Polani, C. L. Nehaniv, Empowerment: A universal agentcentric measure of control, in: 2005 IEEE Congress on Evolutionary Computation, volume 1, IEEE, 2005, pp. 128–135. URL: https://doi.org/ 10.1109/CEC.2005.1554676. doi:10.1109/CEC.2005.1554676. [16] D. P. Kingma, M. Welling, Auto-encoding variational bayes, in: International Conference on Learning Representations, 2014. URL: https: //arxiv.org/abs/1312.6114. doi:10.48550/arXiv.1312.6114. [17] C. Doersch, Tutorial on variational autoencoders, arXiv preprint arXiv:1606.05908 (2016). URL: https://arxiv.org/abs/1606.05908. doi:10.48550/arXiv.1606.05908. [18] J. V. Michalowicz, J. M. Nichols, F. Bucholtz, Handbook of Differential Entropy, Chapman and Hall/CRC, Boca Raton, Florida, 2013. doi:10.1201/ b15991. [19] Y. Du, S. Tiomkin, E. Kiciman, D. Polani, P. Abbeel, A. Dragan, AvE: Assistance via empowerment, 2020. URL: https://arxiv.org/abs/2006. 14796. doi:10.48550/arXiv.2006.14796. arXiv:2006.14796. [20] S. P. Diggle, A. Gardner, S. A. West, A. S. Griffin, Evolutionary theory of bacterial quorum sensing: When is a signal not a signal?, Philosophical Transactions of the Royal Society B: Biological Sciences 362 (2007) 1241–1249. URL: https://pmc.ncbi.nlm.nih.gov/articles/ PMC2435587/. doi:10.1098/rstb.2007.2049. [21] A. Prindle, J. Liu, M. Asally, S. Ly, J. Garcia-Ojalvo, G. M. Süel, Ion channels enable electrical communication in bacterial communities, 47

Nature 527 (2015) 59–63. URL: https://www.nature.com/articles/ nature15709. doi:10.1038/nature15709. [22] J. Bódis, J. Berke, B. Nagy, I. Gulyas, P. Hersics, Á. Várnagy, K. Kovács, Ultra-weak photon emission: From oxidative metabolism to DNA-based communication: A review of biochemical, biophysical and quantum biological perspectives, Frontiers in Endocrinology 17 (2026) 1861061. URL: https://www.frontiersin.org/journals/ endocrinology/articles/10.3389/fendo.2026.1861061/full. doi:10.3389/fendo.2026.1861061. [23] M. I. Belghazi, A. Baratin, S. Rajeshwar, S. Ozair, Y. Bengio, A. Courville, R. D. Hjelm, Mutual information neural estimation, in: Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, PMLR, 2018, pp. 531–540. URL: https://proceedings.mlr.press/v80/belghazi18a.html. [24] A. van den Oord, Y. Li, O. Vinyals, Representation learning with contrastive predictive coding, arXiv preprint arXiv:1807.03748 (2018). URL: https: //arxiv.org/abs/1807.03748. doi:10.48550/arXiv.1807.03748. [25] D. Hafner, J. Pasukonis, J. Ba, T. Lillicrap, Mastering diverse domains through world models, arXiv preprint arXiv:2301.04104 (2023). URL: https://arxiv.org/abs/2301.04104. doi:10.48550/ arXiv.2301.04104. [26] M. Burchi, R. Timofte, Learning transformer-based world models with contrastive predictive coding, arXiv preprint arXiv:2503.04416 (2025). URL: https://arxiv.org/abs/2503.04416. doi:10.48550/ arXiv.2503.04416. [27] B. Eysenbach, A. Gupta, J. Ibarz, S. Levine, Diversity is all you need: Learning skills without a reward function, arXiv preprint arXiv:1802.06070 (2018). URL: https://arxiv.org/abs/1802.06070. doi:10.48550/arXiv.1802.06070. [28] J. Schmidhuber, Formal theory of creativity & fun & intrinsic motivation (1990–2010), IDSIA webpage, 2010. URL: https://people.idsia.ch/ ~juergen/creativity.html, accessed September 15, 2026.

48

[29] J. Schmidhuber, Formal theory of creativity, fun, and intrinsic motivation (1990–2010), IEEE Transactions on Autonomous Mental Development 2 (2010) 230–247. URL: https://doi.org/10.1109/TAMD.2010. 2056368. doi:10.1109/TAMD.2010.2056368. [30] L. Wiskott, T. J. Sejnowski, Slow feature analysis: Unsupervised learning of invariances, Neural Computation 14 (2002) 715–770. doi:10.1162/ 089976602317318938. [31] A. Adams, H. Zenil, P. C. W. Davies, S. I. Walker, Formal definitions of unbounded evolution and innovation reveal universal mechanisms for open-ended evolution in dynamical systems, Scientific Reports 7 (2017) 997. URL: https://www.nature.com/articles/ s41598-017-00810-8. doi:10.1038/s41598-017-00810-8.

49

Record · ID 919439 · SHA-256 9bb08d9760fd362d
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.