ConceptioArchivearXiv CS
arXiv CSopen access

SKooP: Symmetric Koopman Predictions for Faster and More Generalizable Legged Robot Locomotion with Reinforcement Learning

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

SKooP: Symmetric Koopman Predictions for Faster and More Generalizable Legged Robot Locomotion with Reinforcement Learning

arXiv:2607.11624v1 [cs.RO] 13 Jul 2026

Evelyn D’Elia1,2 , Weishu Zhan2 , Giulio Turrisi3 , Giulio Romualdi4 , Giuseppe L’Erario4 , Raffaello Camoriano5,6 , Wei Pan7 , and Daniele Pucci4 Abstract— Reinforcement learning (RL) algorithms classically suffer from poor sample efficiency. In robotics, a recent line of work has emerged addressing this problem by encoding physics priors in the learning process. However, most of these approaches are validated on well-defined, low-dimensional benchmark systems rather than high-dimensional robots with complex nonlinear dynamics. In this paper, we introduce SKooP (Symmetric Koopman Predictions), an approach combining the advantages of morphological symmetries with those of a Koopman model learned via autoencoder to enhance policy learning. SKooP learns a Koopman model of the system dynamics alongside the policy. The resulting Koopman predictions are used as privileged observations for the critic, allowing the agent to learn based on smoother, more informative features. We also incorporate group symmetries into the actor, critic, encoder and decoder networks to produce a highly equivariant policy. The SKooP approach is validated via in-depth analysis of the learned Koopman models and symmetric policies to showcase how each of these influences the agent’s performance. We also show that the learned policies are transferable to different simulation environments. Our results show that SKooP consistently reduces convergence time and increases the learned reward for multiple challenging bipedal locomotion tasks on a quadruped robot. Project page: https://evelyd.github. io/SymmetricKoopmanPredictions/

I. INTRODUCTION The use of reinforcement learning approaches for robotic control has skyrocketed in recent years, mainly thanks to advances in computational efficiency and simulation fidelity. However, for systems such as legged robots which have complex, nonlinear dynamics, learning effective policies is still a challenging and active research area. Although powerful, purely data-driven model-free RL approaches rely on costly *This study was carried out within the FAIR - Future Artificial Intelligence Research and received funding from the European Union Next-GenerationEU (PIANO NAZIONALE DI RIPRESA E RESILIENZA (PNRR) – MISSIONE 4 COMPONENTE 2, INVESTIMENTO 1.3 – D.D. 1555 11/10/2022, PE00000013). This manuscript reflects only the authors’ views and opinions, neither the European Union nor the European Commission can be considered responsible for them. 1 IIT@MIT, Italian Institute of Technology (IIT), 16163 Genoa, Italy

[email protected] 2 Machine Learning and Optimisation, University of Manchester, M13 9PL Manchester, U.K. 3 Dynamic Legged Systems Laboratory, IIT, 16163 Genoa, Italy 4 Generative Bionics S.R.L, 16163 Genoa, Italy 5 DAUIN, Politecnico di Torino, 10129 Turin, Italy 6 Rehab Technologies Lab, IIT, 16163 Genoa, Italy 7 School of Engineering, Newcastle University, NE1 7RU Newcastle upon Tyne, U.K. E. D’Elia, G. Romualdi, G. L’Erario, and D. Pucci contributed to this work while at the Artificial and Mechanical Intelligence lab, IIT, Italy. W. Pan contributed to this work while at the Machine Learning and Optimisation group, University of Manchester, U.K.

Fig. 1. Comparison of SKooP performance on trained vs. mirrored push door task. Top row: right-opening door (training task). Bottom row: leftopening door (unseen symmetric task).

trial and error. Conversely, more traditional model-based control approaches exploit the known physics of the system, but typically require expert knowledge and extensive manual tuning. Fusing model-based control with RL holds promise for overcoming their respective limitations and devising more efficient and effective robot control methods. Model-based control methods employ a complete or reduced description of robot dynamics. Using the complete dynamics model requires explicit design and precise modeling, but produces a high-fidelity result. Instead, reduced or simplified model-based approaches, e.g., with inverted pendulum [1] or centroidal dynamics [2], [3] models, improve computational complexity while sacrificing fidelity. Reinforcement learning, conversely, can be model-free thanks to the parallelized, high-fidelity simulators that are nowadays available, and recent works show its success in learning legged robot tasks [4], [5]. Nonetheless, model-free strategies suffer from low sample efficiency, lack of generalization, and convergence to suboptimal policies. In recent years, injecting physics priors into the RL training process has emerged as a strategy to mitigate the drawbacks of the two separate approaches. Here, we utilize two specific types of priors: symmetry information and linearized dynamics in the Koopman-lifted space [6]. The use of symmetries as priors was first introduced for supervised deep learning, namely through the imposition of symmetry constraints on learning architectures [7], [8], [9]. In robot learning, a handful of works incorporate symmetries into the RL process via data augmentation [10], [11]. The RL framework proposed in [11] constrains the learned policy to be equivariant. Encoding symmetry information in this way reduces the number of samples required for training

and yields more generalizable policies, but introducing such constraints can negatively affect performance. The other type of prior we consider is based on Koopman theory [6], which states that for any nonlinear system there exists an infinite-dimensional function space in which the system dynamics are globally, not just locally, linear. Early data-driven methods found finite approximations of the Koopman operator, such as via dynamic mode decomposition (DMD) [12] and DMD with control (DMDc), which extends the theory to controlled systems to produce a controlled Koopman model [13]. More recent works approximate both the lifting functions and the linear embedding with kernel methods or autoencoders [14], [15]. It has been shown that the controlled Koopman system can be effectively regulated with linear model predictive control (MPC) [16], [17]. Some recent works also utilize Koopman priors for RL. For instance, [18], [19] use the Koopman operator to facilitate learning optimal controllers. [20] presents two strategies in which the value function constitutes a linear mapping of the Koopman observables, and shows these methods to have state-of-the-art policy-learning performance on benchmark systems such as the Lorenz model [21]. Another approach, proposed in [22], trains the Koopman autoencoder in parallel with Proximal Policy Optimization (PPO) and learns the policy directly from the encoded Koopman state, achieving modest performance improvements on RL benchmark tasks. Although the results are relevant from a theoretical perspective, they are validated only on well-defined and lowdimensional systems. In robotics, high-dimensional state spaces can be lifted and linearized at relatively low cost with an autoencoder, bypassing the hand-design of lifting functions. Many of the Koopman-based approaches for robotics specifically focus on characterizing the dynamics [23] or tailor the linearization of the system for use in optimal control frameworks such as MPC [24]. Some approaches use the Koopman operator with RL to learn the optimal control tuning parameters rather than the actions directly, e.g., [25]. However, none of these robotic control approaches directly inject the Koopman prediction into the policy learning process. A key motivation for our work is that Koopman theory can facilitate the identification and exploitation of symmetries. For example, [26] improves the sample efficiency of an offline RL algorithm by using the Koopman linearization to more effectively enforce symmetries in the system via data augmentation, validating the approach on RL benchmark tasks. In [23], a system’s morphological symmetries are predefined and embedded in an equivariant Koopman autoencoder to more efficiently learn a global linear model. The authors present a harmonic analysis of quadruped locomotion modes, but do not address the control of the robot. In this work, we propose SKooP (Symmetric Koopman Predictions) an approach which combines the advantages of a controlled Koopman embedding and morphological symmetry priors to enable faster initial policy convergence and improve learned behavior quality and generalization. By learning a symmetric linearized latent space concurrently

with the policy, we ensure that the lifting function is trained on a relevant state space. Then, by providing the critic with information about the future in the form of Koopmanpredicted latent states, we enable it to learn higher-quality motions. Specifically, our contributions are as follows: • A control- and symmetry-informed Koopman autoencoder to predict the future state, trained concurrently with the RL policy. • A modular, symmetric Koopman-based RL architecture, SKooP, which learns the Koopman embedding and employs the lifted state prediction as a prior for learning the value function. • Detailed ablation analysis and performance validation on multiple challenging two-legged locomotion tasks for a quadruped robot, in two different simulation environments. Our results show that SKooP improves initial reward convergence rate and symmetry of learned motions compared with standard PPO, and that it generalizes better than a stateof-the-art symmetric RL method to unseen, mirrored tasks. II. BACKGROUND A. Notation g ∈ G denotes a group element within a group representation. Specifically, C2 := {e, gs |gs2 = e} is a reflection group; gs , e are the reflection and identity transformations. • ρX , ρU , ρY are matrix representations of a symmetry ′ group acting on the state space X ⊂ Rn , the control ′ m space U ⊂ R , and the output space Y ⊂ Rn ×n , respectively. • K is the Koopman operator, which acts on observables z := ψ(x) := [ψ1 (x), . . . , ψn (x)], where ψi (x) ∈ F(X ), to predict the system dynamics F (xk , uk ) in the lifted space. ′ n×n • A ∈ R , B ∈ Rn×m , C ∈ Rn ×n are linear system matrices and H ∈ N is the look-ahead horizon length. • q, q̇ are joint positions and velocities. • θ, ξ, ϕ represent learned network parameters. • Lsr , Llp , Lsp are the standard autoencoder losses: state reconstruction, latent state prediction, and state prediction, respectively. •

B. Markov Decision Processes and Bellman Equation Markov Decision Processes (MDPs) are used in optimal control and RL for modeling transition-based decisions. An MDP is represented by the tuple (X , U , P, r, γ) where P : X × U × X 7→ [0, 1] is the transition probability, r : X × U × X 7→ R is the reward function, and γ ∈ [0, 1] is the discount factor. RL aims to find an optimal policy π ∗ that maximizes the Bellman equation: X Vπ (x) = P(x′ |x, π(x)) [r(x, π(x), x′ ) + γVπ (x′ )] , x′ ∈X

(1) where Vπ (x) is the value of state x using policy π.

C. Morphological Symmetries In this work, we take advantage of the inherent symmetries of the state and action spaces of the robot. Most legged robots possess C2 symmetry over the sagittal plane, allowing information such as joint positions and poses to be mirrored over this plane. A symmetry acting on a point x ∈ X can be interpreted as a matrix-vector multiplication with g ▷ x := ρX (g)x ∈ X . A map f is considered group invariant, or G-invariant, if its output does not change under any group transformation of the input, i.e., f (ρX (g)x) = f (x), ∀g ∈ G, while it is Gequivariant if applying a group transformation to the input and then evaluating f yields the same result as applying the corresponding group transformation to the output, i.e., f (ρX (g)x) = ρY (g)f (x), ∀g ∈ G. In the context of MDPs, the state and action group representations ρX and ρU can be defined as symmetric representations given the morphological symmetries of the system. This definition proves useful in the context of RL, where it makes the policy π : X 7→ U G-equivariant and the reward function r : X × U 7→ R G-invariant. Thus, the value function Vπ : X 7→ R is also G-invariant [27]. An unknown target function to approximate via a neural network can be expressed as fϕ ∈ F : X 7→ Y. To impose symmetry constraints, we can require that the learned function fϕ be G-invariant or G-equivariant: FGinv := {fϕ ∈ F|y = fϕ (g ▷ x), ∀g ∈ G}, FGeq := {fϕ ∈ F|g ▷ y = fϕ (g ▷ x), ∀g ∈ G}.

Actor

Prioritized Replay Buffer

Environment

Critic

Actor Critic with Privileged Koopman Observations

DAE Loss Symmetric Koopman Dynamics Autoencoder

Fig. 2.

Schematic representation of SKooP.

utilizes a symmetric Koopman dynamics autoencoder to learn a lifted space in which the dynamics are linear and system symmetries are respected. The inclusion of a Koopman prediction as part of the critic input in an actor-critic framework enables the critic to more easily learn a precise value function. This is thanks to the linear representation of the system dynamics that the Koopman model provides, which simplifies the critic’s learning problem of predicting the expected future return. The overall scheme of the proposed approach is displayed in Figure 2.

(2) A. Equivariant Networks

D. Koopman Operator The Koopman operator offers a linear representation of the nonlinear system dynamics xk+1 = F (xk ): Kψ(xk ) := ψ(F (xk )) = ψ(xk+1 ), zk := ψ(xk ).

(3)

A key challenge in finding a Koopman model is to design the lifting function ψ(x). Given that we must approximate an infinite-dimensional Koopman operator, the latent space F(X ) is much higher-dimensional than the state space X [28]. Moreover, [16] proposes an extension of the Koopman operator approximation for controlled systems xk+1 = F (xk , uk ): zk+1 = Azk + Buk , (4) xk = Czk . For a controlled system, the linear representation is supported on the policy control set. As such, finding an approximation of the Koopman operator K amounts to finding [A, B]. Extending Koopman theory to the controlled case opens up new possibilities for its use in optimal control and RL to provide predictive priors.

Our approach implements symmetry priors in two ways. The first way is via the structure of the actor and critic networks. We constrain the actor to be G-equivariant and the critic to be G-invariant, following the assertions made for MDPs in Section II-C. We do so by predefining the symmetric state and action representations ρX and ρU . For example, in the robot’s joint space, due to the C2 symmetry, a state (q, q̇) can be reflected to obtain the symmetric state (gs ▷ q, gs ▷ q̇) [27]. This actor-critic architecture is similar to that proposed in [11], but we extend it by defining the symmetries for the privileged latent-state observations. We also apply symmetry priors to the latent state. The Koopman autoencoder learns to encode and decode the Koopman state using the symmetry information for the state xk and action uk . This is made possible by constraining the encoder fϕ and decoder fϕ−1 to be G-equivariant. Using the notation from Section II-C, the symmetry rules for the actor, critic, encoder, and decoder are summarized as: πθ (g ▷ x) = g ▷ π(x), Vξ (g ▷ x) = Vξ (x),

III. METHODS

fϕ (g ▷ x) = g ▷ z,

In this Section, we present our novel, modular, and online SKooP approach that exploits symmetries and Koopman theory to enhance the sample efficiency and improve the performance of actor-critic RL algorithms. Our approach

fϕ−1 (g ▷ z) = g ▷ x.

(5)

Structurally, all these networks are termed equivariant multilayer perceptrons, or EMLPs.

B. Koopman Autoencoder with Symmetries To learn a lifting function, we train an equivariant autoencoder concurrently with the agent. This simultaneous learning strategy avoids pretraining and ensures training data are relevant to the state and action space regions explored by the agent since the control policy changes over the course of the training. We adapt the Equivariant Dynamics Auto-Encoder (eDAE) structure proposed in [23], which is based on the original DAE [15], to learn the dynamics of a controlled system, which we term an Equivariant Controlled Dynamics Auto-Encoder (ecDAE). Thus, the linearized dynamics take the form of Equation (4). We now represent the lifting function ψ(x) as fϕ and the state reconstruction matrix C as fϕ−1 to show that they are parameterized as neural networks in the autoencoder framework. We also employ a multi-step Koopman prediction, so we can adapt Equation (4) to reflect these specifics of our approach, as follows: H−1 X zk+H = AH fϕ (xk ) + AH−1−i Buk+i , (6) i=0 xk = fϕ−1 (zk ). We use standard losses for training the autoencoders: Lsr := ∥xk − fϕ−1 (fϕ (xk )) ∥22 , Llp := ∥fϕ (xk+H ) − KH fϕ (xk )∥22 ,  Lsp := ∥xk+H − fϕ−1 KH fϕ (xk ) ∥22 .

(7)

Equation (6) shows how the control input fits into the autoencoder structure, but does not address the symmetries. fϕ (x) = z and πθ (x) = u are constrained to be Gequivariant by ensuring that the basis functions are Gequivariant. Thus F (X ) becomes a symmetric function space, which Lp can be decomposed into isotypic subspaces F(X ) := i=1 Fi (X ), where p is the number of distinct irreducible representations of G [27]. This means that A and B are block-diagonal with p blocks. Essentially, Equation (7) learns a Koopman operator Ki for each of these subspaces, and symmetry constraints are preserved throughout. C. Latent State Privileged Observations In standard actor-critic frameworks, the state x is used as input to both the actor and the critic. Our actor-critic method is asymmetric (e.g., [29]), so the critic receives a single privileged Koopman state prediction zk+1 = Afϕ (xk )+Buk learned by the ecDAE. Since the actor only requires xk as input, this architectural choice informs training without increasing computational overhead at deployment time. The autoencoder is trained to predict the linear dynamics of the state z over an H-step horizon. This property means it provides the critic with an informative predictive prior. D. Online DAE Training To avoid costly pretraining and to ensure that the explored state space is represented in the learned Koopman model, we train the ecDAE concurrently with the RL policy. At the beginning of training, rollouts of random actions are

performed to fill a replay buffer with training data for the DAE. This buffer is distinct from the buffer used for training the actor and critic in PPO. At each RL training step, an ecDAE training step is also executed. We choose to store and select training data from a Prioritized Experience Replay (PER) [30] buffer D(xk , uk , xk+1 ) for each iteration. PER prioritizes keeping buffer data which results in high loss. IV. EXPERIMENTAL SETUP We now show that SKooP is modular and succeeds in improving performance, sample efficiency, and generalization to unseen scenarios. We present the results of training SKooP on multiple challenging bipedal locomotion tasks to empirically evaluate its capabilities. In our experiments, we adopt vanilla PPO as an actor-critic algorithm due to its reliability and widespread use. However, note that SKooP is compatible with any actor-critic algorithm. We test on multiple bipedal locomotion tasks inspired by [31] for the Cyberdog 2 quadruped [32]. We consider 3 tasks, all of which require the quadruped to transition into an upright, bipedal position and walk on its hind legs: stand dance, walk slope, and push door. Details on task parameters are available in Table VI (Appendix). We employ the massively parallel Isaac Gym [33] physics simulator and execute our experiments on an NVIDIA A100 GPU. Averaged over 100 iterations, a single training iteration of SKooP, PPO, and PPOeqic respectively for the stand dance task takes 2.651 s, 1.978 s, and 2.245 s. PPOeqic takes 13.5% longer than baseline PPO, and SKooP takes 34.0% longer. To apply our symmetry and Koopman techniques, we utilize the open source libraries DynamicsHarmonicsAnalysis [23], MorphoSymm [27], and ESCNN [34]. We employ a horizon of H = 5 for our ecDAE training. We use a PER buffer with 10000 data points to train the autoencoder. Each task is trained with 8192 parallel environments, over 5 random seeds. The dimensionality ratio of the latent state z to the state x is 3. The stand dance and walk slope tasks are trained for 30000 iterations and have a state dimension of 47, while push door is trained for 18000 iterations and has a state dimension of 43. All tasks use an action space of size 12, equal to the number of controllable joints. For further details on the task, refer to [11]. V. RESULTS We define and adopt several standard metrics to evaluate our approach, including the mean episode return, training value loss, success rate (SR), symmetry index (SI), stability, and controllability. We also evaluate SR and SI when starting in out-of-distribution (OOD) initial states. Additionally, we provide an ablation study to isolate the effects of the symmetry constraints, latent state critic input, and Koopman prediction prior. Examples of SKooP policy performance on the considered tasks are available in the accompanying video. The nomenclature we use for the ablation is as follows: PPO refers to vanilla PPO, PPOeqic denotes equivariant actor and invariant critic without using the autoencoder (presented in [11]), and SKooP denotes our complete approach.

TABLE I

Mean Episode Return

P USH DOOR ABLATION STUDY, SHOWING (% MEAN ± STD ). Method

Right SR ↑

Left SR ↑

SI ↓

OOD Right SR ↑

OOD Left SR ↑

OOD SI ↓

PPO PPOeqic

70.01 ± 6.65 39.65 ± 5.07

0.00 ± 0.00 39.12 ± 6.31

199.99 ± 0.02 13.82 ± 6.88

42.27 ± 7.65 21.43 ± 4.44

0.01 ± 0.02 21.72 ± 4.10

199.94 ± 0.12 17.68 ± 13.44

SKooP-NoSym-NoPred SKooP-NoSym SKooP-NoPred

84.18 ± 4.31 83.36 ± 3.86 74.37 ± 12.88

0.02 ± 0.02 0.05 ± 0.02 68.45 ± 11.32

199.91 ± 0.10 199.74 ± 0.09 8.03 ± 6.89

45.95 ± 6.56 41.23 ± 4.84 44.94 ± 7.61

0.02 ± 0.02 0.01 ± 0.01 43.43 ± 5.90

199.82 ± 0.16 199.89 ± 0.08 9.14 ± 10.57

SKooP

82.77 ± 2.96

77.70 ± 6.66

8.49 ± 4.26

47.65 ± 2.59

46.18 ± 4.37

9.06 ± 9.24

Stand Dance

Walk Slope

30

6

15

10

5000

10000

15000

20000

25000

30000

10

4

5

2

0 0

Training Iterations PPO SKooP-NoSym-NoPred

Fig. 3.

8

20

20

0 0

Push Door

5000

10000

15000

20000

25000

30000

Training Iterations SKoop-NoSym PPOeqic

0 0

2500

5000

7500 10000 12500 15000 17500

Training Iterations SKooP-NoPred SKooP

Ablation of training reward applied to the following bipedal tasks: dancing, climbing a slope, and opening a door, each averaged over 5 seeds.

SKooP-NoSym-NoPred is our approach without symmetries (using the controlled autoencoder, cDAE) and using zk instead of the prediction zk+1 , SKooP-NoPred also uses zk , but is with symmetries, while SKooP-NoSym uses the prediction zk+1 and no symmetries. A. Policy Symmetry Evaluation We select the push door task as a means of evaluating the generalization of the policy learned by our method. To do this, we collect the SR, SI, OOD SR, and OOD SI metrics over 10000 episodes, averaged over 5 seeds per setting. We define a successful episode as one in which the agent manages to open the door to at least 60◦ and walk through it. The SR is the percentage of episodes in which the agent succeeds. The SI is defined using the SR. R −XL | Specifically, SI := 2|X XR +XL × 100%, where XR is the SR for a scenario where the door opens on the right, while XL is the SR for the mirrored task which differs only in that the door opens on the left. To evaluate OOD performance, the initial base orientation of the robot is sampled uniformly at random within ±15◦ around each axis compared to the trained initial orientation. These metrics are similar to those used in [11]. All models are trained only in the scenario with the right-opening door, while they are tested with both right- and left-opening doors to evaluate generalization to symmetric tasks. The results are reported in Table I. 1) Success Rate (SR) and Symmetry Index (SI): We find that all approaches that do not employ symmetry constraints, i.e., PPO, SKooP-NoSym-NoPred, and SKooP-NoSym, exhibit a very high SI, meaning that they do not generalize at all to the mirrored task (left-opening door). SKooP-NoSym and SKooP-NoSym-NoPred display the highest average SR for the right-side training task, although not significantly higher than SKooP. However, they systematically fail on the left-side task. Conversely, PPOeqic, which constrains the actor and critic with symmetry priors, achieves a diminished right-side SR compared to the baseline PPO, but generalizes well to

TABLE II stand dance STEADY- STATE GAIT METRICS AVERAGED OVER 5 SEEDS . Method

Mean Phase MAE (rad)

Mean ∆θ (rad)

Mean ∆θdiff (rad)

PPO PPOeqic SKooP-NoSym SKooP

0.216 ± 0.006 0.136 ± 0.012 0.521 ± 0.179 0.170 ± 0.015

0.761 ± 0.027 0.800 ± 0.061 0.854 ± 0.222 0.813 ± 0.049

0.100 ± 0.042 0.103 ± 0.000 0.237 ± 0.142 0.137 ± 0.054

TABLE III S UMMARY OF LEARNED KOOPMAN MODEL STABILITY AND CONTROLLABILITY METRICS , AVERAGED OVER 5 SEEDS . Task

DAE Type

|λ(A)|max

κ(C)

rank(C)/Total

Stand Dance

cDAE ecDAE

0.557 ± 0.012 0.613 ± 0.019

(1.61 ± 1.16)e+09 (1.99 ± 0.48)e+09

(69.4 ± 2.9)/141 (55.8 ± 1.8)/141

Walk Slope

cDAE ecDAE

0.551 ± 0.023 1.013 ± 0.018

(1.77 ± 0.89)e+09 (1.26 ± 2.41)e+11

(68.6 ± 3.1)/141 (45.2 ± 13.5)/141

Push Door

cDAE ecDAE

0.572 ± 0.017 1.000 ± 0.003

(3.96 ± 0.70)e+08 (8.17 ± 9.07)e+09

(71.0 ± 1.4)/129 (54.6 ± 1.5)/129

the mirrored task. Both SKooP-NoPred and SKooP enforce symmetry priors on the actor and critic in the same way that PPOeqic does. SKooP-NoPred and SKooP additionally enforce the symmetry constraints in the ecDAE, leading to a latent space with embedded symmetries. The results show that the use of symmetry priors in the privileged zk or zk+1 input enables significantly higher SRs on both the right and left sides, and lower SI, with respect to PPOeqic. Compared to SKooP-NoPred, SKooP yields stronger performance across the board and similarly low SI, demonstrating the benefit of the Koopman prediction zk+1 . We note that the results for SKooP also exhibit lower variance compared to the other methods. Figure 1 visually compares the SKooP method on the original and mirrored door-opening tasks, showing similar behavior in both. The takeaways from these results are that, by providing a Koopman state prediction to the critic, and by constraining the actor, critic, and autoencoder with symmetry information,

PPO - Rear Hip Pitch

PPO - Rear Knee 2

−1 −2

1

PPOeqic - Rear Hip Pitch

PPOeqic - Rear Knee

Joint Pos (rad)

−1

2

−2

1

SKooP-NoSym - Rear Hip Pitch

0

SKooP-NoSym - Rear Knee 2

−2

1

SKooP - Rear Hip Pitch

SKooP - Rear Knee

−1

Left Rear Knee

2

−2

Right Rear Knee

1 0

100

200

300

400

500

0

100

200

Timestep

Fig. 4.

300

400

500

Timestep

Comparison of stand dance rear leg joint angles for PPO, PPOeqic, SKooP-NoSym, and SKooP.

we enable the agent to learn not just a more successful policy, but also one that can generalize to a previously unseen, mirrored task. 2) OOD Performance: When tested on the OOD initial states that we define, we observe that while SKooP-NoSymNoPred and SKooP-NoSym still perform comparably with PPO on the right-hand opening task, they also similarly fail on the mirrored task. In contrast, PPOeqic, SKooP-NoPred, and SKooP consistently perform highly symmetrically even on OOD tasks, as revealed by the low SI values. Notably, in the OOD case SKooP significantly outperforms PPOeqic for both SR and SI, also achieving stronger mean performance and lower variance with respect to SKoop-NoPred. 3) Joint Trajectory Analysis: We also examine a sample of joint trajectories for the rear legs in the stand dance task (see Figure 4). Symmetric motion of the rear legs indicates greater stability of the robot. In the plots, the motion is symmetric if the right and left joint trajectories oscillate in similar ranges. The results show that SKooP produces the most symmetric behavior. Instead, PPO learns to command larger angles for the right rear knee than for the left. Table II reports the phase mean absolute error (MAE), the mean step amplitude ∆θ, and the amplitude difference between left and right joints ∆θdiff , P for the considered methods. T The phase MAE is defined as T1 k=0 |qL (k) − qR (k + τ )|, where T is the total number of timesteps, qL , qR are the left and right joint positions of a given joint, and τ = 1/(2fd ∆t) is the expected phase offset assuming that the ideal is a 180◦ offset, given the desired gait frequency fd = 2.5Hz and timestep ∆t. SKooP-NoSym learns significantly different ranges for the left and right rear hip pitch joints, which is expected since it does not have access to symmetry priors. Table II shows that SKooP rectifies this with the use of symmetry priors, reducing the mean phase MAE by 96.7% and ∆θdiff by 42.3% in comparison. While PPOeqic learns more visually symmetric right and left joint ranges, the left rear knee consistently has a larger maximum angle than the right one. Furthermore, SKooP learns joint trajectories that oscillate within the same range for both the left and right joints. Finally, Table II also shows that SKooP achieves

larger and more consistent ∆θ across seeds than PPOeqic. B. Training Performance Across the three tasks, Figure 3 shows that the SKooP, SKooP-NoPred, and PPOeqic approaches incorporating symmetries have steeper initial convergence rates than competitors. We note that across the tasks, the -NoPred methods which learn with zk converge slightly faster than their counterparts which learn with the Koopman prediction zk+1 . This is explained by the fact that during the initial learning phase the autoencoder is exposed to a wide variety of states and actions and has not yet converged. Using only the lifting function to get zk introduces less error in this phase than using both the lifting function and the Koopman model prediction to get zk+1 . For the walk slope task, incorporating the latent state (predicted or not) is sufficient to improve the convergence rate without the use of symmetries, while symmetries provide further improvement. Overall, SKooP achieves higher returns than PPO and PPOeqic. The easiest task to learn, stand dance, exhibits comparably small improvements from the use of Koopman prediction and symmetry priors. Instead, on the most challenging, asymmetric task, push door, PPOeqic provides faster initial convergence than PPO while achieving similar reward. Interestingly, for push door, SKooP-NoSym and SKooPNoSym-NoPred learn to maximize rewards and become highly specialized. However, as discussed in Section V-A.1, they fail to generalize to the mirrored task. On push door, SKooP significantly outperforms PPO and PPOeqic in terms of reward values and converges faster than SKooP-NoPred. C. Stability and Controllability of the Koopman Model A discrete-time linear system with eigenvalues λ(A) closer to the center of the unit circle has dynamics that quickly decay to an equilibrium point. Conversely, at the edge of the unit circle, eigenvalues are marginally stable, meaning the system has dynamics that are time invariant. Remarkably, Table III and Figure 5 show that the use of symmetry constraints always produces a larger |λ(A)|max , with low variance across seeds. In fact, in two out of three

Imag

cDAE λ(A) for stand dance

ecDAE λ(A) for stand dance

1.0

1.0

0.5

0.5

0.0

0.0

−0.5

−0.5

−1.0

−1.0 −1

0

1

−1

Imag

cDAE λ(A) for walk slope 1.0

1.0

0.5

0.5

0.0

0.0

−0.5

−0.5

−1.0 0

1

−1

cDAE λ(A) for push door

Imag

1

−1.0 −1

1.0

0.5

0.5

0.0

0.0

−0.5

−0.5

−1.0

−1.0 −1

0 Real

0

1

ecDAE λ(A) for push door

1.0

Fig. 5.

0

ecDAE λ(A) for walk slope

1

D. Sim-to-sim −1

0 Real

1

Eigenvalues of selected learned Koopman state matrices A.

TABLE IV MSE COMPARISON OF C DAE AND THE SYMMETRY- CONSTRAINED EC DAE FOR push door, WITH ABLATION OF INVARIANT MODE , AVERAGED OVER 5 SEEDS .

Metric Lsr MSE 1-Step Llp Latent MSE 5-Step Lsp MSE (Right) 5-Step Lsp MSE (Left) Invariant λ

cDAE

ecDAE

1.503 ± 0.153 0.283 ± 0.014 2.144 ± 0.078 3.150 ± 0.143 N/A

1.245 ± 0.152 0.274 ± 0.017 1.778 ± 0.153 1.848 ± 0.196 1.000 ± 0.003

Ablation Ablated Lsp MSE (Right) Ablated Lsp MSE (Left) Only Inv |λ| Lsp MSE (Right) Only Inv |λ| Lsp MSE (Left)

autoencoder state reconstruction, 1-step latent prediction, and H-step state prediction errors for a horizon H = 5 (Equation (7)) for cDAE and ecDAE trained for push door. The table also shows the value of the invariant mode (if any) and compares the state prediction error on the trained (right) and mirrored task (left) both with the invariant mode enabled, disabled, and with only the invariant mode enabled. This analysis shows that while the symmetry constraints may have no significant benefit for latent state prediction, they improve state reconstruction and state prediction. The use of symmetry priors also enables the ecDAE to achieve almost identical state prediction on the mirrored task as it does on the trained task, whereas the cDAE’s state prediction is around 50% worse on the mirrored task. Furthermore, removing the ecDAE’s invariant mode significantly increases Lsp MSE, as does disabling all the other modes except for the invariant one. Since only the ecDAE learns this |λ| ≈ 1, it must help represent the invariance of the system dynamics across symmetry groups. Unlike the other tasks, stand dance does not require significant base displacement, which may explain why for that task the ecDAE does not identify an invariant mode to represent the symmetry.

N/A N/A N/A N/A

3.054 ± 1.047 3.524 ± 1.403 3.186 ± 1.265 3.555 ± 1.549

tasks, the ecDAE learns exactly one eigenvalue |λ(A)| ≈ 1.0. Table III further reports numerical values related to the controllability, which measures how well the controller can reach goal states in the state space of the Koopman model. This can be evaluated using  the singular values σ of the controllability matrix C = B AB · · · An−1 B . Across all tasks and seeds, the ecDAE has a larger condition number κ(C) = σmax (C)/σmin (C) and lower controllability rank rank(C) than its cDAE counterpart. This suggests that the ecDAE learns a more compact representation, by recognizing that symmetric modes cannot be controlled independently. The larger condition number shows that the ecDAE is forcing redundant control directions towards 0. In order to better understand the performance and behavior of the cDAE vs. the ecDAE, Table IV reports the average

To validate the robustness of the learned SKooP policies, we evaluate sim-to-sim transfer from the training simulator, Isaac Gym, to Mujoco. Unlike Isaac Gym, which trades some physical accuracy for massively parallelized training efficiency, Mujoco’s high-fidelity solver better models realworld joint and contact physics. Table V compares the command tracking MSE for the stand dance task across both simulators. Because this task is governed by linear x, y and angular z velocity commands, we report the linear vxy and angular ωz tracking errors. To test robustness, we evaluate conditions at the edge of the agent’s training distribution. The training ranges are: friction coefficient [1.0, 3.0], added mass [−0.5, 0.5] kg, velocity pushes [0.0, 0.2] m/s, and rotor inertia 0.0 kg m2 . Unless stated otherwise, the default test settings are: friction 1.0, added mass 0.0 kg, velocity push 0.0 m/s, and rotor inertia 0.005 kg m2 . Linear velocity commands vxy are set to 0.0. We note that Isaac Gym by default assigns the rotor inertia (armature) to 0.0 kg m2 during training, and in Mujoco this causes failure. None of the rollouts summarized in Table V result in failure. However, they show that while the policy produces similar linear velocity MSE across test conditions in both simulators, it suffers in the yaw velocity tracking. These results demonstrate that SKooP is robust to perturbations and transferable across simulated settings. VI. CONCLUSIONS We introduce SKooP, a unified learning approach which accelerates convergence and improves reward performance. Imposing symmetry constraints enables SKooP to generalize beyond training scenarios, while ecDAE-trained Koopman predictions for value function learning improves total return and qualitative learned behavior. We also demonstrate that

TABLE V S TUDY OF SKOO P S IM - TO -S IM ROBUSTNESS ON stand dance L INEAR vxy AND A NGULAR ωz C OMMAND T RACKING P ERFORMANCE (MSE). Test Condition Friction µ = 0.4 Friction µ = 1.0 Mass Offset 0.5 kg Push Vel 0.3 m/s Armature 0.002 Armature 0.005

vxy MSE ωz MSE Isaac Gym MuJoCo Isaac Gym MuJoCo 0.0341 0.0271 0.0273 0.0320 0.0308 0.0406

0.0245 0.0231 0.0257 0.0257 0.0229 0.0227

0.2795 0.1719 0.1656 0.1981 0.1494 0.2339

0.6017 0.5048 0.2904 0.5325 0.4826 0.4986

SKooP is modular and produces consistent advantages across tasks. Future work will address alternative Koopman learning strategies and transfer to other robots and tasks. A PPENDIX Here, we include a table listing some of the parameter settings we use for the quadruped training tasks we consider. TABLE VI S UMMARY OF TASK SETTINGS . R EMAINING PARAMETERS AS IN [11]. Task

Parameter

Value

Stand dance Stand dance Stand dance Stand dance

Lin. vel. tracking weight Ang. vel. tracking weight Upright weight Lift up weight

0.8 0.5 1.2 0.8

Push door

Lin. vel. x command range

[0.15, 0.3]

R EFERENCES [1] S. Kajita, F. Kanehiro, K. Kaneko, K. Yokoi, and H. Hirukawa, “The 3d linear inverted pendulum mode: a simple modeling for a biped walking pattern generation,” in Proceedings 2001 IEEE/RSJ International Conference on Intelligent Robots and Systems. IEEE. [2] J. Englsberger, C. Ott, and A. Albu-Schaffer, “Three-dimensional bipedal walking control based on divergent component of motion,” IEEE Transactions on Robotics, vol. 31, no. 2, pp. 355–368, 2015. [3] H. Dai, A. Valenzuela, and R. Tedrake, “Whole-body motion planning with centroidal dynamics and full kinematics,” in 2014 IEEE-RAS International Conference on Humanoid Robots. IEEE, Nov. 2014. [4] I. Radosavovic, T. Xiao, B. Zhang, T. Darrell, J. Malik, and K. Sreenath, “Real-world humanoid locomotion with reinforcement learning,” Science Robotics, vol. 9, no. 89, p. eadi9579, 2024. [5] J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V. Tsounis, V. Koltun, and M. Hutter, “Learning agile and dynamic motor skills for legged robots,” Science Robotics, vol. 4, no. 26, 2019. [6] B. O. Koopman, “Hamiltonian systems and transformation in hilbert space,” Proceedings of the National Academy of Sciences, vol. 17, no. 5, pp. 315–318, 1931. [7] M. Weiler, P. Forré, E. Verlinde, and M. Welling, Equivariant and Coordinate Independent Convolutional Networks, 2023. [8] J. E. Gerken, J. Aronsson, O. Carlsson, H. Linander, F. Ohlsson, C. Petersson, and D. Persson, “Geometric deep learning and equivariant neural networks,” Artificial Intelligence Review, vol. 56, no. 12, pp. 14 605–14 662, Jun. 2023. [9] P. de Haan, T. Cohen, and J. Brehmer, “Euclidean, projective, conformal: Choosing a geometric algebra for equivariant transformers,” in Proceedings of The 27th International Conference on Artificial Intelligence and Statistics, vol. 238. PMLR, 02–04 May 2024. [10] M. Mittal, N. Rudin, V. Klemm, A. Allshire, and M. Hutter, “Symmetry considerations for learning task symmetric robot policies,” in International Conference on Robotics and Automation. IEEE, 2024. [11] Z. Su, X. Huang, D. Ordoñez-Apraez, Y. Li, Z. Li, Q. Liao, G. Turrisi, M. Pontil, C. Semini, Y. Wu, and K. Sreenath, “Leveraging symmetry in rl-based legged locomotion control,” in 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, Oct. 2024, pp. 6899–6906.

[12] J. H. Tu, C. W. Rowley, D. M. Luchtenburg, S. L. Brunton, and J. Nathan Kutz, “On dynamic mode decomposition: Theory and applications,” Journal of Computational Dynamics, vol. 1, no. 2, pp. 391–421, 2014. [13] J. L. Proctor, S. L. Brunton, and J. N. Kutz, “Generalizing koopman theory to allow for inputs and control,” SIAM Journal on Applied Dynamical Systems, vol. 17, no. 1, pp. 909–930, 2018. [14] M. O. Williams, C. W. Rowley, and I. G. Kevrekidis, “A kernelbased method for data-driven koopman spectral analysis,” Journal of Computational Dynamics, vol. 2, no. 2, pp. 247–265, 2015. [15] B. Lusch, J. N. Kutz, and S. L. Brunton, “Deep learning for universal linear embeddings of nonlinear dynamics,” Nature Communications, vol. 9, no. 1, Nov. 2018. [16] M. Korda and I. Mezić, “Linear predictors for nonlinear dynamical systems: Koopman operator meets model predictive control,” Automatica, vol. 93, pp. 149–160, 2018. [17] M. Korda and I. Mezić, “Optimal construction of koopman eigenfunctions for prediction and control,” IEEE Transactions on Automatic Control, vol. 65, no. 12, pp. 5114–5129, 2020. [18] L. Song, J. Wang, and J. Xu, “A data-efficient reinforcement learning method based on local koopman operators,” in 2021 20th IEEE International Conference on Machine Learning and Applications (ICMLA), 2021, pp. 515–520. [19] D. Mayfrank, M. Velioglu, A. Mitsos, and M. Dahmen, “Sampleefficient reinforcement learning of koopman enmpc,” Computers & Chemical Engineering, vol. 201, p. 109240, 2025. [20] P. Rozwood, E. Mehrez, L. Paehler, W. Sun, and S. Brunton, “Koopman-assisted reinforcement learning,” in NeurIPS 2023 AI for Science Workshop, 2023. [21] E. N. Lorenz, “Deterministic nonperiodic flow.” Journal of the Atmospheric Sciences, vol. 20, no. 2, pp. 130–148, 1963. [22] A. Cozma, L. Harris, and H. Qi, “Kippo: Koopman-inspired proximal policy optimization,” in Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, IJCAI-25, 8 2025. [23] D. Ordonez-Apraez, G. Turrisi, V. R. Kostic, P. Novelli, C. Semini, C. Mastalli, and M. Pontil, “Dynamics harmonic analysis of robotic systems: Application in data-driven koopman modeling,” in 6th Annual Learning for Dynamics & Control Conference, vol. 242, 2024. [24] F. Li, A. Abuduweili, Y. Sun, R. Chen, W. Zhao, and C. Liu, “Continual learning and lifting of koopman dynamics for linear control of legged robots,” in 7th Annual Learning for Dynamics & Control Conference, vol. 283, 2025. [25] S. Martini, S. Sönmez, M. Stefanovic, M. J. Rutherford, and K. P. Valavanis, “Koopman-based reinforcement learning for lq control gains estimation of quadrotors,” in 2025 International Conference on Unmanned Aircraft Systems (ICUAS), 2025, pp. 465–472. [26] M. Weissenbacher, S. Sinha, A. Garg, and K. Yoshinobu, “Koopman qlearning: Offline reinforcement learning via symmetries of dynamics,” in Proceedings of the 39th International Conference on Machine Learning, vol. 162. PMLR, 17–23 Jul 2022. [27] D. Ordoñez-Apraez, G. Turrisi, V. Kostic, M. Martin, A. Agudo, F. Moreno-Noguer, M. Pontil, C. Semini, and C. Mastalli, “Morphological symmetries in robotics,” The International Journal of Robotics Research, vol. 44, no. 10-11, 2025. [28] S. L. Brunton, M. Budišić, E. Kaiser, and J. N. Kutz, “Modern koopman theory for dynamical systems,” SIAM Review, vol. 64, no. 2, pp. 229–340, 2022. [29] L. Pinto, M. Andrychowicz, P. Welinder, W. Zaremba, and P. Abbeel, “Asymmetric actor critic for image-based robot learning,” RSS, 2018. [30] T. Schaul, J. Quan, I. Antonoglou, and D. Silver, “Prioritized experience replay,” CoRR, vol. abs/1511.05952, 2015. [31] Y. Li, J. Li, W. Fu, and Y. Wu, “Learning agile bipedal motions on a quadrupedal robot,” in 2024 IEEE International Conference on Robotics and Automation (ICRA), 2024, pp. 9735–9742. [32] Xiaomi, https://www.mi.com/cyberdog2, [Accessed: Sept. 2025]. [33] V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, and G. State, “Isaac gym: High performance gpu based physics simulation for robot learning,” in Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, vol. 1, 2021. [34] G. Cesa, L. Lang, and M. Weiler, “A program to build e(n)-equivariant steerable CNNs,” in International Conference on Learning Representations, 2022.

Record · ID 363260 · SHA-256 bce32eabb47a3cba
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.