ConceptioArchivearXiv CS
arXiv CSopen access

Directional Constraints for Efficient Exploration in Safe Reinforcement Learning

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

Directional Constraints for Efficient Exploration in Safe Reinforcement Learning

arXiv:2607.12784v1 [cs.RO] 14 Jul 2026

Paolo Magliano1 , Puze Liu2,3 , Jan Peters4,3,5 , Davide Tateo6,4,† , Raffaello Camoriano1,7,† Abstract— Reinforcement Learning has revolutionized the landscape of robotic research, allowing robust learning of complex robotic skills in simulation. However, real-world deployment in open-ended environments requires strong safety guarantees to prevent dangerous or harmful behaviors. Safe Reinforcement Learning methods address this requirement by enforcing safety constraints. Nevertheless, learning under constraints often reduces learning speed and could lead to suboptimal task performance, as the agent must solve a more complex constrained optimization problem compared to unconstrained settings. To tackle this issue, in this work, we propose an extension of the ATACOM framework, a state-of-the-art reliable safety layer that can be integrated with existing Reinforcement Learning algorithms to enforce constraints derived from prior knowledge of the system or learned directly from data. Our proposed method, named ATACOM Directional Constraints (ATACOM-DC), significantly improves the safety-performance trade-off by introducing directional constraints that distinguish between actions approaching and moving away from constraint boundaries, activating constraint enforcement only when necessary. We evaluate our method across a range of challenging robotic control tasks in simulation, analyzing both constraint-violation costs and achieved task performance. Code and additional material at https://atacom-dc.robot-learning.net.

I. I NTRODUCTION In recent years, Reinforcement Learning (RL) has become the dominant technique for learning complex, dynamic robotic skills, both in the area of locomotion [1], [2] and manipulation [3], [4], [5]. Most of the proposed approaches rely on the concept of Domain Randomization [6], [1]. The key idea is to generate policies that can be robustly deployed in the real world, reacting to different environmental conditions by training on a family of simulators offline. However, when we deploy policies in the real world, policy robustness may not be sufficient due to safety-critical requirements in many applications. Indeed, while safety violations in simulation are not problematic, in the real world they may damage the This study was carried out within the FAIR - Future Artificial Intelligence Research and received funding from the European Union NextGenerationEU (PIANO NAZIONALE DI RIPRESA E RESILIENZA (PNRR) – MISSIONE 4 COMPONENTE 2, INVESTIMENTO 1.3 – D.D. 1555 11/10/2022, PE00000013). This manuscript reflects only the authors’ views and opinions, neither the European Union nor the European Commission can be considered responsible for them. 1 Dipartimento di Automatica e Informatica, Politecnico di Torino, Turin, Italy. [email protected] 2 Tongji University, Shanghai Research Institute for Intelligent Autonomous Systems. puze [email protected] 3 German Research Center for AI (DFKI). 4 Intelligent Autonomous Systems Group, TU Darmstadt, Germany. [email protected] 5 Hessian.AI. 6 Lund University, Sweden. [email protected] 7 Istituto Italiano di Tecnologia, Genoa, Italy. † Co-last authors.

(a) Unconstrained

(b) ATACOM [7]

(c) Directional (ours)

Fig. 1: Illustration of four actions near a constraint (red line) under three handling strategies: (a) raw agent actions without modification; (b) scaling of the orthogonal component for all actions while preserving tangential components; (c) selective scaling of only actions directed toward the constraint (a1, a2), leaving safe actions (a3, a4) unchanged. robot, the environment, or harm people. Relying only on disturbance rejection capabilities from a random distribution of the environment may not be sufficient and definitely lacks sound theoretical guarantees on the safety of the system. To tackle this problem, both data-driven Safe Reinforcement Learning (SafeRL) [8] methods based on a constraint budget, and model-based Safe Exploration (SafeExp) [9], [10] approaches based on state constraints have been developed. While each of these methods has different strengths, weaknesses, and application domains, there is a common issue to both settings: learning under safety constraints is considerably more challenging than the standard unconstrained RL problem. Indeed, trying to satisfy safety requirements while optimizing a policy may be challenging, either because the policy objective could push the system towards constraint violations—e.g., when extreme motions are more effective than slower and safer ones—or because the safety constraints restrict exploration too much. Learning is even more problematic when constraint functions are unknown and must be estimated from data, since the additional functionapproximation process may further restrict exploration.

In this paper, we address the issue of efficient exploration under safety constraints. Our work is based on one of the state-of-the-art approaches for learning under safety constraints, the Acting on the TAngent Space of the COnstraint Manifold (ATACOM) [11] framework, which has been recently [7] shown to be capable of performing real-world finetuning of RL policies on complex, dynamic, contactrich manipulation tasks such as the Robot Air Hockey. This approach has originally been developed in the context of SafeExp, being closely related to Control Barrier Function (CBF) [12], [7] and later extended in the context of SafeRL [13] to include distributional critics [14], [15] and Feasibility Value Function (FVF) [16]. While the ATACOM approach is effective in maintaining safety and has a low computational overhead, only requiring simple matrix operations and no optimization, the safety layer hampers exploration, making learning slower and more inefficient, particularly during the first training epochs. To address this issue, we introduce the concept of Directional Constraints, which morphs the action through the safety layer only when approaching the constraints, while leaving it unchanged when moving away from them, thereby promoting safe exploration. A comparison between the original [7] and our Directional Constraints approach is illustrated in Figure 1. Both (b) and (c) ensure constraint satisfaction, unlike case (a), where ATACOM is not applied. Additionally, Directional Constraints (c) selectively scale down only the actions that move towards the constraint. In contrast, the original ATACOM (b) also modifies outbound actions, forcing the agent to stay closer to the boundary even when unnecessary. The contributions of this work are summarized as follows: We introduce ATACOM with Directional Constraints (ATACOM-DC), a novel approach to generate safe actions. This is a simple, yet very effective modification of the ATACOM safety layer, which allows significantly faster learning and, in some cases, achieve better final performances in multiple challenging robotic control benchmark tasks. • Furthermore, we systematically analyze the sensitivity of the parameters of ATACOM and ATACOM-DC, highlighting the performance trade-off between the policy’s safety and performance.

II. R ELATED W ORK Many approaches have been developed to deal with real-world safety-critical settings in learning-based control. Machine learning-based SafeRL methods are built on the Constrained Markov Decision Processes (CMDP) framework [17]. The key idea of SafeRL is to impose safety constraints, represented as a threshold on a cumulative constraint cost during the learning process. These approaches are generally based either on the idea of safety filters [18], [19] or on Constrained optimization [20], [21], [22], among which the family of Lagrangian-based methods [23], [24], [14], [15] is the most represented.

In the fields of robotics and control, major efforts have been made to advance the SafeExp area, due to its closeness to real-world applications. To tackle the SafeExp problem, a variety of approaches have been developed based on the theory of CBF [12], [25], [26], [27], [28], Lyapunov Stability [29], [30], Reachability Analysis [31], [32], [33], and Model Predictive Control [34], [35], [36]. Furthermore, data-driven solutions [10], [37] tackle the problem by exploiting the regularity of the environment, such as Lipschitzcontinuity and regularity of the unknown dynamics. Machine learning and robot control approaches have both strengths and weaknesses. Machine learning-based SafeRL approaches are unable to impose safety at every timestep of the interaction with the environment, as safety is guaranteed (in probability) only at convergence. Instead, controlbased SafeExp methods are usually not plug-and-play, since they rely on prior knowledge in the form of accurate dynamics models or hand-crafted functions. Moreover, step-by-step safety guarantees may come at the cost of increased computations and overly conservative exploration. In this work, we successfully tackle some of these limitations by directionally activating constraints, speeding up safe exploration. III. P RELIMINARIES A. Constrained Markov Decision Process In SafeRL, the environment is modeled as a CMDP. A CMDP is defined by the tuple ⟨S, A, R, P, ι, γ, K⟩, where ⟨S, A, R, P, ι, γ⟩ corresponds to a standard MDP, and K is the set of constraint functions: K := {ki : S → R | i ∈ {1, . . . , K}}. In this paper, we will focus on SafeExp, where we must ensure safety at every timestep of the agent-environment interaction. To ensure safe exploration, the objective is to avoid constraint violations throughout the learning process. This is achieved by solving the following constrained optimization problem: # " T X ∗ t γ r(st , at ) , π = arg max Eτ ∼π π

t=1

s.t. ki (st ) ≤ 0, ∀t, ∀i ∈ {1, . . . , K}. B. Acting on the Tangent Space of the Constraint Manifold The ATACOM framework [11], [38], [7] provides a method that achieves safe exploration by leveraging knowledge of the system’s dynamics and constraints. At its core, ATACOM introduces the notion of constraint manifold, transforming the original constrained optimization problem into an unconstrained one defined over this manifold. To maintain safety, the action space is redefined as the tangent space of the constraint manifold. This ensures that any selected action satisfies the system constraints, as it lies within the safe set, i.e., the tangent space. Based on the definition of the constraints, the safe region within the state space is the set C, C := {s ∈ S ⊂ Rn | k(s) ≤ 0} .

To build the constraint manifold, the state space is augmented with a vector of slack variables µ ∈ [0, +∞)K . This leads to the augmented constraint function: c(s, µ) := k(s) + µ. The corresponding constraint manifold is then defined as the set of all state-slack pairs that satisfy the equality condition: M := {(s, µ) ∈ D | c(s, µ) = 0} . ATACOM constructs a safety controller that ensures that the system evolves within the tangent space of the constraint manifold. To achieve this, it assumes that the system can be modeled as a control-affine dynamical system ṡ = f (s) + G(s)us

(1)

and that the constraints k(s), along with their Jacobians Jk (s), are defined analytically. ATACOM also defines perelement dynamics for the slack variables, using a class K function αi (·) which is locally Lipschitz continuous. In this paper, we use exponential slack dynamics µi = αi (µi )uµ,i ,

αi (µi ) = exp(βµi ) − 1.

The β parameter is used by ATACOM to perform a tradeoff between safety and performance, with lower values of β representing constraints with wider margin. The controller is designed to adjust the augmented state along directions that remain tangent to the constraint  ⊤manifold. In other words, the velocity vector ṡ µ̇ must belong to the tangent space T(s,µ) M. Under this condition, the safe control input is defined as:   us = −Ju† ψ − λJu† c + Bu u, (2) uµ where ψ = Jk (s)f (s) is the constraint drift that describes † how the constraints evolve without  control action; Judenotes the pseudoinverse of Ju = Jk (s)G(s) A(µ) . Here, A(µ) : RK → RK×K is a diagonal matrix with entries Aii = αi (µi ), while Bu (s, µ) is the tangent space basis in matrix form, such that the following equation holds Ju (s, µ)Bu (s, µ) = 0.

evaluates the expected constraint violation under the policy. To incorporate the uncertainty arising from various sources of stochasticity in the environment, the FVF is learned using distributional RL. The algorithm exploits the uncertainty estimate to improve safety through the concept of Conditional Value-at-Risk (CVaR). Furthermore, the algorithm introduces a dynamic threshold δ to counteract constraint function fitting errors, ensuring that the system reaches a predefined value of safety while allowing for meaningful exploration in the initial phase of training. While ATACOM performs effectively on a wide range of tasks, it can produce suboptimal policies in scenarios where multiple constraints significantly restrict the action space. In such settings, the agent is forced to operate conservatively, which can negatively affect task execution and learning efficiency. By design, ATACOM suppresses the action component along the orthogonal direction of the constraint, thereby limiting motion toward constraint boundaries while allowing the RL agent to freely explore the tangent space. While this mechanism effectively prevents constraint violations, it also unintentionally suppresses actions that would move the agent away from the constraint boundary. As a consequence, once the agent approaches a constraint, escaping toward safer regions of the state space becomes difficult. This limitation arises from the symmetric nature of the action space morphing near constraint boundaries, which penalizes both approaching and escaping motions. D. Key idea To address the issue of symmetric morphing, we implement the Directional Constraints mechanism. The key idea is to scale down only the actions that are moving the robot towards the constraints. To decide which actions to scale, we evaluate the constraint derivatives, which describe how the constraint values evolve based on the current action. If the derivative is positive, the constraint value is increasing; otherwise, it is decreasing. To evaluate the effect of an action on the i-th constraint, we consider the constraint derivative ċi , as follows: ċi (s) = ψ i (s) + Jki (s)G(s)us ,

The first term Ju† ψ of the ATACOM controller represents the drift compensation term, which counteracts the natural system drift. The second term λJu† c corresponds to the contraction term, driving the system state back toward the constraint manifold when the constraint is violated. The final component Bu u is the tangential term, generating a vector field that lies in the tangent space of the constraint manifold. Overall, the controller maps the task-specific control input u into a safe action us .

where Jki (s) is the Jacobian of the i-th constraint, G(s) is the function that models the system, and us is the sampled action. Once the constraints derivatives are computed, we can select the actions to scale based on the derivative’s sign, only keeping the one with positive sign. Thus, Directional Constraints apply the action morphing only if the action sampled from the RL policy would move the robot closer to the constraint, otherwise leaving the action unaffected.

C. Distributional ATACOM

Although the idea described above is intuitive for single constraints, extending it to multiple constraints is not straightforward. In the following, we exploit the geometry of the problem to make the extension viable for an arbitrary number of constraints.

Distributional ATACOM (D-ATACOM) is an extension [13] of the ATACOM framework, removing the necessity of the analytical form of the constraints function. Instead, D-ATACOM introduces the concept of FVF, which

E. Multiple constraints

̇ u) c(s,

̇ u) c(s,

ψ

̇ u) c(s,

Algorithm 1 ATACOM-DC ut

u

u

u

(a)

(b)

(c)

Fig. 2: A representation of the key idea of Directional Constraints with a one-dimensional action. Plots show the constraint derivatives ċ(s, u) for multiple constraints as a function of the sampled action. The system drift affects the constraint values for states close to the boundary, increasing the derivative even with zero control action. This effect is illustrated by the orange line, which intersects the y-axis at a nonzero value. After drift compensation, we sample the (residual) action ut and select constraints based on the sign of their derivative, disabling constraints with ċ(s, u) < 0. First of all, we notice that all constraint derivatives are linear w.r.t. us . This means that each constraint derivative is a hyperplane in the space of actions, i.e., a line in the scalar-action setting, as shown in Figure 2a as an illustrative example. As we are assuming a nonlinear affine system (eq. (1)), most hyperplanes (lines) will pass through the origin. In the ATACOM setting, we assume that this linear system is solvable, as prescribed in Assumption 4 from [7]   u ψ(s) + Ju (s, µ) s = 0. uµ This means that we can always compensate for the constraint drift. Notice that, while this assumption may seem restrictive, we are indeed restricting the drift compensation on the constraint manifold, which does not require complete compensation for the drift of the system. This means that the system drift is compensated only when reaching constraint boundaries; otherwise, the drift can be redirected into an increase of the slack variables. For many practical scenarios, this assumption boils down to not having more active constraints, i.e., those with zero slack variable, than action degrees of freedom. Therefore, in our setting, we can always compensate first for the drift of the system, and only learn the residual part of the controller, representing the “tangent” part of the action. With the drift compensated, all hyperplanes of the residual action will pass through the center of the plane, as illustrated in Figure 2b. We further assume that the null-space basis Bu does not invert the sign of the actions of the actual system, which holds in most scenarios when using the smooth basis algorithm introduced in [7]. Notice that this assumption only prescribes us to avoid the use of an inverting basis, i.e., a basis that changes the meaning of the action drastically. Given that the null-space basis can be chosen arbitrarily among all possible bases, this is not a strict assumption, as we can always use a basis that preserves the action’s meaning. A possible problematic case is when there are equality constraints, and the Jacobian rank degenerates, e.g., when the constraint manifold has sphere-like topologies. In

Require: s, u ▷ At each step 1: Determine the slack variable µ ← max(−k(s), tol) 2: Compute the Jacobians and the drift JG ← Jk (s)G(s)   Ju (s, µ) ← JG (s) A(µ) ψ(s) ← Jk (s)f (s) 3: Compute the constraint value and derivative assuming drift compensation c(s, µ) ← k(s) + µ ċ(s, u) ← JG (s)u 4: Disable constraints with negative ċ(s, u) A(µ)dir ← A(µ)[i|ċ>0] 5: Compute the directional Jacobian   Ju (s, µ)dir ← JG (s) A(µ)dir 6: Compute the tangent space basis Bu ← SmoothBasis(Judir ) 7: Compute us compensating the drift ▷ Eq. (2) 8: Output: us

this scenario, the selection of a proper base is impossible due to the hairy-ball theorem [39]. However, these are edge cases, as most robotics systems either do not have such complex constraints or can be designed to circumvent the issue. With these two assumptions, we see that the sampled “tangential” action prescribes a quadrant where our morphed action will also lie. Here, only the constraints active in this quadrant are relevant for the ATACOM morphing. Thus, these active constraints are the only constraints that we need to consider to compute the morphing, effectively turning off the others. Intuitively, if the action pushes away from the constraints, we can avoid considering them. The complete ATACOM-DC algorithm consists of the following steps: 1) Compensate for the drift; 2) Sample a residual action; 3) Compute the active set of constraints by computing the constraint derivatives; 4) Compute the morphing of the action using the ATACOM controller only on the active set of constraints. Notice that this approach retains the same safety guarantees as the original algorithm under the assumption of a perfect model and constraints, since it modifies the algorithm’s behavior only when taking safe actions. F. Practical Implementation The implementation of Directional Constraints follows Algorithm 1 and is conceptually illustrated in Figure 2c. Given the sampled action at time step t, the derivative of each constraint is computed and evaluated to disable those constraints whose value would not increase if the action were applied. Once the constraint derivatives ċ(s, u) are computed, the key step consists of removing the rows of the matrix A(µ) corresponding to the constraints with negative derivative, under the assumption that drift compensation is handled by the ATACOM controller. The modified Jacobian Judir (s, µ) is then constructed accordingly, together with the corresponding smooth basis Bu .

(a) Kuka iiwa air hockey.

(b) Planar air hockey.

(c) Quadrotor.

Fig. 3: The task used in our experimental evaluation Success rate 1.0

1.0

0.5

0.5

Success rate

Distance from target 4

0.0

2

0.0 0.0

0.5

1.0 Steps

1.5

2.0 1e6

0 0.0

0.5

1.0 Steps

1.5

2.0 1e6

0.0

0.2

0.4

Episodic cost

Episodic cost

10 100

0.0

300

1000

0

0

0 0.5

1.0

1.5

Length of episodes Steps

(a) Planar air hockey (velocity)

0.8

1.0 1e6

0

0.5

1.0 Steps

1.5

2.0 1e6

0.0

(b) Planar air hockey (acceleration)

200

1.0 1e6

0 0.0

2.0 1e6

0.8

200

1 0

0.6

Episodic cost

200 100

Steps

ATACOM-DC

0.2

0.4 0.6 Steps

(c) Quadrotor

SAC

100 4: Comparison of our method against unconstrained SAC in the planar air hockey and quadrotor tasks, both in terms Fig. of task metric (top row) and safety (bottom row) 0.0

0.5

1.0 Steps

1.5

2.0 1e6

For computational efficiency, especially in configurations with parallel environments where the matrix A(µ) is computed in batch form, explicitly removing rows becomes impractical. Different environments may require removing a different number of rows, leading to batched matrices with inconsistent dimensions. An equivalent and more practical solution consists in setting the corresponding diagonal entries to the upper bound value, i.e., αi (µi ) = µη for all i such that ċi (s, u) < 0, where µη < ∞ is a sufficiently large constant such that µ ≤ µη . IV. E XPERIMENTAL E VALUATION For the experimental evaluation of the ATACOM-DC approach, we focus on three simulated tasks that are slightly modified versions of those proposed in [7] and [13]. Policy rollouts and qualitative comparisons on the considered tasks are available in the supplementary video. Air hockey task: In the air hockey task, a KUKA iiwa14 robotic manipulator is trained to strike a puck, initialized at random positions on the table, toward the opponent’s goal. The robot observes both its own state and the puck state, including joint configurations and puck motion. Based on these observations, the policy outputs desired joint velocities, which are executed by a low-level controller to generate smooth torque commands. The objective of the task is to maximize performance by effectively hitting the puck, increasing its forward velocity, and ultimately scoring a goal. To ensure feasibility and safety, constraints are imposed on joint limits, workspace boundaries defined by multiple planes around the table, and collision avoidance by enforcing minimum height limits on selected robot links.

Planar air hockey task: In the planar air hockey task, the agent controls a 3-DoF planar robotic arm equipped with a mallet at the end-effector that aims to strike a puck toward the opponent’s goal using the same constrained setup as the air hockey task. This task can be controlled either in acceleration or in velocity. Quadrotor navigation task: The quadrotor navigation task involves a quadrotor drone required to track a moving target that follows an eight-shaped trajectory, while avoiding a large cylindrical obstacle placed in the center of a confined environment. The control input consists of the total thrust generated by the propellers and the torques applied along the roll, pitch, and yaw axes. The agent observes the robot’s proprioceptive state and the target’s position and velocity. To guarantee safe and stable navigation, several constraints are enforced. These include collision-avoidance constraints to prevent impacts with obstacles, workspace constraints that confine the quadrotor within boundaries defined along the x, y, and z axes, and an additional constraint limiting the angular velocity around the z-axis to improve flight stability. A. Experimental setup The following results compare the proposed ATACOMDC with the unconstrained approach, the original ATACOM, and the D-ATACOM method. Each experiment is repeated over 15 independent random seeds. All approaches employ SAC [40] as the underlying RL policy optimization algorithm. The policy is an MLP with 2 layers parameterizing a state-dependent Gaussian distribution over actions. While we log the cumulative discounted return, for presentation reasons we only show task-specific performance metrics. For the air

Success rate

0.5

0.00

0.0 0.0

0.5

1.0 Steps

175

1.5

2.0 1e6

0.0

0.5

1.0 Steps

ATACOM-DC

150

0.000

0.02

Length of episodes

0.00

200

0.005

0.04

0.50 0.25

Episodic cost

Puck velocity

1.0

0.75

1.5

2.0 1e6

0.0

0.5

1.0 Steps

1.5

2.0 1e6

ATACOM

Fig. 5: Impact of directional constraints in the kuka iiwa air Hockey task

125

Success rate

0.0

1.0

0.5

1.0 Steps

Success rate

1.5

Distance from target

2.0 1e6

0.8

0.5

0.5

0.0

0.0

0.6 0.4

0.0

0.5

1.0 Steps

1.5

2.0 1e6

0.2 0.0

0.5

2.0 1e6

0.0

0.2

0.4

0.005

0.0

0.5

1.0

1.5

Length of episodes Steps

(a) Planar air hockey (velocity)

0.5

1.0 Steps

1.5

2.0 1e6

(b) Planar air hockey (acceleration)

2000

0.8

1.0 1e6

0.8

1.0 1e6

0.000

0.00 0.0

2.0 1e6

0.6

0.005 0.05

0.0000

0.000

0.00

Steps

Episodic cost

0.10

0.0001

0.02

2100

1.5

Episodic cost

Episodic cost 0.04

2200

1.0 Steps

ATACOM-DC

0.0

0.2

0.4 0.6 Steps

(c) Quadrotor

ATACOM

1900 Fig. 6: Impact of directional constraints in the planar air hockey tasks and in the quadrotor task. The figure presents task

performance (top row) and the safety violations (bottom row). 1800 0.0 0.2 0.4 0.6 0.8 1.0 1e6 Steps hockey setups (both KUKA and planar), these include the goal success rate and the puck velocity. For the quadrotor task, the performance metric is the distance from the moving target. For safety evaluation, we consider the episodic cost, defined as the cumulative cost over an episode, where the cost at timestep t is max(k(st ), 0), as in [13]. B. Performance against unconstrained methods In Figure 4, we present a comparison of our approach in all the tasks against the unconstrained Soft Actor-Critic (SAC) algorithm. In this setting, we do not report results for the kuka iiwa air hockey task, as SAC is unable to learn in such a complex setup without the possibility of exploiting the constraint information. As a result of this experiment, we demonstrate that the optimal policy is safe for some tasks, and is not safe for others. In fact, in quadrotor and planar air hockey controlled in acceleration, the SAC final policy still violates constraints, thus achieving higher task performance. Instead, ATACOM-DC maintains safety throughout the entire training process, with similar final performance and generally better or comparable learning curves. C. Impact of Directional Constraints In Figures 5 and 6, we analyze the impact of Directional Constraints against the vanilla ATACOM algorithm. Our results show that in all tasks Directional Constraints consistently lead to faster learning and overall improved final performance, while constraint violations are generally reduced across tasks or at least comparable, except during the first epochs. Indeed, Directional Constraints are less restrictive, allowing for more movement in the initial phases

of learning. However, when the task objective is not to push against the constraints, this behavior allows the system to learn to leave the unsafe area, leading to lower long-term constraint violations. In any case, all the approaches show very low, close to zero, constraint violations. The benefit of enhanced exploration is particularly impactful in the air hockey setting, as highlighted in Figure 5. In this complex scenario, the less restrictive exploration allows us to learn faster and achieve higher-speed and more precise policies, which results in higher success rate and puck velocities. However, the faster motion may cause small violations due to the model inaccuracies, as it does not include the torque model and the low-level controller. D. D-ATACOM Improvements The methodology can be easily extended to the DATACOM approach. Here, we report results only for the planar air hockey environment, both in velocity and acceleration. In the planar velocity scenario, the result shows that, while having comparable performance in terms of safety as the vanilla D-ATACOM, directional constraints boost the overall learning performances. However, directional constraints have no clear statistical impact in the planar air hockey acceleration scenario. Furthermore, we investigate whether it is possible to learn without using a FVF, only using the constraints and their uncertainty. Indeed, learning a constraint function is more stable and faster than learning a FVF, which requires temporal difference learning. We investigate this in the setting of planar air hockey in velocity, which is a first-order system,

Success rate

0.5

0.5

0.0

0.0 0.00

0.25

0.50

0.75 Steps

1.00

1.25

1.50 1e6

Episodic cost 40

0.00

0.25

0.50

0.75 Steps

1.00

1.25

1.50 1e6

1.25

1.50 1e6

Episodic cost

150

2

10 100

20

0

0 0.00

0.25

0.50

0.75 Steps

1.00

1.25

(a) Planar air hockey (velocity) DATACOM-DC

1.50 1e6

0.00

0.25

0.50

0.75 Steps

1.00

(c) Planar air hockey - δ analysis

(b) Planar air hockey (acceleration) DATACOM-DC-constraint

DATACOM

DATACOM-constraint

Fig. 7: Impact of Directional constraints in the D-ATACOM algorithm. The figure presents task performance (top row) and the safety violations (bottom row). The last column presents the results for direct constraint learning.

1.00

1.25

1.50 1e6

where vanilla constraints are sufficient to impose safety to the system. Unfortunately, directly learning the constraints and avoiding learning the FVF produces comparable performance, at the cost of increased constraint violations. Notably, if we also fix the dynamic threshold δ, tuned automatically by D-ATACOM to trade off exploration and exploitation, and Beta analysis remove warm-up trajectories, we can achieve a performance 1.8 1.0 2.2 2.2 3.0 1.4 0.8shown boost in terms of learning speed and safety, as clearly (b) Quadrotor 3.0 2.6 (a) Planar air hockey 2.6 0.6 in Figure. 7a, where the DATACOM-DC-constraint line, 1.0 1.8 0.6 ATACOM ATACOM-DC 1.4 which represents this version of the algorithm, outperforms 0.4 Fig. 8: Analysis of the effects of the β parameter in the all other approaches. In particular, the safety improvement 0.6 0.2phase, planar air hockey and quadrotor tasks is mostly due to the reduced unconstrained warm-up 0.00 0.05 which limits early unsafe exploration while still allowingViolation rate without degrading performance. In practice, ATACOM-DC the agent to learn the constraints. This is due both to the improves the performance–safety trade-off, resulting in a removal of warm-up trajectories and the faster convergence Pareto-superior behavior compared to the baseline. to the desired constraint, allowed by the lack of temporal difference learning. V. C ONCLUSION Furthermore, we analyze the sensitivity of the effect of the In this paper, we introduced ATACOM-DC, an improved δ parameter over 5 different seeds for each value, comparing version of the ATACOM safety layer that boosts exploration D-ATACOM-DC with D-ATACOM. Here, having a higher capabilities simply and effectively through the concept of δ results in higher constraint violations. However, looser Directional Constraints. The approach only modifies safe constraints allow for more exploration, particularly in the actions, retaining the safety guarantees of the original frameinitial episodes, when the constraint is not yet correctly work. Our simulated experiments show that the learning approximated. Our results, presented in Figure 7c, show that speed is increased while not compromising safety. On the the Directional Constraints yield a higher success rate, while contrary, in some settings, the safety guarantees are imkeeping the cost lower for all values of δ, showing that the proved, as the agent policy can move away from the conproposed method is Pareto-optimal w.r.t. the baseline. straint quickly, moving the policy state distribution towards safer states. Using Directional Constraints, in settings where E. Beta Analysis: Performance–Safety Trade-off the long-term safety is not a concern, it is not necessary Finally, we perform an ablation on the role of the β to use TD-learning and automatic tuning of safety margin, parameter of the original ATACOM safety layer for slack dy- allowing the algorithm to directly learn a constraint, which is namics. Each value is evaluated across 5 independent seeds. faster and more accurate than learning a FVF. Furthermore, Results for different betas are reported in Figure 8 for both we show that the new approach is less sensitive to the safety the planar air hockey and the quadrotor tasks. Results show parameters and allows us to obtain a better performancethat Directional Constraints allow effective operation even safety trade-off, which is Pareto-dominant compared to when a larger safety margin from the constraints is required, vanilla ATACOM, at least in our experimental setting. Success rate

0.75 Steps

0

50

0

ngth of episodes

.50

Success rate

1.0

In future work, we aim to bring these exploration advances together with more modern off-policy actor-critic algorithms to rapidly learn from scratch in complex realworld environments, e.g., the Robot Air Hockey, overcoming the limitations of the previous approach. R EFERENCES [1] N. Rudin, D. Hoeller, P. Reist, and M. Hutter, “Learning to walk in minutes using massively parallel deep reinforcement learning,” in Conference on robot learning. PMLR, 2022, pp. 91–100. [2] Z. Zhuang, Z. Fu, J. Wang, C. Atkeson, S. Schwertfeger, C. Finn, and H. Zhao, “Robot parkour learning,” in Conference on Robot Learning (CoRL), 2023. [3] B. Huang, Y. Chen, T. Wang, Y. Qin, Y. Yang, N. Atanasov, and X. Wang, “Dynamic handover: Throw and catch with bimanual hands,” in 7th Annual Conference on Robot Learning, 2023. [4] T. Lin, Z.-H. Yin, H. Qi, P. Abbeel, and J. Malik, “Twisting lids off with two hands,” in 8th Annual Conference on Robot Learning, 2024. [Online]. Available: https://openreview.net/forum?id=3wBqoPfoeJ [5] Z. Li, Y. Jin, D. Ordonez-Apraez, C. Semini, P. Liu, and G. Chalvatzaki, “Morphologically symmetric reinforcement learning for ambidextrous bimanual manipulation,” in 9th Annual Conference on Robot Learning, 2025. [6] J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel, “Domain randomization for transferring deep neural networks from simulation to the real world,” in 2017 IEEE/RSJ international conference on intelligent robots and systems (IROS). IEEE, 2017, pp. 23–30. [7] P. Liu, H. Bou-Ammar, J. Peters, and D. Tateo, “Safe reinforcement learning on the constraint manifold: Theory and applications,” IEEE Transactions on Robotics, 2025. [8] S. Gu, L. Yang, Y. Du, G. Chen, F. Walter, J. Wang, and A. Knoll, “A review of safe reinforcement learning: Methods, theories, and applications,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 12, pp. 11 216–11 235, 2024. [9] J. Garcia and F. Fernández, “Safe exploration of state and action spaces in reinforcement learning,” Journal of Artificial Intelligence Research, vol. 45, pp. 515–564, 2012. [10] F. Berkenkamp, M. Turchetta, A. P. Schoellig, and A. Krause, “Safe Model-based Reinforcement Learning with Stability Guarantees,” in Conference on Neural Information Processing Systems (NIPS), 2017. [11] P. Liu, D. Tateo, H. B. Ammar, and J. Peters, “Robot reinforcement learning on the constraint manifold,” in 5th Conference on Robot Learning (CoRL). PMLR, 2021. [12] A. D. Ames, S. Coogan, M. Egerstedt, G. Notomista, K. Sreenath, and P. Tabuada, “Control barrier functions: Theory and applications,” in 2019 18th European control conference (ECC). IEEE, 2019, pp. 3420–3431. [13] J. Günster, P. Liu, J. Peters, and D. Tateo, “Handling long-term safety and uncertainty in safe reinforcement learning,” in Conference on Robot Learning (CoRL), 2024. [14] Q. Yang, T. D. Simão, S. H. Tindemans, and M. T. Spaan, “Wcsac: Worst-case soft actor critic for safety-constrained reinforcement learning,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35 (12), 2021, pp. 10 639–10 646. [15] Q. Yang, T. D. Simao, S. H. Tindemans, and M. T. Spaan, “Safetyconstrained reinforcement learning with a distributional safety critic,” Machine Learning, vol. 112, no. 3, pp. 859–887, 2023. [16] Y. Yang, Z. Zheng, and S. E. Li, “Feasible policy iteration,” arXiv preprint arXiv:2304.08845, 2023. [17] E. Altman, “Constrained Markov Decision Processes with Total Cost Criteria: Lagrangian Approach and Dual Linear Program,” Mathematical methods of operations research, vol. 48, no. 3, pp. 387–417, 1998. [18] G. Dalal, K. Dvijotham, M. Vecerik, T. Hester, C. Paduraru, and Y. Tassa, “Safe exploration in continuous action spaces,” arXiv preprint arXiv:1801.08757, 2018. [19] D. P. Nguyen, K.-C. Hsu, W. Yu, J. Tan, and J. F. Fisac, “Gameplay filters: Robust zero-shot safety through adversarial imagination,” in 8th Annual Conference on Robot Learning, 2024. [20] J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained Policy Optimization,” in International Conference on Machine Learning (ICML), 2017.

[21] Y. Liu, J. Ding, and X. Liu, “Ipo: Interior-point policy optimization under constraints,” in AAAI Conference on Artificial Intelligence (AAAI), vol. 34(04), 2020, pp. 4940–4947. [22] D. Ding, X. Wei, Z. Yang, Z. Wang, and M. R. Jovanovic, “Provably Efficient Safe Exploration via Primal-Dual Policy Optimization,” in International Conference on Artificial Intelligence and Statistics (AISTATS), vol. 130, 2021. [23] A. Ray, J. Achiam, and D. Amodei, “Benchmarking safe exploration in deep reinforcement learning.(2019),” URL https://cdn. openai. com/safexp-short. pdf, pp. 1–25, 2019. [24] S. Ha, P. Xu, Z. Tan, S. Levine, and J. Tan, “Learning to walk in the real world with minimal human effort,” in Proceedings of the 2020 Conference on Robot Learning, ser. Proceedings of Machine Learning Research, vol. 155. PMLR, 16–18 Nov 2021, pp. 1110–1120. [25] A. Taylor, A. Singletary, Y. Yue, and A. Ames, “Learning for safetycritical control with control barrier functions,” in Learning for Dynamics and Control. PMLR, 2020, pp. 708–717. [26] W. Xiao and C. Belta, “High-Order Control Barrier Functions,” IEEE Transactions on Automatic Control, vol. 67, no. 7, pp. 3655–3662, 2022. [27] D. C. Tan, F. Acero, R. McCarthy, D. Kanoulas, and Z. A. Li, “Your value function is a control barrier function: Verification of learned policies using control theory,” 2nd Workshop on Formal Verification of Machine Learning in the 40th International Conference on Machine Learning, 2023. [28] Y. Yang, Y. Jiang, Y. Liu, J. Chen, and S. E. Li, “Model-free safe reinforcement learning through neural barrier certificate,” IEEE Robotics and Automation Letters, vol. 8, no. 3, pp. 1295–1302, 2023. [29] Y. Chow, O. Nachum, E. Duenez-Guzman, and M. Ghavamzadeh, “A Lyapunov-based Approach to Safe Reinforcement Learning,” in Conference on Neural Information Processing Systems (NIPS), 2018. [30] Y. Chow, O. Nachum, A. Faust, E. Duenez-Guzman, and M. Ghavamzadeh, “Lyapunov-based Safe Policy Optimization for Continuous Control,” in RL4RealLife Workshop in the 36 th International Conference on Machine Learning, 2019. [31] A. K. Akametalu, J. F. Fisac, J. H. Gillula, S. Kaynama, M. N. Zeilinger, and C. J. Tomlin, “Reachability-based safe learning with gaussian processes,” in 53rd IEEE conference on decision and control. IEEE, 2014, pp. 1424–1431. [32] J. F. Fisac, A. K. Akametalu, M. N. Zeilinger, S. Kaynama, J. Gillula, and C. J. Tomlin, “A general safety framework for learning-based control in uncertain robotic systems,” IEEE Transactions on Automatic Control, vol. 64, no. 7, pp. 2737–2752, 2018. [33] Y. S. Shao, C. Chen, S. Kousik, and R. Vasudevan, “Reachability-based trajectory safeguard (rts): A safe and fast reinforcement learning safety layer for continuous control,” IEEE Robotics and Automation Letters, vol. 6, no. 2, pp. 3663–3670, 2021. [34] L. Hewing, K. P. Wabersich, M. Menner, and M. N. Zeilinger, “Learning-based model predictive control: Toward safe learning in control,” Annual Review of Control, Robotics, and Autonomous Systems, vol. 3, no. 1, pp. 269–296, 2020. [35] M. Prajapat, A. Lahr, J. Köhler, A. Krause, and M. N. Zeilinger, “Towards safe and tractable gaussian process-based mpc: Efficient sampling within a sequential quadratic programming framework,” in 2024 IEEE 63rd Conference on Decision and Control (CDC). IEEE, 2024, pp. 7458–7465. [36] M. Prajapat, J. Köhler, M. Turchetta, A. Krause, and M. N. Zeilinger, “Safe guaranteed exploration for non-linear systems,” IEEE Transactions on Automatic Control, 2025. [37] M. Wendl, Y. As, M. Prajapat, A. Pollak, S. Coros, and A. Krause, “Safe exploration via policy priors,” in The Fourteenth International Conference on Learning Representations, 2026. [38] P. Liu, K. Zhang, D. Tateo, S. Jauhri, Z. Hu, J. Peters, and G. Chalvatzaki, “Safe reinforcement learning of dynamic high-dimensional robotic tasks: navigation, manipulation, interaction,” in 2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2023, pp. 9449–9456. [39] M. Eisenberg and R. Guy, “A proof of the hairy ball theorem,” The American Mathematical Monthly, vol. 86, no. 7, pp. 571–574, 1979. [40] T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Offpolicy maximum entropy deep reinforcement learning with a stochastic actor,” in International conference on machine learning. Pmlr, 2018, pp. 1861–1870.

Record · ID 366256 · SHA-256 7c2e8cf1a42280e5
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.