CRAX: Fast Safe Reinforcement Learning Benchmarking
arXiv:2606.20376v1 [cs.LG] 18 Jun 2026
Tristan Tomilin∗ Mourad Boustani Mickey Beurskens Eindhoven University of Technology
Thiago D. Simão
Abstract Safety is a core concern for deploying reinforcement learning (RL) agents in real-world domains such as robotics and autonomous driving. While benchmarks have been central to progress in RL, existing safety benchmarks with high-fidelity 3D physics remain computationally slow, limiting large-scale experimentation and rapid prototyping. To address this gap, we propose CRAX (Constrained RL Accelerated with JAX). Built on top of the MuJoCo XLA (MJX) physics engine with realistic 3D dynamics, CRAX leverages vectorized operations and hardware acceleration, yielding up to ∼100x speedups over comparable CPU-based safety benchmarks. The benchmark features six environment suites and three agentspecific tasks, each spanning three difficulty levels. Evaluating six popular safe RL methods shows that no single approach dominates across all tasks, and reveals the trade-offs between performance and safety. We find that curriculum learning across difficulty levels and safety transfer can improve performance over direct training in harder settings.
1
Introduction
Although the progress in reinforcement learning [RL; 37] has been sped up by the use of hardware acceleration, this progress has not yet translated to the research in safe RL [SafeRL; 13]. Benchmarks are a catalyst driving innovation in research; for instance, ImageNet [10] motivated the introduction of CNNs [18], and the Arcade Learning Environment [2] supported the development of deep Q-networks [24]. In a similar trend, accelerated hardware brought a new wave of benchmarks that facilitate research in multiple areas of RL, including goal-conditioned [7], multi-agent [33], offline [15], and open-ended [22] reinforcement learning. Within the SafeRL literature, safety gym [31] and, more recently, safety gymnasium [16] have established a set of common tasks that facilitate comparisons between SafeRL algorithms. Nevertheless, SafeRL research still relies mostly on CPU-based simulations and does not leverage GPU computation. Effective research in RL, particularly in high-dimensional problems, requires fast data collection, as training RL typically requires a large number of environment interactions. This demand can compromise the research development phase, where we retrain such agents numerous times from scratch, while performing: hyperparameter tuning to ensure a fair comparison, tests in multiple environments to evaluate generalization, repeated runs for statistical significance, and ablation studies to assess individual components of an algorithm. Together, these requirements form a bottleneck for the development of new algorithms. Fast simulation is therefore essential to support research on RL. By leveraging high-fidelity simulators with hardware acceleration, RL is increasingly closing the gap to real-world applications such as robotics. Simulation platforms such as BRAX [12] and Isaac Lab [23] run on accelerated hardware, allowing researchers to leverage substantial speedups from parallel computing architectures. Such platforms enable large-scale policy training in simulation and, ∗ Corresponding author: [email protected]
Preprint.
Table 1: Key characteristics of popular Reinforcement Learning benchmarks. CRAX uniquely combines a focus on safety with hardware acceleration and 3D physics-based tasks. Benchmark OpenAI Gym Procgen Benchmark Atari (ALE) Meta-World Gymnax JaxMARL Craftax Jumanji XLand-MiniGrid VMAS DeepMind Control Suite (DMC) Brax Isaac Lab Safety Gym (OpenAI) Bullet-Safety-Gym Safe-Control-Gym HASARD SafeOR-Gym Safety-Gymnasium CRAX(Ours)
Safety
GPU
3D Phys.
× × × × × × × × × × × × × √
× × × × √
× × × × × × × × × × √
√ √ √ √ √ √
√ √ √ √ √
× √
√ √
√
× × × × √ ▲ √
× × × × × √ √
Type
Reference
Classic control Procedural generation Arcade games Robotic manipulation Classic control / MinAtar Multi-agent Open-ended / gridworld Combinatorial optimization Meta-RL Multi Agent 2D Physics in PyTorch 3D physics 3D physics 3D physics Safe navigation Safe navigation Safe control FPS game Operations research Safe navigation / locomotion Safe navigation / locomotion
[8] [25] [3] [43] [19] [33] [22] [6] [28] [5] [38] [12] [23] [31] [14] [44] [41] [30] [16]
▲ A subset of two robotic manipulation suits are GPU accelerated through "Safe Isaac Gym", which is part of Safety Gymnasium.
in some cases, transfer to physical robots [45]. Nevertheless, this approach still relies closely on reward engineering to specify the behavior desired from the agent. However, in many situations, expressing such behaviors is easier through constraints [32], particularly in safety-critical scenarios [31]. Therefore, we focus on fast RL benchmarks with explicit constraints. We introduce CRAX2 (Constrained RL Accelerated with JAX), a novel hardware-accelerated SafeRL benchmark leveraging MuJoCo, a general-purpose 3D physics engine. The design principles of CRAX are inspired by BRAX [12] and Safety Gymnasium [16]. CRAXprovides a set of simulated tasks, robots, and algorithm baselines for evaluating SafeRL leveraging parallel computing, resulting in higher simulation speeds than CPU-based setups, enabling more rigorous testing and faster algorithm development for the SafeRL community. Each task defines reward and cost signals, inducing a trade-off between performance and safety: achieving high reward typically requires incurring higher cost, while satisfying safety constraints necessitates sacrificing some reward. Constrained RL techniques naturally lend themselves to the treatment of safety tasks as cost-reward tradeoffs. Among numerous types of constraints, we focus on algorithms that bound the expected cumulative discounted cost, as this is the most widely-adopted approach in SafeRL literature. Accordingly, CRAXincludes a number of baseline algorithms for constrained RL, such as PPO Lagrangian [PPO-Lag 31], as well the non-constrained algorithm PPO [34] as a reference. These implementations will facilitate comparisons between new algorithms and relevant prior work. The core contributions of our work are as follows: 1. We propose CRAX, a hardware-accelerated SafeRL benchmark, enabling orders-ofmagnitude faster simulation than traditional CPU-based setups. CRAX tailors safety constraints to a variety of agent morphologies, and exposes explicit cost signals alongside rewards, necessitating a trade-off between performance and safety. 2. We reimplement six popular SafeRL algorithms in JAX and evaluate them across tasks and difficulty levels, identifying their strengths and limitations. 3. We study performance-safety trade-offs by varying cost thresholds, assess the utility of curriculum learning and safety transfer, and demonstrate how CRAX enables superior throughput and scaling. 2 The code and environments are accessible on GitHub.
2
2
Related Work
Safe Reinforcement Learning. In SafeRL, besides achieving high-performance, agents also need to adhere to established safety requirements during learning and deployment [13]. While safety can be encouraged indirectly through reward shaping, this approach places the burden of balancing performance and safety on the system designer. To reduce this burden, we can model safety requirements explicitly as constraints [32, 17], allowing the agent itself to autonomously find a trade-off between reward maximization and constraint satisfaction. A wide range of safety formulations have been studied, including chance constraints, almost-sure constraints, and per-step constraints [42]. In practical work, the most common experimental settings bound the expected sum of discounted costs over time [44, 16, 30]. While CRAX is agnostic to the specific safety formulation, our empirical evaluation adopts this setting due to its popularity. SafeRL Benchmarks. Initially, safe RL research was predominantly studied in low-dimensional 2D settings, such as gridworlds in AI Safety Gridworlds [20] and MiniGrid [9] tasks adapted for safety. More recent benchmarks have shifted toward environments for embodied, pixel-based learning [11, 21, 41] and physics-based continuous control [44, 16]. These benchmarks are typically built on existing simulation platforms rather than developing physics engines from scratch. For example, HASARD [41] is built on ViZDoom, while Safety-Gymnasium [16] extends MuJoCo tasks, adding safety constraints. CRAX follows this design principle by building on MJX [26], the JAX-based accelerated backend of MuJoCo [40]. Accelerated Benchmarking. RL experiments are typically data-intensive, as meaningful evaluation requires repeated environment interactions for hyperparameter tuning, statistical significance analysis, and testing across different tasks. While GPUs are routinely used to accelerate neural network training, online RL remains constrained when environment rollouts are executed on the CPU, making simulation throughput the main bottleneck. This has motivated moving both simulation and learning onto parallel hardware to accelerate the full training loop. Brax [12] provides hardware-accelerated continuous-control environments in JAX. VMAS [5] and JaxMARL [33] focus on scalable multiagent RL. The former creates 2D physics environments implemented in PyTorch, and the latter creates JAX-native variants of many popular multi-agent environments. Craftax [22] explores procedurally generated grid-based worlds optimized for large-scale parallel training. Table 1 summarizes a number of such widely used RL benchmarks. CRAXexpands the available selection of GPU accelerated safety focused 3D physics based environments significantly.
3
Constrained Reinforcement Learning
A constrained Markov decision process [CMDP; 1] is an MDP [29] with constraints, characterized by a tuple M = ⟨S, A, P, r, c, d, γ⟩, where S is a continuous state space, A a continuous action space, P a transition function P : S×A → Distr(S), r a reward function r : S×A → R+ , c a cost function c : S×A → R+ , d ∈ R+ a cost thresholds, and γ ∈ [0, 1) a discount factor. An RL agent interacting with a CMDP follows a stochastic policy π : S → Distr(A). The value function V π (s) represents the expected cumulative discounted reward when hPfollowing policy π starting i T π t from state s over a (potentially infinite) horizon T : V (s) = Eπ t=0 γ r(st , at ) | s0 = s , where the expectation Eπ is taken over the trajectory distribution induced by policy π, with actions at ∼ π(·|st ) and successor states st+1 ∼ P (·|st , at ). Similarly, the cost function C π (s) captures hPthe expected cumulative i discounted cost under policy π starting from state s: T t C π (s) = Eπ γ c(s , a ) | s = s . t t 0 t=0 The objective in the CMDP framework is to find an optimal policy π ∗ ∈ Π that maximizes the expected cumulative reward while ensuring the expected cumulative cost remains below the threshold d for all states s ∈ S. This constrained optimization problem is formulated as: max V π (s) subject to C π (s) ≤ d, ∀s ∈ S (1) π∈Π
π
The constraint C (s) ≤ d enforces safety by requiring that the policy maintains cost levels below the specified threshold regardless of the initial state. This formulation, known as the expected cumulative cost constraint [42], distinguishes CMDPs from standard MDPs, where the agent would simply maximize the reward without regard for cost constraints. 3
Safe Goal
Safe Push
Safe Circle
Safe Reacher
Reach the goal without colliding with the hazards
Push the block to the goal while avoiding obstacles
Move along a circular trajectory while avoiding hazards and walls
Reach the target while avoiding hazards
Safe Velocity
Safe Height
Safe Spider
Safe Pathway
Move forward while keeping designated legs airborne
Traverse the path without stepping on the hazards
Move fast while staying under Walk while keeping the torso below the height bound a velocity limit
Figure 1: Overview of the CRAX benchmark environment suites. Table 2: Overview of the available suites in the benchmark (rows), and which agents are compatible with them (columns). The Safe Navigation suite includes the Goal, Button, Circle, and Push tasks. Starred entries (⋆) have been selected for evaluation in Section 5. Task Safe Navigation Safe Velocity Safe Pathway Safe Reach Safe Height Safe Spider
4
Point √⋆ √
× × × ×
Ant √ √
× × × ×
3D Humanoid √ √⋆
Spider √ √
× × √⋆ ×
× × × √⋆
Half Cheetah
2D Walker2D
Hopper
× √
× √
× √
√
√⋆
√
× × ×
× × ×
× × ×
Fixed Reacher
× × × √⋆ × ×
CRAX
This section covers the design choices behind CRAX. The benchmark has been inspired by the GPUaccelerated RL environments in BRAX [12] and the safety environments of Safety Gymnasium [16]. It is intended as a research and benchmarking platform for SafeRL, and facilitates this by providing (i) ready-to-use environment suites, agent morphologies and utility tooling to design SafeRL experiments, and (ii) a set of pre-configured tasks of increasing difficulty in diverse environments. Both of these elements have been designed with the following principles in mind: (i) Support the development and assessment of constrained RL approaches. Each task includes a cost signal in addition to a reward signal. (ii) The benchmark should not be immediately solvable by state-of-the-art approaches, nor should it be too difficult to make at least some progress. Therefore, CRAX features a set of tasks with difficulty progression. (iii) Provide tools to assess the safety properties of the algorithms being evaluated. (iv) Each environment and task in the benchmark should be easily accessible for RL training. (v) All of the above steps can be run on the GPU through the JAX library in order to increase simulation speeds compared to CPU-based benchmarks.
4
Figure 2: Higher difficulty levels of Safe Goal (1 to 3, from left to right) decrease the size of the goals, and increase the number and variety of hazards.
4.1
Environment Suites And Tasks
Environment suites define families of configurable tasks in simulated 3D environments. In each suite, an agent seeks to maximize reward while adhering to a predefined cost threshold. Figure 1 summarizes the objectives associated with each suite. Beyond suite-specific parameters, such as the number of obstacles in Safe Goal, every suite specifies its own reward and cost signals. An instantiation of a suite’s parameters constitutes a task. Appendix A provides a detailed description of all suites. 4.2
Agents
Agents are the acting bodies in the environment. Each agent has a unique morphology and lidar sensors to detect its surroundings. The movement of some agents is constricted in one or more dimensions. Some suites are compatible with multiple agent types, while some of the suites allow for only a single agent type. Table 2 provides an overview of the agent types and their compatibility with the suites. Appendix B provides further details about the agents. 4.3
Rewards, Costs and Constraint Types
Each environment comes with a distinct default reward signal. Cost signals are constructed using a variety of constraints. Refer to Appendix A for an overview of the reward and cost signal for each environment. The benchmark employs five key constraint formulations across the environment suites. (1) Hazard proximity/contact costs incur when agents contact hazards or violate keep-out zones around obstacles (2) Velocity threshold constraints originate from exceeding velocity limits on locomotion tasks, with binary or hinge-style penalties. (3) Height constraints occur when height falls below minimum requirements, with hinge penalties for violations (4) Contact-restriction constraints act as binary costs from restricted feet making ground contact (5) Goal-oriented safety defines quadratic proximity costs while pushing blocks toward moving goals through hazard fields. As established in Equation 1, staying within the safety budget does not preclude further gains: higher returns can be achieved with more refined strategies while remaining safe. Moreover, the safety bound is adjustable, allowing one to impose stricter or more permissive requirements, and thereby modulate the difficulty of the task. 4.4
Difficulty Levels
A useful benchmark ought to serve two complementary purposes. First, it should have a lenient evaluation setting that allows for comparing and analyzing existing methods. Second, it should pose a significant challenge to remain relevant for more advanced future methods. To this end, we create each CRAX suite in three difficulty levels. The lowest level tasks are designed such that most existing methods are capable of learning a reasonable policy and achieving meaningful performance, while the highest levels leave substantial room for improvement. The difficulty increase between levels depends on the nature of the environment. For example, in Safe Goal, higher levels introduce a greater number and variety of hazards that the agent must avoid (Figure 2). In Safe Spider, each successive level requires the Spider agent to keep one additional leg off the ground to avoid costs. Appendix A provides the exact parameter settings defining each difficulty level, and Figure 9 visually depicts the difficulty levels of the tasks. 5
25
50 0
Reward
40
Cost
Reward
6
20 0
PPO
PPOCost
PPOPID
481 244
980
Safe Height
1e3
4 2 0
PPOLag
60 40 20 0
269
67
0
Safe Circle 100
Cost
50
Cost
0
Safe Push
75
Reward
Cost
10
60 40 20 0
287
Safe Spider 60 40 20 0
Safe Velocity
7.5 5.0 2.5 0.0
Cost
10
20 0
1e3
0
20
Cost
94
99
0
Reward
Cost
Reward
0.0
40
100
20
0.5
Safe Reacher
200
Safe Pathway
1e4 1.0
Reward
60 40 20 0
Reward
30 20 10 0
Cost
Reward
Safe Goal
20 10 0
PPOSaute
P3O
FOCOPS
Threshold
Figure 3: Rewards and costs of baseline methods on Level 1 tasks after 500M environment steps used for training. Error bars denote 95% confidence intervals across five seeds.
Table 3: Algorithm summary results across 8 CRAX environments. Wins: number of environments where the algorithm achieved the highest reward while being safe (cost < 25). Safe%: percentage of environments where the algorithm was safe. Total: sum of wins and average safety percentage across levels. Green indicates 100% safe. Level 1
5
Level 2
Level 3
Total
Algorithm
Wins ↑
Safe% ↑
Wins ↑
Safe% ↑
Wins ↑
Safe% ↑
Wins ↑
Safe% ↑
PPO PPOCost PPOLag PPOPID PPOSaute P3O FOCOPS
0 1 1 2 1 2 1
25 50 62 88 38 62 62
1 1 0 0 0 3 3
12 38 100 50 25 75 88
0 1 0 1 0 3 3
0 25 100 88 12 75 88
1 3 1 3 1 8 7
12 38 88 75 25 71 79
Empirical Evaluation
To assess CRAX, we evaluate several popular SafeRL baselines and one unconstrained RL baseline. (1) We include PPO [34] to serve as an unconstrained reference point, ignoring costs entirely. (2) PPOCost [41] extends PPO by treating costs as negative rewards. (3) PPOLag [31], a primal–dual approach that updates both the policy and a learned Lagrange multiplier to balance return and safety. (4) PPOPID [36] refines PPOLag’s strategy by updating the Lagrange multiplier with a proportional–integral–derivative controller, allowing the reward–safety trade-off to adjust more responsively during training. (5) PPOSauté [35] augments the observed state with a safety budget, treating the constraint as part of the dynamics. (6) P3O [46] progressively increases a cost penalty coefficient when constraints are violated, encouraging the policy to adapt toward feasibility without explicit dual updates. (7) Finally, FOCOPS [47] enforces safety by constraining policy updates through a trust-region formulation, optimizing reward while explicitly bounding expected cost. 6
0
20
75
100
50
50
0
0
mal lum ansfer Nor urricu Tr C 1e3 4
0.5 0.0
mal lum ansfer Nor urricu Tr C
Reward
20
Cost
Reward
mal mal lum ansfer lum ansfer Nor urricu Nor urricu Tr Tr C C Safe Pathway 1.0 1e4 10 0
mal lum ansfer Nor urricu Tr C PPOLag PPOPID
2
Cost
20
40
150
25 0
mal lum ansfer Nor urricu Tr C Safe Height 10
Cost
Cost
Reward
40
Safe Reacher
Reward
Safe Goal 60
5
0
0 al m nsfer u m mal l lum ansfer r u No urric Nor urricu Tra Tr C C P3O FOCOPS Threshold
Figure 4: Curriculum learning and safety transfer in CRAX environments. We compare direct training (Normal), curriculum learning across difficulty levels (Curriculum), and transfer from an unconstrained PPO policy (Transfer) on Level 3 tasks.
Experimental Setup. We run each experiment for 500 million environment steps, repeated over 5 seeds. All experiments are conducted on a dedicated compute node with a 72-core 3.2 GHz AMD EPYC 7F72 CPU and a single NVIDIA H100 GPU. Appendix C provides the exact hyperparameters. 5.1
Baseline Algorithm Analysis
Figure 3 shows the performance of baseline algorithms. PPO focuses entirely on maximizing rewards, providing a rough sense of the return achievable when safety is ignored. PPOCost simply subtracts costs from the reward, settling into a compromise, but offers no guarantee of adhering to the constraint. Some methods show distinct affinities for constraint types. PPOSauté satisfies the cost bound in Reacher and Pathway but otherwise behaves close to unconstrained PPO. FOCOPS performs best on navigation tasks (Goal, Push) and Reacher, but worst on forward-locomotion tasks (Spider, Height, Pathway). Table 3 provides a summary of the evaluations. P3O and FOCOPS are the strongest baselines on CRAX. PPOLag achieves the highest safety percentage, being the only baseline to satisfy all cost bounds on Levels 2 and 3. However, it fails to reach high rewards. PPOPID is less conservative, trading stricter safety adherence for slightly higher performance. Appendix E provides more detailed baseline results and training curves. 5.2
Curriculum Learning and Safety Transfer
As seen in Figure 3, when trained directly on the hardest difficulty level, agents often struggle to discover a good policy, as exploration becomes dominated by constraint violations and sparse progress. Training RL agents on progressively more complex settings has been shown to substantially improve learning efficiency, a paradigm commonly referred to as curriculum learning [4, 27]. In parallel, transfer learning aims to reuse knowledge acquired in one environment to accelerate learning in a related setting [39, 48]. We investigate whether curriculum and transfer learning can improve performance on the most difficult tasks in CRAX. In the curriculum setting, agents are trained sequentially on increasing difficulty levels, carrying over parameters between stages, with the data budget split equally across levels. This exposes the agent to simpler dynamics before confronting denser hazards. In the transfer setting, we first train an unconstrained PPO policy directly on Level 3 and subsequently use its parameters to initialize a safe RL algorithm, which is then trained with the remaining half of the allowed timesteps. 7
Number of Parallel Environments
213 211 29 27 25 23 21
CRAX (Ours) Safety-Gymnasium Ideal Scaling
1 2 4 8 16 32 6 124 258 56 1012 2024 4048 8196 92
103
Speedup Factor
104
1 2 4 8 16 32 6 124 258 56 1012 2024 4048 8196 92
Steps per Second (SPS)
105
CRAX (Ours) Safety-Gymnasium
Number of Parallel Environments
Figure 5: Throughput comparison between CRAX and Safety-Gymnasium. Left: CRAXachieves up to ∼300K steps per second and roughly two orders of magnitude higher throughput than SafetyGymnasium. Right: CRAX closely follows ideal scaling up to hundreds of environments, while Safety-Gymnasium plateaus early due to CPU and memory bottlenecks. As shown in Figure 4, the impact of curriculum learning and transfer is strongly environment- and algorithm-dependent. In Safe Goal, neither curriculum nor transfer improves over direct training. In Safe Reacher, curriculum learning boosts performance for all methods except PPOPID. Transfer proves ineffective, as all methods except FOCOPS violate the threshold after transfer, suggesting that unconstrained policies struggle to adapt for safety in this environment. Curriculum learning improves performance In Safe Pathway and Safe Height, where as safe transfer yield benefits only in Safe Pathway. Here, direct training yields overly conservative policies that underutilize the available safety budget, while safety-transferred agents make better use of it. Interestingly, this is the opposite in Safe Goal. Overall, these results indicate that curriculum learning and transfer can be beneficial in some circumstances for learning challenging safety-constrained tasks. 5.3
Computational Efficiency: A Case Study
As discussed in Section 1, simulation throughput is a core bottleneck in SafeRL. To assess the extent to which this manifests in practice, we conduct a case study on scalability. Safety-Gymnasium (SG) is currently the most widely used benchmark for continuous-control SafeRL, and thus serves as a natural point of comparison. In particular, we evaluate how simulation throughput scales with the number of parallel environments, measuring steps per second (SPS) under identical hardware using CRAX Safe Point Goal Level 1 and SG SafetyPointGoal1-v0. Figure 5 shows that CRAX scales far beyond Safety-Gymnasium, reaching ∼ 300K steps per second (SPS) at around 8192 parallel environments, while Safety-Gymnasium saturates at low concurrency and fails to scale further, reaching only ∼ 3K SPS. Attempts to scale beyond 256 environments were unsuccessful, as the process exhausted available CPU memory. This early saturation and low SPS reflects the inherent limitations of CPUbound physics simulation. The two-order-of-magnitude gap in achievable throughput between CRAX and Safety-Gymnasium demonstrates the importance of hardware-accelerated simulation for largescale safe RL experimentation. To put this in concrete terms, our full evaluation suite (algorithms × environments × difficulty levels × seeds) amounts to hundreds of training runs and trillions of environment steps. On a single H100 GPU, CRAX completes this in 2 weeks. The equivalent run on Safety-Gymnasium would take close to a year. 5.4
Varying Safety Bounds
Safety requirements in real-world applications can vary substantially across use cases and are often accompanied by trade-offs with performance. For instance, autonomous vehicle navigation must contend with unpredictable pedestrian behavior and sensor limitations, making perfect safety unattainable in certain circumstances. The objective instead becomes to minimize unsafe behavior. In contrast, domains such as nuclear reactor control require absolute precision and tolerate no margin 8
0 0.0
0.5
1.0 1e8
Steps
0.0
0.5
Safe Block Push
Steps
1.0 1e8
0
0
2
4
Steps 1e8
Reward
50
Cost
Reward
100
50 0
2
4
100 0
0
2
20
4
0
2
4
2
4
Steps 1e8 Steps 1e8 Safe Point Circle 100
100 50 0
Steps 1e8 Bound 15
40
Cost
25
Safe Reacher
200
Cost
20
Reward
50
Cost
Reward
Safe Point Goal
Bound 25
2
4
Steps 1e8
50 0
Steps 1e8
Bound 35
Figure 6: PPOLag increasingly sacrifices rewards to adhere to tighter safety bounds in all tasks. for error. To support such requirements, CRAX environments expose an adjustable safety bound. We examine how varying this constraint induces safety–performance trade-offs by evaluating PPOLag under stricter and more lenient cost budgets d ∈ {15, 25 (default), 35}. As shown in Figure 6, PPOLag reliably adapts to the varied bounds across all environments, with a non-linear trade-off in performance. Tightening the bound leads to a substantially larger drop in score than the gains obtained by increasing it by the same amount. Moreover, some safe RL methods [42] are explicitly designed to enforce strict safety guarantees rather than negotiate trade-offs, and can therefore also be evaluated under the most stringent safety bounds.
6
Conclusions
We introduced CRAX, a hardware-accelerated benchmark for SafeRL built on high-fidelity 3D physics simulation, enabling large-scale experimentation that is infeasible with existing CPUbound benchmarks. Its diverse environment suites with difficulty progression allow evaluation of safety–performance trade-offs across agent morphologies. We empirically demonstrate the computational advantages of CRAX, the strengths and limitations of popular SafeRL algorithms, and the potential for curriculum learning and safety transfer to improve performance in challenging settings. We hope CRAX serves as a vital tool for developing, analyzing, and benchmarking future safe RL methods.
7
Limitations and Future Work
As of the time of writing, MJX does not yet support all features available in the CPU-based MuJoCo, such as certain rigid-body collision types. This restricts the range of scenarios that can currently be expressed within CRAX. In addition, our evaluation focuses exclusively on on-policy methods. Exploring off-policy and model-based safe RL approaches within CRAX remains an important direction for future work. In the scope of this work, we evaluate safety exclusively through expected cumulative cost constraints, leaving other safety formulations unexplored. Finally, our experiments only consider state-based observations and single-agent settings. Future work could investigate learning solely from pixel observations from an embodied perspective and extend CRAX to multiagent safety scenarios.
9
References [1] Eitan Altman. Constrained Markov Decision Processes. Routledge, 1st edition, 1999. doi: 10.1201/9781315140223. [2] Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling. The arcade learning environment: An evaluation platform for general agents. J. Artif. Intell. Res., 47:253–279, 2013. [3] Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling. The arcade learning environment: An evaluation platform for general agents (extended abstract). In IJCAI, pages 4148–4152. AAAI Press, 2015. [4] Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. Curriculum learning. In ICML, pages 41–48, 2009. [5] Matteo Bettini, Ryan Kortvelesy, Jan Blumenkamp, and Amanda Prorok. VMAS: A vectorized multi-agent simulator for collective robot learning. In DARS, pages 42–56, 2022. [6] Clément Bonnet, Daniel Luo, Donal Byrne, Shikha Surana, Sasha Abramowitz, Paul Duckworth, Vincent Coyette, Laurence Illing Midgley, Elshadai Tegegn, Tristan Kalloniatis, Omayma Mahjoub, Matthew Macfarlane, Andries P. Smit, Nathan Grinsztajn, Raphaël Boige, Cemlyn N. Waters, Mohamed A. Mimouni, Ulrich A. Mbou Sob, Ruan de Kock, Siddarth Singh, Daniel Furelos-Blanco, Victor Le, Arnu Pretorius, and Alexandre Laterre. Jumanji: a diverse suite of scalable reinforcement learning environments in JAX. In ICLR, 2024. [7] Michal Bortkiewicz, Wladyslaw Palucki, Vivek Myers, Tadeusz Dziarmaga, Tomasz Arczewski, Lukasz Kucinski, and Benjamin Eysenbach. Accelerating goal-conditioned reinforcement learning algorithms and research. In ICLR, 2025. [8] Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. OpenAI Gym. arXiv preprint arXiv:1606.01540, 2016. [9] Maxime Chevalier-Boisvert, Bolun Dai, Mark Towers, Rodrigo Perez-Vicente, Lucas Willems, Salem Lahlou, Suman Pal, Pablo Samuel Castro, and Jordan K. Terry. Minigrid & Miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks. In NeurIPS, 2023. [10] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, pages 248–255, 2009. [11] Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun. CARLA: An open urban driving simulator. In CoRL, pages 1–16, 2017. [12] C. Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem. Brax - A differentiable physics engine for large scale rigid body simulation. In NeurIPS Datasets and Benchmarks, 2021. [13] Javier García and Fernando Fernández. A comprehensive survey on safe reinforcement learning. J. Mach. Learn. Res., 16:1437–1480, 2015. [14] Sven Gronauer. Bullet-Safety-Gym: A framework for constrained reinforcement learning. Technical report, mediaTUM, 2022. [15] Matthew Thomas Jackson, Uljad Berdica, Jarek Luca Liesen, Shimon Whiteson, and Jakob Nicolaus Foerster. A clean slate for offline reinforcement learning. In NeurIPS, 2025. [16] Jiaming Ji, Borong Zhang, Jiayi Zhou, Xuehai Pan, Weidong Huang, Ruiyang Sun, Yiran Geng, Yifan Zhong, Josef Dai, and Yaodong Yang. Safety gymnasium: A unified safe reinforcement learning benchmark. In NeurIPS, 2023. [17] Danial Kamran, Thiago D Simão, Qisong Yang, Canmanie T Ponnambalam, Johannes Fischer, Matthijs TJ Spaan, and Martin Lauer. A modern perspective on safe automated driving for different traffic dynamics using constrained reinforcement learning. In ITSC, pages 4017–4023, 2022. 10
[18] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. ImageNet classification with deep convolutional neural networks. In NIPS, pages 1106–1114, 2012. [19] Robert Tjarko Lange. gymnax: A JAX-based reinforcement learning environment library, 2022. URL http://github.com/RobertTLange/gymnax. [20] Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg. AI safety gridworlds. arXiv preprint arXiv:1711.09883, 2017. [21] Quanyi Li, Zhenghao Peng, Lan Feng, Qihang Zhang, Zhenghai Xue, and Bolei Zhou. MetaDrive: Composing diverse driving scenarios for generalizable reinforcement learning. IEEE Trans. Pattern Anal. Mach. Intell., 45(3):3461–3475, 2022. [22] Michael T. Matthews, Michael Beukman, Benjamin Ellis, Mikayel Samvelyan, Matthew Thomas Jackson, Samuel Coward, and Jakob Nicolaus Foerster. Craftax: A lightning-fast benchmark for open-ended reinforcement learning. In ICML, 2024. [23] Mayank Mittal, Pascal Roth, James Tigue, Antoine Richard, Octi Zhang, Peter Du, Antonio Serrano-Muñoz, Xinjie Yao, René Zurbrügg, Nikita Rudin, Lukasz Wawrzyniak, Milad Rakhsha, Alain Denzler, Eric Heiden, Ales Borovicka, Ossama Ahmed, Iretiayo Akinola, Abrar Anwar, Mark T. Carlson, Ji Yuan Feng, Animesh Garg, Renato Gasoto, Lionel Gulich, Yijie Guo, M. Gussert, Alex Hansen, Mihir Kulkarni, Chenran Li, Wei Liu, Viktor Makoviychuk, Grzegorz Malczyk, Hammad Mazhar, Masoud Moghani, Adithyavairavan Murali, Michael Noseworthy, Alexander Poddubny, Nathan Ratliff, Welf Rehberg, Clemens Schwarke, Ritvik Singh, James Latham Smith, Bingjie Tang, Ruchik Thaker, Matthew Trepte, Karl Van Wyk, Fangzhou Yu, Alex Millane, Vikram Ramasamy, Remo Steiner, Sangeeta Subramanian, Clemens Volk, CY Chen, Neel Jawale, Ashwin Varghese Kuruttukulam, Michael A. Lin, Ajay Mandlekar, Karsten Patzwaldt, John Welsh, Huihua Zhao, Fatima Anes, Jean-Francois Lafleche, Nicolas Moënne-Loccoz, Soowan Park, Rob Stepinski, Dirk Van Gelder, Chris Amevor, Jan Carius, Jumyung Chang, Anka He Chen, Pablo de Heras Ciechomski, Gilles Daviet, Mohammad Mohajerani, Julia von Muralt, Viktor Reutskyy, Michael Sauter, Simon Schirm, Eric L. Shi, Pierre Terdiman, Kenny Vilella, Tobias Widmer, Gordon Yeoman, Tiffany Chen, Sergey Grizan, Cathy Li, Lotus Li, Connor Smith, Rafael Wiltz, Kostas Alexis, Yan Chang, David Chu, Linxi "Jim" Fan, Farbod Farshidian, Ankur Handa, Spencer Huang, Marco Hutter, Yashraj Narang, Soha Pouya, Shiwei Sheng, Yuke Zhu, Miles Macklin, Adam Moravanszky, Philipp Reist, Yunrong Guo, David Hoeller, and Gavriel State. Isaac Lab: A GPU-accelerated simulation framework for multi-modal robot learning. arXiv preprint arXiv:2511.04831, 2025. [24] Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. Human-level control through deep reinforcement learning. Nat., 518(7540):529–533, 2015. [25] Sharada P. Mohanty, Jyotish Poonganam, Adrien Gaidon, Andrey Kolobov, Blake Wulfe, Dipam Chakraborty, Grazvydas Semetulskis, João Schapke, Jonas Kubilius, Jurgis Pasukonis, Linas Klimas, Matthew J. Hausknecht, Patrick MacAlpine, Quang Nhat Tran, Thomas Tumiel, Xiaocheng Tang, Xinwei Chen, Christopher Hesse, Jacob Hilton, William Hebgen Guss, Sahika Genc, John Schulman, and Karl Cobbe. Measuring sample efficiency and generalization in reinforcement learning benchmarks: NeurIPS 2020 Procgen benchmark. In NeurIPS (Competition and Demos), pages 361–395, 2020. [26] MuJoCo XLA Authors. MuJoCo XLA (MJX) - MuJoCo documentation. https://mujoco. readthedocs.io/en/stable/mjx.html, 2023. [Accessed 28-01-2026]. [27] Sanmit Narvekar, Bei Peng, Matteo Leonetti, Jivko Sinapov, Matthew E Taylor, and Peter Stone. Curriculum learning for reinforcement learning domains: A framework and survey. J. Mach. Learn. Res., 21(181):1–50, 2020. [28] Alexander Nikulin, Vladislav Kurenkov, Ilya Zisman, Artem Agarkov, Viacheslav Sinii, and Sergey Kolesnikov. XLand-MiniGrid: Scalable meta-reinforcement learning environments in JAX. In NeurIPS, 2024. 11
[29] Martin L Puterman. Markov Decision Processes: Discrete Stochastic Dynamic Programming. John Wiley & Sons, Inc., 1 edition, 1994. [30] Asha Ramanujam, Adam Elyoumi, Hao Chen, Sai Madhukiran Kompalli, Akshdeep Singh Ahluwalia, Shraman Pal, Dimitri J. Papageorgiou, and Can Li. SafeOR-Gym: A benchmark suite for safe reinforcement learning algorithms on practical operations research problems. arXiv preprint arXiv:2506.02255, 2025. [31] Alex Ray, Joshua Achiam, and Dario Amodei. Benchmarking safe exploration in deep reinforcement learning. arXiv preprint arXiv:1910.01708, 2019. [32] Julien Roy, Roger Girgis, Joshua Romoff, Pierre-Luc Bacon, and Christopher J. Pal. Direct behavior specification via constrained reinforcement learning. In ICML, pages 18828–18843, 2022. [33] Alexander Rutherford, Benjamin Ellis, Matteo Gallici, Jonathan Cook, Andrei Lupu, Garðar Ingvarsson, Timon Willi, Ravi Hammond, Akbir Khan, Christian Schröder de Witt, Alexandra Souly, Saptarashmi Bandyopadhyay, Mikayel Samvelyan, Minqi Jiang, Robert T. Lange, Shimon Whiteson, Bruno Lacerda, Nick Hawes, Tim Rocktäschel, Chris Lu, and Jakob N. Foerster. JaxMARL: Multi-agent RL environments and algorithms in JAX. In NeurIPS, 2024. [34] John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017. [35] Aivar Sootla, Alexander I. Cowen-Rivers, Taher Jafferjee, Ziyan Wang, David Henry Mguni, Jun Wang, and Haitham Ammar. Saute RL: almost surely safe reinforcement learning using state augmentation. In ICML, pages 20423–20443, 2022. [36] Adam Stooke, Joshua Achiam, and Pieter Abbeel. Responsive safety in reinforcement learning by PID Lagrangian methods. In ICML, pages 9133–9143, 2020. [37] Richard S. Sutton and Andrew G. Barto. Reinforcement learning - an introduction, 2nd Edition. MIT Press, 2018. [38] Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy P. Lillicrap, and Martin A. Riedmiller. DeepMind control suite. arXiv preprint arXiv:1801.00690, 2018. [39] Matthew E. Taylor and Peter Stone. Transfer learning for reinforcement learning domains: A survey. J. Mach. Learn. Res., 10:1633–1685, 2009. [40] Emanuel Todorov, Tom Erez, and Yuval Tassa. MuJoCo: A physics engine for model-based control. In IROS, pages 5026–5033, 2012. [41] Tristan Tomilin, Meng Fang, and Mykola Pechenizkiy. HASARD: A benchmark for visionbased safe reinforcement learning in embodied agents. In ICLR, 2025. [42] Akifumi Wachi, Xun Shen, and Yanan Sui. A survey of constraint formulations in safe reinforcement learning. In IJCAI, pages 8262–8271, 2024. [43] Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine. Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning. In CoRL, volume 100 of Proceedings of Machine Learning Research, pages 1094– 1100. PMLR, 2019. [44] Zhaocong Yuan, Adam W. Hall, Siqi Zhou, Lukas Brunke, Melissa Greeff, Jacopo Panerati, and Angela P. Schoellig. Safe-control-gym: A unified benchmark suite for safe learning-based control and reinforcement learning in robotics. IEEE Robotics Autom. Lett., 7(4):11142–11149, 2022. [45] Kevin Zakka, Baruch Tabanpour, Qiayuan Liao, Mustafa Haiderbhai, Samuel Holt, Jing Yuan Luo, Arthur Allshire, Erik Frey, Koushil Sreenath, Lueder A Kahrs, et al. MuJoCo playground. arXiv preprint arXiv:2502.08844, 2025. 12
[46] Linrui Zhang, Li Shen, Long Yang, Shixiang Chen, Bo Yuan, Xueqian Wang, and Dacheng Tao. Penalized proximal policy optimization for safe reinforcement learning. arXiv preprint arXiv:2205.11814, 2022. [47] Yiming Zhang, Quan Vuong, and Keith Ross. First order constrained optimization in policy space. In NeurIPS, pages 15338–15349, 2020. [48] Markel Zubia, Thiago D Simão, and Nils Jansen. Robust transfer of safety-constrained reinforcement learning agents. In ICLR, 2025.
13
A
Environment Descriptions
A.1
Overview
Task
Compatible Agents
Constraint Type
Cost Mechanism
Goal Button Circle Push
Point, Ant, Humanoid, Spider Point, Ant, Humanoid, Spider Point, Ant, Humanoid, Spider Point, Ant, Humanoid, Spider
Spatial avoidance Spatial + Selection Spatial + Boundary Spatial avoidance
Contact / Proximity Contact + Wrong button Proximity + Boundary Contact / Proximity
3 3 3 3
Velocity Height
All locomotion agents Humanoid
Speed limit Posture
Binary / Hinge Soft hinge
3 3
Pathway Reach SpiderLegs
Walker2d, HalfCheetah, Hopper Reacher Spider
Foot placement Spatial avoidance Gait constraint
Quadratic penetration Binary intersection Binary contact
3 3 3
A.2
Task Descriptions
A.2.1
Navigation Suite
Goal. Agents: Point, Ant, Humanoid, Spider Navigate to goal regions while avoiding hazards scattered throughout an arena. When the agent reaches a goal, it respawns at a new random location. Hazards may be collidable (blocking, with contact-based cost) or non-collidable (pass-through, with proximity-based cost). This is a continuous task with no terminal success state. Difficulty levels increase the number of hazards and introduce mixed hazard types (cylinders and cubes). Button. Agents: Point, Ant, Humanoid, Spider Navigate to press the correct “active” button among multiple buttons while avoiding hazards and optional moving gremlins. The active button is visually indicated and observable through a compass sensor. Pressing a wrong button can optionally incur a cost. In continual mode, a new button becomes active after each success. Level 1 has 4 hazards and 4 gremlins in a arena. Level 2 increases to 8 hazards and 6 gremlins in a larger arena. Level 3 has 12 hazards and 8 faster gremlins in a smaller arena, making navigation more challenging. Circle. Agents: Point, Ant, Humanoid, Spider Maintain a circular orbit at a target radius around a fixed center point. The agent is rewarded for tangential velocity (moving along the circle) and penalized for deviating from the target radius. Difficulty levels progressively add boundary constraints and hazards: Level 1 has x-boundaries only with no hazards. Level 2 adds a square boundary with 1 cylinder hazard. Level 3 has a smaller boundary with 2 cylinder hazards. Push. Agents: Point, Ant, Humanoid, Spider Push a movable block into goal regions while avoiding hazards. Unlike Goal, the block (not the agent) must reach the goal. The agent must coordinate approaching the block and pushing it in the correct direction. Only agent-hazard interactions incur cost; the block passes through hazards freely. Difficulty levels increase goal movement speed: Level 1 is stationary, Level 2 moves slowly, and Level 3 moves fast. A.2.2
Other Suites
Velocity. Agents: Ant, HalfCheetah, Hopper, Humanoid, Walker2d Standard locomotion with an added velocity constraint. The agent must maximize forward progress while keeping its speed below a threshold. Three difficulty levels progressively tighten the speed limit (100%, 75%, 50% of baseline). Height. Agents: Humanoid The humanoid must move forward while staying below a maximum height threshold, simulating a 14
Levels
Figure 7: In levels 2 and 3 of Push, the goal that the agent must push the block into moves at a fixed velocity. To succeed, the agent ought to anticipate the goal’s trajectory while avoiding hazard zones.
Figure 8: The Pathway agent incurs a per-step cost when its foot contacts a hazard, scaled by the penetration depth into the hazard region.
low ceiling constraint. Cost increases smoothly as height exceeds the threshold, encouraging the agent to crouch. Based on the HumanoidStandup environment with added forward locomotion reward. Difficulty levels lower the maximum height requirement. Pathway. Agents: Walker2d, HalfCheetah, Hopper A bipedal or hopping agent traverses a path with non-collidable hazard zones placed along its route. Cost is incurred when feet step inside hazard regions, with deeper penetration causing higher cost. Hazards are randomly placed with varying gaps and lateral offsets, requiring the agent to time its steps carefully. Falling incurs a terminal cost. Difficulty levels decrease the maximum gap between hazards. Reach. Agents: Reacher (2-link arm) A planar 2-link robotic arm must reach a randomly placed target while avoiding flat hazards scattered on the workspace. Cost is incurred when any part of the arm intersects a hazard. The arm is sampled at discrete points along both links to detect collisions. Reward emphasizes proximity to the target with a bonus for reaching it. Difficulty levels increase the number of hazards: Level 1 (4), Level 2 (7), Level 3 (10). SpiderLegs. Agents: Spider (6-legged) A hexapod spider must walk forward while keeping specified legs off the ground, forcing unusual gaits. Level 1 restricts 2 diagonal legs, Level 2 restricts 3 legs (tripod pattern), and Level 3 restricts 4 legs (only center legs may touch). Cost is incurred each timestep a restricted foot contacts the floor. A.3 A.3.1
Technical Reference Notation
rt ct d(a, b) 1[·]
reward at timestep t cost at timestep t Euclidean distance indicator function 15
A.3.2
Reward and Cost Components
Across all environments, CRAX constructs the per-timestep reward Rt and cost Ct from a small set of common components. Each environment instantiates a subset of these terms with environmentspecific weights. Reward Components. • Forward progress: rforward = (xt − xt−1 )/∆t, where xt denotes the agent’s position along the forward axis at timestep t, and ∆t is the control timestep. • Survival bonus: A constant rhealthy while the agent remains in a valid state (e.g., not fallen). • Goal reward: A sparse reward rgoal for reaching goal locations, plus an optional dense distance-shaping term rdist = dt−1 − dt . • Control penalty: rctrl = −wctrl
2 m i ai , where a ∈ R is the action vector.
P
Cost Components. Let pt ∈ R2 denote the agent’s position in the horizontal plane, H the set of hazards, and h ∈ H an individual hazard. • Hazard proximity/contact (Goal, Push, Button, Circle, Reach): Binary penalty upon collision with collidable hazards, or continuous penalty when inside non-collidable keep-out zones scaling with penetration depth. • Velocity threshold (Velocity): Costs for exceeding speed limits, with binary penalties 1[v > τ ] or hinge-style penalties max(0, v − τ ) that increase with violation magnitude. • Height/posture (Height): Soft hinge costs when the agent’s center of mass exceeds a maximum height threshold, encouraging crouching. Penalty scales smoothly with violation degree. • Gait restriction (SpiderLegs): Binary costs when restricted feet make ground contact, forcing constrained locomotion patterns. • Foot placement (Pathway): Quadratic penetration costs when feet step inside hazard regions, with deeper penetration causing higher cost. Includes a terminal penalty for falling. A.3.3
Reward Functions
Task
Reward Formula
Goal Button
rt = αdist (dt−1 − dt ) + αgoal · nreached rt = αdist (dt−1 − dt ) + αgoal · 1[pressed active]
Circle
rt = vtan · (1 + |ractual − rtarget |)−1 · α
Push Velocity
bg ab ab rt = αbg (dbg t−1 − dt ) + αgoal · 1[reached] + αab (dt−1 − dt ) rt = α · rbase
Height Pathway Reach
rt = vforward + 1.0 − 0.01∥a∥2 rt = α(vforward + rhealthy ) rt = α(1 − dt /dmax )γ + rb · 1[dt < ϵ]
SpiderLegs
rt = vforward + rhealthy − 0.5∥a∥2 16
A.3.4
Cost Functions
Task
Type
Goal, Push
Height
Contact Proximity (cyl) Proximity (cube) Hazard + Wrong Prox + Boundary Binary Hinge Soft hinge
Pathway Reach SpiderLegs
Quadratic Binary Binary
Button Circle Velocity
Cost Formula P ct = i ccol · 1[contact(a, hi )] P ct = i cprox · max(0, 1 − di /ri ) P ct = i cprox · 1[|dxi | ≤ s ∧ |dyi | ≤ s] ct = chazard + cwrong · 1[wrong pressed] ct = chazard + cb · 1[outside boundary] ct = w · 1[vt > τ ] ct = w · max(0, vt − τ ) ct = w · max(0, ht − hmax )/δ P (i) ct = β i maxf [max(0, 1 − df /ri )]2 + cterm · 1[fell] P ct = β i 1[arm ∩ hi ] P ct = β f ∈Frestr 1[contact(f, floor)]
A.3.5
Default Parameters
Task
Parameter
Default
Description
Goal
reward_goal cost_scale collision_cost
1.0 2.0 3.0
Reward per goal reached Proximity cost multiplier Contact cost per hazard
Button
button_count wrong_button_cost
4 1.0
Number of buttons Cost for wrong press
Circle
circle_radius boundary_cost
1.5 1.0
Target orbit radius Cost for boundary violation
Push
goal_velocity agent_block_scale
0.0 0.1
Goal movement speed Agent-to-block reward weight
Velocity
level cost_mode reward_scaler
1 binary 0.01
Difficulty (1, 2, or 3) Cost type (binary/hinge) Reward scaling
Height
max_height hinge_margin
1.15 0.08
Maximum CoM height Soft hinge width
Pathway
num_hazards hazard_radius terminal_cost
100 0.25 5.0
Hazards along path Cylinder radius Cost for falling
Reach
num_hazards samples_per_link
10 5
Hazards in workspace Collision check density
SpiderLegs
restricted_feet cost_scale
(varies) 1.0
Legs that must stay up Cost per violation
A.4
Difficulty Levels
Figure 9 provides a visual overview of how environments change across difficulty levels. 17
Level 2
Level 3
Height
Circle
SpiderLegs
Pathway
Reach
Goal
Level 1
Figure 9: Visual comparison of difficulty levels across environments. Increasing difficulty generally adds more hazards, tightens constraints, or reduces margins for error.
18
A.4.1
Velocity Thresholds by Level
Agent
Level 1 (1.0×)
Level 2 (0.75×)
Level 3 (0.5×)
2.62 3.21 0.74 1.41 2.34
1.97 2.41 0.56 1.06 1.76
1.31 1.60 0.37 0.71 1.17
Ant HalfCheetah Hopper Humanoid Walker2d A.4.2 Level 1 2 3 A.4.3 Level 1 2 3 A.4.4
SpiderLegs Difficulty Levels Restricted Feet
Gait Pattern
front-left, back-right front-left, mid-right, back-left front-left, front-right, back-left, back-right
Diagonal constraint Alternating tripod Center legs only
Goal Difficulty Levels Hazard Configuration 12 cylinders 8 + 8 cylinders 6 cubes + 6 cylinders (prox) + 4 cubes + 4 cylinders (col) Circle Difficulty Levels
Level
X Boundary
Y Boundary
Hazards
1 2 3
±1.125 ±1.05 ±0.975
None ±1.05 ±0.975
0 1 2
A.4.5
Push Difficulty Levels
Level
Goal Velocity
Description
1 2 3
0.0 0.3 0.6
Stationary goal Slow moving goal Fast moving goal
A.4.6
Button Difficulty Levels
Level
Hazards
Gremlins
Arena Extents
Gremlin Travel
1 2 3
4 8 12
4 6 8
±1.5 ±1.8 ±1.2
0.35 0.35 0.45
A.4.7
Pathway Difficulty Levels
Level
Max Gap (m)
Description
1 2 3
6.0 4.0 2.0
Wide gaps Medium gaps Narrow gaps 19
Proximity
Collidable
12 8 12
0 8 8
A.4.8
Reach Difficulty Levels
Level
Number of Hazards
1 2 3
4 7 10
A.4.9
Height Difficulty Levels
Level
Max Height (m)
Description
1 2 3
1.30 1.15 1.00
Slight crouch Medium crouch Deep crouch
B
Agent Descriptions
B.1
Overview
Agent
Type
Actions
Observations
DoF
Point Ant Humanoid Spider
Holonomic sphere 3D quadruped 3D bipedal 3D hexapod
2 8 17 12
62 27 376 35
3 15 23 19
HalfCheetah Walker2d Hopper
2D planar runner 2D bipedal 2D one-legged
6 6 3
18 17 11
9 9 6
Reacher
2-link arm
2
11
4
B.2 B.2.1
Agent Descriptions 3D Navigation Agents
Point. Actions: 2 Observations: 62 A simple holonomic sphere that can move in any direction on a 2D plane. Controlled via forward thrust and angular velocity. Includes accelerometer, velocimeter, gyro, and magnetometer sensors plus configurable lidar and compass observations. Used in Goal, Button, Circle, and Push tasks. Ant. Actions: 8 Observations: 27 A four-legged 3D robot with torque-controlled joints. Each leg has two actuated joints (hip and ankle), totaling 8 actuators. Observations include joint positions and velocities plus contact forces. Compatible with navigation tasks and velocity constraints. Humanoid. Actions: 17 Observations: 376 A complex 3D bipedal robot with 17 actuated joints spanning the legs, arms, and torso. The large observation space includes body inertia, velocity, and actuator forces. Used in navigation tasks, velocity constraints, and the Height constraint task. Spider. Actions: 12 Observations: 35 A six-legged 3D hexapod robot. Each leg has a hip joint and ankle joint, totaling 12 actuators. Observations include joint angles (excluding root position) and joint velocities. Compatible with navigation tasks and the SpiderLegs gait constraint task. B.2.2
2D Planar Locomotion Agents
HalfCheetah. Actions: 6 Observations: 18 A fast planar running robot with two legs optimized for forward velocity. Each leg has three joints 20
(hip, knee, ankle), totaling 6 actuators. Observations include joint angles and angular velocities. Used in Velocity and SkipHop tasks.
Walker2d. Actions: 6 Observations: 17 A 2D bipedal walker that must balance while moving forward. Each leg has three joints (hip, knee, ankle), totaling 6 actuators. Used in Velocity and SkipHop tasks.
Hopper. Actions: 3 Observations: 11 A single-legged hopping robot in 2D. Three actuators control the hip, knee, and ankle joints. Must hop forward while maintaining balance. Used in Velocity and SkipHop tasks.
B.2.3
Static Base Agents
Reacher. Actions: 2 Observations: 11 A 2-link planar robotic arm with a fixed base. Two rotational joints control the shoulder and elbow. Observations include joint angles, angular velocities, fingertip position, and target location. Used exclusively in the Reach task with spatial hazard avoidance.
B.3
Technical Reference
B.3.1
Healthy Bounds and Termination
Agent
Healthy Height Range
Healthy Angle Range
Point Ant Humanoid Spider HalfCheetah Walker2d Hopper Reacher
0.05–0.3 0.2–1.0 1.0–2.0 0.2–1.0 – 0.8–2.0 0.7–∞ –
– – – – – |θ| < 1.0 rad |θ| < 0.2 rad –
B.3.2
Terminates Yes Yes Yes Yes No Yes Yes No
Action Spaces
Agent
Dim
Actuator Description
Point Ant Humanoid Spider HalfCheetah Walker2d Hopper Reacher
2 8 17 12 6 6 3 2
Forward thrust (x-axis motor), angular velocity (z-axis) Hip and ankle torques for 4 legs Torques for legs (6), arms (6), abdomen (3), pelvis (2) Hip and ankle torques for 6 legs Back hip, back knee, back ankle, front hip, front knee, front ankle Right hip, right knee, right ankle, left hip, left knee, left ankle Hip, knee, ankle torques Shoulder and elbow torques 21
B.3.3
Observation Spaces
Agent
Dim
Observation Components
Point
62
Ant Humanoid
27 376
Spider HalfCheetah Walker2d Hopper Reacher
35 18 17 11 11
Sensors (12), goal lidar (16), hazard lidar (16), goal compass (2), hazard compasses (16) qpos (13), qvel (14), excluding root x,y qpos, qvel, cinert (body inertias), cvel (body velocities), qfrc_actuator qpos (17, excluding x,y), qvel (18) qpos (8, excluding root x), qvel (9), root z qpos (8, excluding root x), qvel (9) qpos (5, excluding root x), qvel (6) cos(θ), sin(θ) for joints, target position, fingertip-target distance, angular velocities
B.3.4
Physical Properties
Agent Point Ant Humanoid Spider HalfCheetah Walker2d Hopper Reacher
C
Total DoF
Bodies
Control Timestep
3 (x, y, θ) 15 (free root + 8 joints) 23 (free root + 17 joints) 19 (free root + 12 joints) 9 (root + 6 joints) 9 (root + 6 joints) 6 (root + 3 joints) 4 (2 joints + target)
1 13 13 13 8 7 4 3
0.008s 0.05s 0.015s 0.05s 0.05s 0.008s 0.008s 0.02s
Hyperparameters
Table 4 lists the configuration we use for our experiments.
D
Use of Large Language Models
LLM-based coding assistants were used during development to help implement the CRAX environments and the JAX reimplementations of the safe RL baselines. All generated code was reviewed, tested, and validated by the authors against reference implementations and the reported empirical results. LLMs were not used as part of any agent, policy, reward model, or evaluation procedure in this work.
E
Extended Results
In this section, we provide additional experimental results that complement the main findings and offer deeper insight into the behavior of the evaluated methods. E.1
Detailed Baseline Performance
Table 5 provides a comprehensive overview of the baseline results across all difficulty levels and tasks. E.2
Training curves
Figures 10–12 show training curves for all baseline methods across difficulty levels, illustrating differences in learning dynamics and convergence behavior.
22
Table 4: Fixed hyper-parameters used for all experiments in this paper unless stated otherwise. Parameter
Value Optimization / PPO core
Optimizer Learning rate η Entropy coef. αent Discount γ Reward scaling GAE λ PPO clip ϵ
Adam (Optax) 5 × 10−4 5 × 10−3 0.99 0.1 0.95 0.3 Network architecture
Actor network 4-layer MLP (32×4) Value network 5-layer MLP (256×5) Activation function Swish Observation normalization Running mean/variance Layer/Spectral normalization False Total parameters ∼2–3×105 Training scale / rollout Total env. steps N Episode length Parallel envs Unroll length Batch size Minibatches per update SGD updates per batch Eval passes during training Eval parallel envs Safety bound (episodic cost) Logging interval
108 2000 steps 2048 8 1024 32 6 5 128 25.0 106 env. steps
PPO-Cost 1.0
Cost weight
PPO-Lagrange 3.0 0.0
Lagrangian LR coef. Initial λlagr
PPO-PID PID gains (Kp , Ki , Kd ) PID integral clip PID λ clip PID derivative EMA β
(10.0, 0.01, 0.01) 1.0 106 0.95
PPO-Saute Budget discount factor Terminal violation penalty Normalize budget observation
0.99 (same as γ) −1.0 True
P3O Initial cost penalty κ κ increase factor Max κ
0.01 1.1 50.0 FOCOPS
Initial ν ν learning rate Max ν KL penalty coef. λfocops Advantage norm. temp. ηfocops
23
0.1 1.0 100.0 1.5 0.02
Table 5: Detailed algorithm comparison across environments and difficulty levels. R: Reward (↑ higher is better), C: Cost (↓ lower is better). Green indicates safe (cost < 25). Bold indicates best safe result. Safe Goal Safe Reacher Level 1
Level 2
Level 3
Level 1
Level 2
Level 3
AlgorithmR ↑
C↓
R↑
C↓
R↑
C↓
R↑
C↓
R↑
C↓
R↑
C↓
PPO 36.9 PPOCost 32.5 PPOLag 34.8 PPOPID 34.5 PPOSaute33.9 P3O 34.1 FOCOPS 37.2
99.3 27.5 25.1 24.6 94.0 26.0 25.4
30.0 23.5 16.7 25.6 61.3 64.1 73.3
97.3 50.2 7.3 25.3 126.9 23.4 24.6
25.7 14.8 7.0 8.6 52.7 56.0 63.6
192.7 87.5 18.8 26.0 191.0 23.5 26.6
206.7 188.2 197.2 168.1 176.6 196.2 204.7
41.5 32.3 25.3 21.4 6.2 24.4 25.5
203.6 144.9 66.6 129.9 104.2 139.3 177.5
74.1 51.4 24.9 26.1 13.1 24.7 24.6
206.2 78.6 18.5 42.7 21.8 47.5 125.4
105.6 36.4 23.1 23.6 11.6 20.8 24.2
Safe Pathway Level 1 AlgorithmR ↑
C↓
PPO 9015.7 14.9 PPOCost 8456.9 8.6 PPOLag 8960.2 16.3 PPOPID 4504.1 12.3 PPOSaute9095.2 17.7 P3O 9352.9 14.3 FOCOPS 830.0 6.1
Level 2 R↑
C↓
10433.821.8 5128.9 7.8 6823.0 20.0 4186.2 13.3 6978.6 19.5 7460.7 19.2 856.3 6.8
Safe Velocity Level 3 R↑
C↓
11024.536.7 4311.7 8.9 2076.2 14.8 1147.7 10.4 9304.0 35.1 2098.4 13.3 769.1 6.3
Level 1 R↑
C↓
4682.5 979.7 1988.9 0.9 2563.7 9.9 2821.7 16.7 2866.5 480.5 1205.4 15.1 188.7 3.0
Safe Spider Level 1
Level 2
Level 2 R↑
C↓
Level 3 R↑
C↓
5966.2 1244.7 2583.0 606.5 1542.8 0.4 888.8 1.5 1594.2 13.9 844.1 14.5 1258.6 9.0 902.5 4.2 1372.0 267.3 972.3 220.5 1785.6 18.1 856.4 19.7 119.8 3.3 79.1 4.5 Safe Push
Level 3
Level 1
Level 2
Level 3
AlgorithmR ↑
C↓
R↑
C↓
R↑
C↓
R↑
C↓
R↑
C↓
R↑
C↓
PPO 17.1 PPOCost 16.9 PPOLag 18.3 PPOPID 16.2 PPOSaute11.7 P3O 16.2 FOCOPS 0.4
67.3 45.9 11.0 11.9 49.8 13.4 1.5
17.0 16.3 4.9 10.2 13.0 13.6 0.1
59.1 56.7 8.8 14.7 76.1 14.5 1.0
15.6 17.5 0.0 -0.1 17.0 0.1 0.0
135.9 115.8 1.4 1.6 107.2 1.4 1.3
65.2 51.1 11.4 1.7 41.7 41.0 57.2
287.4 268.7 10.5 15.0 244.3 27.9 24.7
54.6 40.1 10.9 7.4 34.1 14.1 40.5
368.1 333.0 24.4 25.2 279.7 26.3 24.6
49.5 39.5 5.8 5.9 25.0 13.9 40.8
461.3 418.0 24.9 24.5 293.6 25.6 24.3
Safe Circle Level 1
Level 2
Safe Height Level 3
Level 1
AlgorithmR ↑
C↓
R↑
C↓
R↑
C↓
R↑
PPO 129.5 PPOCost 107.0 PPOLag 86.7 PPOPID 102.2 PPOSaute120.2 P3O 94.7 FOCOPS 98.5
45.8 10.9 25.9 25.0 46.0 30.8 30.8
127.2 82.8 64.5 68.9 128.6 60.1 62.2
98.8 9.3 24.2 26.5 97.3 28.3 25.1
127.5 80.8 53.5 39.9 123.1 61.3 57.7
99.4 38.1 18.7 21.6 96.8 27.5 24.0
4709.3 22.8 4229.0 18.6 2978.8 7.3 4392.2 7.0 4847.5 17.3 2293.2 6.4 2528.7 6.0
24
C↓
Level 2 R↑
C↓
3532.3 52.1 3691.0 33.8 1703.3 5.5 1087.3 5.4 5652.4 60.9 2310.9 4.5 1244.1 6.9
Level 3 R↑
C↓
3868.4 53.8 3811.3 57.5 1983.7 6.2 2657.9 6.0 5910.1 35.7 2753.3 5.4 376.4 7.7
Safe Reacher
0 2
Steps
4
0
1e8
2
Steps
4
Safe Push
100
200
0 0
2
Steps
4
0 1e8
Steps
4
0 1e8
0
2
Steps
4
20 0
2
Steps
4
1e8 1.0
2
Steps
PPO
4
1e8
PPOCost
Reward
Cost
20
0
2
Steps
4
10 2
Steps
PPOLag
4
1e8
PPOPID
0
0
4
2
4
2
4
Steps
1e8
Safe Height
2
Steps
4
40 30 20 10 0
1e8
0
Steps
1e8
Safe Velocity 1e3
1e4
2
0.5
1 0
0
2
Steps
PPOSaute
Figure 10: Level 1 training curves.
25
2
1e8
50
1e8
0.0 0
4
Steps
100
0
2
Safe Pathway
0
0
4 0
2
150
6
40
1e8
0
1e8
Safe Spider
1e3
Reward
Cost
Reward Reward
2
0
1e4 1.00 0.75 0.50 0.25 0.00
0
60
50
4
50
80
100
Steps
100
Safe Circle
150
0
Reward
Cost
Reward
200
2
150
400
300
Cost 0
1e8
50 25
Cost
0
75
Cost
0
20
100
Cost
100
Reward
40
Cost
Reward
200
Safe Goal 80 60 40 20 0
4
P3O
1e8
0
FOCOPS
Steps
1e8
Threshold
Safe Reacher 75
0
25
50
2
Steps
4
0
1e8
2
Steps
4
0 1e8
0
2
Steps
4
Safe Push
Steps
4
0 1e8
50 0
2
Steps
4
0 1e8
Cost
0.5
0
2
Steps
PPO
2
Steps
4
1e8 1.00
4
20 10 0
1e8
PPOCost
Cost Steps
4
4
2
4
2
4
2
4
Steps
1e8
0
0 1e8
2
Steps
4
2
Steps
PPOLag
4
1e8
PPOPID
0
Steps
1e8
0
Steps
1e8
Safe Velocity 1e3
1e4
2
0.75 0.50 0.25
1 0
0
2
Steps
PPOSaute
Figure 11: Level 2 training curves.
26
100 75 50 25 0
1e8
0.00 0
50
Safe Height
0.5 0.0
30
1.0
2
1.0
0
1e8
0 1e4
100 75 50 25 0
Safe Pathway
1e4
Reward
Steps
4
Reward
Cost
Reward
100
0.0
2
25
Safe Circle
150
0
0
50
Cost
2
2
100
75
Cost
0
200
Reward
0
Reward
Cost
Reward
100
0
100
400
200
0 1e8
Safe Spider
400 300
100 50
25
0 0
150
Cost
100
50
200
75
Reward
Cost
Reward
200
Safe Goal
100
4
P3O
1e8
0
FOCOPS
Steps
1e8
Threshold
Safe Reacher 100
0 2
Steps
4
1e8
2
Steps
4
0
1e8
Safe Push
0
2
Steps
4
0 1e8
2
Steps
4
0
1e8
75 50 25
0
2
Steps
4
0 1e8
Cost
0.5 0.0 2
Steps
PPO
2
Steps
4
4
1e8
PPOCost
Steps
4
0
2
4
2
4
2
4
2
4
Steps
1e8
100 0
1e8
2
Steps
PPOLag
4
1e8
PPOPID
2
Steps
75 50
4
0 1e8
0
Steps
1e8
Safe Velocity 1e3 2
0.5
1 0
0
2
Steps
PPOSaute
Figure 12: Level 3 training curves.
27
1e8
25
0.0 0
Steps
100
1.0
20
0
Safe Height
1e4
40
0
2
1e3
0
1e8
60
1.0
0
0
8 6 4 2 0
Safe Pathway
1e4
1e8
200
0
Reward
Cost
50 0
Reward
0
100
100
0
Safe Spider
1
Safe Circle
150
Reward
200
Reward
0
Steps
4
2
Reward
200
2
1e6
400
Cost
Reward
400
100
0 0
Cost
0
Cost
50
Cost
0
50
200
Cost
100
100
Reward
Cost
Reward
200
Safe Goal
4
P3O
1e8
0
FOCOPS
Steps
1e8
Threshold