Beyond Kinematics: Benchmarking Simulation Fidelity for Muscle-Driven Imitation Learning
arXiv:2609.21909v1 [cs.LG] 18 Sep 2026
Ayah G. Ahmad† , Claire E. Borden† , and Maegan Tucker Abstract— In this work, we conduct a systematic comparison of two state-of-the-art motion-imitation reinforcement learning (MIRL) pipelines, one built on SCONE/HyFyDy and one built on MuJoCo/MyoSim. HyFyDy emphasizes physiological realism through detailed musculotendon modeling, while MuJoCo prioritizes computational efficiency and scalable policy learning. While recent work has demonstrated that both pipelines reproduce human kinematics with high fidelity, it remains unclear if they accurately capture the underlying neuromuscular behavior that produced the movement. This limitation is particularly important for robotic assistive-device design and control, where outcome measures such as muscle activation patterns and metabolic cost are often used as optimization targets. To conduct a systematic comparison, our work compares both pipelines using a common set of human motion-capture and electromyography (EMG) measurements. The results find that while both pipelines produce similar kinematics with relative accuracy, the muscle activations from HyFyDy are more aligned with the experimental EMG, as supported by the average pooled (RMSE, r) values for muscle activations from HyFyDy and MuJoCo: (0.164, 0.4) and (0.344, 0.11), respectively. While we conclude that the more advanced physiological realism of HyFyDy currently makes it more suitable for musculoskeletal modeling, both require further development to bring physiological realism to GPU-parallelizable simulation environments and advance robotic assistive device design.
I. I NTRODUCTION Robotic assistive devices, including exoskeletons and prostheses, have the potential to improve mobility and independence for individuals with motor impairments [1], [2]. Developing controllers that generalize across users, tasks, and environments, however, remains a persistent challenge [3], [4]. Since human experiments are costly, time-intensive, and particularly burdensome for clinical populations, physiologically plausible musculoskeletal simulation has been proposed as a scalable alternative for accelerating controller development and device design [5]. Recent advances in reinforcement learning (RL) have further increased interest in musculoskeletal simulation as a platform for training and evaluating control policies [6]–[9]. However, the extent to which current simulation pipelines accurately reproduce the underlying neuromuscular behavior of humans remains unclear. This limitation is particularly important for assistive robotics, where measures such as muscle activation patterns, muscle recruitment strategies, and metabolic cost are often used to evaluate device efficacy [10]–[12]. Consequently, before musculoskeletal simulation can be relied upon as † Denotes equal contribution This work is supported by the Georgia Institute of Technology. Authors are with the Dynamic Mobility Lab at Georgia Tech, Atlanta, U.S. {ayah, cborden9, mtucker}@gatech.edu
Fig. 1. This work benchmarks the predictive capabilities of two stateof-the-art musculoskeletal modeling simulation environments intended for motion-imitation-style reinforcement learning: HyFyDy/SCONE (middle) and MyoSim/MuJoCo (bottom), against real-world human data modeled in OpenSim (top). The gait tiles show each musculoskeletal model completing one gait cycle with the average speed-binned muscle activations for the medial gastrocnemius throughout it. Darker and lighter colors represent slower and faster speeds, respectively.
a tool for assistive-device design, its ability to predict physiologically meaningful muscle-level outcomes must be benchmarked against experimental human data. To address this gap, we benchmark two state-of-the-art musculoskeletal simulation pipelines that employ motionimitation reinforcement learning (MIRL) to reproduce human locomotion from motion-capture demonstrations. In these approaches, human motion trajectories serve as reference signals for RL objectives, enabling the learned controller to reproduce observed movement while remaining physically consistent with the underlying musculoskeletal model. Our comparison moves beyond the evaluation of human kinematics alone and addresses whether these state-of-the-art pipelines are capable yet of accurately predicting the muscle activation patterns associated with those movements. The first pipeline we benchmark is the framework proposed by Choi et al. [9], which combines MIRL with the High-Fidelity Dynamics (HyFyDy) [13] musculoskeletal simulator, featuring detailed musculotendon dynamics. The second is MuscleMimic [14], a MuJoCo-based [15] MIRL framework designed for scalable training of large muscledriven humanoid models. Both approaches are trained from human motion demonstrations and have demonstrated strong kinematic tracking performance. However, neither framework has been systematically validated against experimental electromyography (EMG) data, and no direct comparison currently exists between them. In this work, we apply both MIRL frameworks in three comparative studies. First, a direct comparison of the two frameworks is performed using equivalent two-dimensional musculoskeletal models. Second, we examine the role of per-
Fig. 2. Pipelines for MuscleMimic and HyFyDy-IL with modifications for this comparative study. The first models correspond to the comparison between the frameworks with similar standard models; the second model is the scaled HyFyDy model for the second comparison; and the third models are the 3D models for the final comparison.
sonalization by evaluating both generic and subject-specific musculoskeletal models. Lastly, the effect of model complexity is explored by comparing two-dimensional and threedimensional implementations of each simulator. Together, these studies provide insight into how simulator fidelity, subject-specific modeling, and model dimensionality influence the physiological accuracy of MIRL. Our paper makes the following contributions: 1) We introduce a benchmark for evaluating the muscle activation prediction accuracy of MIRL pipelines using synchronized motion-capture and surface EMG measurements. 2) We provide the first direct comparison, to our knowledge, of the HyFyDy and MuscleMimic musculoskeletal simulation environments for MIRL. 3) We systematically evaluate the effects of simulator fidelity, subject-specific personalization, and model dimensionality on the prediction of muscle activation patterns during human locomotion. 4) We provide an OpenSim-to-reference data for MuscleMimic and a trained subject-specific HyFyDy model to facilitate future benchmarking and reproducible evaluation of musculoskeletal simulation frameworks. II. BACKGROUND A. Predictive Simulation Predictive simulation generates novel motions without necessitating the use of experimental data [16]. In the context of robotic assistive devices, this can enable testing hypotheses, device designs, and treatment strategies before hardware validation. OpenSim [17] is a dominant platform for physicsbased musculoskeletal simulation, which contains Hill-type muscle-tendon dynamics, inverse kinematics and dynamics, and a full-body model library. Additionally, OpenSim can be used with the Neuromusculoskeletal Modeling (NMSM) Pipeline to allow for patient-specific personalization [18]. Personalizing with NMSM, however, is time-consuming, taking over 25 hours to generate a model for an example
patient [18]. For RL, these costs compound. Each new patient requires hours of offline fitting before training can begin, and training itself requires millions of environment interactions. SCONE is an open-source program built for human and animal neuromuscular modeling that requires an additional musculoskeletal modeling software package, such as OpenSim [17], or HyFyDy [19]. HyFyDy includes the musculoskeletal models, packages, and scripts to simulate human motion, using a complete Millard musculotendon model with elastic tendons, variable pennation angles, and an errorcontrolled integrator running at approximately 7000 Hz [13]. Alternatively, MuJoCo is a different physics engine that is popular within the robotics community due to its high speed and soft contact models [15]. MuJoCo is widely used for developing controllers for robots such as bipeds and quadrupeds [20], [21]. It can also be paired with MyoSim [22], an open-source library with human musculoskeletal models. MuJoCo represents MyoSim’s musculotendon units as a simplified Hill-type model, with completely inelastic tendons and a fixed timestep of 1000 Hz [15]. Both HyFyDy and MuJoCo have machine learning integration capabilities, making them suitable for training predictive musculoskeletal models. In our work, we compare HyFyDy and MuJoCo directly, using two state-of-the-art MIRL frameworks: for HyFyDy, we use an implementation inspired by Choi et al. [9] (we term this framework HyFyDyIL); for MuJoCo, we use the MuscleMimic framework [14]. While both simulate human anatomy and locomotion, and have published impressive kinematic reference tracking, the degree to which they can predict human muscle activations patterns remains unclear. This is the outcome we target in our comparisons. B. Reinforcement Learning Comparison Among the limited literature studying HyFyDy and MuJoCo in the context of predictive human modeling, Schumacher et al., compare the two simulators with the same RL framework [23]. It was concluded that HyFyDy produces
useful for learning with limited data [25]. Alternatively, MuscleMimic is a GPU-based JAX implementation that operates on the MuJoCo physics engine. MuscleMimic’s policy includes a custom Proximal Policy Optimizer (PPO) that updates at each epoch, instead of after every certain number of epochs, to maintain on-policy learning and produce optimal results with parallel environment training [14], [26]. These underlying frameworks are kept constant to their originally published architectures to preserve consistency during our comparative study. B. Preprocessing
Fig. 3. Two-dimensional models MM240-2D (left), H1090 (middle), and H-AB08 (right). Not shown: MM240-3D and H2190 (visually identical to MM240-2D and H1090).
improved gait kinematics and ground reaction force predictions compared to MuJoCo when using models of similar complexity. The work also noted the limitations of both simulators in predicting muscle activation and their inability to produce accurate results, particularly for three-dimensional models with a similar muscle count. The inability of RL frameworks to accurately replicate muscle activations led some researchers to implement motion-imitation and attempt to model complex biological actuation in humans [9], [14]. Motion-imitation allows policies to move musculoskeletal models to mimic motion-capture data, decreasing the amount of necessary data to reproduce a motion and providing a method to potentially match poses and movements with muscle activations. C. State-of-the-art RL Frameworks As previously mentioned, our work directly compares two state-of-the-art MIRL frameworks for musculoskeletal modeling: HyFyDy-IL and MuscleMimic [9], [14]. Among the musculoskeletal modeling research, there exist other frameworks, such as Simos’s KINESIS motion-imitation architecture [7], yet these train for significantly longer times and increase computational expense drastically [14]. HyFyDy-IL and MuscleMimic published kinematic and muscle activation results, showcasing their capabilities. However, additional research must be conducted to identify the most suitable simulator for musculoskeletal modeling, determine if MIRL is adequate for predicting human locomotion, and investigate limitations and potential advancements to improve these frameworks and simulators. The goal of our work is to contribute to the gap in comparative studies and provide further insight into musculoskeletal modeling frameworks and simulators. III. S YSTEMATIC C OMPARISON A. Framework HyFyDy-IL is a CPU-based implementation that uses the SCONE physics engine, which incorporates Soft Actor-Critic (SAC) from Stable Baselines 3 [24]. Due to its off-policy nature, SAC is highly sample efficient, making it particularly
In the original architecture, MuscleMimic uses the AMASS [27] and KIT Motion-Language [28] datasets. The AMASS dataset provides SMPL pose and body parameters, which are fit to 17 anatomical mimic sites on the MyoFullBody model. The motion can then be retargeted using two methods: 1) Mocap-Body, which uses MuJoCo’s inverse kinematics, or 2) GMR-fit, which uses General Motion Retargeting’s [29] kinematic solver and applies joint and equality constraints to find joint configurations. Alternatively, HyFyDy-IL selects a reference walking motion obtained from the Scherpereel et al. open-sourced biomechanics dataset [30], which includes five different walking speeds from 0.6 m/s to 2.2 m/s for approximately three minutes. The treadmill motion is converted to an overground walking motion and post-processed in OpenSim to compute corresponding joint angles using inverse kinematics with retained positions for body segments. These become the targets for MIRL. To directly compare the two frameworks’ ability to learn human locomotion, we train both HyFyDy-IL and MuscleMimic on the same data: a 140-second clip from subject AB08 in the Scherpereel et al. dataset [30]. For HyFyDyIL, since the original framework uses a longer version of this data, no additional retargeting or preprocessing steps are performed. For MuscleMimic, however, the data is retargeted using GMR-fit to extract kinematic information for the mimic sites. This alignment of reference data preserves consistency in the preprocessing stage without affecting either policy’s architecture and allows both simulators to be benchmarked against real-world patient data. C. Rewards Due to their unique architectures, HyFyDy-IL and MuscleMimic have different reward structures. HyFyDy-IL has a DeepMimic-style reward [9], [31], focused on imitation tracking and effort penalties. The imitation rewards include joint position and velocity tracking and pelvic position tracking. Effort penalties impact the reward when 1) the muscle activations are too high and 2) their changes in activation are too significant, attempting to simulate real-life energy minimization while walking [32]. The total reward is: Rimit,H = wpos Rpos + wvel Rvel + wroot Rroot − wef f Cef f − w∆ef f C∆ef f
(1)
2 X Rpos = exp −βpos θ̂i − θi
! (2)
i
2 X˙ Rvel = exp −βvel θ̂ − θ̇i
!
reward structures are maintained for our comparative study, as changing either one could lead to additional changes that would alter the architecture and affect performance.
(3)
TABLE I
i 2
Rroot = exp −βroot ∥p̂root − proot ∥2 Ceff =
C∆eff =
1
M X e2i
µef f i=1
M X (ei,curr − ei,prev )2
1
µ∆ef f i=1
(4) (5)
Term k
Tracked quantity
βk
wk
(6)
pos vel root eff ∆eff
Kinematic pose error Kinematic velocity error End-effector position error Muscle activation magnitude Muscle excitation rate
1.0 1.0 80.0 1.0 1.0
0.60 0.15 0.25 −2.00 −2.00
The β terms are temperature parameters, M is the number of muscles, µ is the mean activation across all muscles, and the terms with a hat represent the reference motion targets, whereas those without are the predicted values from the simulation. Additionally, “eff” refers to the muscle effort. MuscleMimic’s rewards also have a DeepMimic-style [14]. The rewards are described below, with q and q̇ as the respective joint angle and angular velocity; p, θ, ω, and v as the respective mimic site position, orientation, angular velocity; and linear velocity, and r and ṙ are the respective root (pelvis) position and velocity. Rimit,MM = wq Rq + wq̇ Rq̇ + wp Rp + wθ Rθ + wω Rω + wv Rv !# " N 2 1 X ∗ qi,lin − qi,lin + θ̄ Rq = exp − βq N i=1
(7)
(8)
Nq
θ̄ =
1 X ∗ θ qj,quat , qj,quat Nq j=1 "
Nr
1 X Rr = exp −βr (ri − ri∗ )2 Nr i=1 " Rṙ = exp −βṙ
Nṙ
1 X (ṙi − ṙi∗ )2 Nṙ i=1
(9) # (10) # (11)
"
# K−1 X 1 ∗ Rp = exp −βp ∥pi − pi ∥2 K − 1 i=1 " # K−1 X 1 ∗ Rθ = exp −βθ ∥ϕi − ϕi ∥2 K − 1 i=1 # " K−1 X 1 ∗ Rω = exp −βv ∥ωi − ωi ∥2 K − 1 i=1 " # K−1 X 1 ∗ Rv = exp −βv ∥vi − vi ∥2 K − 1 i=1
O RIGINAL COEFFICIENTS FOR H Y F Y DY-IL ( TOP ) AND M USCLE M IMIC ( BOTTOM ) REWARD FUNCTIONS .
(12)
(13)
(14)
(15)
Here, N is the number of joints, K is the number of mimic sites, ϕ represents the rotation as described in [14], D is the dimension of the action space, and each quantity in the norms are computed relative to the reference frame (the pelvis). Despite their differences, both frameworks rely on imitating position and velocity and choose to weigh position more heavily than other factors, as shown in Table I. The
Total (
P
k |wk |)
5.00
Term k
Tracked quantity
βk
wk
q q̇ r ṙ p θ ω v
Joint positions Joint velocities Root coordinates (x, z, θ) Root twist Relative site positions Relative site orientations Relative site angular vel. Relative site linear vel.
10.0 2.0 10.0 10.0 100.0 10.0 0.1 0.1
0.10 0.10 0.10 0.10 0.60 0.01 0.10 0.10
Total(
P
k |wk |)
1.21
D. Musculoskeletal Models 1) Aligning Base Musculature and Bone Structure: The original HyFyDy-IL framework uses the three-dimensional H2190 model, with 21 degrees of freedom, 90 muscles, and no bone structure for the arms [9], while MuscleMimic uses MyoFullBody, a model with a complete bone structure, 72 degrees of freedom, 123 joints, and 416 muscles [14]. Specific changes are made to align the musculature and bone structure of the two models. The models differ in how the individual muscles are subdivided (eg, for each leg, MyoFullBody represents the gluteus maximus as three separate actuators, glmax1, glmax2, glmax3, while H2190 uses a single muscle). Thus, instead of restricting the MyoFullBody model to 90 muscle actuators, the primary muscle groups are aligned. The resultant model MM240 thus has 240 muscles, representing the same muscles as those in H2190. Additionally, for MM240, the arm structures are removed, and the bones of the torso are fused together to represent and match HyFyDy’s rigid torso in two dimensions. 2) Comparison Setup: We train five models: 1) H1090 (HyFyDy, 2D, 10 DoF, 90 muscle actuators): The H2190 model constrained to two-dimensions by removing 11 degrees of freedom. 2) H-AB08 (HyFyDy, 2D, 10 DoF, 90 muscle actuators): The H1090 model scaled to match patient AB08’s parameters from the Scherpereel et al. dataset [30] and represent the subject. It maintains the same number of muscles and DoFs, but the body geometric, mass, and inertia scale factors from the OpenSim model replace the default values in H1090. The muscle actuators are unaltered as it is assumed that muscle size and force do not change significantly between subjects or models.
Notably, we choose not to use HyFyDy’s OpenSimto-HyFyDy conversion tool, which does not include contact geometries needed for ground reaction forces or an architectural complexity similar to the other HyFyDy models (it contains 54 muscles). 3) H2190 (HyFyDy, 3D, 21 DoF, 90 muscle actuators): The default 3D model provided by HyFyDy. 4) MM240-2D (MuscleMimic, 2D, 10 DoF, 240 muscle actuators): The MM240 model constrained to two dimensions to align with H1090. 5) MM240-3D (MuscleMimic, 3D, 23 DoF, 240 muscle actuators): The MM240 model constrained to three dimensions. IV. R ESULTS Kinematic and muscle activation results are shown in Figures 4 and 5, respectively. While the subject’s data has speeds reaching over 2.0 m/s, none of the policies successfully reach these faster speeds. Eight major lower body muscles are used in our analysis, with plots of the additional muscles included in the appendix. We compute two metrics for analyzing the prediction accuracy: Root Mean Squared Error (RMSE) and the Pearson r correlation coefficient. RMSE quantifies the magnitude of disagreement of the model from the human; a lower RMSE indicates more agreement. Pearson r quantifies shape agreement between the model and the human data; -1 indicates perfect negative correlation, 0 indicates no correlation, and 1 indicates perfect positive correlation. For each joint and muscle, RMSE and Pearson r were computed between the human and model waveforms pooled across all gait-cycle points at matched speed bins. A. HyFyDy-IL versus MuscleMimic with Similar Models The first comparison evaluates the performance of MM240-2D and H1090, two-dimensional models similar to those used in the original MIRL frameworks. In terms of kinematics, both policies produce fairly accurate hip flexion results but similarly poor ankle flexion data. While their kinematic performance is similar, HyFyDy-IL produces lower RMSE values for six out of the eight recorded muscles and higher r values for five. While HyFyDy-IL is able to replicate the EMG data better than MuscleMimic, neither produce highly accurate results, as H1090’s pooled values are 0.164 and 0.40, and MM240-2D’s are 0.312 and 0.16 for the RMSE and r statistics, respectively. B. Standard versus Personalized Models The second experiment explores how model personalization influences predictive accuracy by comparing the two-dimensional model from the first comparison, H1090, with the scaled model, H-AB08. MuscleMimic is excluded as available conversion tools, such as MyoConverter, are incapable of accurately scaling musculoskeletal models in MuJoCo while being compatible with MuscleMimic [33]. Both policies (H1090 and H-AB08) show similar hip and knee angle trends, supported by relatively high r values,
Fig. 4. Hip, knee, and ankle angle data from the human data (top) from the policies’ rollouts over identical 140-second walking data, normalized over a gait cycle. Models H1090 (second row), H-AB08 (third row), and H2190 (bottom row), MM240-2D (fourth row), MM240-3D (fifth row).
with H1090 having slightly better results. Additionally, the H1090 produces a greater number of low-value RMSE and high-value r results, such as for the medial gastrocnemius and rectus femoris muscles. The data demonstrates that the change in model kinematics alone may not be significant enough to influence the predictive capabilities of HyFyDyIL. However, these results may change for subject models who are more dissimilar from the H1090 model. C. 2D versus 3D Models The final study evaluates the effect of model complexity on the prediction accuracy by comparing the two- and three-dimensional models. Specifically, we compare H1090 with H2190 and MM240-2D with MM240-3D. Notably, the MM240-3D and H2190 policies are not able to walk stably through the entire reference motion despite training for 36
TABLE II R ESULTING J OINT K INEMATICS AND M USCLE ACTIVATIONS : RMSE, Pearson r H1090
H-AB08
H2190
MM240-2D
MM240-3D1
RMSE
r
RMSE
r
RMSE
r
RMSE
r
RMSE
r
Joint Kinematics (◦ ) Hip Knee Ankle
11.26 12.64 11.69
0.96 0.96 0.82
14.07 21.76 12.77
0.93 0.79 0.57
9.86 12.42 15.60
0.93 0.87 0.53
11.04 18.50 13.42
0.76 0.70 −0.07
n/a n/a n/a
n/a n/a n/a
Pooled
11.88
0.89
16.68
0.80
12.85
0.88
14.65
0.76
n/a
n/a
Muscle Activation (activation) Tibialis Anterior 0.224 Rectus Femoris 0.128 Biceps Femoris 0.238 Gluteus Medius 0.085 Medial Gastrocnemius 0.131 Vastus Lateralis 0.169 Adductor Longus 0.189 Gluteus Maximus 0.052
0.38 0.23 −0.27 0.53 0.81 0.61 −0.08 0.44
0.320 0.165 0.248 0.097 0.221 0.077 0.185 0.083
0.08 −0.23 −0.30 0.31 0.42 −0.01 −0.01 0.36
0.261 0.209 0.254 0.129 0.278 0.261 0.298 0.187
0.33 0.00 0.20 0.40 0.60 0.47 −0.10 0.05
0.284 0.630 0.417 0.187 0.232 0.177 0.366 0.201
0.36 0.15 0.23 0.44 −0.38 -
n/a n/a n/a n/a n/a n/a n/a n/a
n/a n/a n/a n/a n/a n/a n/a n/a
0.164
0.40
0.192
0.20
0.240
0.26
0.344
0.11
n/a
n/a
Pooled
hours on an NVIDIA A100 Tensorcore GPU and 32 hours on an AMD Ryzen Threadripper PRO 5945WX CPU (12 cores / 24 threads), respectively, indicating the increased difficulty of the training task. For HyFyDy-IL, both policies (H1090 and H2190) produce similar overall kinematics, with pooled RMSE and r values of 11.88 and 0.89 for H1090 and 12.85 and 0.88 for H2190, respectively. In terms of muscle activations, H1090 yields lower RMSE values for all muscles and higher r values for seven of the eight muscle actuators compared to H2190. Due to MM240-3D’s inability to complete the walking motion for a significant portion of the reference data, it is excluded from the statistical results. Visually comparing the muscle activation plots MM240-2D and MM240-2D, it is clear that the three-dimensional model struggles to consistently actuate specific models. Overall, increased model complexity resulted in worse predictions across both frameworks. D. MuscleMimic with Modified Reward Structure Three similar MuscleMimic policies are trained to improve the first comparison by adding muscle activation penalties to align the reward function with HyFyDy-IL’s. Since modifying the rewards themselves would significantly affect the structure of the MIRL framework, only the weight values are altered using three different methods. 1) In HyFyDy-IL, the ratios of the rewards to the total sum of all |wk | is 20%, while each muscle activation penalty weight makes up 40% of the total weight, which are used to calculate the activation penalty terms in MuscleMimic. 2) In HyFyDy-IL, the ratios of the individual weights to the total sum of all |wk | are 12% for “position,” 3% for “velocity,” 5% for “root,” and 40% for both muscle activation penalties. MuscleMimic rewards are grouped into the same reward categories for HyFyDy-IL. The largest sum between the categories is maintained; other weight sums and their individual weights are scaled to match HyFyDy-IL’s ratios, and the muscle activation penalties are
calculated. 3) Weight values are chosen to be small (0.05 for each penalty term) and similar to the range of weights in the reward equation. However, after training the three policies for 24 hours, the MM240-2D fails to take even a single step. The poor results may be due to the calculated weights being too large, forcing the policy to learn more from activations than kinematics. While an ablation study would be valuable to tune these rewards, the long training time makes this infeasible without large amounts of compute. Although MuscleMimic produces undesirable muscle activations with a similar reward structure, the results do not signify its inability to train musculoskeletal models. V. C ONCLUSION Our comparative study evaluates the kinematic and muscle activation prediction capabilities of HyFyDy-IL and MuscleMimic with two-dimensional, personalized, and threedimensional models. Each framework produces kinematics and muscle activations that best align with the experimental data using policies trained with their standard twodimensional models, H1090 and MM240-2D. The increase in dimensional complexity in the H2190 and MM240-3D models led to higher pooled RMSE values and lower pooled Pearson r values for both kinematics and muscle activations. Higher dimensionality leads the model to be more unstable during gait simulation, suggesting the need for longer training periods or implementation of termination conditions and reward functions that help stabilize the musculoskeletal model and produce symmetric motions in three-dimensional space. Additionally, the scaled two-dimensional model in HyFyDy (H-AB08) also underperformed despite having the same complexity as H1090. This may be due to small musculotendon unit displacements after scaling the bone structure of the H1090 model, affecting the muscle actuator’s performance. Based on the overall results of the three comparison
Fig. 5. The top row shows EMG experimental data recorded from subject AB08, which serves as the ground-truth benchmark for comparison. The rows below show muscle activations, normalized over a gait cycle, which are produced by each policy’s rollouts on subject AB08’s walking data. Eight muscles crucial for walking were analyzed: tibialis anterior, rectus femoris, biceps femoris, gluteus medius, medial gastrocnemius, vastus lateralis, adductor longus, and gluteus maximus. Models H1090 (second row), H-AB08 (third row), and H2190 (fourth row), MM240-2D (fifth row), MM240-3D (sixth row). Colors indicate the speeds the models were able to achieve. Note that none of the policies were able to reach the peak speed of 2.5 m/s. Also note that MM240-2D learned not to activate several muscles: the gluteus medius, vastus lateralis, and gluteus maximus, which are necessary for walking.
studies, HyFyDy-IL is currently more suitable for predicting human simulation. We hypothesize that this is due to its physiological realism. However, HyFyDy is not currently GPU-parallelizable, resulting in long training times. Thus, our research does not conclude that one framework or simulator is better than another. The work demonstrates that both frameworks and simulators are capable of simulating muscleactuated human locomotion and require further development to improve their predictive capabilities. To improve human locomotion simulation and prediction, this work can be extended to evaluate the simulation accuracy across individuals with diverse gaits. This would be beneficial for testing the robustness of the frameworks and provide insight into simulating the diversity seen in clinical populations. Along these lines, implementing assistive technology with the musculoskeletal model has significant potential and interest among the prosthetic and exoskeleton research
community [34], [35]. Incorporating devices with accurate human simulation models could provide predicted human feedback to improve mechanical design and controller algorithms, advancing development without expensive, tiresome, and time-consuming physical prototyping and clinical testing [36], [37]. Introducing robotic devices that interact with the musculoskeletal model would be a significant contribution to advance the biomechanics and assistive device engineering communities. VI. ACKNOWLEDGEMENTS The authors would like to acknowledge and thank Ilseung Park and Inseung Kang for sharing their knowledge about HyFyDy and their assistance in setting up our environments, as well as the many researchers behind HyFyDy and MuscleMimic for developing the frameworks studied in this work. Models from Anthropic and OpenAI were used to assist in writing code for this project.
R EFERENCES [1] M. K. Shepherd, E. B. Schonhaut, K. L. Scherpereel, F. M. Tourk, and A. J. Young, “A roadmap for end-to-end task-agnostic exoskeleton control,” Nature Machine Intelligence, pp. 1–11, 2026. [2] R. Gehlhar, M. Tucker, A. J. Young, and A. D. Ames, “A review of current state-of-the-art control methods for lower-limb powered prostheses,” Annual reviews in control, vol. 55, pp. 142–164, 2023. [3] R. Gehlhar et al., “A review of current state-of-the-art control methods for lower-limb powered prostheses,” Annual Reviews in Control, vol. 55, pp. 142–164, 2023. [4] C. Siviy et al., “Opportunities and challenges in the development of exoskeletons for locomotor assistance,” Nature Biomedical Engineering, vol. 7, no. 4, pp. 456–472, 2023. [5] P. Slade, C. Atkeson, J. M. Donelan, H. Houdijk, K. A. Ingraham, M. Kim, K. Kong, K. L. Poggensee, R. Riener, M. Steinert et al., “On human-in-the-loop optimization of human–robot interaction,” Nature, vol. 633, no. 8031, pp. 779–788, 2024. [6] S. Song, Ł. Kidziński, X. B. Peng, C. Ong, J. Hicks, S. Levine, C. G. Atkeson, and S. L. Delp, “Deep reinforcement learning for modeling human locomotion control in neuromechanical simulation,” Journal of neuroengineering and rehabilitation, vol. 18, no. 1, p. 126, 2021. [7] M. Simos, A. S. Chiappa, and A. Mathis, “Kinesis: Motion imitation for human musculoskeletal locomotion,” arXiv preprint arXiv:2503.14637, 2025. [Online]. Available: https://arxiv.org/abs/2503.14637 [8] B. Denizdurduran, H. Markram, and M.-O. Gewaltig, “Optimum trajectory learning in musculoskeletal systems with model predictive control and deep reinforcement learning,” Biological cybernetics, vol. 116, no. 5, pp. 711–726, 2022. [9] I. Choi, I. Park, E. Halilaj, and I. Kang, “Musculoskeletal motion imitation for learning personalized exoskeleton control policy in impaired gait,” arXiv preprint arXiv:2604.09431, 2026. [10] C. Nuesslein, K. Bhakta, J. Fernandez, F. Davenport, J. Leestma, R. Kim, D. Lee, A. Mazumdar, G. Sawicki, and A. Young, “Comparing metabolic cost and muscle activation for knee and back exoskeletons in lifting,” IEEE Transactions on Medical Robotics and Bionics, vol. 6, no. 1, pp. 224–234, 2023. [11] A. Fleming et al., “Myoelectric control of robotic lower limb prostheses: A review of electromyography interfaces, control paradigms, challenges and future directions,” Journal of Neural Engineering, vol. 18, no. 4, 2021. [12] H. Han et al., “Selection of muscle-activity-based cost function in human-in-the-loop optimization of multi-gait ankle exoskeleton assistance,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol. 29, pp. 944–952, 2021. [13] T. Geijtenbeek, “The Hyfydy simulation software,” 11 2021, https://hyfydy.com. [Online]. Available: https://hyfydy.com [14] C. Li, C. Wang, B. Ziliotto, M. Simos, J. Kovecses, G. Durandau, and A. Mathis, “Towards embodied ai with musclemimic: Unlocking full-body musculoskeletal motor learning at scale,” arXiv preprint arXiv:2603.25544, 2026. [15] E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems. IEEE, 2012, pp. 5026–5033. [16] M. Febrer-Nafrı́a, A. Nasr, M. Ezati, P. Brown, J. M. Font-Llagunes, and J. McPhee, “Predictive multibody dynamic simulation of human neuromusculoskeletal systems: a review,” Multibody System Dynamics, vol. 58, no. 3, pp. 299–339, 2023. [17] A. Seth, J. L. Hicks, T. K. Uchida, A. Habib, C. L. Dembia, J. J. Dunne, C. F. Ong, M. S. DeMers, A. Rajagopal, M. Millard et al., “Opensim: Simulating musculoskeletal dynamics and neuromuscular control to study human and animal movement,” PLoS computational biology, vol. 14, no. 7, p. e1006223, 2018. [18] C. V. Hammond, S. T. Williams, M. M. Vega, D. Ao, G. Li, R. M. Salati, K. M. Pariser, M. S. Shourijeh, A. W. Habib, C. Patten et al., “The neuromusculoskeletal modeling pipeline: Matlab-based model personalization and treatment optimization functionality for opensim,” Journal of NeuroEngineering and Rehabilitation, vol. 22, no. 1, p. 112, 2025. [19] T. Geijtenbeek, “Scone: Open source software for predictive simulation of biological motion,” Journal of Open Source Software, vol. 4, no. 38, p. 1421, 2019. [Online]. Available: https://doi.org/10.21105/joss.01421
[20] J. Embley-Riches, J. Liu, S. Julier, and D. Kanoulas, “Unreal robotics lab: A high-fidelity robotics simulator with advanced physics and rendering,” arXiv preprint arXiv:2504.14135, 2026. [21] S. Wei, Z. Ni, J. Liu, Z. Zhao, J. Ye, H. Jing, J. Xia, X. Liu, M. Leong, L. Heng, D. Huang, and Y. Wang, “Simple: Simulation-based policy learning and evaluation for humanoid loco-manipulation,” arXiv preprint arXiv:2606.08278, 2026. [Online]. Available: https://arxiv.org/abs/2606.08278 [22] H. Wang, V. Caggiano, G. Durandau, M. Sartori, and V. Kumar, “Myosim: Fast and physiologically realistic mujoco models for musculoskeletal and exoskeletal studies,” in 2022 International Conference on Robotics and Automation (ICRA), Philadelphia, PA, USA, 2022, pp. 8104–8111. [23] P. Schumacher, T. Geijtenbeek, V. Caggiano, V. Kumar, S. Schmitt, G. Martius, and D. F. B. Haeufle, “Emergence of natural and robust bipedal walking by learning from biologically plausible objectives,” iScience, 2025. [Online]. Available: https://doi.org/10.1016/j.isci.2025.112203 [24] A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann, “Stable-baselines3: Reliable reinforcement learning implementations,” Journal of Machine Learning Research, vol. 22, no. 268, pp. 1–8, 2021. [Online]. Available: http://jmlr.org/papers/v22/201364.html [25] T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Offpolicy maximum entropy deep reinforcement learning with a stochastic actor,” in International conference on machine learning. Pmlr, 2018, pp. 1861–1870. [26] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017. [27] N. Mahmood, N. Ghorbani, N. F. Troje, G. Pons-Moll, and M. J. Black, “Amass: Archive of motion capture as surface shapes,” arXiv preprint arXiv:1904.03278, 2019. [Online]. Available: https://arxiv.org/abs/1904.03278 [28] M. Plappert, C. Mandery, and T. Asfour, “The KIT motion-language dataset,” Big Data, vol. 4, no. 4, pp. 236–252, 2016. [29] J. P. Araujo, Y. Ze, P. Xu, J. Wu, and C. K. Liu, “Retargeting matters: General motion retargeting for humanoid motion tracking,” arXiv preprint arXiv:2510.02252, 2025. [30] K. Scherpereel, D. Molinaro, O. T. Inan, M. Shepherd, and A. J. Young, “A human lower-limb biomechanics and wearable sensors dataset during cyclic and non-cyclic activities,” Scientific Data, vol. 10, p. 924, 2023. [Online]. Available: https://doi.org/10.1038/s41597-02302840-6 [31] X. B. Peng, P. Abbeel, S. Levine, and M. van de Panne, “Deepmimic: Example-guided deep reinforcement learning of physics-based character skills,” ACM Transactions on Graphics, vol. 37, no. 4, pp. 1–14, 2018. [32] J. C. Selinger, S. M. O’Connor, J. D. Wong, and J. M. Donelan, “Humans can continuously optimize energetic cost during walking,” Current Biology, vol. 25, no. 18, pp. 2452–2456, 2015. [Online]. Available: https://doi.org/10.1016/j.cub.2015.08.016 [33] A. Ikkala and P. Hämäläinen, “Converting biomechanical models from opensim to mujoco,” in Converging Clinical and Engineering Research on Neurorehabilitation IV: Proceedings of the 5th International Conference on Neurorehabilitation (ICNR2020). Springer, 2022, pp. 277– 281. [34] C. Zuo, J. Xu, M. Q. Vergnolle, and Y. Sui, “Embodied human simulation for quantitative design and analysis of interactive robotics,” 2026. [Online]. Available: https://arxiv.org/abs/2603.09218 [35] H. Ryu, W. Hong, and P. Hur, “Towards realistic prosthetic gait simulations: Enhancing the accuracy of OpenSim analysis by integrating the transfemoral prosthesis model,” in 2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN). IEEE, 2023. [Online]. Available: https://doi.org/10.1109/RO-MAN57019.2023.10309521 [36] D. Le, R. Posh, S. Cheng, M. Ghaffari, and R. D. Gregg, “A replay-constrained simulation framework for personalization of powered knee–ankle prosthesis controllers,” 2026. [Online]. Available: https://doi.org/10.48550/arXiv.2607.22858 [37] A. Mahmoudi et al., “Design optimization platform for assistive wearable devices applied to a knee damper exoskeleton,” Wearable Technologies, vol. 6, p. e30, 2025. [Online]. Available: https://doi.org/10.1017/wtc.2025.10016