GUIDE · 7 MIN READ
MuJoCo for humanoids
How humanoid teams use MuJoCo: robot models, GPU training with MuJoCo Warp and mjlab, sim-to-sim checks, planning in the loop and ROS 2.
REVIEWED BY WBH · · DRAFTED WITH AI ASSISTANCE FROM THE CITED SOURCES
For: People who build, program or evaluate humanoid robots.
What MuJoCo is
MuJoCo, short for Multi-Joint dynamics with Contact, is a general-purpose physics engine maintained by Google DeepMind, aimed at fields that need fast and accurate simulation of articulated structures interacting with their environment, robotics among them [1]. It has a C API, its runtime works on low-level data structures that a built-in XML compiler preallocates, and it ships with a native interactive viewer rendered in OpenGL; the Python bindings install with pip [1].
Robots are described in MJCF, an XML format: in the MuJoCo Menagerie collection, each robot's XML file holds the MJCF definition of the model, and a separate scene file adds a plane and a light source [2].
For humanoid work, MuJoCo is also used to check policies trained elsewhere. The SIMPLE authors observe that locomotion policies trained with massively parallel reinforcement learning in Isaac Gym and Isaac Lab are frequently deployed and evaluated in MuJoCo because of its contact fidelity [14].
Finding a robot model
Start with MuJoCo Menagerie, a collection of models for MuJoCo curated by Google DeepMind [2]. Its maintainers warn that a simulator is only as good as the model it simulates, and that MuJoCo's many modelling options make it easy to create models that do not behave as expected [2]. The humanoid section lists, among others, the Unitree H1, Unitree G1, Booster T1, PAL TALOS, ROBOTIS OP3, Apptronik Apollo, Fourier N1, Berkeley Humanoid and ToddlerBot, each with its DoF count and licence [2].
Model quality varies: the maintainers say the current state of many models is not necessarily as good as it could be, and they define grades from A+, for values that come from proper system identification, down to C, for models that are only conditionally stable [2]. Licences vary as well, because the XML and asset files in each model directory have their own licence terms [2].
The mujoco-menagerie Python package downloads each model into a per-user cache the first time it is loaded, and the Menagerie README notes that version 2026.9.0 of the package pins every model to that release [2].
Menagerie models are also available through robot_descriptions.py, a third-party package that imports open-source robot descriptions as Python modules and downloads and caches them on first import [2] [6]. Its maintainers state that all of its descriptions load successfully in MuJoCo (MJCF) or in Pinocchio, iDynTree, PyBullet and yourdfpy (URDF) [6]. Its humanoid table shows how far licences differ: Apache-2.0 for the Booster T1 descriptions, GPL-3.0 for the Fourier GR-1 URDF and CC-BY-NC-4.0 for the GENE.01 [6].
Berkeley Humanoid Lite releases its robot description assets in URDF, MJCF and USD [7]. The Open-X-Humanoid xSIM_MUJOCO repository is a MuJoCo simulation environment for the Tiangong 3.0 robot, meant for developing and testing control algorithms, and includes a tool that converts URDF to MuJoCo XML [8].
Training on the GPU: MuJoCo Warp and mjlab
The MuJoCo README covers a multithreaded rollout module and MuJoCo XLA (MJX), a branch of MuJoCo written in JAX [1].
MuJoCo Warp (MJWarp) is a GPU-accelerated version of MuJoCo designed for NVIDIA hardware, maintained by Google DeepMind and NVIDIA as part of the Newton project [3]. It needs an NVIDIA GPU for fast simulation but supports the CPU for development and debugging, and its quick-start example simulates a Unitree G1 [3]. Newton itself is a GPU-accelerated physics engine built on NVIDIA Warp that integrates MuJoCo Warp as its primary backend and emphasises OpenUSD support and differentiability; it is a Linux Foundation project, initiated by Disney Research, Google DeepMind and NVIDIA [4].
mjlab combines Isaac Lab's manager-based API with MuJoCo Warp and gives direct access to native MuJoCo data structures [5]. Training needs an NVIDIA GPU, while macOS is supported for evaluation only [5]. Its examples train a Unitree G1 to follow velocity commands on flat terrain and to imitate reference motions with 4096 environments, and training can be spread over multiple GPUs [5]. Before training a new task, the README suggests checking the MDP with built-in agents that send zero actions or uniform random actions [5].
YAHMP, a general motion-tracking policy for the Unitree G1, has a training pipeline built on mjlab, trains with 8192 environments, includes a teacher-student example, and runs its pre-trained policy as an ONNX model in MuJoCo [12]. Its authors used it for an empirical study of design choices in humanoid motion tracking, and deployed the resulting policies zero-shot on the real Unitree G1 [13].
Planning with MuJoCo in the loop
MuJoCo also serves model-based control. Sumo, from the RAI Institute and collaborators, steers a pre-trained whole-body control policy with a sample-based planner; for the Unitree G1 it uses the standard velocity-tracking policy from mjlab and overrides the arm commands with targets from the planner [15]. The planner's parallel rollouts run that policy inside the CPU-based MuJoCo engine, extended in C++ with a thread pool, and the authors report total rollout times faster than the planner's update rate of 20 Hz, or 50 ms [15].
The authors also explain the trade-off behind running rollouts on the CPU: CPU parallelisation offers low latency at lower throughput, GPU batching offers high throughput at the cost of higher latency, and real-time control favours low latency [15]. In their simulated experiments, the Unitree G1 solved loco-manipulation tasks such as pushing a box and opening a door [15]. The released code needs a local g1_extensions build for its G1 tasks, and has a headless MPC runner for benchmarking and data collection [16].
Checking policies from other simulators
Tokyo Robotics' torobo_isaac_lab README refers users who want to run a trained policy in MuJoCo to the torobo_mujoco repository [10]. The torobo_mujoco examples for playing trained reinforcement-learning policies, listed under Sim2Sim, include a bipedal walk for the leg_v1 model and the reorientation of a sphere and a cube in the hand; its README notes that the leg_v1 model is under research and development and not currently scheduled for sale [9].
The torobo_mujoco README also records practical limits: in its pitching example, the timing of the ball release differs depending on the computer environment, and the reorientation scripts have to be restarted when the object gets stuck in the hand or falls from it [9].
MuJoCo can also be paired with a separate renderer. SIMPLE, a simulation testbed for humanoid loco-manipulation, handles rigid-body dynamics, contact and robot control in MuJoCo, with the contact simulation running at 500 Hz, and synchronises the states to Isaac Sim for photorealistic rendering; its whole-body controllers output low-level joint position commands through MuJoCo [14]. For VR teleoperation the SIMPLE authors switch Isaac Sim rendering off and stream stereo video from MuJoCo's native renderer to a PICO XR headset, and every object mesh goes through convex decomposition with CoACD to keep the MuJoCo physics stable [14].
MuJoCo and ROS 2
mujoco_ros2_control provides a ros2_control system interface for MuJoCo, so the ros2_control stack, with its controller manager and controllers, can run against simulated robots described in MJCF or generated from URDF [11]. It includes MJCF and URDF conversion utilities that generate MuJoCo models, and its support matrix lists the Humble, Jazzy, Kilted, Lyrical and Rolling distributions [11].
xSIM_MUJOCO takes its own route: an asynchronous simulator script integrates with ROS 2 Jazzy or Humble, and the Tiangong 3.0 model supports position and torque control, a hybrid position-velocity-torque controller, and IMU and joint sensors [8]. Its built-in terrain course has stairs with step heights of 8, 12 and 18 cm, slopes of 10, 20 and 30 degrees, stepping stones and random rough terrain [8]. Its README also notes that simulation speed depends on system performance and can be tuned with the timestep parameter [8].
A practical checklist
- Read a model's README for how its MJCF file was generated: each Menagerie model directory has one [2].
- Check the licence file in the directory of each model you use, since licence terms differ per model directory [2].
- Pin the mujoco-menagerie version you load models from: version 2026.9.0, for example, pins every model to that release [2].
- Sanity-check a new task with zero and random actions before you train [5].
- Choose CPU or GPU simulation by the need: low latency for real-time planning, throughput for training [15].
- Check timing-sensitive behaviour on the computer you will use: in the torobo_mujoco pitching example, the timing of the ball release differs depending on the computer environment [9].