TOPIC HUB · /topics/reinforcement-learning
Reinforcement Learning
Reinforcement learning for humanoids: PPO, massively parallel simulation, Isaac Lab and MuJoCo, and the humanoids with official RL environments.
15 ROBOTS · 2 UPDATES
Introduction
Reinforcement learning (RL) trains a control policy from a reward signal instead of from demonstrations. For humanoid control the workhorse is on-policy policy-gradient training such as PPO, which was introduced with simulated robotic locomotion among its first benchmarks [1].
What made RL practical for legged robots is scale: training on thousands of simulated robots in parallel on a single workstation GPU produces walking policies in minutes rather than days [2]. Isaac Lab, built on NVIDIA Isaac Sim, packages this for robot learning [3] and ships humanoid tasks out of the box [4]; MuJoCo is the other widely used physics engine [5].
Several manufacturers publish their own RL environments, which is the fastest way to a policy that runs on real hardware; Unitree's unitree_rl_gym, for example, covers training through deployment on the robot [6]. The development tables of the robots below show which publish official Isaac or MuJoCo environments.
CONCEPTS · 4
Key concepts
- Policy
- The learned controller that maps observations (joint states, IMU, commands) to actions (joint targets).
- PPO
- Proximal Policy Optimization, a policy-gradient method that allows several epochs of minibatch updates per batch of experience [1].
- Parallel environments
- Thousands of simulated copies of the robot stepped together on one GPU to gather experience quickly [2].
- Reward shaping
- Designing the reward terms (velocity tracking, energy, foot contact, posture) the policy maximises; the environment defines them [6].
KNOWLEDGE GRAPH · 15 ROBOTS
Robots and Reinforcement Learning
| ROBOT | ISAAC | MUJOCO |
|---|---|---|
| AGIBOT A3AGIBOT | Not published | MJCF files per model in the same repo. |
| AGIBOT Genie G2AGIBOT | Not published | Not published |
| AGIBOT X2AGIBOT | Not published | Official MJCF models (scene.xml, x2_ultra.xml / X2-Ultra.xml / X2-EDU.xml, plus omnihand/omnipicker variants) in the agibot_x2_urdf repository, and an official 'X2 MuJoCo Motion-Control Simulation' guide with local (ROS 2 Humble) or Docker deployment on Ubuntu 22.04. |
| Agility Digit 5Agility | Not published | Not published |
| Booster K1Booster | Not published | Not published |
| Booster T1Booster | Booster Gym trains T1 locomotion policies in NVIDIA Isaac Gym (legacy Isaac Gym, not Isaac Sim/Lab) and cross-validates them in MuJoCo before deployment. Booster's newer Isaac-Lab-based framework, Booster Train, currently targets the K1 robot rather than T1. | Officially supported for sim-to-sim testing and deployment (booster_gym's play_mujoco.py, booster_deploy's --mujoco flag with a T1 mjcf_path). Also has a community-maintained Booster T1 model in Google DeepMind's MuJoCo Menagerie and a T1 joystick locomotion environment in MuJoCo Playground. |
| Boston Dynamics AtlasBoston Dynamics | Not published | Not published |
| HMND 01 Alpha BipedalHumanoid | Not published | Not published |
| HMND 01 Alpha WheeledHumanoid | Not published | Not published |
| Noetix N2Noetix | Not published | Not published |
| PAL Robotics KANGAROOPAL | Not published | kangaroo_simulation repository (default branch humble-devel) states MuJoCo (via mujoco_ros2_control, in the kangaroo_mujoco package) is 'the primary supported simulator', officially released as kangaroo_mujoco 2.7.0 in the ROS 2 Humble rosdistro. The product page and datasheet additionally list 'MuJoCo and mjlab' under Simulation and RL Tools; PAL also maintains a public mujoco_vendor ROS 2 vendor package and a pal_mjlab repository. |
| ROBOTERA L7ROBOTERA | roboterax/humanoid-lab: RL locomotion and single-motion tracking pipeline for L7 on Isaac Lab 2.3.2 / Isaac Sim 5.1.0, sim2sim and sim2real. | Not published |
| ROBOTERA M7ROBOTERA | Not published | Not published |
| Unitree G1Unitree | unitree_rl_lab (Isaac Lab 2.3.0) · unitree_rl_gym (Isaac Gym) | unitree_mujoco · MuJoCo Menagerie model |
| Unitree H1Unitree | unitree_rl_lab (Isaac Lab 2.3.0 / Isaac Sim 5.1.0): dedicated h1 locomotion task folder, official support statement names Go2, H1 and G1-29dof. unitree_rl_gym (Isaac Gym): h1 and h1_2 tasks, with Sim2Sim/Sim2Real deployment. NVIDIA's own Isaac Lab (third-party) additionally ships public Isaac-Velocity-Flat/Rough-H1-v0 environments and an H1_CFG asset. | unitree_mujoco ships MJCF for both h1 and h1_2 (with DDS idl notes: H1 uses unitree_go, H1-2 uses unitree_hg); unitree_rl_gym documents Sim2Sim (MuJoCo) for h1 and h1_2. MuJoCo Menagerie (third-party, Google DeepMind) additionally provides a derived 19-joint H1 model under BSD-3-Clause. |
Values come from each robot's record, where every fact links to its source.
INTELLIGENCE
Latest updates
RESEARCH
Research
0 PAPERS
No research has been published on this topic yet.
Papers are reviewed in weekly research digests before anything is published.
LEARN