Skip to content
WE BUILDHUMANOIDS

TOPIC HUB · /topics/reinforcement-learning

Reinforcement Learning

Reinforcement learning for humanoids: PPO, massively parallel simulation, Isaac Lab and MuJoCo, and the humanoids with official RL environments.

15 ROBOTS · 2 UPDATES

Introduction

Reinforcement learning (RL) trains a control policy from a reward signal instead of from demonstrations. For humanoid control the workhorse is on-policy policy-gradient training such as PPO, which was introduced with simulated robotic locomotion among its first benchmarks [1].

What made RL practical for legged robots is scale: training on thousands of simulated robots in parallel on a single workstation GPU produces walking policies in minutes rather than days [2]. Isaac Lab, built on NVIDIA Isaac Sim, packages this for robot learning [3] and ships humanoid tasks out of the box [4]; MuJoCo is the other widely used physics engine [5].

Several manufacturers publish their own RL environments, which is the fastest way to a policy that runs on real hardware; Unitree's unitree_rl_gym, for example, covers training through deployment on the robot [6]. The development tables of the robots below show which publish official Isaac or MuJoCo environments.

CONCEPTS · 4

Key concepts

Policy
The learned controller that maps observations (joint states, IMU, commands) to actions (joint targets).
PPO
Proximal Policy Optimization, a policy-gradient method that allows several epochs of minibatch updates per batch of experience [1].
Parallel environments
Thousands of simulated copies of the robot stepped together on one GPU to gather experience quickly [2].
Reward shaping
Designing the reward terms (velocity tracking, energy, foot contact, posture) the policy maximises; the environment defines them [6].

KNOWLEDGE GRAPH · 15 ROBOTS

Robots and Reinforcement Learning

All robots
Robots tagged Reinforcement Learning, with the relevant facts from each robot's sourced record
ROBOTISAACMUJOCO
AGIBOT A3AGIBOTNot publishedMJCF files per model in the same repo.
AGIBOT Genie G2AGIBOTNot publishedNot published
AGIBOT X2AGIBOTNot publishedOfficial MJCF models (scene.xml, x2_ultra.xml / X2-Ultra.xml / X2-EDU.xml, plus omnihand/omnipicker variants) in the agibot_x2_urdf repository, and an official 'X2 MuJoCo Motion-Control Simulation' guide with local (ROS 2 Humble) or Docker deployment on Ubuntu 22.04.
Agility Digit 5AgilityNot publishedNot published
Booster K1BoosterNot publishedNot published
Booster T1BoosterBooster Gym trains T1 locomotion policies in NVIDIA Isaac Gym (legacy Isaac Gym, not Isaac Sim/Lab) and cross-validates them in MuJoCo before deployment. Booster's newer Isaac-Lab-based framework, Booster Train, currently targets the K1 robot rather than T1.Officially supported for sim-to-sim testing and deployment (booster_gym's play_mujoco.py, booster_deploy's --mujoco flag with a T1 mjcf_path). Also has a community-maintained Booster T1 model in Google DeepMind's MuJoCo Menagerie and a T1 joystick locomotion environment in MuJoCo Playground.
Boston Dynamics AtlasBoston DynamicsNot publishedNot published
HMND 01 Alpha BipedalHumanoidNot publishedNot published
HMND 01 Alpha WheeledHumanoidNot publishedNot published
Noetix N2NoetixNot publishedNot published
PAL Robotics KANGAROOPALNot publishedkangaroo_simulation repository (default branch humble-devel) states MuJoCo (via mujoco_ros2_control, in the kangaroo_mujoco package) is 'the primary supported simulator', officially released as kangaroo_mujoco 2.7.0 in the ROS 2 Humble rosdistro. The product page and datasheet additionally list 'MuJoCo and mjlab' under Simulation and RL Tools; PAL also maintains a public mujoco_vendor ROS 2 vendor package and a pal_mjlab repository.
ROBOTERA L7ROBOTERAroboterax/humanoid-lab: RL locomotion and single-motion tracking pipeline for L7 on Isaac Lab 2.3.2 / Isaac Sim 5.1.0, sim2sim and sim2real.Not published
ROBOTERA M7ROBOTERANot publishedNot published
Unitree G1Unitreeunitree_rl_lab (Isaac Lab 2.3.0) · unitree_rl_gym (Isaac Gym)unitree_mujoco · MuJoCo Menagerie model
Unitree H1Unitreeunitree_rl_lab (Isaac Lab 2.3.0 / Isaac Sim 5.1.0): dedicated h1 locomotion task folder, official support statement names Go2, H1 and G1-29dof. unitree_rl_gym (Isaac Gym): h1 and h1_2 tasks, with Sim2Sim/Sim2Real deployment. NVIDIA's own Isaac Lab (third-party) additionally ships public Isaac-Velocity-Flat/Rough-H1-v0 environments and an H1_CFG asset.unitree_mujoco ships MJCF for both h1 and h1_2 (with DDS idl notes: H1 uses unitree_go, H1-2 uses unitree_hg); unitree_rl_gym documents Sim2Sim (MuJoCo) for h1 and h1_2. MuJoCo Menagerie (third-party, Google DeepMind) additionally provides a derived 19-joint H1 model under BSD-3-Clause.

Values come from each robot's record, where every fact links to its source.

INTELLIGENCE

Latest updates

All updates
  1. Isaac Lab 3.0 Early Access separates physics, rendering and visualisation backendsRELEASE · GITHUB · OFFICIAL
  2. Figure says BotQ has delivered 350+ Figure 03 units and reached one robot per hourHISTORICAL RECORD · PRODUCT · FIGURE AI · OFFICIAL

RESEARCH

Research

0 PAPERS

No research has been published on this topic yet.

Papers are reviewed in weekly research digests before anything is published.

LEARN

Guides