GUIDE · 4 MIN READ
Reinforcement learning for humanoids
How humanoid makers train walking policies with reinforcement learning in simulation, and how those policies are checked before they reach a real robot.
REVIEWED BY WBH · · DRAFTED WITH AI ASSISTANCE FROM THE CITED SOURCES
For: People who build, program or evaluate humanoid robots.
What reinforcement learning does for a humanoid
In reinforcement learning (RL) a controller is not programmed step by step. The robot acts in a simulated environment, and training searches for the policy that earns the most reward under a reward design the developer writes [1]. For humanoids the method is used above all for locomotion: Booster Gym and Humanoid-Gym both present themselves as RL frameworks for training humanoid walking [3] [2], and Unitree RL GYM applies it to motion control on Unitree robots [1].
The optimisation itself is done by policy-gradient algorithms. The authors of proximal policy optimization (PPO) describe a family of such methods that alternate between collecting data through interaction with the environment and optimising a surrogate objective, with several epochs of minibatch updates per batch of data; among their benchmarks was simulated robotic locomotion [4]. Humanoid-Gym keeps all of its training parameters in a configuration class named LeggedRobotCfgPPO [2].
The reward is where the developer's intent enters the process. In Humanoid-Gym every non-zero reward scale in the configuration adds a function of the same name to the total reward, and the framework ships rewards written specifically for humanoids to ease the move to hardware [2].
Why training runs in simulation, many robots at once
Simulators supply data in quantity and remove some of the safety concerns of learning on a real machine [6]. Current frameworks also simulate many robots at the same time: Unitree RL GYM takes the number of parallel environments as a training parameter [1], and Booster Gym trains in Isaac Gym with parallelised environments [3].
How much this matters was shown by Rudin and colleagues, who trained thousands of simulated robots in parallel on a single workstation GPU. Policies for flat terrain took under four minutes and policies for uneven terrain twenty minutes, several orders of magnitude faster than earlier work, and they transferred to the real robot [5]. Their robot was the quadruped ANYmal rather than a humanoid [5], but the humanoid frameworks above rely on the same idea of many simulated robots learning in parallel [3] [1].
From simulator to robot: the usual pipeline
The open frameworks follow much the same sequence. Unitree names its stages Train, Play, Sim2Sim and Sim2Real [1]:
- Train: the robot interacts with the simulated environment while training looks for the policy with the highest designed reward; Unitree advises against watching training live, because visualisation slows it down [1].
- Play: the trained policy is run to confirm it behaves as intended, and the network is exported, as policy_1.pt for a standard MLP or policy_lstm_1.pt for a recurrent network [1].
- Sim2Sim: the policy runs in a second simulator, MuJoCo, to show that it does not rely on peculiarities of the training simulator [1].
- Sim2Real: the policy is deployed on the physical robot, which must first be put in debug mode [1].
Booster Gym inserts one more check. After testing in the training environment and cross-checking in MuJoCo, the policy is exported from a .pth file to a JIT-optimised .pt file and run through the SDK in Webots; the same Webots deployment script then drives the physical robot [3].
Closing the gap to reality
A policy that works in simulation can still fail on hardware. Peng and colleagues explain why: behaviour learned in a simulator is often specific to that simulator, and because of modelling error a strategy that succeeds there may not carry over [6]. Their answer was to randomise the simulator's dynamics during training. The resulting policies adapted to dynamics they had not seen, including the real system, without any training on it; they demonstrated this on an object-pushing task with a robot arm [6].
Humanoid frameworks attack the same gap from several sides. Humanoid-Gym checks policies by moving them from Isaac Gym to MuJoCo, whose settings its authors have tuned to resemble real-world behaviour, and reports zero-shot sim-to-real transfer on RobotEra's XBot-S, a 1.2-metre humanoid, and XBot-L, which is 1.65 metres tall [2]. Booster lists settings and techniques that reduce the sim-to-real gap among the features of Booster Gym [3].
Frameworks for specific robots
- Unitree RL GYM supports the Go2, H1, H1_2 and G1, with Isaac Gym for training and MuJoCo for the sim-to-sim step [1]; see Unitree H1 and Unitree G1.
- Booster Gym is pre-configured for the Booster T1. Booster now also offers a training pipeline in Isaac Lab that supports the Booster K1, together with robot descriptions and example motion data [3].
- Humanoid-Gym uses RobotEra's XBot-L as its main example and says it can be used for other robots with minimal adjustments [2].
Makers also mention RL in product descriptions. AGIBOT says the A2-W combines 3D model synthesis training with RL so that adapting to a new object takes hours [7], and NEURA lists reinforcement learning among the capabilities of the 4NE1 Mini [8]. Statements like these describe how the maker develops its robot; on their own they do not mean that a buyer receives a training framework.
What to check as a builder
- Is there an open training framework for the robot, and which simulator does it use: Isaac Gym, Isaac Lab or MuJoCo [3] [1]?
- Does the maker publish robot descriptions for simulation [3]?
- Is there a second simulator for a sim-to-sim check before the policy reaches hardware [2]?
- How is a trained policy deployed, and in which safety mode [1]?
- Has transfer to real hardware been shown, and on which robot [2]?