GUIDE · 4 MIN READ
Sim-to-real
How policies trained in simulation reach real humanoids: the reality gap, domain randomization, sim-to-sim checks and correcting with real-world data.
REVIEWED BY WBH · · DRAFTED WITH AI ASSISTANCE FROM THE CITED SOURCES
For: People who train control policies for humanoid robots and deploy them on hardware.
The reality gap
The reality gap is what separates simulated robotics from experiments on hardware, and Tobin et al. argue that bridging it could accelerate robotics research through improved data availability [1].
Unitree describes the workflow of its reinforcement learning examples in four steps: train, play, sim-to-sim and sim-to-real [2]. Training lets the robot interact with a simulated environment to find a policy that maximises the designed rewards, and the play step verifies that the trained policy meets expectations [2]. Unitree advises against real-time visualisation during training, because it reduces efficiency [2]. The sim-to-sim step deploys the trained policy in other simulators to make sure it is not overly specific to the training simulator, and the last step deploys it on the physical robot [2].
Domain randomization
Tobin et al. describe domain randomization as a simple technique for training models on simulated images that transfer to real images, by randomising rendering in the simulator [1]. The idea is that, with enough variability in the simulator, the real world may appear to the model as just another variation [1]. Using only simulated data with non-realistic random textures, the authors trained a real-world object detector accurate to 1.5 cm [1]. To the authors' knowledge, it was the first successful transfer of a deep neural network trained only on simulated RGB images to the real world for robotic control [1].
Isaac Lab lists domain randomization for improving robustness and adaptability among its key features, next to fast and accurate physics simulation provided by PhysX and tiled rendering APIs for vectorised rendering [3]. Its documentation also lists the Unitree H1 and the Unitree G1 among the humanoids included with the platform [3].
Checking a policy in a second simulator
HumanoidVerse supports multiple simulators and tasks for humanoid sim-to-real learning; its design separates simulators, tasks and algorithms, so that you can switch between simulators and tasks [5].
The ASAP codebase is built on top of HumanoidVerse, and its README lists two deployment paths among the released components: sim-to-sim in MuJoCo, and sim-to-real with the Unitree SDK [6].
MuJoCo, short for Multi-Joint dynamics with Contact, is a general purpose physics engine meant to support research and development in fields such as robotics and machine learning that need fast and accurate simulation of articulated structures interacting with their environment [4]. The engine's runtime simulation module is tuned to maximise performance [4]. Models are described in MuJoCo's native MJCF scene description language, an XML format, and URDF model files can also be loaded [4].
Correcting with real-world data
ASAP aligns simulation and real-world physics for learning agile humanoid whole-body skills [6]. Besides motion-tracking training, it trains a delta action model and then fine-tunes the policy with that model [6]. According to its documentation, the delta action model needs motion files that also record the actions, so that the policy can use it to match real-world or sim-to-sim motions [6].
The MOSAIC authors observe that generalist humanoid motion trackers have reached strong simulation metrics by scaling data and training, yet often remain brittle on hardware during sustained teleoperation, because of errors induced by the operator interface and by dynamics [7]. MOSAIC first learns a teleoperation-oriented general motion tracker with reinforcement learning on a multi-source motion bank [7]. To bridge the interface gap it then uses rapid residual adaptation: a policy for one operator interface is trained on a small amount of interface-specific data and then distilled into the general tracker through an additive residual module [7].
Moving a policy to another robot
Any2Any addresses a related transfer: reusing a whole-body tracking model trained for one humanoid on a new humanoid, with only a small amount of data and compute [8]. It first aligns the kinematics of the source and target robots, then adapts to the target's dynamics with lightweight parameter-efficient fine-tuning of selected modules [8]. The authors transfer Sonic models pre-trained on the Unitree G1 to the LimX Oli and LimX Luna [8].
Before deploying on hardware
Before running a policy on a real robot, check:
- that it was trained with randomisation of the properties in which your robot and environment differ from the simulator;
- that it has been run in a second simulator;
- how the deployment code talks to the robot, for example through the maker's SDK;
- whether real-world data has been used to correct the policy or the simulator;
- which simulators the training code supports: HumanoidVerse, for example, currently supports IsaacGym, Genesis and IsaacLab [5].