GUIDE · 7 MIN READ
Humanoid robot development explained
How humanoid software fits together: robot models, ROS 2, control, simulation, learned controllers, sim-to-real transfer and deployment.
REVIEWED BY WBH · · DRAFTED WITH AI ASSISTANCE FROM THE CITED SOURCES
For: People who build, program or evaluate humanoid robots.
The development loop
The authors of AGILE, a workflow for humanoid reinforcement learning, argue that in many real deployments the main obstacle is no longer simulation speed or the choice of algorithm, but the lack of infrastructure that connects environment checks, training, evaluation and deployment into one coherent loop [15]. Their workflow is split into four stages: interactive verification of the environment, reproducible training, a unified evaluation, and deployment driven by configuration descriptors for the robot and the task [15].
The ToddlerBot team makes a related point about hardware: learning-based research needs a robot that both executes policies and serves as a tool for collecting the embodied data those policies are trained on [18].
Describing the robot: URDF, MJCF and USD
LimX Dynamics publishes descriptions of its full-size humanoid robots in three formats, URDF, MuJoCo MJCF and USD, so that one robot can be loaded in ROS, MuJoCo and NVIDIA Isaac Sim [6]. Not every variant comes in every format: the gripper and dexterous-hand versions of its current humanoid have URDF and USD files but no MJCF, which is provided for the base model only [6].
UBTECH follows the same split for the Walker S2: a URDF with STL meshes for modelling, visualisation and simulation under ROS and ROS 2, and a USD scene for NVIDIA Omniverse and Isaac Sim [13]. The URDF defines the links, the joints, the joint motion limits and the mass and inertia parameters [13].
UBTECH states that the model parameters are for reference only and should be adjusted to the physical robot in real applications [13]. Treat a published description as a starting point for identification, not as a measured model.
ROS 2 as the middleware
Open Robotics describes ROS 2 as middleware in which processes exchange messages through a strongly typed, anonymous publish/subscribe mechanism [1]. The nodes of a system and the connections they communicate over together form the ROS graph [1].
Underneath, ROS 2 runs on DDS or its RTPS wire protocol, which take care of discovery, serialisation and transport; each vendor implementation is connected through an RMW (ROS middleware interface) package [2]. In Humble, eProsima Fast DDS is the default implementation, and Eclipse Cyclone DDS is also fully supported and shipped with the binary releases [2]. Nodes using different RMW implementations can often communicate, but that is not guaranteed, so the documentation recommends running every part of a distributed system on the same ROS version and the same RMW implementation [2].
Humble Hawksbill runs from May 2022 to May 2027 with Tier 1 support on Ubuntu Jammy (22.04), and Jazzy Jalisco runs from May 2024 to May 2029 on Ubuntu Noble (24.04) [3].
Control and motion planning
ros2_control is a framework for real-time robot control with ROS 2 [4]. Its packages are a rewrite of the ros_control packages from the original ROS, and the project's stated goal is to make integrating new hardware simpler [4]. The framework is spread over several repositories under the ros-controls organisation, among them ros2_controllers with widely used controllers such as a joint trajectory controller and a forward command controller, realtime_tools for real-time support, and simulator plugins for Gazebo (gz_ros2_control) and MuJoCo (mujoco_ros2_control) [4].
Tokyo Robotics bases the software of its Torobo humanoid on ROS, and lists state visualisation in RViz, trajectory planning with MoveIt and logging of sensor data such as camera images, joint angles and joint torques [12]. Because Torobo is ROS-compatible, the same program can operate the robot in Gazebo and the real robot, which Tokyo Robotics says makes it possible to verify the robot's behaviour safely [12].
Choosing simulation backends
Gazebo ties each ROS distribution to one official Gazebo release, which is integrated, tested and supported for the life of that distribution; for ROS 2 Humble this is Gazebo Fortress [5]. For new users, the Gazebo documentation currently recommends Ubuntu Noble 24.04 with ROS 2 Jazzy Jalisco and Gazebo Harmonic [5].
TienKung-Lab is built on Isaac Lab and supports Sim2Sim transfer to MuJoCo [7]. AGILE is also built on NVIDIA Isaac Lab and includes a generic framework for validating policies across simulators in MuJoCo [9].
Astribot's simulation takes a different approach: one abstraction layer over MuJoCo, Genesis, ManiSkill and Isaac Lab, where every environment inherits from a base class compatible with gym.Env and the backend is selected in a YAML configuration [11].
Astribot's backends do not all offer the same sensors: in its sensor table, MuJoCo provides RGB, depth, point clouds, force/torque and IMU, ManiSkill provides RGB only, and Genesis none of these [11]. Astribot also notes that the force-control mode of its joint-space commands is mainly for simulation and does not guarantee sim-to-real accuracy [11].
Learned controllers and sim-to-real transfer
TienKung-Lab is a reinforcement-learning locomotion system for the full-size TienKung humanoid that combines AMP-style rewards with periodic gait rewards for walking and running [7]. Trained policies are evaluated in MuJoCo as a cross-simulation check, and the project reports that the framework has been validated on the real TienKung robot [7].
AGILE trains whole-body control policies with privileged observations and then distils them into student policies that can be deployed [9]. Its evaluation framework combines random rollouts, deterministic scenarios and motion metrics [9]. The AGILE paper reports five humanoid skills, spanning locomotion, recovery, motion imitation and loco-manipulation, on two platforms, the Unitree G1 and the Booster T1, with consistent sim-to-real transfer [15].
The ToddlerBot authors say that its zero-point calibration and transferable motor system identification give a high-fidelity digital twin, which enables zero-shot policy transfer from simulation to the real world [18]. Booster Lab adds real-to-sim model adaptation to a pipeline that ends in sim-to-real deployment, validated on the Booster T1 with preliminary cross-platform validation on the Booster K1 [14]. ZEST is trained entirely in simulation with moderate domain randomisation, and its authors give a procedure for selecting joint-level gains from approximate analytical armature values for closed-chain actuators, along with a refined actuator model [16].
Motion data and imitation
The Booster Lab authors note that motion data a robot can physically execute is often scarce: raw human demonstrations may not suit the robot's morphology, open-source clips vary in quality, and trajectories collected in simulation still need feasibility checks [14]. Their pipeline therefore curates motion data before AMP-based reinforcement learning and sim-to-real deployment [14].
TienKung-Lab prepares the reference motions for its AMP training by motion retargeting with GMR, which it currently supports only for SMPLX data such as AMASS and OMOMO [7]. The retargeted motion is then converted in two steps: into data for playback, used to check the correctness and quality of the motion, and into the expert reference data that AMP training uses [7].
ZEST takes imitation further: it trains policies with reinforcement learning from motion capture, monocular video and animation without physics constraints, and deploys them to hardware zero-shot [16]. It avoids contact labels, state estimators and extensive reward shaping, and uses adaptive sampling of difficult motion segments plus an automatic curriculum based on a model-based assistive wrench [16]. The authors report transferring dance and scene-interaction skills such as box-climbing from videos to Atlas and the Unitree G1 [16].
ToddlerBot includes a teleoperation interface for collecting real-world data to learn motor skills from human demonstrations [18]. The Sprout platform from Fauna Robotics integrates virtual-reality teleoperation with whole-body control and manipulation in one hardware-software stack [17].
Deployment, safety and licences
Deploy_Tienkung provides a ROS2-based reinforcement-learning control library for the TienKung series that supports both simulation and the real robot, plus an SDK with a finite state machine and standardised robot interfaces [10].
TienKung-Lab warns that an RL policy may cause unexpected or violent motions, and asks users to have accident insurance in place and to make sure the emergency stop works [7].
Platform design can also reduce the risk. The Sprout paper describes safety as one of the platform's main emphases, pursued through compliant control, soft exteriors and limits on joint torque, so that the robot can work in spaces shared with people [17].
Licences can differ within one project: the ToddlerBot codebase, including its documentation, is MIT-licensed, while the ToddlerBot design (the Onshape document, STL files and so on) is released under CC BY-NC-SA, which allows use and building on the work non-commercially [8]. AGILE likewise contains code under two licences, with most of it under the Apache License 2.0 and an RSL-RL compatibility patch under the BSD 3-Clause License [9]. LimX Dynamics publishes its humanoid descriptions under the Apache License 2.0 [6], and Astribot's simulation platform is released under the BSD 3-Clause License [11].