GUIDE · 4 MIN READ
Teleoperation
How humanoid teleoperation works: input devices, motion retargeting, learned whole-body control, latency, and teleoperation as a data source.
REVIEWED BY WBH · · DRAFTED WITH AI ASSISTANCE FROM THE CITED SOURCES
For: People who build, program or evaluate humanoid robots and their control software.
Why teleoperate a humanoid
Teleoperation puts a person in the control loop of a humanoid, and it is also an effective way to collect data for training robot policies [1]. The authors of ExtremControl state that a low-latency teleoperation system is essential for collecting diverse reactive and dynamic demonstrations [1].
The iCub3 avatar system lets a human operator embody a humanoid robot remotely, covering locomotion, manipulation, voice and face expressions, with visual, auditory, haptic, weight and touch feedback [2]. It was validated on iCub3, a humanoid developed at the Istituto Italiano di Tecnologia [2]. In one of its tests the operator was in Genoa and the robot in Venice, about 290 km away [2].
Operator input devices
The ExtremControl authors describe the limits of teleoperation input modalities: optical motion capture is restricted to a limited capture volume, video-based pose estimators give noisy results, standard VR headsets do not track the feet, and exoskeletons are inconvenient and expensive [1]. Their own system accepts either optical motion capture or VR-based motion tracking as operator input [1]. A VR headset with two hand controllers and three trackers on the waist and both feet provides six tracked poses, which the authors map onto matching links of a Unitree G1 [1].
Unitree publishes its own teleoperation software, xr_teleoperate, which controls its humanoids from XR devices such as Apple Vision Pro, PICO 4 Ultra Enterprise and Meta Quest 3 [3]. Its list of supported robots includes the G1 in 29 and 23 degree-of-freedom versions as well as the H1, H1_2 and H2 [3].
Retargeting: mapping a person onto a robot
Humanoid motion tracking, on which teleoperation pipelines are built, has to bridge the embodiment gap between humans and humanoid robots [4]. The usual approach retargets human motion data to the humanoid and then trains a reinforcement learning policy to imitate the resulting reference trajectories [4]. Retargeting can introduce artefacts such as foot sliding, self-penetration and physically infeasible motion, which are often left in the reference for the policy to correct [4]. In the GMR study, such artefacts significantly reduced policy robustness, particularly for dynamic or long sequences [4].
The GMR repository lists Retargeting Matters among its related papers and provides real-time, high-quality retargeting that supports multiple humanoid robots and multiple human motion data formats [5].
The dex-retargeting library, which originates from the AnyTeleop project, provides retargeting optimisers that translate human hand motion to robot hand motion [6]. Its documentation names teleoperation as one application of retargeting from human hand video [6]. Because URDF parsers in ROS, simulators and robot drivers may order joints differently, it advises handling joint order explicitly by joint name [6].
Learned whole-body control
For OmniH2O, a privileged teacher policy is trained first and a student policy is distilled from it as the policy to deploy; the H2O policy is trained for sim-to-real with reinforcement learning directly [7].
The HumanPlus repository implements a Humanoid Shadowing Transformer, whose reinforcement learning in simulation builds on legged_gym and rsl_rl, and a Humanoid Imitation Transformer, whose imitation learning in the real world builds on the ACT and Mobile ALOHA code [8]. The repository also includes instructions for whole-body pose estimation, and its hardware code is based on Unitree's unitree_ros2 package [8].
X-WBC trains one shared whole-body controller across multiple humanoid robots, and its command tokens align full human motion, robot reference motion and sparse VR observations [9]. The authors deploy it on the Unitree G1, R1, H1-2 and H2 with the same policy architecture and sparse-VR interface [9].
Latency
Because teleoperation runs in a closed loop with a human in it, system latency determines how well the operator can perform responsive tasks [1]. The ExtremControl authors observe that most real-time humanoid teleoperation systems show end-to-end latencies around 200 ms, which they estimated by running optical flow on the videos those systems reported [1]. They describe these latencies as largely independent of the robot, the retargeting strategy and the length of future motion used [1].
ExtremControl moves beyond position-only PD control: a velocity feedforward term reduces the low-level control response time by approximately 100 ms [1]. It also maps human motion in Cartesian space directly to targets for selected links of the robot, primarily its extremities, so that full-body retargeting is avoided [1]. The resulting system supports optical motion capture and VR-based tracking and reaches end-to-end latency as low as 50 ms [1].
Putting a teleoperation stack together
Genesis Humanoid is an all-in-one research platform built on the Genesis simulator: it supports real-time human-to-humanoid retargeting, includes motion datasets in a unified format, provides an end-to-end learning pipeline and offers a modular low-level control testbed [10].
Before choosing a stack, check:
- which input devices it supports, and whether they track the operator's feet;
- whether its retargeter supports your robot model and your motion data format;
- how joint order is handled between the retargeter, the simulator and the robot driver;
- the measured end-to-end latency, and how it was measured;
- whether the licence allows your use: the H2O and OmniH2O code is under a CC BY-NC 4.0 licence that does not allow commercial use [7].