Who has run a vision-language-action (VLA) or robot foundation model on a real humanoid, on which tasks, with how much fine-tuning data, and where did it break?
On WeBuildHumanoids: Unitree open-sources UnifoLM-WLA and releases the UnifoLM-ER models · Unitree publishes UnifoLM-ER-1 · Unitree releases UnifoLM-ER-Flow · Figure introduces Helix 2.5