DISCUSSIONSTARTED BY WBH
Vision-language-action models on real humanoids: what works today?
Who has run a vision-language-action (VLA) or robot foundation model on a real humanoid, on which tasks, with how much fine-tuning data, and where did it break? On WeBuildHumanoids: Unitree open-sources UnifoLM-WLA and…
Research & papers#foundation-models#imitation-learning#vlaWBH Editorial ·