Skip to content
WE BUILDHUMANOIDS

OPEN SOURCE✓ OFFICIAL SOURCE

Unitree publishes UnifoLM-ER-1, a 4B vision-language model for spatial reasoning in embodied settings

Unitree Robotics
ORIGINAL SOURCE · Web pageRead on Unitree Robotics (Hugging Face) huggingface.co/unitreerobotics/UnifoLM-ER-1

SUMMARY

Unitree Robotics has published the weights of UnifoLM-ER-1 on Hugging Face. The model card describes a 4B-parameter vision-language model built on Qwen3-VL-4B, distributed as BF16 Safetensors under the Apache 2.0 licence. The repository was created on 11 September 2026 and belongs to Unitree's UnifoLM-WLA-1.0 collection.

According to Unitree, the model saw more than 5 million training samples covering point prediction in images, object detection, reasoning across several images, 2D trajectory prediction, 3D object detection and spatial question answering over multiple images. Those embodied tasks were mixed with general image-text data during training, which Unitree says keeps the model's general vision-language ability while strengthening spatial understanding and reasoning for robots.

The card reports scores on 16 perception and understanding benchmarks, and Unitree says the model leads open-source models on seven of them. Set against its own base, Qwen3-VL-4B, the table shows RoboVQA rising from 47.7 to 62.4, Where2Place from 63.0 to 82.0 and ERQA from 41.3 to 50.0. The base model still scores higher on VSI-Bench, 59.3 against 54.2, and on RealWorldQA, MME and MMMU_VAL.

The footnotes set out how the comparison was assembled: figures for other open models come from their technical reports or papers, proprietary models were tested through their official APIs, and BLINK scores are averaged over the Relative Depth and Spatial Relation subtasks only. No inference provider serves the model yet; one Hugging Face Space uses it for a demo.

Drafted with AI assistance from the source and reviewed by WBH. Follow the source link for the full text.

WHY IT MATTERS

EDITORIAL

UnifoLM-ER-1 gives robot developers an openly licensed 4B embodied-reasoning model they can download and test, and Unitree's benchmark table shows both where it improves on its Qwen3-VL-4B base in robot-oriented spatial tasks and where the base model still scores higher.

COMMUNITY

Discuss this update

Tried it, or have a question about it? Start the discussion; it stays linked to this update.

Start the discussion