What is a VLM?
A vision-language model connects visual information with language. It can describe scenes, answer questions or reason over images without directly controlling a robot.
What adds the action layer?
A vision-language-action system extends the interface toward physical behavior. It must map multimodal context to an action representation that a robot can execute.
Why embodiment changes the problem
Actions depend on the robot. A model that predicts actions for one arm, hand or humanoid cannot assume another body has the same action space. VLA research therefore intersects directly with embodiment and transfer.
Explore the field
Continue through the research map, model directory, robot directory and company directory.