Guide / VLA
Vision. Language. Action.
Vision-language-action models connect what a robot sees and what a person asks for with actions the robot can execute.
Vision
Encode the scene, objects, geometry and state from visual or multimodal observations.
Language
Ground natural-language goals and instructions in the observed physical context.
Action
Generate an action representation suitable for downstream robot control.
Why VLA matters
Robots need more than perception or language understanding in isolation. VLA research explores a shared pathway from multimodal context to physical behavior, with the goal of making robots more adaptable across tasks.
Open research questions
- How much robot data is required for robust generalization?
- How should high-level reasoning connect to low-level control?
- How can safety and uncertainty be handled during real-world action?
- How well do policies transfer across robot embodiments?