Guide / VLA

Vision. Language. Action.

Vision-language-action models connect what a robot sees and what a person asks for with actions the robot can execute.

Vision

Encode the scene, objects, geometry and state from visual or multimodal observations.

Language

Ground natural-language goals and instructions in the observed physical context.

Action

Generate an action representation suitable for downstream robot control.

Why VLA matters

Robots need more than perception or language understanding in isolation. VLA research explores a shared pathway from multimodal context to physical behavior, with the goal of making robots more adaptable across tasks.

Open research questions

Related

Robot foundation models →   Embodied AI →