Guide / Models

Robot foundation models.

Models designed to connect multimodal understanding with physical action are becoming an important layer in modern robotics.

What makes a model a “foundation model” for robots?

The term is used for models intended to support multiple tasks, environments or robot embodiments rather than a single narrowly programmed behavior. The exact scope differs between projects.

Multimodal input

Vision, language, proprioception and other observations can provide context about the task and environment.

Action output

VLA systems connect observations and instructions to robot actions, trajectories or intermediate control representations.

Generalization

A central research question is whether skills and knowledge transfer across objects, tasks, environments and robot bodies.

The stack

A practical mental model is: multimodal perception → reasoning/planning → action generation → robot-specific control → physical feedback. Different architectures collapse or separate these layers in different ways.

Examples

See the model reference for source-linked entries covering Gemini Robotics and NVIDIA GR00T.

Editorial guide · Updated September 2026