A model that combines seeing, understanding words, and controlling real actions.
VLA is a kitchen helper with eyes and hands. Say “pass the spatula,” and it reaches for the spatula, not the TV remote.
Robots use it to see a room. They use it to understand a request. Then they move. You meet it in robots doing real-world jobs.
Multimodal AI
VLA adds action output to multimodal understanding.
Embodied AI
VLA is a core model for many embodied systems.
Computer use
VLA extends computer use from screens into the real world.
World model
A world model can help VLA predict what its moves may cause.