AI Rookies

VLA — Vision-Language-Action model

Fact

A model that combines seeing, understanding words, and controlling real actions.

In Plain Words

VLA is a kitchen helper with eyes and hands. Say “pass the spatula,” and it reaches for the spatula, not the TV remote.

Robots use it to see a room. They use it to understand a request. Then they move. You meet it in robots doing real-world jobs.

Related Concepts

Multimodal AI
VLA adds action output to multimodal understanding.

Embodied AI
VLA is a core model for many embodied systems.

Computer use
VLA extends computer use from screens into the real world.

World model
A world model can help VLA predict what its moves may cause.