A 9B multimodal base model for spatial reasoning and robot tasks.
It is like a chess player at the board. It sees where every piece sits. Then it plans the next move.
It reads words, pictures, and video. It helps robots and Agents plan several steps.
Multimodal AI
It understands text, single images, many images, and video together.
Spatial Intelligence
It focuses on spatial reasoning and understanding 3D scenes.
Embodied AI
It helps embodied systems sense their surroundings and plan actions.
Agent
It helps Agents use tools over several steps.