Gemini Robotics ER 2

Gemini Robotics ER is a vision-language model (VLM) that brings advanced reasoning about the physical world to robotics. The model interprets visual data, performs spatial and temporal reasoning, plans multi-step tasks, and orchestrates robots and tools.

Early access

Documentation for Gemini Robotics ER 2 on Agent Platform is available to participants in the early access program. Contact your Google Cloud account team to request access.

For the publicly available documentation, including spatial reasoning, agentic capabilities, and task orchestration guides, see the Gemini API robotics documentation.