Thinking models generate an internal "thinking" process before returning a response. This capability helps the model perform complex multi-step planning, solve mathematical problems, and generate accurate code.
This page explains how to:
- Configure thinking levels and token budgets
- View thought summaries
- Preserve reasoning state across multi-turn conversations using thought signatures
Supported models
Thinking is supported in the following models:
Click to expand supported models
- Gemini Omni Flash
- Gemini Omni 1.1 Flash
- Gemini 3.8 Flash Cyber
- Gemini 3.8 Flash
- Gemini 3.7 Flash
- Gemini 3.6 Flash
- Gemini 3.5 Flash-Lite
- Gemini 3.5 Flash
- Gemini 3.1 Pro
- Gemini 3.1 Flash-Lite Image (Nano Banana 2 Lite)
- Gemini 3.1 Flash-Lite
- Gemini 3.1 Flash Image
- Gemini 3 Pro Image
- Gemini 3 Flash
- Gemini 2.5 Pro
- Gemini 2.5 Flash-Lite
- Gemini 2.5 Flash
Control model thinking
Thinking is enabled by default in supported Gemini models. In Agent Studio, you can inspect the full thinking process alongside the generated response.
How you configure thinking depends on the model version:
Gemini 3 and later models
Gemini 3 models use the thinking_level parameter. This
parameter sets discrete reasoning tiers so you can optimize between latency and
reasoning depth.
Console
-
Go to Agent Studio and select New > Chat and expand the model panel.
- In the Model settings panel, select a supported model from the Model menu.
- Select a value from the Thinking level drop-down menu.
Python
from google import genai
from google.genai import types
client = genai.Client()
response = client.models.generate_content(
model="gemini-3.5-flash",
contents="How does AI work?",
config=types.GenerateContentConfig(
thinking_config=types.ThinkingConfig(
# Options: MINIMAL, LOW, MEDIUM, HIGH
thinking_level=types.ThinkingLevel.THINKING_LEVEL_VALUE
)
),
)
print(response.text)Thinking level values
You can set thinking_level to one of the following values:
MINIMAL: Uses the fewest possible tokens for thinking. Best for straightforward tasks that do not require extended reasoning.MINIMALrequires thought signatures in multi-turn conversations; if omitted, the model returns a400: INVALID_ARGUMENTerror.LOW: Uses fewer thinking tokens for faster responses. Best for high-throughput applications with low task complexity.MEDIUM: Balances reasoning quality and latency. Suitable for tasks with moderate complexity that benefit from intermediate reasoning steps.HIGH: Uses the maximum thinking capacity. Best for complex prompts requiring deep reasoning, multi-step problem solving, formal code verification, or multi-turn tool execution.
Supported thinking levels by model
The following table lists supported thinking_level values and default configurations by model:
| Model | Supported thinking_level values |
Default |
|---|---|---|
| Gemini 3.8 Flash Cyber | LOW, MEDIUM, HIGH |
MEDIUM |
| Gemini 3.8 Flash | LOW, MEDIUM, HIGH |
MEDIUM |
| Gemini 3.7 Flash | LOW, MEDIUM, HIGH |
MEDIUM |
| Gemini 3.6 Flash | MINIMAL, LOW, MEDIUM, HIGH |
MEDIUM |
| Gemini 3.5 Flash-Lite | MINIMAL, LOW, MEDIUM, HIGH |
MINIMAL |
| Gemini 3.5 Flash | MINIMAL, LOW, MEDIUM, HIGH |
MEDIUM |
| Gemini 3.1 Pro | LOW, MEDIUM, HIGH |
HIGH |
| Gemini 3.1 Flash-Lite Image (Nano Banana 2 Lite) | MINIMAL, HIGH |
MINIMAL |
| Gemini 3.1 Flash-Lite | MINIMAL, LOW, MEDIUM, HIGH |
MINIMAL |
| Gemini 3.1 Flash Image | MINIMAL, HIGH |
MINIMAL |
| Gemini 3 Pro Image | HIGH |
HIGH |
| Gemini 3 Flash | MINIMAL, LOW, MEDIUM, HIGH |
HIGH |
Gemini 2.5 and earlier models
For Gemini 2.5 and earlier models, configure thinking using the
thinking_budget parameter. This parameter sets a soft limit on the number of
tokens the model can use during internal reasoning.
Console
-
Go to Agent Studio and select New > Chat.
- In the Model settings panel, select a supported model from the Model menu.
- In the Thinking budget selector, select Manual and use the slider to adjust the token limit.
Python
Install
pip install --upgrade google-genai
To learn more, see the SDK reference documentation.
Set environment variables to use the Google Gen AI SDK with Vertex AI:
# Replace the `GOOGLE_CLOUD_PROJECT` and `GOOGLE_CLOUD_LOCATION` values # with appropriate values for your project. export GOOGLE_CLOUD_PROJECT=GOOGLE_CLOUD_PROJECT export GOOGLE_CLOUD_LOCATION=global export GOOGLE_GENAI_USE_ENTERPRISE=True
Node.js
Install
npm install @google/genai
To learn more, see the SDK reference documentation.
Set environment variables to use the Google Gen AI SDK with Vertex AI:
# Replace the `GOOGLE_CLOUD_PROJECT` and `GOOGLE_CLOUD_LOCATION` values # with appropriate values for your project. export GOOGLE_CLOUD_PROJECT=GOOGLE_CLOUD_PROJECT export GOOGLE_CLOUD_LOCATION=global export GOOGLE_GENAI_USE_ENTERPRISE=True
Go
Learn how to install or update the Go.
To learn more, see the SDK reference documentation.
Set environment variables to use the Google Gen AI SDK with Vertex AI:
# Replace the `GOOGLE_CLOUD_PROJECT` and `GOOGLE_CLOUD_LOCATION` values # with appropriate values for your project. export GOOGLE_CLOUD_PROJECT=GOOGLE_CLOUD_PROJECT export GOOGLE_CLOUD_LOCATION=global export GOOGLE_GENAI_USE_ENTERPRISE=True
Java
Learn how to install or update the Java.
To learn more, see the SDK reference documentation.
Set environment variables to use the Google Gen AI SDK with Vertex AI:
# Replace the `GOOGLE_CLOUD_PROJECT` and `GOOGLE_CLOUD_LOCATION` values # with appropriate values for your project. export GOOGLE_CLOUD_PROJECT=GOOGLE_CLOUD_PROJECT export GOOGLE_CLOUD_LOCATION=global export GOOGLE_GENAI_USE_ENTERPRISE=True
If you do not specify a budget, the model sets the token budget dynamically up
to 8,192 tokens. To explicitly enable dynamic budgeting in the API, set
thinking_budget to -1.
Supported thinking budgets by model
The following table lists minimum, maximum, and default token limits for each model:
| Model | Minimum tokens | Maximum tokens | Default |
|---|---|---|---|
| Gemini 2.5 Flash | 1 | 24,576 | Auto (up to 8,192 tokens) |
| Gemini 2.5 Pro | 128 | 32,768 | Auto (up to 8,192 tokens) |
| Gemini 2.5 Flash-Lite | 512 | 24,576 | Auto (up to 8,192 tokens) |
Turn off thinking
You can turn off thinking for Gemini 2.5 Flash and
Gemini 2.5 Flash-Lite by setting thinking_budget to 0. Although
thinking content is not returned in the response, the generated text might still
show reasoning-style output.
You can't turn off thinking for Gemini 2.5 Pro.
View thought summaries
Thought summaries display intermediate reasoning steps alongside the model's final response. Thought summaries are supported in Gemini 2.5 and later models.
Console
Thought summaries are enabled by default in Agent Studio. To view the summarized reasoning steps, expand the Thoughts panel.
Python
Install
pip install --upgrade google-genai
To learn more, see the SDK reference documentation.
Set environment variables to use the Google Gen AI SDK with Vertex AI:
# Replace the `GOOGLE_CLOUD_PROJECT` and `GOOGLE_CLOUD_LOCATION` values # with appropriate values for your project. export GOOGLE_CLOUD_PROJECT=GOOGLE_CLOUD_PROJECT export GOOGLE_CLOUD_LOCATION=global export GOOGLE_GENAI_USE_ENTERPRISE=True
Node.js
Install
npm install @google/genai
To learn more, see the SDK reference documentation.
Set environment variables to use the Google Gen AI SDK with Vertex AI:
# Replace the `GOOGLE_CLOUD_PROJECT` and `GOOGLE_CLOUD_LOCATION` values # with appropriate values for your project. export GOOGLE_CLOUD_PROJECT=GOOGLE_CLOUD_PROJECT export GOOGLE_CLOUD_LOCATION=global export GOOGLE_GENAI_USE_ENTERPRISE=True
Go
Learn how to install or update the Go.
To learn more, see the SDK reference documentation.
Set environment variables to use the Google Gen AI SDK with Vertex AI:
# Replace the `GOOGLE_CLOUD_PROJECT` and `GOOGLE_CLOUD_LOCATION` values # with appropriate values for your project. export GOOGLE_CLOUD_PROJECT=GOOGLE_CLOUD_PROJECT export GOOGLE_CLOUD_LOCATION=global export GOOGLE_GENAI_USE_ENTERPRISE=True
Java
Learn how to install or update the Java.
To learn more, see the SDK reference documentation.
Set environment variables to use the Google Gen AI SDK with Vertex AI:
# Replace the `GOOGLE_CLOUD_PROJECT` and `GOOGLE_CLOUD_LOCATION` values # with appropriate values for your project. export GOOGLE_CLOUD_PROJECT=GOOGLE_CLOUD_PROJECT export GOOGLE_CLOUD_LOCATION=global export GOOGLE_GENAI_USE_ENTERPRISE=True
A response might return a thought signature without thought summary text in the following situations:
- Low-complexity requests: The request required minimal reasoning steps.
- Disabled summaries: Thought summaries were not requested or were turned off.
- Non-text modalities: Reasoning over certain modalities (such as image analysis) might not produce text summaries.
Your application must handle empty or missing thought summary fields gracefully while preserving any accompanying thought signatures.
Thought signatures
Thought signatures are encrypted representations of the model's internal reasoning state. They maintain context across multi-turn conversations, especially when using function calling.
To preserve the reasoning context across multi-turn interactions, pass the thought signatures returned in earlier responses back into subsequent requests, regardless of the thinking level configured.
If you use the official Google Google Gen AI SDK (Python, Node.js, Go, or Java), thought signatures are managed automatically when using standard chat sessions or when appending complete response objects to your message history.
For implementation patterns, requirements, and examples, see Thought signatures.
Pricing
You are billed for tokens generated during the thinking process. For models where thinking is enabled by default—such as Gemini 3 Pro and Gemini 2.5 Pro—these thinking tokens are included in your billable usage.
For full rate details, see Pricing.
What's next
Thought signatures
Learn how to preserve the Gemini reasoning state during multi-turn and multi-step conversations using thought signatures.
Thinking prompting guide
Explore prompt engineering techniques and best practices tailored for Gemini thinking models.