Interactions API 提供了一个统一的有状态接口,用于在 Gemini Enterprise Agent Platform 上使用 Gemini 模型和自主代理构建生成式 AI 应用。使用 Interactions API 可运行多轮对话、实时传输回答、强制执行结构化输出、执行函数调用,以及编排长时间运行的后台任务。
本指南将向您展示如何安装 Google Gen AI SDK、对客户端进行身份验证,以及实现常见的互动工作流。如需了解互动生命周期的概念性详情,请参阅 Interactions API 概览。
准备工作
在向 Interactions API 发送请求之前,请设置您的 Google Cloud项目和开发环境:
- 登录您的 Google Cloud 账号。如果您是 Google Cloud新手,请 创建一个账号来评估我们的产品在实际场景中的表现。新客户还可获享 $300 赠金,用于运行、测试和部署工作负载。
-
In the Google Cloud console, on the project selector page, select or create a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Verify that billing is enabled for your Google Cloud project.
Enable the Agent Platform API, if it is not already enabled.
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.-
Make sure that you have the following role or roles on the project: Agent Platform User (
roles/aiplatform.user)Check for the roles
-
In the Google Cloud console, go to the IAM page.
Go to IAM - Select the project.
-
In the Principal column, find all rows that identify you or a group that you're included in. To learn which groups you're included in, contact your administrator.
- For all rows that specify or include you, check the Role column to see whether the list of roles includes the required roles.
Grant the roles
-
In the Google Cloud console, go to the IAM page.
Go to IAM - Select the project.
- Click Grant access.
-
In the New principals field, enter your user identifier. This is typically the email address for a Google Account.
- Click Select a role, then search for the role.
- To grant additional roles, click Add another role and add each additional role.
- Click Save.
-
-
In the Google Cloud console, on the project selector page, select or create a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Verify that billing is enabled for your Google Cloud project.
Enable the Agent Platform API, if it is not already enabled.
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.-
Make sure that you have the following role or roles on the project: Agent Platform User (
roles/aiplatform.user)Check for the roles
-
In the Google Cloud console, go to the IAM page.
Go to IAM - Select the project.
-
In the Principal column, find all rows that identify you or a group that you're included in. To learn which groups you're included in, contact your administrator.
- For all rows that specify or include you, check the Role column to see whether the list of roles includes the required roles.
Grant the roles
-
In the Google Cloud console, go to the IAM page.
Go to IAM - Select the project.
- Click Grant access.
-
In the New principals field, enter your user identifier. This is typically the email address for a Google Account.
- Click Select a role, then search for the role.
- To grant additional roles, click Add another role and add each additional role.
- Click Save.
-
主要概念
请查看以下概念,了解 Interactions API 如何管理状态和响应:
Interaction:Interactions API 围绕一个核心资源展开:Interaction。Interaction表示对话或任务中的完整轮次,用于跟踪模型想法、工具调用和最终输出的时间顺序。它为提示-回答互动和复杂的多步骤代理工作流提供统一的封装。- 有状态保留:默认情况下,互动存储在服务器端(Python 中为
store=True,TypeScript/JavaScript 中为store: true)。存储的互动数据会保留 7 天,之后会自动删除。 将store=False设置为启用无状态模式,该模式会停用服务器端保留功能,并符合零数据保留 (ZDR) 标准。无状态模式还会停用previous_interaction_id链式调用和异步执行 (background=True)。 - 回答辅助程序:Google Gen AI SDK 版本
2.3.0及更高版本在互动回答中提供了便捷的属性,包括interaction.output_text、interaction.output_image和interaction.output_audio。使用interaction.output_text读取文本响应,而不是手动对步骤数组(例如interaction.steps[-1].content[0].text)进行索引。
要求
在与 Interactions API 集成之前,请确保您的环境和请求符合以下要求:
SDK 版本支持:使用统一的 Google Gen AI SDK(
>= 2.3.0[适用于 Python] 或@google/genai >= 2.3.0[适用于 TypeScript 和 JavaScript])。- 响应辅助属性和代理功能需要版本
2.3.0或更高版本,而版本 2.0.0 支持基本steps架构。 - 旧版 SDK(
google-cloud-aiplatform、@google-cloud/vertexai和google-generativeai)不支持 Interactions API。
- 响应辅助属性和代理功能需要版本
支持的模型:使用受支持的 Gemini 3 模型或更高版本。较早的模型系列不支持此 API。如需查看支持的模型的完整列表,请参阅支持的模型和迁移到最新模型版本。
回合范围参数:
tools、system_instruction和generation_config等配置参数仅适用于当前回合。如果您的工作流需要在多轮对话中用到这些参数,请在后续的每次互动轮次中传递这些参数。
安装 Google Gen AI SDK
为首选语言安装或升级 Google Gen AI SDK (>= 2.3.0):
Python
pip install --upgrade "google-genai>=2.3.0"
TypeScript / JavaScript
npm install "@google/genai>=2.3.0"
对客户端进行身份验证
您可以使用以下任一身份验证方法连接到 Agent Platform 上的 Interactions API:
使用具有应用默认凭据 (ADC) 的 Google Cloud 项目建立连接
我们建议将此身份验证方法用于 Google Cloud上的企业工作负载和生产部署。如需使用应用默认凭据 (ADC) 进行身份验证,请使用以下属性初始化客户端:
enterprise=Trueproject= Google Cloud project IDlocation="global"
如果您尚未配置本地凭据,请运行 gcloud auth application-default login。
在以下代码示例中,将 PROJECT_ID 替换为您的Google Cloud 项目 ID。
Python
from google import genai
client = genai.Client(
enterprise=True,
project="PROJECT_ID",
location="global",
)
interaction = client.interactions.create(
model="gemini-3.8-flash",
input="Explain serverless computing in one sentence.",
)
print(interaction.output_text)
TypeScript / JavaScript
import { GoogleGenAI } from "@google/genai";
const ai = new GoogleGenAI({
enterprise: true,
project: "PROJECT_ID",
location: "global",
});
const interaction = await ai.interactions.create({
model: "gemini-3.8-flash",
input: "Explain serverless computing in one sentence.",
});
console.log(interaction.output_text);
REST
curl -X POST "https://aiplatform.googleapis.com/v1beta1/projects/PROJECT_ID/locations/global/interactions" \
-H "Authorization: Bearer $(gcloud auth application-default print-access-token)" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.8-flash",
"input": [{
"role": "user",
"content": [{
"type": "text",
"text": "Explain serverless computing in one sentence."
}]
}]
}'
使用快速模式(API 密钥)建立连接
我们建议使用此身份验证方法进行快速原型设计、编写轻量级脚本,或在通过 API 密钥进行身份验证的环境中使用此方法。在初始化客户端时或在 x-goog-api-key HTTP 标头中传递您的 API 密钥。
在以下代码示例中,将 API_KEY 替换为您的 API 密钥。
Python
from google import genai
client = genai.Client(
enterprise=True,
api_key="API_KEY",
)
interaction = client.interactions.create(
model="gemini-3.8-flash",
input="Explain serverless computing in one sentence.",
)
print(interaction.output_text)
TypeScript / JavaScript
import { GoogleGenAI } from "@google/genai";
const ai = new GoogleGenAI({
enterprise: true,
apiKey: "API_KEY",
});
const interaction = await ai.interactions.create({
model: "gemini-3.8-flash",
input: "Explain serverless computing in one sentence.",
});
console.log(interaction.output_text);
REST
curl -X POST "https://aiplatform.googleapis.com/v1beta1/locations/global/interactions" \
-H "x-goog-api-key: API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.8-flash",
"input": [{
"role": "user",
"content": [{
"type": "text",
"text": "Explain serverless computing in one sentence."
}]
}]
}'
常见互动工作流
配置客户端后,您可以使用 interactions.create 方法来构建多轮对话、实时流式传输输出令牌、生成经过架构验证的 JSON、调用外部函数以及运行自主智能体。
管理有状态多轮对话
与无状态聊天 API 不同,无状态聊天 API 要求您在每次请求时重新发送完整的消息历史记录,而 Interactions API 默认在服务器上管理对话状态(在 Python 中为 store=True,在 TypeScript/JavaScript 中为 store: true)。
如需继续现有对话,请将前一次互动的 id 传递给 previous_interaction_id 参数。代理平台会自动检索存储的对话上下文并附加新对话轮次。如果您设置了 store=False(在 TypeScript/JavaScript 中为 store: false),则服务器端持久性会被停用,并且您无法使用 previous_interaction_id 链接后续对话轮次。
Python
# Turn 1: Start a conversation (store=True by default)
turn1 = client.interactions.create(
model="gemini-3.8-flash",
input="Hi! My name is John. I am working on AI agents.",
store=True,
)
print(f"Turn 1: {turn1.output_text}")
# Turn 2: Reference the stored conversation state using previous_interaction_id
turn2 = client.interactions.create(
model="gemini-3.8-flash",
input="What is my name?",
previous_interaction_id=turn1.id,
)
print(f"Turn 2: {turn2.output_text}")
TypeScript / JavaScript
// Turn 1: Start a conversation (store: true by default)
const turn1 = await ai.interactions.create({
model: "gemini-3.8-flash",
input: "Hi! My name is John. I am working on AI agents.",
store: true,
});
console.log(`Turn 1: ${turn1.output_text}`);
// Turn 2: Reference the stored conversation state using previous_interaction_id
const turn2 = await ai.interactions.create({
model: "gemini-3.8-flash",
input: "What is my name?",
previous_interaction_id: turn1.id,
});
console.log(`Turn 2: ${turn2.output_text}`);
实时显示回答内容
为了缩短互动应用的感知延迟时间,您可以流式传输模型在生成时的回答。调用 interactions.create 时,设置 stream=True(在 TypeScript/JavaScript 中为 stream: true),以接收服务器发送事件的可迭代流。过滤 step.delta 事件,以便在增量文本块到达时进行渲染:
Python
response = client.interactions.create(
model="gemini-3.8-flash",
input="Write a short poem about debugging.",
stream=True,
)
for event in response:
if event.event_type == "step.delta" and hasattr(event.delta, "text"):
print(event.delta.text, end="", flush=True)
print()
TypeScript / JavaScript
const responseStream = await ai.interactions.create({
model: "gemini-3.8-flash",
input: "Write a short poem about debugging.",
stream: true,
});
for await (const event of responseStream) {
if (event.event_type === "step.delta" && event.delta && "text" in event.delta) {
process.stdout.write(event.delta.text);
}
}
console.log();
生成结构化输出
如果您的应用需要以可预测的机器可读格式生成回答,您可以限制模型输出以匹配特定的 JSON 架构。将目标架构(例如 Python 中的 Pydantic 模型 JSON 架构或 TypeScript/JavaScript 中的 Type 架构对象)直接传递给多态 response_format 参数:
Python
from pydantic import BaseModel, Field
class Book(BaseModel):
title: str = Field(description="The title of the book")
author: str = Field(description="The book's author")
year_published: int
interaction = client.interactions.create(
model="gemini-3.8-flash",
input="Recommend one famous sci-fi book.",
response_format=Book.model_json_schema(),
)
# The output text is valid JSON matching the Book schema
print(interaction.output_text)
TypeScript / JavaScript
import { Type } from "@google/genai";
const BookSchema = {
type: Type.OBJECT,
properties: {
title: { type: Type.STRING, description: "The title of the book" },
author: { type: Type.STRING, description: "The book's author" },
yearPublished: { type: Type.INTEGER },
},
required: ["title", "author", "yearPublished"],
};
const interaction = await ai.interactions.create({
model: "gemini-3.8-flash",
input: "Recommend one famous sci-fi book.",
response_format: BookSchema,
});
console.log(interaction.output_text);
使用函数调用(工具使用)
借助函数调用,模型可以请求执行自定义函数或外部 API 来收集信息,然后再制定最终回答。在有状态的交互工作流中,函数调用遵循两轮模式:
- 声明并传递工具:在初始请求中,通过
tools参数提供函数声明。 - 执行并返回结果:检查
function_call步骤的响应步骤 (interaction.steps),使用模型提供的arguments运行本地函数,并发送后续互动,其中包含由call_id和previous_interaction_id关联的function_result项。
Python
# Define a declarative function tool schema
stock_tool = {
"type": "function",
"name": "get_stock_price",
"description": "Gets the stock price for a given ticker symbol.",
"parameters": {
"type": "object",
"properties": {
"ticker": {"type": "string", "description": "The stock ticker symbol"}
},
"required": ["ticker"],
},
}
def get_stock_price(ticker: str) -> float:
"""Executes the local tool function."""
if ticker.upper() == "GOOG":
return 175.50
return 100.0
# Turn 1: Pass the tool declaration to the model
interaction = client.interactions.create(
model="gemini-3.8-flash",
input="What is the stock price of GOOG?",
tools=[stock_tool],
)
# Inspect the interaction steps for function call requests
for step in interaction.steps:
if step.type == "function_call" and step.name == "get_stock_price":
ticker_arg = step.arguments.get("ticker")
price = get_stock_price(ticker_arg)
# Turn 2: Submit the function execution result to the conversation
final_turn = client.interactions.create(
model="gemini-3.8-flash",
input=[{
"type": "function_result",
"call_id": step.id,
"result": {"price": price},
}],
previous_interaction_id=interaction.id,
)
print(final_turn.output_text)
TypeScript / JavaScript
// Define a declarative function tool schema
const stockTool = {
type: "function",
name: "getStockPrice",
description: "Gets the stock price for a given ticker symbol.",
parameters: {
type: "object",
properties: {
ticker: { type: "string", description: "The stock ticker symbol" },
},
required: ["ticker"],
},
};
function getStockPrice({ ticker }: { ticker: string }): number {
if (ticker.toUpperCase() === "GOOG") return 175.50;
return 100.00;
}
// Turn 1: Pass the tool declaration to the model
const interaction = await ai.interactions.create({
model: "gemini-3.8-flash",
input: "What is the stock price of GOOG?",
tools: [stockTool],
});
// Inspect the interaction steps for function call requests
for (const step of interaction.steps ?? []) {
if (step.type === "function_call" && step.name === "getStockPrice") {
const tickerArg = step.arguments.ticker as string;
const price = getStockPrice({ ticker: tickerArg });
// Turn 2: Submit the function execution result to the conversation
const finalTurn = await ai.interactions.create({
model: "gemini-3.8-flash",
input: [{
type: "function_result",
call_id: step.id,
result: { price },
}],
previous_interaction_id: interaction.id,
});
console.log(finalTurn.output_text);
}
}
运行代理和长时间运行的后台任务
除了基础模型之外,您还可以使用 Interactions API 通过 agent 参数调用专业自主代理:
antigravity-preview-05-2026:通用托管式智能体,可在安全的沙盒化 Linux 环境中执行代码、管理文件和浏览网页。如需了解详情,请参阅与代理互动。deep-research-preview-04-2026:Gemini Deep Research 代理,可规划和执行多步网络研究任务,并将多个来源的发现结果整合为全面的报告。如需了解详情,请参阅使用 Gemini Deep Research Agent。- 自定义代理:使用
client.agents.create()配置和预配的自定义代理资源。
由于代理工作流通常需要几分钟才能完成,因此请通过设置 background=True 在后台异步运行它们。该 API 会立即返回一个 Interaction 对象,其中包含一个 id,您可以使用 client.interactions.get() 轮询该对象,直到 interaction.status 转换为 completed:
在尝试此示例之前,请将 PROJECT_ID 替换为您的Google Cloud 项目 ID。
import time
from google import genai
client = genai.Client(
enterprise=True,
project="PROJECT_ID",
location="global",
)
interaction = client.interactions.create(
input="Analyze competitive positioning for solar energy providers.",
agent="deep-research-preview-04-2026",
background=True,
)
print(f"Research started: {interaction.id}")
while True:
interaction = client.interactions.get(interaction.id)
if interaction.status == "completed":
print(interaction.output_text)
break
elif interaction.status in ("failed", "cancelled"):
print(f"Research ended with status: {interaction.status}")
break
time.sleep(10)
访问已上传的 Cloud Storage 文件
您可以使用 Interactions API 访问已上传的 Cloud Storage 文件。请参阅以下示例:
from google import genai
# Credentials must belong to an identity with storage.objects.get permissions
client = genai.Client()
interaction = client.interactions.create(
model="gemini-3.8-flash",
input=[
{"type": "text", "text": "Summarize the attached document:"},
{
"type": "document",
"uri": "gs://my-secure-bucket/quarterly_report.pdf",
"mime_type": "application/pdf"
}
],
)
print(interaction.output_text)
将 Cloud Storage URI(例如 gs://bucket-name/path/to/file)传递给 Interactions API 时,系统会使用最终用户凭据 (EUC) 来评估请求。该 API 使用经过身份验证的调用者的身份(而非后台项目服务代理)提取 Cloud Storage 对象。
如需在互动请求中传递 Cloud Storage 文件,调用方主体(用户账号、服务账号或联合身份)必须拥有对所有引用对象的 storage.objects.get 权限。
配置用于访问 Cloud Storage 文件的 IAM 角色
授予包含 storage.objects.get 权限的标准预定义角色之一:
- Storage Object Viewer (
roles/storage.objectViewer):对对象的读取访问权限(推荐)。 - Storage Object User (
roles/storage.objectUser):拥有对对象的读写权限。
如需使用 Google Cloud CLI 向用户账号授予访问权限,请使用以下命令:
gcloud storage buckets add-iam-policy-binding gs://BUCKET_NAME \
--member="user:user-email@example.com" \
--role="roles/storage.objectViewer"
如需向特定调用服务账号授予访问权限,请使用以下命令:
gcloud storage buckets add-iam-policy-binding gs://BUCKET_NAME \
--member="serviceAccount:sa-name@PROJECT_ID.iam.gserviceaccount.com" \
--role="roles/storage.objectViewer"
排查 Cloud Storage 文件访问问题
如果调用主体缺少足够的权限,Interactions API 会返回类似于以下内容的 403 Forbidden 错误:
Access error:
PERMISSION_DENIED - 403 Forbidden: Calling principal lacks
storage.objects.get on one or more GCS URIs.
如需解决此问题,请向经过身份验证的调用者授予存储桶或对象的 Storage Object Viewer 角色 (roles/storage.objectViewer)。
如果指定的对象不存在,或者存储桶权限阻止调用方查看对象是否存在,则 Interactions API 会返回类似于以下内容的 404 Not Found 错误:
Access error:
NOT_FOUND - 404 Not Found: The object does not exist, or bucket
permissions prevent revealing object existence.
如需解决此问题,请验证 Cloud Storage URI 是否正确,并确认经过身份验证的调用方是否具有对相应存储桶的读取权限。
高级 REST 工作流
对于基于 shell 的自动化、CI/CD 流水线或没有 Python 或 TypeScript/JavaScript 运行时的环境,您可以使用 curl 通过 HTTP 直接调用 Interactions API。
REST 端点
向以下 Interactions API 端点发送 POST 请求:
POST https://aiplatform.googleapis.com/v1beta1/projects/PROJECT_ID/locations/LOCATION/interactions
替换请求中的以下变量:
- PROJECT_ID:您的 Google Cloud 项目 ID。
- LOCATION:设置为
global(或您的配置所需的受支持的自定义区域)。
设置环境变量和身份验证
在运行以下部分中的 curl 示例之前,请导出您的项目 ID、目标模型或代理 ID,以及通过应用默认凭据生成的 OAuth 2.0 访问令牌:
PROJECT_ID="PROJECT_ID"
MODEL_ID="gemini-3.8-flash"
AGENT_ID="deep-research-preview-04-2026"
ACCESS_TOKEN=$(gcloud auth print-access-token)
同步响应格式
同步 POST 请求会返回一个 JSON interaction 对象,其中包含唯一的互动 id、执行 status、对话 steps 和令牌 usage 元数据:
{
"id": "your-interaction-id",
"status": "completed",
"steps": [
{
"type": "model_output",
"content": [
{
"type": "text",
"text": "Serverless computing is a cloud execution model where the cloud provider dynamically manages the allocation and provisioning of servers, charging customers based on actual usage rather than pre-purchased capacity."
}
]
}
],
"usage": {
"total_tokens": 24751,
"total_input_tokens": 23894,
"total_output_tokens": 857
},
"created": "2026-05-08T10:44:43Z",
"updated": "2026-05-08T10:44:43Z",
"environment_id": "your-environment-id",
"object": "interaction"
}
继续进行多轮有状态交互
如需通过 REST 继续已存储的对话,请在 JSON 请求正文的 previous_interaction_id 字段中传递之前响应中的 id。
在试用此示例之前,请将 PREVIOUS_INTERACTION_ID 替换为之前互动返回的 id。
curl -X POST "https://aiplatform.googleapis.com/v1beta1/projects/${PROJECT_ID}/locations/global/interactions" \
-H "Authorization: Bearer ${ACCESS_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"model": "'"${MODEL_ID}"'",
"store": true,
"previous_interaction_id": "PREVIOUS_INTERACTION_ID",
"input": [{
"role": "user",
"content": [{
"type": "text",
"text": "Can you elaborate on that?"
}]
}]
}'
通过服务器发送的事件流式传输输出
如需通过 REST 流式传输增量更新,请在 JSON 请求正文中添加 "stream": true:
curl -X POST "https://aiplatform.googleapis.com/v1beta1/projects/${PROJECT_ID}/locations/global/interactions" \
-H "Authorization: Bearer ${ACCESS_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"model": "'"${MODEL_ID}"'",
"stream": true,
"input": [{
"role": "user",
"content": [{
"type": "text",
"text": "Write a long story about space travel."
}]
}]
}'
如果设置了 "stream": true,服务器会返回 Transfer-Encoding: chunked 和 Content-Type: text/event-stream(服务器发送的事件)。数据流中的每个事件都包含一个 data: 前缀,其中包含一个 JSON 载荷,该载荷具有 event_type 和步数增量内容。curl 会自动保持 HTTP 连接处于打开状态,并将传入的块实时写入 stdout,直到互动完成。
在后台运行受管理的代理
如需通过 REST 异步启动长时间运行的受管代理任务,请指定目标 agent、设置 "background": true 并配置 "environment": "remote":
curl -X POST "https://aiplatform.googleapis.com/v1beta1/projects/${PROJECT_ID}/locations/global/interactions" \
-H "Authorization: Bearer ${ACCESS_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"agent": "'"${AGENT_ID}"'",
"environment": "remote",
"background": true,
"input": [{
"role": "user",
"content": [{
"type": "text",
"text": "Analyze competitive positioning for commercial solar energy providers."
}]
}]
}'
后续步骤
- 如需详细了解相关关键概念,请参阅 Interactions API 概览。
- 在 Interactions API 参考文档中探索请求和响应架构。
- 了解如何与受管理的代理互动以及使用 Gemini Deep Research 代理。