Gemini Enterprise Agent Platform の Mistral AI モデルは、API としてフルマネージド モデルとサーバーレス モデルを提供します。Agent Platform で Mistral AI モデルを使用するには、Agent Platform API エンドポイントにリクエストを直接送信します。Mistral AI モデルはマネージド API を使用します。インフラストラクチャをプロビジョニングしたり、管理する必要はありません。
レスポンスをストリーミングして、エンドユーザーのレイテンシを軽減できます。レスポンスをストリーミングする際には、サーバー送信イベント(SSE)を使用してレスポンスを段階的にストリーミングします。
Mistral AI モデルは従量課金制です。従量課金制の料金については、Gemini Enterprise Agent Platform の料金ページで Mistral AI モデルの料金をご覧ください。
利用可能な Mistral AI モデル
Gemini Enterprise Agent Platform で使用できる Mistral AI のモデルは次のとおりです。Mistral AI モデルにアクセスするには、Model Garden のモデルカードに移動します。
Mistral AI モデルを使用する
curl コマンドを使用すると、次のモデル名を使用して Gemini Enterprise Agent Platform エンドポイントにリクエストを送信できます。
- Mistral Medium 3 の場合は
mistral-medium-3を使用します - Mistral OCR(25.05)の場合は、
mistral-ocr-2505を使用します - Mistral Small 3.1(25.03)の場合は、
mistral-small-2503を使用します - Codestral 2 の場合は
codestral-2を使用します
Mistral AI SDK の使用方法については、 Mistral AI Gemini Enterprise Agent Platform のドキュメントをご覧ください。
始める前に
Gemini Enterprise Agent Platform で Mistral AI モデルを使用するには、次の操作を行う必要があります。Gemini Enterprise Agent Platform を使用するには、Agent Platform API(aiplatform.googleapis.com)を有効にする必要があります。既存のプロジェクトで Agent Platform API が有効になっている場合は、新しいプロジェクトを作成する代わりに、そのプロジェクトを使用できます。
- Google Cloud アカウントにログインします。 Google Cloudを初めて使用する場合は、 アカウントを作成して、実際のシナリオで Google プロダクトのパフォーマンスを評価してください。新規のお客様には、ワークロードの実行、テスト、デプロイができる無料クレジット $300 分も差し上げます。
-
In the Google Cloud console, on the project selector page, select or create a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Verify that billing is enabled for your Google Cloud project.
Enable the Gemini Enterprise Agent Platform API, if it is not already enabled.
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.-
In the Google Cloud console, on the project selector page, select or create a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Verify that billing is enabled for your Google Cloud project.
Enable the Gemini Enterprise Agent Platform API, if it is not already enabled.
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.- 次のいずれかの Model Garden モデルカードに移動し、[有効にする] をクリックします。
Mistral AI モデルにストリーミング呼び出しを行う
次のサンプルでは、Mistral AI モデルへのストリーミング呼び出しを行います。
REST
環境をセットアップしたら、REST を使用してテキスト プロンプトをテストできます。次のサンプルは、パブリッシャー モデルのエンドポイントにリクエストを送信します。
リクエストのデータを使用する前に、次のように置き換えます。
- LOCATION: Mistral AI モデルをサポートするリージョン。
- MODEL: 使用するモデル名。リクエスト本文で、
@モデルのバージョン番号を除外します。 - ROLE: メッセージに関連付けられたロール。
userまたはassistantを指定できます。最初のメッセージでは、userロールを使用する必要があります。Claude モデルはuserとassistantのターンを交互に操作します。最後のメッセージがassistantロールを使用する場合、そのメッセージのコンテンツの直後にレスポンス コンテンツが続きます。これを使用して、モデルの回答の一部を制限できます。 - STREAM: レスポンスがストリーミングされるかどうかを指定するブール値。レスポンスのストリーミングを行うことで、エンドユーザーが認識するレイテンシを短縮できます。レスポンスをストリーミングする場合は
true、すべてのレスポンスを一度に戻すにはfalseに設定します。 - CONTENT:
userまたはassistantのメッセージの内容(テキストなど)。 - MAX_OUTPUT_TOKENS: レスポンスで生成できるトークンの最大数。トークンは約 3.5 文字です。100 トークンは約 60~80 語に相当します。
回答を短くしたい場合は小さい値を、長くしたい場合は大きい値を指定します。
HTTP メソッドと URL:
POST https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION/publishers/mistralai/models/MODEL:streamRawPredict
リクエストの本文(JSON):
{
"model": MODEL,
"messages": [
{
"role": "ROLE",
"content": "CONTENT"
}],
"max_tokens": MAX_TOKENS,
"stream": true
}
リクエストを送信するには、次のいずれかのオプションを選択します。
curl
リクエスト本文を request.json という名前のファイルに保存して、次のコマンドを実行します。
curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json; charset=utf-8" \
-d @request.json \
"https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION/publishers/mistralai/models/MODEL:streamRawPredict"
PowerShell
リクエスト本文を request.json という名前のファイルに保存して、次のコマンドを実行します。
$cred = gcloud auth print-access-token
$headers = @{ "Authorization" = "Bearer $cred" }
Invoke-WebRequest `
-Method POST `
-Headers $headers `
-ContentType: "application/json; charset=utf-8" `
-InFile request.json `
-Uri "https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION/publishers/mistralai/models/MODEL:streamRawPredict" | Select-Object -Expand Content
次のような JSON レスポンスが返されます。
Mistral AI モデルに単一呼び出しを行う
次のサンプルでは、Mistral AI モデルへの単一呼び出しを行います。
REST
環境をセットアップしたら、REST を使用してテキスト プロンプトをテストできます。次のサンプルは、パブリッシャー モデルのエンドポイントにリクエストを送信します。
リクエストのデータを使用する前に、次のように置き換えます。
- LOCATION: Mistral AI モデルをサポートするリージョン。
- MODEL: 使用するモデル名。リクエスト本文で、
@モデルのバージョン番号を除外します。 - ROLE: メッセージに関連付けられたロール。
userまたはassistantを指定できます。最初のメッセージでは、userロールを使用する必要があります。Claude モデルはuserとassistantのターンを交互に操作します。最後のメッセージがassistantロールを使用する場合、そのメッセージのコンテンツの直後にレスポンス コンテンツが続きます。これを使用して、モデルの回答の一部を制限できます。 - STREAM: レスポンスがストリーミングされるかどうかを指定するブール値。レスポンスのストリーミングを行うことで、エンドユーザーが認識するレイテンシを短縮できます。レスポンスをストリーミングする場合は
true、すべてのレスポンスを一度に戻すにはfalseに設定します。 - CONTENT:
userまたはassistantのメッセージの内容(テキストなど)。 - MAX_OUTPUT_TOKENS: レスポンスで生成できるトークンの最大数。トークンは約 3.5 文字です。100 トークンは約 60~80 語に相当します。
回答を短くしたい場合は小さい値を、長くしたい場合は大きい値を指定します。
HTTP メソッドと URL:
POST https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION/publishers/mistralai/models/MODEL:rawPredict
リクエストの本文(JSON):
{
"model": MODEL,
"messages": [
{
"role": "ROLE",
"content": "CONTENT"
}],
"max_tokens": MAX_TOKENS,
"stream": false
}
リクエストを送信するには、次のいずれかのオプションを選択します。
curl
リクエスト本文を request.json という名前のファイルに保存して、次のコマンドを実行します。
curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json; charset=utf-8" \
-d @request.json \
"https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION/publishers/mistralai/models/MODEL:rawPredict"
PowerShell
リクエスト本文を request.json という名前のファイルに保存して、次のコマンドを実行します。
$cred = gcloud auth print-access-token
$headers = @{ "Authorization" = "Bearer $cred" }
Invoke-WebRequest `
-Method POST `
-Headers $headers `
-ContentType: "application/json; charset=utf-8" `
-InFile request.json `
-Uri "https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION/publishers/mistralai/models/MODEL:rawPredict" | Select-Object -Expand Content
次のような JSON レスポンスが返されます。
Mistral AI モデルで利用可能なリージョンと割り当て
Mistral AI モデルの場合、モデルが使用可能なリージョンごとに割り当てが適用されます。割り当ては、1 分あたりのクエリ数(QPM)と 1 分あたりのトークン数(TPM)で指定されます。TPM には、入力トークンと出力トークンの両方が含まれます。
| モデル | リージョン | 割り当て | コンテキストの長さ |
|---|---|---|---|
| Mistral Medium 3 | |||
us-central1 |
|
128,000 | |
europe-west4 |
|
128,000 | |
| Mistral OCR(25.05) | |||
us-central1 |
|
30 ページ | |
europe-west4 |
|
30 ページ | |
| Mistral Small 3.1(25.03) | |||
us-central1 |
|
128,000 | |
europe-west4 |
|
128,000 | |
| Codestral 2 | |||
us-central1 |
|
128,000 トークン | |
europe-west4 |
|
128,000 トークン |
Agent Platform の割り当てを引き上げる場合は、 Google Cloud コンソールを使用して割り当ての引き上げをリクエストできます。割り当ての詳細については、クラウド割り当ての概要をご覧ください。