פריסה והיקש של Gemma באמצעות Model Garden ונקודות קצה של Gemini Enterprise Agent Platform עם תמיכה ב-GPU

במדריך הזה נשתמש ב-Model Garden כדי לפרוס את המודל הפתוח Gemma 3 1B לנקודת קצה של Gemini Enterprise Agent Platform עם תמיכה ב-GPU. כדי להשתמש במודל כדי להציג תחזיות אונליין, צריך לפרוס את המודל לנקודת קצה. פריסת מודל משייכת למודל משאבים פיזיים כדי שהוא יוכל לספק תחזיות אונליין עם זמן אחזור נמוך.

אחרי שפורסים את מודל Gemma 3 1B, מסיקים מסקנות מהמודל שאומן באמצעות PredictionServiceClient כדי לקבל חיזויים אונליין. תחזיות אונליין הן בקשות סנכרוניות שנשלחות למודל שפריסתו מתבצעת לנקודת קצה (endpoint).

מטרות

במדריך הזה מוסבר איך לבצע את הפעולות הבאות:

  • פריסת המודל הפתוח Gemma 3 1B לנקודת קצה עם GPU באמצעות Model Garden
  • משתמשים ב-PredictionServiceClient כדי לקבל תחזיות אונליין

עלויות

במסמך הזה משתמשים ברכיבים הבאים של Google Cloud, והשימוש בהם כרוך בתשלום:

כדי להעריך את ההוצאות בהתאם לתחזית השימוש שלכם, אתם יכולים להיעזר במחשבון העלויות.

משתמשים חדשים של Google Cloud ? יכול להיות שאתם זכאים לתקופת ניסיון בחינם.

כשמסיימים את המשימות שמתוארות במסמך הזה אפשר למחוק את המשאבים שיצרתם כדי להימנע מחיובים נוספים. מידע נוסף זמין בקטע הסרת המשאבים.

לפני שמתחילים

כדי לבצע את הפעולות במדריך הזה, צריך:

  • הגדרת פרויקט והפעלת Agent Platform API Google Cloud
  • במכונה המקומית:
    • התקנה, הפעלה ואימות באמצעות Google Cloud CLI
    • התקנה של ה-SDK בשפה שלכם

הגדרת Google Cloud פרויקט

מגדירים את הפרויקט ב- Google Cloud ומפעילים את Agent Platform API.

  1. נכנסים לחשבון Google Cloud . אם אתם משתמשים חדשים ב- Google Cloud, צרו חשבון כדי שתוכלו להעריך את הביצועים של המוצרים שלנו בתרחישים מהעולם האמיתי. לקוחות חדשים מקבלים בחינם גם קרדיט בשווי 300$ להרצה, לבדיקה ולפריסה של עומסי העבודה.
  2. In the Google Cloud console, on the project selector page, select or create a Google Cloud project.

    Roles required to select or create a project

    • Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
    • Create a project: To create a project, you need the Project Creator role (roles/resourcemanager.projectCreator), which contains the resourcemanager.projects.create permission. Learn how to grant roles.

    Go to project selector

  3. Verify that billing is enabled for your Google Cloud project.

  4. Enable the Agent Platform API.

    Roles required to enable APIs

    To enable APIs, you need the serviceusage.services.enable permission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.

    Enable the API

  5. In the Google Cloud console, on the project selector page, select or create a Google Cloud project.

    Roles required to select or create a project

    • Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
    • Create a project: To create a project, you need the Project Creator role (roles/resourcemanager.projectCreator), which contains the resourcemanager.projects.create permission. Learn how to grant roles.

    Go to project selector

  6. Verify that billing is enabled for your Google Cloud project.

  7. Enable the Agent Platform API.

    Roles required to enable APIs

    To enable APIs, you need the serviceusage.services.enable permission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.

    Enable the API

הגדרת Google Cloud CLI

במחשב המקומי, מגדירים את Google Cloud CLI.

  1. מתקינים ומפעילים את Google Cloud CLI.

  2. אם התקנתם בעבר את ה-CLI של gcloud, מריצים את הפקודה הזו כדי לוודא שהרכיבים של gcloud מעודכנים.

    gcloud components update
  3. כדי לבצע אימות באמצעות ה-CLI של gcloud, מריצים את הפקודה הזו כדי ליצור קובץ מקומי של Application Default Credentials ‏ (ADC). תהליך האינטרנט שמופעל על ידי הפקודה משמש להזנת פרטי הכניסה שלכם.

    gcloud auth application-default login

    למידע נוסף, ראו הגדרת אימות ב-CLI של gcloud והגדרת ADC.

הגדרת ה-SDK לשפת התכנות

כדי להגדיר את הסביבה שבה נעשה שימוש במדריך הזה, צריך להתקין את Gemini Enterprise Agent Platform SDK בשפה שלכם ואת ספריית Protocol Buffers. בדוגמאות הקוד נעשה שימוש בפונקציות מהספרייה של מאגרי אחסון לפרוטוקולים כדי להמיר את מילון הקלט לפורמט JSON שה-API מצפה לו.

במחשב המקומי, לוחצים על אחת מהכרטיסיות הבאות כדי להתקין את ה-SDK עבור שפת התכנות שלכם.

Python

במחשב המקומי, לוחצים על אחת מהכרטיסיות הבאות כדי להתקין את ה-SDK בשפת התכנות הרצויה.

  • כדי להתקין ולעדכן את Agent Platform SDK for Python, מריצים את הפקודה הבאה.

    pip3 install --upgrade "google-cloud-aiplatform>=1.64"
  • מריצים את הפקודה הבאה כדי להתקין את ספריית Protocol Buffers ל-Python.

    pip3 install --upgrade "protobuf>=5.28"

Node.js

כדי להתקין או לעדכן את aiplatform SDK for Node.js, מריצים את הפקודה הבאה.

npm install @google-cloud/aiplatform

Java

כדי להוסיף את google-cloud-aiplatform כתלות, מוסיפים את הקוד המתאים לסביבה שלכם.

‫Maven עם BOM

מוסיפים את קוד ה-HTML הבא לקובץ pom.xml:

<dependencyManagement>
<dependencies>
  <dependency>
    <artifactId>libraries-bom</artifactId>
    <groupId>com.google.cloud</groupId>
    <scope>import</scope>
    <type>pom</type>
    <version>26.34.0</version>
  </dependency>
</dependencies>
</dependencyManagement>
<dependencies>
<dependency>
  <groupId>com.google.cloud</groupId>
  <artifactId>google-cloud-aiplatform</artifactId>
</dependency>
<dependency>
  <groupId>com.google.protobuf</groupId>
  <artifactId>protobuf-java-util</artifactId>
</dependency>
<dependency>
  <groupId>com.google.code.gson</groupId>
  <artifactId>gson</artifactId>
</dependency>
</dependencies>

‫Maven בלי BOM

מוסיפים את השורות הבאות לקובץ pom.xml:

<dependency>
  <groupId>com.google.cloud</groupId>
  <artifactId>google-cloud-aiplatform</artifactId>
  <version>1.1.0</version>
</dependency>
<dependency>
  <groupId>com.google.protobuf</groupId>
  <artifactId>protobuf-java-util</artifactId>
  <version>5.28</version>
</dependency>
<dependency>
  <groupId>com.google.code.gson</groupId>
  <artifactId>gson</artifactId>
  <version>2.11.0</version>
</dependency>

‫Gradle ללא BOM

מוסיפים את הנתונים הבאים אל build.gradle:

implementation 'com.google.cloud:google-cloud-aiplatform:1.1.0'

המשך

כדי להתקין את חבילות Go האלה, מריצים את הפקודות הבאות.

go get cloud.google.com/go/aiplatform
go get google.golang.org/protobuf
go get github.com/googleapis/gax-go/v2

פריסת Gemma באמצעות Model Garden

אפשר לפרוס את Gemma 3 1B באמצעות כרטיס המודל שלה במסוף Google Cloud או באופן פרוגרמטי.

מידע נוסף על הגדרת Google Gen AI SDK או Google Cloud CLI זמין במאמרים סקירה כללית על Google Gen AI SDK והתקנת Google Cloud CLI.

Python

במאמר התקנת Vertex AI SDK ל-Python מוסבר איך להתקין או לעדכן את Vertex AI SDK ל-Python. מידע נוסף מופיע ב מאמרי העזרה של Python API.

  1. מציגים את רשימת המודלים שאפשר לפרוס ורושמים את מזהה המודל לפריסה. אפשר גם לראות את רשימת המודלים הנתמכים של Hugging Face ב-Model Garden, ואפילו לסנן אותם לפי שמות המודלים. הפלט לא כולל מודלים שעברו התאמה.

    
    import vertexai
    from vertexai import model_garden
    
    # TODO(developer): Update and un-comment below lines
    # PROJECT_ID = "your-project-id"
    vertexai.init(project=PROJECT_ID, location="us-central1")
    
    # List deployable models, optionally list Hugging Face models only or filter by model name.
    deployable_models = model_garden.list_deployable_models(list_hf_models=False, model_filter="gemma")
    print(deployable_models)
    # Example response:
    # ['google/gemma2@gemma-2-27b','google/gemma2@gemma-2-27b-it', ...]
    
  2. כדי לראות את מפרטי הפריסה של מודל, משתמשים במזהה המודל מהשלב הקודם. אפשר לראות את סוג המכונה, סוג המאיץ ו-URI של קובץ אימג' של קונטיינר ש-Model Garden אימת עבור מודל מסוים.

    
    import vertexai
    from vertexai import model_garden
    
    # TODO(developer): Update and un-comment below lines
    # PROJECT_ID = "your-project-id"
    # model = "google/gemma3@gemma-3-1b-it"
    vertexai.init(project=PROJECT_ID, location="us-central1")
    
    # For Hugging Face modelsm the format is the Hugging Face model name, as in
    # "meta-llama/Llama-3.3-70B-Instruct".
    # Go to https://console.cloud.google.com/vertex-ai/model-garden to find all deployable
    # model names.
    
    model = model_garden.OpenModel(model)
    deploy_options = model.list_deploy_options()
    print(deploy_options)
    # Example response:
    # [
    #   dedicated_resources {
    #     machine_spec {
    #       machine_type: "g2-standard-12"
    #       accelerator_type: NVIDIA_L4
    #       accelerator_count: 1
    #     }
    #   }
    #   container_spec {
    #     ...
    #   }
    #   ...
    # ]
    
  3. פורסים מודל בנקודת קצה. ב-Model Garden נעשה שימוש בהגדרת הפריסה שמוגדרת כברירת מחדל, אלא אם מציינים ארגומנטים וערכים נוספים.

    
    import vertexai
    from vertexai import model_garden
    
    # TODO(developer): Update and un-comment below lines
    # PROJECT_ID = "your-project-id"
    vertexai.init(project=PROJECT_ID, location="us-central1")
    
    open_model = model_garden.OpenModel("google/gemma3@gemma-3-12b-it")
    endpoint = open_model.deploy(
        machine_type="g2-standard-48",
        accelerator_type="NVIDIA_L4",
        accelerator_count=4,
        accept_eula=True,
    )
    
    # Optional. Run predictions on the deployed endoint.
    # endpoint.predict(instances=[{"prompt": "What is Generative AI?"}])
    

gcloud

לפני שמתחילים, צריך לציין פרויקט מכסה כדי להריץ את הפקודות הבאות. הפקודות שמריצים נספרות כחלק מהמכסות של הפרויקט. מידע נוסף זמין במאמר הגדרת פרויקט לצורכי מכסה.

  1. מריצים את הפקודה gcloud ai model-garden models list כדי לראות את רשימת המודלים שאפשר לפרוס. הפקודה הזו מציגה רשימה של כל מזהי המודלים, ומציינת אילו מודלים אפשר לפרוס באופן עצמאי.

    gcloud ai model-garden models list --model-filter=gemma
    

    בפלט, מאתרים את מזהה המודל לפריסה. בדוגמה הבאה מוצג פלט מקוצר.

    MODEL_ID                                      CAN_DEPLOY  CAN_PREDICT
    google/gemma2@gemma-2-27b                     Yes         No
    google/gemma2@gemma-2-27b-it                  Yes         No
    google/gemma2@gemma-2-2b                      Yes         No
    google/gemma2@gemma-2-2b-it                   Yes         No
    google/gemma2@gemma-2-9b                      Yes         No
    google/gemma2@gemma-2-9b-it                   Yes         No
    google/gemma3@gemma-3-12b-it                  Yes         No
    google/gemma3@gemma-3-12b-pt                  Yes         No
    google/gemma3@gemma-3-1b-it                   Yes         No
    google/gemma3@gemma-3-1b-pt                   Yes         No
    google/gemma3@gemma-3-27b-it                  Yes         No
    google/gemma3@gemma-3-27b-pt                  Yes         No
    google/gemma3@gemma-3-4b-it                   Yes         No
    google/gemma3@gemma-3-4b-pt                   Yes         No
    google/gemma3n@gemma-3n-e2b                   Yes         No
    google/gemma3n@gemma-3n-e2b-it                Yes         No
    google/gemma3n@gemma-3n-e4b                   Yes         No
    google/gemma3n@gemma-3n-e4b-it                Yes         No
    google/gemma@gemma-1.1-2b-it                  Yes         No
    google/gemma@gemma-1.1-2b-it-gg-hf            Yes         No
    google/gemma@gemma-1.1-7b-it                  Yes         No
    google/gemma@gemma-1.1-7b-it-gg-hf            Yes         No
    google/gemma@gemma-2b                         Yes         No
    google/gemma@gemma-2b-gg-hf                   Yes         No
    google/gemma@gemma-2b-it                      Yes         No
    google/gemma@gemma-2b-it-gg-hf                Yes         No
    google/gemma@gemma-7b                         Yes         No
    google/gemma@gemma-7b-gg-hf                   Yes         No
    google/gemma@gemma-7b-it                      Yes         No
    google/gemma@gemma-7b-it-gg-hf                Yes         No
    

    הפלט לא כולל מודלים שעברו כוונון או מודלים של Hugging Face. כדי לראות אילו מודלים של Hugging Face נתמכים, מוסיפים את הדגל --can-deploy-hugging-face-models.

  2. כדי לראות את מפרטי הפריסה של מודל, מריצים את הפקודה gcloud ai model-garden models list-deployment-config. אפשר לראות את סוג המכונה, סוג המאיץ ו-URI של קובץ אימג' של קונטיינר ש-Model Garden תומך בהם עבור מודל מסוים.

    gcloud ai model-garden models list-deployment-config \
        --model=MODEL_ID
    

    מחליפים את MODEL_ID במזהה המודל מהפקודה הקודמת של רשימת המודלים, כמו google/gemma@gemma-2b או stabilityai/stable-diffusion-xl-base-1.0.

  3. כדי לפרוס מודל לנקודת קצה, מריצים את הפקודה gcloud ai model-garden models deploy. ‫Model Garden יוצר שם מוצג לנקודת הקצה ומשתמש בהגדרות הפריסה שמוגדרות כברירת מחדל, אלא אם מציינים ארגומנטים וערכים נוספים.

    כדי להריץ את הפקודה באופן אסינכרוני, כוללים את הדגל --asynchronous.

    gcloud ai model-garden models deploy \
        --model=MODEL_ID \
        [--machine-type=MACHINE_TYPE] \
        [--accelerator-type=ACCELERATOR_TYPE] \
        [--endpoint-display-name=ENDPOINT_NAME] \
        [--hugging-face-access-token=HF_ACCESS_TOKEN] \
        [--reservation-affinity reservation-affinity-type=any-reservation] \
        [--reservation-affinity reservation-affinity-type=specific-reservation, key="compute.googleapis.com/reservation-name", values=RESERVATION_RESOURCE_NAME] \
        [--asynchronous]
    

    מחליפים את ה-placeholders הבאים:

    • MODEL_ID: מזהה המודל מפקודת הרשימה הקודמת. במקרה של מודלים של Hugging Face, צריך להשתמש בפורמט של כתובת ה-URL של המודל של Hugging Face, כמו stabilityai/stable-diffusion-xl-base-1.0.
    • MACHINE_TYPE: הגדרת קבוצת המשאבים לפריסה של המודל, כמו g2-standard-4.
    • ACCELERATOR_TYPE: מציינים מאיצים להוספה לפריסה כדי לשפר את הביצועים כשעובדים עם עומסי עבודה אינטנסיביים, כמו NVIDIA_L4.
    • ENDPOINT_NAME: שם לנקודת הקצה של Gemini Enterprise Agent Platform שהופעלה.
    • HF_ACCESS_TOKEN: במודלים של Hugging Face, אם המודל מוגבל, צריך לספק אסימון גישה.
    • RESERVATION_RESOURCE_NAME: כדי להשתמש במקום שמור ספציפי ב-Compute Engine, מציינים את שם המקום השמור. אם מציינים הזמנה ספציפית, אי אפשר לציין את any-reservation.

    הפלט כולל את הגדרת הפריסה שבה נעשה שימוש ב-Model Garden, את מזהה נקודת הקצה ואת מזהה פעולת הפריסה, שבהם אפשר להשתמש כדי לבדוק את סטטוס הפריסה.

    Using the default deployment configuration:
     Machine type: g2-standard-12
     Accelerator type: NVIDIA_L4
     Accelerator count: 1
    
    The project has enough quota. The current usage of quota for accelerator type NVIDIA_L4 in region us-central1 is 0 out of 28.
    
    Deploying the model to the endpoint. To check the deployment status, you can try one of the following methods:
    1) Look for endpoint `ENDPOINT_DISPLAY_NAME` at the [Agent Platform] -> [Online prediction] tab in Cloud Console
    2) Use `gcloud ai operations describe OPERATION_ID --region=LOCATION` to find the status of the deployment long-running operation
    
  4. כדי לראות פרטים על הפריסה, מריצים את הפקודה gcloud ai endpoints list --list-model-garden-endpoints-only:

    gcloud ai endpoints list --list-model-garden-endpoints-only \
        --region=LOCATION_ID
    

    מחליפים את LOCATION_ID באזור שבו פרסתם את המודל.

    הפלט כולל את כל נקודות הקצה שנוצרו מ-Model Garden, וכולל מידע כמו מזהה נקודת הקצה, שם נקודת הקצה והאם נקודת הקצה משויכת למודל שנפרס. כדי למצוא את הפריסה, מחפשים את שם נקודת הקצה שהוחזר מהפקודה הקודמת.

REST

מציגים רשימה של כל המודלים שאפשר לפרוס, ואז מקבלים את המזהה של המודל שרוצים לפרוס. אחר כך תוכלו לפרוס את המודל עם הגדרות ברירת המחדל ונקודת הקצה שלו. אפשר גם לבחור להתאים אישית את הפריסה, למשל להגדיר סוג מכונה ספציפי או להשתמש בנקודת קצה ייעודית.

הצגת רשימת המודלים שאפשר לפרוס

לפני שמשתמשים בנתוני הבקשה, צריך להחליף את הנתונים הבאים:

  • PROJECT_ID: מזהה הפרויקט ב- Google Cloud .
  • QUERY_PARAMETERS: כדי להציג רשימה של מודלים ב-Model Garden, מוסיפים את פרמטרי השאילתה הבאים listAllVersions=True&filter=can_deploy(true). כדי להציג רשימה של מודלים של Hugging Face, מגדירים את המסנן ל-alt=json&is_hf_wildcard(true)+AND+labels.VERIFIED_DEPLOYMENT_CONFIG%3DVERIFIED_DEPLOYMENT_SUCCEED&listAllVersions=True.

ה-method של ה-HTTP וכתובת ה-URL:

GET https://us-central1-aiplatform.googleapis.com/v1/publishers/*/models?QUERY_PARAMETERS

כדי לשלוח את הבקשה אתם צריכים לבחור אחת מהאפשרויות הבאות:

curl

מריצים את הפקודה הבאה:

curl -X GET \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "x-goog-user-project: PROJECT_ID" \
"https://us-central1-aiplatform.googleapis.com/v1/publishers/*/models?QUERY_PARAMETERS"

PowerShell

מריצים את הפקודה הבאה:

$cred = gcloud auth print-access-token
$headers = @{ "Authorization" = "Bearer $cred"; "x-goog-user-project" = "PROJECT_ID" }

Invoke-WebRequest `
-Method GET `
-Headers $headers `
-Uri "https://us-central1-aiplatform.googleapis.com/v1/publishers/*/models?QUERY_PARAMETERS" | Select-Object -Expand Content

מקבלים תגובת JSON שדומה לזו.

{
  "publisherModels": [
    {
      "name": "publishers/google/models/gemma3",
      "versionId": "gemma-3-1b-it",
      "openSourceCategory": "GOOGLE_OWNED_OSS_WITH_GOOGLE_CHECKPOINT",
      "supportedActions": {
        "openNotebook": {
          "references": {
            "us-central1": {
              "uri": "https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/model_garden/model_garden_gradio_streaming_chat_completions.ipynb"
            }
          },
          "resourceTitle": "Notebook",
          "resourceUseCase": "Chat Completion Playground",
          "resourceDescription": "Chat with deployed Gemma 2 endpoints via Gradio UI."
        },
        "deploy": {
          "modelDisplayName": "gemma-3-1b-it",
          "containerSpec": {
            "imageUri": "us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20250312_0916_RC01",
            "args": [
              "python",
              "-m",
              "vllm.entrypoints.api_server",
              "--host=0.0.0.0",
              "--port=8080",
              "--model=gs://vertex-model-garden-restricted-us/gemma3/gemma-3-1b-it",
              "--tensor-parallel-size=1",
              "--swap-space=16",
              "--gpu-memory-utilization=0.95",
              "--disable-log-stats"
            ],
            "env": [
              {
                "name": "MODEL_ID",
                "value": "google/gemma-3-1b-it"
              },
              {
                "name": "DEPLOY_SOURCE",
                "value": "UI_NATIVE_MODEL"
              }
            ],
            "ports": [
              {
                "containerPort": 8080
              }
            ],
            "predictRoute": "/generate",
            "healthRoute": "/ping"
          },
          "dedicatedResources": {
            "machineSpec": {
              "machineType": "g2-standard-12",
              "acceleratorType": "NVIDIA_L4",
              "acceleratorCount": 1
            }
          },
          "publicArtifactUri": "gs://vertex-model-garden-restricted-us/gemma3/gemma3.tar.gz",
          "deployTaskName": "vLLM 128K context",
          "deployMetadata": {
            "sampleRequest": "{\n    \"instances\": [\n        {\n          \"@requestFormat\": \"chatCompletions\",\n          \"messages\": [\n              {\n                  \"role\": \"user\",\n                  \"content\": \"What is machine learning?\"\n              }\n          ],\n          \"max_tokens\": 100\n        }\n    ]\n}\n"
          }
        },
        ...

פריסת מודל

פריסת מודל מ-Model Garden או מ-Hugging Face. אפשר גם להתאים אישית את הפריסה על ידי ציון שדות JSON נוספים.

פריסת מודל עם הגדרות ברירת המחדל שלו.

לפני שמשתמשים בנתוני הבקשה, צריך להחליף את הנתונים הבאים:

  • LOCATION: אזור שבו המודל פרוס.
  • PROJECT_ID: מזהה הפרויקט ב- Google Cloud .
  • MODEL_ID: המזהה של המודל שרוצים לפרוס. אפשר לקבל אותו מרשימת כל המודלים שאפשר לפרוס. המזהה הוא בפורמט הבא: publishers/PUBLISHER_NAME/models/ MODEL_NAME@MODEL_VERSION.

ה-method של ה-HTTP וכתובת ה-URL:

POST https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION:deploy

גוף בקשת JSON:

{
  "publisher_model_name": "MODEL_ID",
  "model_config": {
    "accept_eula": "true"
  }
}

כדי לשלוח את הבקשה עליכם לבחור אחת מהאפשרויות הבאות:

curl

שומרים את גוף הבקשה בקובץ בשם request.json. כדי ליצור או להחליף את הקובץ הזה בספרייה הנוכחית, מריצים את הפקודה הבאה בטרמינל:

cat > request.json << 'EOF'
{
  "publisher_model_name": "MODEL_ID",
  "model_config": {
    "accept_eula": "true"
  }
}
EOF

לאחר מכן מבצעים את הפקודה הבאה כדי לשלוח את בקשת ה-REST:

curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json; charset=utf-8" \
-d @request.json \
"https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION:deploy"

PowerShell

שומרים את גוף הבקשה בקובץ בשם request.json. כדי ליצור או להחליף את הקובץ הזה בספרייה הנוכחית, מריצים את הפקודה הבאה בטרמינל:

@'
{
  "publisher_model_name": "MODEL_ID",
  "model_config": {
    "accept_eula": "true"
  }
}
'@  | Out-File -FilePath request.json -Encoding utf8

לאחר מכן מבצעים את הפקודה הבאה כדי לשלוח את בקשת ה-REST:

$cred = gcloud auth print-access-token
$headers = @{ "Authorization" = "Bearer $cred" }

Invoke-WebRequest `
-Method POST `
-Headers $headers `
-ContentType: "application/json; charset=utf-8" `
-InFile request.json `
-Uri "https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION:deploy" | Select-Object -Expand Content

מקבלים תגובת JSON שדומה לזו.

{
  "name": "projects/PROJECT_ID/locations/LOCATION/operations/OPERATION_ID",
  "metadata": {
    "@type": "type.googleapis.com/google.cloud.aiplatform.v1.DeployOperationMetadata",
    "genericMetadata": {
      "createTime": "2025-03-13T21:44:44.538780Z",
      "updateTime": "2025-03-13T21:44:44.538780Z"
    },
    "publisherModel": "publishers/google/models/gemma3@gemma-3-1b-it",
    "destination": "projects/PROJECT_ID/locations/LOCATION",
    "projectNumber": "PROJECT_ID"
  }
}

פריסת מודל של Hugging Face

לפני שמשתמשים בנתוני הבקשה, צריך להחליף את הנתונים הבאים:

  • LOCATION: אזור שבו המודל פרוס.
  • PROJECT_ID: מזהה הפרויקט ב- Google Cloud .
  • MODEL_ID: מזהה המודל של Hugging Face שרוצים לפרוס. אפשר לקבל אותו מתוך רשימת כל המודלים שניתנים לפריסה. המזהה הוא בפורמט הבא: PUBLISHER_NAME/MODEL_NAME.
  • ACCESS_TOKEN: אם המודל מוגבל, צריך לספק אסימון גישה.

ה-method של ה-HTTP וכתובת ה-URL:

POST https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION:deploy

גוף בקשת JSON:

{
  "hugging_face_model_id": "MODEL_ID",
  "hugging_face_access_token": "ACCESS_TOKEN",
  "model_config": {
    "accept_eula": "true"
  }
}

כדי לשלוח את הבקשה עליכם לבחור אחת מהאפשרויות הבאות:

curl

שומרים את גוף הבקשה בקובץ בשם request.json. כדי ליצור או להחליף את הקובץ הזה בספרייה הנוכחית, מריצים את הפקודה הבאה בטרמינל:

cat > request.json << 'EOF'
{
  "hugging_face_model_id": "MODEL_ID",
  "hugging_face_access_token": "ACCESS_TOKEN",
  "model_config": {
    "accept_eula": "true"
  }
}
EOF

לאחר מכן מבצעים את הפקודה הבאה כדי לשלוח את בקשת ה-REST:

curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json; charset=utf-8" \
-d @request.json \
"https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION:deploy"

PowerShell

שומרים את גוף הבקשה בקובץ בשם request.json. כדי ליצור או להחליף את הקובץ הזה בספרייה הנוכחית, מריצים את הפקודה הבאה בטרמינל:

@'
{
  "hugging_face_model_id": "MODEL_ID",
  "hugging_face_access_token": "ACCESS_TOKEN",
  "model_config": {
    "accept_eula": "true"
  }
}
'@  | Out-File -FilePath request.json -Encoding utf8

לאחר מכן מבצעים את הפקודה הבאה כדי לשלוח את בקשת ה-REST:

$cred = gcloud auth print-access-token
$headers = @{ "Authorization" = "Bearer $cred" }

Invoke-WebRequest `
-Method POST `
-Headers $headers `
-ContentType: "application/json; charset=utf-8" `
-InFile request.json `
-Uri "https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION:deploy" | Select-Object -Expand Content

מקבלים תגובת JSON שדומה לזו.

{
  "name": "projects/PROJECT_ID/locations/us-central1LOCATION/operations/OPERATION_ID",
  "metadata": {
    "@type": "type.googleapis.com/google.cloud.aiplatform.v1.DeployOperationMetadata",
    "genericMetadata": {
      "createTime": "2025-03-13T21:44:44.538780Z",
      "updateTime": "2025-03-13T21:44:44.538780Z"
    },
    "publisherModel": "publishers/PUBLISHER_NAME/model/MODEL_NAME",
    "destination": "projects/PROJECT_ID/locations/LOCATION",
    "projectNumber": "PROJECT_ID"
  }
}

פריסת מודל עם התאמות אישיות

לפני שמשתמשים בנתוני הבקשה, צריך להחליף את הנתונים הבאים:

  • LOCATION: אזור שבו המודל פרוס.
  • PROJECT_ID: מזהה הפרויקט ב- Google Cloud .
  • MODEL_ID: המזהה של המודל שרוצים לפרוס. אפשר לקבל אותו מרשימת כל המודלים שאפשר לפרוס. המזהה הוא בפורמט הבא: publishers/PUBLISHER_NAME/models/ MODEL_NAME@MODEL_VERSION, לדוגמה: google/gemma@gemma-2b או stabilityai/stable-diffusion-xl-base-1.0.
  • MACHINE_TYPE: הגדרת קבוצת המשאבים לפריסה של המודל, כמו g2-standard-4.
  • ACCELERATOR_TYPE: מציינים מאיצים להוספה לפריסה כדי לשפר את הביצועים כשעובדים עם עומסי עבודה אינטנסיביים, כמו NVIDIA_L4
  • ACCELERATOR_COUNT: מספר המאיצים לשימוש בפריסה.
  • reservation_affinity_type: כדי להשתמש בהזמנה קיימת של Compute Engine לפריסה, מציינים הזמנה כלשהי או הזמנה ספציפית. אם מציינים את הערך הזה, לא מציינים את הערך spot.
  • spot: האם להשתמש ב-VM במודל Spot לפריסה.
  • IMAGE_URI: המיקום של קובץ אימג' של קונטיינר שבו רוצים להשתמש, למשל us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20241016_0916_RC00_maas
  • CONTAINER_ARGS: ארגומנטים להעברה לקונטיינר במהלך הפריסה.
  • CONTAINER_PORT: מספר היציאה של הקונטיינר.
  • fast_tryout_enabled: כשבודקים מודל, אפשר לבחור להשתמש בפריסה מהירה יותר. האפשרות הזו זמינה רק למודלים בשימוש נרחב עם סוגים מסוימים של מכונות. אם האפשרות הזו מופעלת, אי אפשר לציין מודל או תצורות פריסה.

ה-method של ה-HTTP וכתובת ה-URL:

POST https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION:deploy

גוף בקשת JSON:

{
  "publisher_model_name": "MODEL_ID",
  "deploy_config": {
    "dedicated_resources": {
      "machine_spec": {
        "machine_type": "MACHINE_TYPE",
        "accelerator_type": "ACCELERATOR_TYPE",
        "accelerator_count": ACCELERATOR_COUNT,
        "reservation_affinity": {
          "reservation_affinity_type": "ANY_RESERVATION"
        }
      },
      "spot": "false"
    }
  },
  "model_config": {
    "accept_eula": "true",
    "container_spec": {
      "image_uri": "IMAGE_URI",
      "args": [CONTAINER_ARGS ],
      "ports": [
        {
          "container_port": CONTAINER_PORT
        }
      ]
    }
  },
  "deploy_config": {
    "fast_tryout_enabled": false
  },
}

כדי לשלוח את הבקשה עליכם לבחור אחת מהאפשרויות הבאות:

curl

שומרים את גוף הבקשה בקובץ בשם request.json. כדי ליצור או להחליף את הקובץ הזה בספרייה הנוכחית, מריצים את הפקודה הבאה בטרמינל:

cat > request.json << 'EOF'
{
  "publisher_model_name": "MODEL_ID",
  "deploy_config": {
    "dedicated_resources": {
      "machine_spec": {
        "machine_type": "MACHINE_TYPE",
        "accelerator_type": "ACCELERATOR_TYPE",
        "accelerator_count": ACCELERATOR_COUNT,
        "reservation_affinity": {
          "reservation_affinity_type": "ANY_RESERVATION"
        }
      },
      "spot": "false"
    }
  },
  "model_config": {
    "accept_eula": "true",
    "container_spec": {
      "image_uri": "IMAGE_URI",
      "args": [CONTAINER_ARGS ],
      "ports": [
        {
          "container_port": CONTAINER_PORT
        }
      ]
    }
  },
  "deploy_config": {
    "fast_tryout_enabled": false
  },
}
EOF

לאחר מכן מבצעים את הפקודה הבאה כדי לשלוח את בקשת ה-REST:

curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json; charset=utf-8" \
-d @request.json \
"https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION:deploy"

PowerShell

שומרים את גוף הבקשה בקובץ בשם request.json. כדי ליצור או להחליף את הקובץ הזה בספרייה הנוכחית, מריצים את הפקודה הבאה בטרמינל:

@'
{
  "publisher_model_name": "MODEL_ID",
  "deploy_config": {
    "dedicated_resources": {
      "machine_spec": {
        "machine_type": "MACHINE_TYPE",
        "accelerator_type": "ACCELERATOR_TYPE",
        "accelerator_count": ACCELERATOR_COUNT,
        "reservation_affinity": {
          "reservation_affinity_type": "ANY_RESERVATION"
        }
      },
      "spot": "false"
    }
  },
  "model_config": {
    "accept_eula": "true",
    "container_spec": {
      "image_uri": "IMAGE_URI",
      "args": [CONTAINER_ARGS ],
      "ports": [
        {
          "container_port": CONTAINER_PORT
        }
      ]
    }
  },
  "deploy_config": {
    "fast_tryout_enabled": false
  },
}
'@  | Out-File -FilePath request.json -Encoding utf8

לאחר מכן מבצעים את הפקודה הבאה כדי לשלוח את בקשת ה-REST:

$cred = gcloud auth print-access-token
$headers = @{ "Authorization" = "Bearer $cred" }

Invoke-WebRequest `
-Method POST `
-Headers $headers `
-ContentType: "application/json; charset=utf-8" `
-InFile request.json `
-Uri "https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION:deploy" | Select-Object -Expand Content

מקבלים תגובת JSON שדומה לזו.

{
  "name": "projects/PROJECT_ID/locations/LOCATION/operations/OPERATION_ID",
  "metadata": {
    "@type": "type.googleapis.com/google.cloud.aiplatform.v1.DeployOperationMetadata",
    "genericMetadata": {
      "createTime": "2025-03-13T21:44:44.538780Z",
      "updateTime": "2025-03-13T21:44:44.538780Z"
    },
    "publisherModel": "publishers/google/models/gemma3@gemma-3-1b-it",
    "destination": "projects/PROJECT_ID/locations/LOCATION",
    "projectNumber": "PROJECT_ID"
  }
}

המסוף

  1. נכנסים לדף Model Garden במסוף Google Cloud .

    כניסה ל-Model Garden

  2. מוצאים מודל נתמך שרוצים לפרוס ולוחצים על כרטיס המודל שלו.

  3. לוחצים על Deploy (פריסה) כדי לפתוח את החלונית Deploy model (פריסת מודל).

  4. בחלונית Deploy model (פריסת המודל), מציינים את פרטי הפריסה.

    1. משתמשים בשמות של המודל ונקודת הקצה שנוצרו או משנים אותם.
    2. בוחרים מיקום שבו רוצים ליצור את נקודת הקצה של המודל.
    3. בוחרים את סוג המכונה שבה רוצים להשתמש לכל צומת בפריסה.
    4. כדי להשתמש בהזמנה ב-Compute Engine, בקטע Deployment settings (הגדרות פריסה), בוחרים באפשרות Advanced (מתקדם).

      בשדה סוג ההזמנה, בוחרים סוג הזמנה. ההזמנה צריכה להתאים למפרט המכונה שציינתם.

      • שימוש אוטומטי במקום שמור שנוצר: Gemini Enterprise Agent Platform בוחרת באופן אוטומטי מקום שמור מותר עם מאפיינים תואמים. אם אין קיבולת בהזמנה שנבחרה אוטומטית, Gemini Enterprise Agent Platform משתמש במאגר הכללי של Google Cloud משאבים.
      • בחירת הזמנות ספציפיות: Gemini Enterprise Agent Platform משתמש בהזמנה ספציפית. אם אין קיבולת להזמנה שבחרתם, תופיע שגיאה.
      • לא להשתמש (ברירת מחדל): Gemini Enterprise Agent Platform משתמש במאגר הכללי של משאביGoogle Cloud . הערך הזה זהה למצב שבו לא מציינים הזמנה.
  5. לוחצים על פריסה.

Terraform

כדי ללמוד איך להחיל הגדרות ב-Terraform או להסיר אותן, ראו פקודות בסיסיות ב-Terraform. למידע נוסף, ראו את מאמרי העזרה לספקים של Terraform.

פריסת מודל

בדוגמה הבאה, מפריסים את מודל gemma-3-1b-it לנקודת קצה חדשה של Agent Platform ב-us-central1 באמצעות הגדרות ברירת מחדל.

terraform {
  required_providers {
    google = {
      source = "hashicorp/google"
      version = "6.45.0"
    }
  }
}

provider "google" {
  region  = "us-central1"
}

resource "google_vertex_ai_endpoint_with_model_garden_deployment" "gemma_deployment" {
  publisher_model_name = "publishers/google/models/gemma3@gemma-3-1b-it"
  location = "us-central1"
  model_config {
    accept_eula = True
  }
}

פרטים נוספים על פריסת מודל עם התאמה אישית זמינים במאמר נקודת קצה של Agent Platform עם פריסה של Model Garden.

החלת ההגדרה

terraform init
terraform plan
terraform apply

אחרי שמחילים את ההגדרות, Terraform מקצה נקודת קצה חדשה של Agent Platform ופורס את המודל הפתוח שצוין.

מחיקה

כדי למחוק את נקודת הקצה ואת פריסת המודל, מריצים את הפקודה הבאה:

terraform destroy

הסקת מסקנות של Gemma 3 1B באמצעות PredictionServiceClient

אחרי שמפעילים את Gemma 3 1B, משתמשים ב-PredictionServiceClient כדי לקבל תחזיות אונליין להנחיה: "למה השמיים כחולים?"

פרמטרים של קוד

בדוגמאות הקוד של PredictionServiceClient צריך לעדכן את הפרטים הבאים.

  • PROJECT_ID: כדי למצוא את מזהה הפרויקט, פועלים לפי השלבים הבאים.

    1. נכנסים לדף Welcome במסוף Google Cloud .

      מעבר לדף Welcome

    2. בוחרים את הפרויקט מתוך כלי לבחירת פרויקטים בחלק העליון של הדף.

      שם הפרויקט, מספר הפרויקט ומזהה הפרויקט מופיעים אחרי הכותרת Welcome.

  • ENDPOINT_REGION: האזור שבו פרסתם את נקודת הקצה.

  • ENDPOINT_ID: כדי למצוא את מזהה נקודת הקצה, אפשר לראות אותו במסוף או להריץ את הפקודה gcloud ai endpoints list. תצטרכו את שם נקודת הקצה והאזור מהחלונית Deploy model.

    המסוף

    כדי לראות את הפרטים של נקודת הקצה, לוחצים על Online prediction (חיזוי אונליין) > Endpoints (נקודות קצה) ובוחרים את האזור. שימו לב למספר שמופיע בעמודה ID.

    לדף Endpoints

    gcloud

    כדי לראות את פרטי נקודת הקצה, מריצים את הפקודה gcloud ai endpoints list.

    gcloud ai endpoints list \
      --region=ENDPOINT_REGION \
      --filter=display_name=ENDPOINT_NAME
    

    הפלט אמור להיראות כך:

    Using endpoint [https://us-central1-aiplatform.googleapis.com/]
    ENDPOINT_ID: 1234567891234567891
    DISPLAY_NAME: gemma2-2b-it-mg-one-click-deploy
    

קוד לדוגמה

בקוד לדוגמה בשפה שלכם, מעדכנים את PROJECT_ID, ENDPOINT_REGION ו-ENDPOINT_ID. לאחר מכן מריצים את הקוד.

Python

במאמר התקנת Vertex AI SDK ל-Python מוסבר איך להתקין או לעדכן את Vertex AI SDK ל-Python. מידע נוסף מופיע ב מאמרי העזרה של Python API.

"""
Sample to run inference on a Gemma2 model deployed to a Vertex AI endpoint with GPU accellerators.
"""

from google.cloud import aiplatform
from google.protobuf import json_format
from google.protobuf.struct_pb2 import Value

# TODO(developer): Update & uncomment lines below
# PROJECT_ID = "your-project-id"
# ENDPOINT_REGION = "your-vertex-endpoint-region"
# ENDPOINT_ID = "your-vertex-endpoint-id"

# Default configuration
config = {"max_tokens": 1024, "temperature": 0.9, "top_p": 1.0, "top_k": 1}

# Prompt used in the prediction
prompt = "Why is the sky blue?"

# Encapsulate the prompt in a correct format for GPUs
# Example format: [{'inputs': 'Why is the sky blue?', 'parameters': {'temperature': 0.9}}]
input = {"inputs": prompt, "parameters": config}

# Convert input message to a list of GAPIC instances for model input
instances = [json_format.ParseDict(input, Value())]

# Create a client
api_endpoint = f"{ENDPOINT_REGION}-aiplatform.googleapis.com"
client = aiplatform.gapic.PredictionServiceClient(
    client_options={"api_endpoint": api_endpoint}
)

# Call the Gemma2 endpoint
gemma2_end_point = (
    f"projects/{PROJECT_ID}/locations/{ENDPOINT_REGION}/endpoints/{ENDPOINT_ID}"
)
response = client.predict(
    endpoint=gemma2_end_point,
    instances=instances,
)
text_responses = response.predictions
print(text_responses[0])

Node.js

לפני שמנסים את הדוגמה הזו, צריך לפעול לפי Node.jsההוראות להגדרה במאמר מדריך למתחילים של Agent Platform באמצעות ספריות לקוח.

כדי לבצע אימות ב-Agent Platform, צריך להגדיר את Application Default Credentials. מידע נוסף זמין במאמר הגדרת אימות לסביבת פיתוח מקומית.

async function gemma2PredictGpu(predictionServiceClient) {
  // Imports the Google Cloud Prediction Service Client library
  const {
    // TODO(developer): Uncomment PredictionServiceClient before running the sample.
    // PredictionServiceClient,
    helpers,
  } = require('@google-cloud/aiplatform');
  /**
   * TODO(developer): Update these variables before running the sample.
   */
  const projectId = 'your-project-id';
  const endpointRegion = 'your-vertex-endpoint-region';
  const endpointId = 'your-vertex-endpoint-id';

  // Default configuration
  const config = {maxOutputTokens: 1024, temperature: 0.9, topP: 1.0, topK: 1};
  // Prompt used in the prediction
  const prompt = 'Why is the sky blue?';

  // Encapsulate the prompt in a correct format for GPUs
  // Example format: [{inputs: 'Why is the sky blue?', parameters: {temperature: 0.9}}]
  const input = {
    inputs: prompt,
    parameters: config,
  };

  // Convert input message to a list of GAPIC instances for model input
  const instances = [helpers.toValue(input)];

  // TODO(developer): Uncomment apiEndpoint and predictionServiceClient before running the sample.
  // const apiEndpoint = `${endpointRegion}-aiplatform.googleapis.com`;

  // Create a client
  // predictionServiceClient = new PredictionServiceClient({apiEndpoint});

  // Call the Gemma2 endpoint
  const gemma2Endpoint = `projects/${projectId}/locations/${endpointRegion}/endpoints/${endpointId}`;

  const [response] = await predictionServiceClient.predict({
    endpoint: gemma2Endpoint,
    instances,
  });

  const predictions = response.predictions;
  const text = predictions[0].stringValue;

  console.log('Predictions:', text);
  return text;
}

module.exports = gemma2PredictGpu;

// TODO(developer): Uncomment below lines before running the sample.
// gemma2PredictGpu(...process.argv.slice(2)).catch(err => {
//   console.error(err.message);
//   process.exitCode = 1;
// });

Java

לפני שמנסים את הדוגמה הזו, צריך לפעול לפי Javaההוראות להגדרה במאמר מדריך למתחילים של Agent Platform באמצעות ספריות לקוח.

כדי לבצע אימות ב-Agent Platform, צריך להגדיר את Application Default Credentials. מידע נוסף זמין במאמר הגדרת אימות לסביבת פיתוח מקומית.


import com.google.cloud.aiplatform.v1.EndpointName;
import com.google.cloud.aiplatform.v1.PredictResponse;
import com.google.cloud.aiplatform.v1.PredictionServiceClient;
import com.google.cloud.aiplatform.v1.PredictionServiceSettings;
import com.google.gson.Gson;
import com.google.protobuf.InvalidProtocolBufferException;
import com.google.protobuf.Value;
import com.google.protobuf.util.JsonFormat;
import java.io.IOException;
import java.util.ArrayList;
import java.util.HashMap;
import java.util.List;
import java.util.Map;

public class Gemma2PredictGpu {

  private final PredictionServiceClient predictionServiceClient;

  // Constructor to inject the PredictionServiceClient
  public Gemma2PredictGpu(PredictionServiceClient predictionServiceClient) {
    this.predictionServiceClient = predictionServiceClient;
  }

  public static void main(String[] args) throws IOException {
    // TODO(developer): Replace these variables before running the sample.
    String projectId = "YOUR_PROJECT_ID";
    String endpointRegion = "us-east4";
    String endpointId = "YOUR_ENDPOINT_ID";

    PredictionServiceSettings predictionServiceSettings =
        PredictionServiceSettings.newBuilder()
            .setEndpoint(String.format("%s-aiplatform.googleapis.com:443", endpointRegion))
            .build();
    PredictionServiceClient predictionServiceClient =
        PredictionServiceClient.create(predictionServiceSettings);
    Gemma2PredictGpu creator = new Gemma2PredictGpu(predictionServiceClient);

    creator.gemma2PredictGpu(projectId, endpointRegion, endpointId);
  }

  // Demonstrates how to run inference on a Gemma2 model
  // deployed to a Vertex AI endpoint with GPU accelerators.
  public String gemma2PredictGpu(String projectId, String region,
               String endpointId) throws IOException {
    Map<String, Object> paramsMap = new HashMap<>();
    paramsMap.put("temperature", 0.9);
    paramsMap.put("maxOutputTokens", 1024);
    paramsMap.put("topP", 1.0);
    paramsMap.put("topK", 1);
    Value parameters = mapToValue(paramsMap);

    // Prompt used in the prediction
    String instance = "{ \"inputs\": \"Why is the sky blue?\"}";
    Value.Builder instanceValue = Value.newBuilder();
    JsonFormat.parser().merge(instance, instanceValue);
    // Encapsulate the prompt in a correct format for GPUs
    // Example format: [{'inputs': 'Why is the sky blue?', 'parameters': {'temperature': 0.8}}]
    List<Value> instances = new ArrayList<>();
    instances.add(instanceValue.build());

    EndpointName endpointName = EndpointName.of(projectId, region, endpointId);

    PredictResponse predictResponse = this.predictionServiceClient
        .predict(endpointName, instances, parameters);
    String textResponse = predictResponse.getPredictions(0).getStringValue();
    System.out.println(textResponse);
    return textResponse;
  }

  private static Value mapToValue(Map<String, Object> map) throws InvalidProtocolBufferException {
    Gson gson = new Gson();
    String json = gson.toJson(map);
    Value.Builder builder = Value.newBuilder();
    JsonFormat.parser().merge(json, builder);
    return builder.build();
  }
}

Go

לפני שמנסים את הדוגמה הזו, צריך לפעול לפי Goההוראות להגדרה במאמר מדריך למתחילים של Agent Platform באמצעות ספריות לקוח.

כדי לבצע אימות ב-Agent Platform, צריך להגדיר את Application Default Credentials. מידע נוסף זמין במאמר הגדרת אימות לסביבת פיתוח מקומית.

import (
	"context"
	"fmt"
	"io"

	"cloud.google.com/go/aiplatform/apiv1/aiplatformpb"

	"google.golang.org/protobuf/types/known/structpb"
)

// predictGPU demonstrates how to run interference on a Gemma2 model deployed to a Vertex AI endpoint with GPU accelerators.
func predictGPU(w io.Writer, client PredictionsClient, projectID, location, endpointID string) error {
	ctx := context.Background()

	// Note: client can be initialized in the following way:
	// apiEndpoint := fmt.Sprintf("%s-aiplatform.googleapis.com:443", location)
	// client, err := aiplatform.NewPredictionClient(ctx, option.WithEndpoint(apiEndpoint))
	// if err != nil {
	// 	return fmt.Errorf("unable to create prediction client: %v", err)
	// }
	// defer client.Close()

	gemma2Endpoint := fmt.Sprintf("projects/%s/locations/%s/endpoints/%s", projectID, location, endpointID)
	prompt := "Why is the sky blue?"
	parameters := map[string]interface{}{
		"temperature":     0.9,
		"maxOutputTokens": 1024,
		"topP":            1.0,
		"topK":            1,
	}

	// Encapsulate the prompt in a correct format for TPUs.
	// Pay attention that prompt should be set in "inputs" field.
	// Example format: [{'inputs': 'Why is the sky blue?', 'parameters': {'temperature': 0.9}}]
	promptValue, err := structpb.NewValue(map[string]interface{}{
		"inputs":     prompt,
		"parameters": parameters,
	})
	if err != nil {
		fmt.Fprintf(w, "unable to convert prompt to Value: %v", err)
		return err
	}

	req := &aiplatformpb.PredictRequest{
		Endpoint:  gemma2Endpoint,
		Instances: []*structpb.Value{promptValue},
	}

	resp, err := client.Predict(ctx, req)
	if err != nil {
		return err
	}

	prediction := resp.GetPredictions()
	value := prediction[0].GetStringValue()
	fmt.Fprintf(w, "%v", value)

	return nil
}

הסרת המשאבים

כדי להימנע מחיובים בחשבון Google Cloud בגלל השימוש במשאבים שנעשה במסגרת המדריך הזה, אפשר למחוק את הפרויקט שמכיל את המשאבים, או להשאיר את הפרויקט ולמחוק את המשאבים בנפרד.

מחיקת הפרויקט

  1. במסוף Google Cloud , נכנסים לדף Manage resources.

    כניסה לדף Manage resources

  2. ברשימת הפרויקטים, בוחרים את הפרויקט שרוצים למחוק ולוחצים על Delete.
  3. כדי למחוק את הפרויקט, כותבים את מזהה הפרויקט בתיבת הדו-שיח ולוחצים על Shut down.

מחיקת משאבים בודדים

אם אתם רוצים לשמור את הפרויקט, אתם צריכים למחוק את המשאבים שבהם השתמשתם במדריך הזה:

  • ביטול הפריסה של המודל ומחיקת נקודת הקצה
  • מחיקת המודל ממרשם המודלים

ביטול הפריסה של המודל ומחיקת נקודת הקצה

כדי לבטל את הפריסה של מודל ולמחוק את נקודת הקצה, אפשר להשתמש באחת מהשיטות הבאות.

המסוף

  1. במסוף Google Cloud , לוחצים על חיזוי מיידי ואז על נקודות קצה.

    כניסה לדף Endpoints

  2. בתפריט הנפתח Region, בוחרים את האזור שבו פרסתם את נקודת הקצה.

  3. לוחצים על שם נקודת הקצה כדי לפתוח את דף הפרטים. לדוגמה: gemma2-2b-it-mg-one-click-deploy.

  4. בשורה של מודל Gemma 2 (Version 1), לוחצים על Actions ואז על Undeploy model from endpoint.

  5. בתיבת הדו-שיח Undeploy model from endpoint (ביטול הפריסה של המודל מנקודת הקצה), לוחצים על Undeploy (ביטול הפריסה).

  6. לוחצים על הלחצן הקודם כדי לחזור לדף Endpoints.

    כניסה לדף Endpoints

  7. בסוף השורה gemma2-2b-it-mg-one-click-deploy, לוחצים על Actions (פעולות) ואז בוחרים באפשרות Delete endpoint (מחיקת נקודת קצה).

  8. בהודעת האישור, לוחצים על אישור.

gcloud

כדי לבטל את הפריסה של המודל ולמחוק את נקודת הקצה באמצעות Google Cloud CLI, פועלים לפי השלבים הבאים.

בפקודות האלה, מחליפים את:

  • PROJECT_ID בשם הפרויקט
  • LOCATION_ID עם האזור שבו פרסתם את המודל ואת נקודת הקצה
  • ENDPOINT_ID עם מזהה נקודת הקצה
  • DEPLOYED_MODEL_NAME עם השם המוצג של המודל
  • DEPLOYED_MODEL_ID עם מזהה המודל
  1. מריצים את הפקודה gcloud ai endpoints list כדי לקבל את מזהה נקודת הקצה. הפקודה הזו מציגה רשימה של מזהי נקודות הקצה של כל נקודות הקצה בפרויקט. חשוב לשים לב למזהה של נקודת הקצה שבה נעשה שימוש במדריך הזה.

    gcloud ai endpoints list \
        --project=PROJECT_ID \
        --region=LOCATION_ID
    

    הפלט אמור להיראות כך: בפלט, המזהה נקרא ENDPOINT_ID.

    Using endpoint [https://us-central1-aiplatform.googleapis.com/]
    ENDPOINT_ID: 1234567891234567891
    DISPLAY_NAME: gemma2-2b-it-mg-one-click-deploy
    
  2. מריצים את הפקודה gcloud ai models describe כדי לקבל את מזהה המודל. רושמים בצד את המזהה של המודל שפרסתם במדריך הזה.

    gcloud ai models describe DEPLOYED_MODEL_NAME \
        --project=PROJECT_ID \
        --region=LOCATION_ID
    

    הפלט המקוצר נראה כך. בפלט, המזהה נקרא deployedModelId.

    Using endpoint [https://us-central1-aiplatform.googleapis.com/]
    artifactUri: [URI removed]
    baseModelSource:
      modelGardenSource:
        publicModelName: publishers/google/models/gemma2
    ...
    deployedModels:
    - deployedModelId: '1234567891234567891'
      endpoint: projects/12345678912/locations/us-central1/endpoints/12345678912345
    displayName: gemma2-2b-it-12345678912345
    etag: [ETag removed]
    modelSourceInfo:
      sourceType: MODEL_GARDEN
    name: projects/123456789123/locations/us-central1/models/gemma2-2b-it-12345678912345
    ...
    
  3. מבטלים את הפריסה של המודל מנקודת הקצה. תצטרכו את מזהה נקודת הקצה ואת מזהה המודל מהפקודות הקודמות.

    gcloud ai endpoints undeploy-model ENDPOINT_ID \
        --project=PROJECT_ID \
        --region=LOCATION_ID \
        --deployed-model-id=DEPLOYED_MODEL_ID
    

    הפקודה הזו לא יוצרת פלט.

  4. מריצים את הפקודה gcloud ai endpoints delete כדי למחוק את נקודת הקצה.

    gcloud ai endpoints delete ENDPOINT_ID \
        --project=PROJECT_ID \
        --region=LOCATION_ID
    

    כשמופיעה בקשה, מקלידים y כדי לאשר. הפקודה הזו לא יוצרת פלט.

מחיקת המודל

המסוף

  1. עוברים לדף מרשם המודלים בקטע Agent Platform במסוף Google Cloud .

    כניסה לדף Model Registry

  2. בתפריט הנפתח Region, בוחרים את האזור שבו פרסתם את המודל.

  3. בסוף השורה gemma2-2b-it-1234567891234, לוחצים על פעולות.

  4. בוחרים באפשרות מחיקת המודל.

    כשמוחקים את המודל, כל הגרסאות וההערכות שמשויכות אליו נמחקות מהפרויקט Google Cloud .

  5. בהודעת האישור, לוחצים על מחיקה.

gcloud

כדי למחוק את המודל באמצעות Google Cloud CLI, צריך להזין את שם התצוגה והאזור של המודל בפקודה gcloud ai models delete.

gcloud ai models delete DEPLOYED_MODEL_NAME \
    --project=PROJECT_ID \
    --region=LOCATION_ID

מחליפים את DEPLOYED_MODEL_NAME בשם לתצוגה של המודל. מחליפים את PROJECT_ID בשם הפרויקט. מחליפים את LOCATION_ID באזור שבו פרסתם את המודל.

המאמרים הבאים