במדריך הזה מוסבר איך לפרוס מודל שפה גדול (LLM) ב-Google Kubernetes Engine (GKE) באמצעות GKE Inference Gateway. ההדרכה כוללת שלבים להגדרת אשכול, לפריסת מודל, להגדרת GKE Inference Gateway ולטיפול בבקשות של LLM.
המדריך הזה מיועד למהנדסי למידת מכונה (ML), לאדמינים ולמפעילים של פלטפורמות ולמומחים בתחום הנתונים וה-AI שרוצים לפרוס ולנהל אפליקציות LLM ב-GKE באמצעות GKE Inference Gateway.
לפני שקוראים את הדף הזה, כדאי להכיר את המושגים הבאים:
- מידע על הסקת מסקנות ממודלים ב-GKE
- הרצת מסקנות לפי שיטות מומלצות באמצעות מתכונים למתחילים של GKE Inference
- מצב Autopilot ומצב רגיל
- יחידות GPU ב-GKE
רקע
בקטע הזה מתוארות הטכנולוגיות העיקריות שבהן נעשה שימוש במדריך הזה. מידע נוסף על מושגים ומינוחים שקשורים להצגת מודלים, ועל האופן שבו יכולות ה-AI הגנרטיבי של GKE יכולות לשפר את הביצועים של הצגת המודלים ולתמוך בהם, זמין במאמר מידע על הסקת מסקנות ממודלים ב-GKE.
vLLM
vLLM הוא מודל קוד פתוח לאירוח מודלים גדולים של שפה (LLM) שעבר אופטימיזציה גבוהה, ומגדיל את קצב העברת הנתונים (throughput) של האירוח ביחידות GPU. התכונות העיקריות כוללות:
- הטמעה אופטימלית של טרנספורמציה באמצעות PagedAttention
- הוספת תכונה של אצווה מתמשכת שמשפרת את התפוקה הכוללת של הצגת המודעות
- מקביליות טנסורים והגשה מבוזרת בכמה מעבדי GPU
מידע נוסף מופיע במאמרי העזרה בנושא vLLM.
GKE Inference Gateway
GKE Inference Gateway משפר את היכולות של GKE להפעלת מודלים גדולים של שפה (LLM). הוא מבצע אופטימיזציה של עומסי עבודה של הסקת מסקנות באמצעות תכונות כמו:
- איזון עומסים שעבר אופטימיזציה להסקת מסקנות על סמך מדדי עומס.
- תמיכה בהצגת מודלים של מתאמי LoRA בצפיפות גבוהה של עומסי עבודה מרובים.
- ניתוב מודע-מודל לפעולות פשוטות.
מידע נוסף זמין במאמר מידע על GKE Inference Gateway.
מטרות
לפני שמתחילים
- נכנסים לחשבון Google Cloud . אם אתם משתמשים חדשים ב- Google Cloud, צרו חשבון כדי שתוכלו להעריך את הביצועים של המוצרים שלנו בתרחישים מהעולם האמיתי. לקוחות חדשים מקבלים בחינם גם קרדיט בשווי 300$ להרצה, לבדיקה ולפריסה של עומסי העבודה.
-
In the Google Cloud console, on the project selector page, select or create a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Verify that billing is enabled for your Google Cloud project.
Enable the required API.
Roles required to enable APIs
To enable APIs, you need the Service Usage Admin IAM role (
roles/serviceusage.serviceUsageAdmin), which contains theserviceusage.services.enablepermission. Learn how to grant roles.-
In the Google Cloud console, on the project selector page, select or create a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Verify that billing is enabled for your Google Cloud project.
Enable the required API.
Roles required to enable APIs
To enable APIs, you need the Service Usage Admin IAM role (
roles/serviceusage.serviceUsageAdmin), which contains theserviceusage.services.enablepermission. Learn how to grant roles.-
צריך לוודא שיש לכם בפרויקט את התפקיד או התפקידים הבאים: roles/container.admin, roles/iam.serviceAccountAdmin
בדיקת התפקידים
-
נכנסים לדף IAM במסוף Google Cloud .
כניסה לדף IAM - בוחרים את הפרויקט.
-
בעמודה Principal (חשבון המשתמש), מוצאים את כל השורות שבהן מופיע השם שלכם או של קבוצה שאתם נכללים בה. כדי לברר באילו קבוצות אתם נכללים, פנו לאדמין.
- בודקים את העמודה Role בכל השורות שבהן מצוין או מופיע השם שלכם, כדי לראות אם רשימת התפקידים כוללת את התפקידים הנדרשים.
מתן התפקידים
-
נכנסים לדף IAM במסוף Google Cloud .
כניסה לדף IAM - בוחרים את הפרויקט.
- לוחצים על Grant access.
-
בשדה New principals, מזינים את מזהה המשתמש. בדרך כלל מזהה המשתמש הוא כתובת האימייל של חשבון Google.
- לוחצים על Select a role ומחפשים את התפקיד.
- כדי לתת עוד תפקידים, לוחצים על Add another role ומוסיפים אותם.
- לוחצים על Save.
-
- יוצרים חשבון ב-Hugging Face, אם עדיין אין לכם חשבון.
- מוודאים שיש בפרויקט מכסה מספקת לשימוש במעבדי GPU מדגם H100. מידע נוסף זמין במאמרים בנושא מכסת GPU בתוכנית ומכסות הקצאה.
גישה למודל
כדי לפרוס את מודל Llama3.1 ב-GKE, צריך לחתום על הסכם ההסכמה לרישיון וליצור טוקן גישה של Hugging Face.
חתימה על הסכם הסכמה לרישיון
כדי להשתמש במודל Llama3.1, צריך לחתום על הסכם ההסכמה. פועלים לפי ההוראות הבאות:
- עוברים לדף ההסכמה ומאשרים את ההסכמה לשימוש בחשבון Hugging Face.
- מאשרים את התנאים של המודל.
יצירת אסימון גישה
כדי לגשת למודל דרך Hugging Face, צריך טוקן של Hugging Face.
אם עדיין אין לכם אסימון, אתם יכולים ליצור אסימון חדש באמצעות השלבים הבאים:
- לוחצים על הפרופיל שלך > הגדרות > טוקנים של גישה.
- בוחרים באפשרות New Token (טוקן חדש).
- מציינים שם לבחירתכם ותפקיד ברמת
Readלפחות. - לוחצים על יצירת אסימון.
- מעתיקים את הטוקן שנוצר ללוח.
הכנת הסביבה
במדריך הזה משתמשים ב-Cloud Shell כדי לנהל משאבים שמתארחים ב-Google Cloud. ב-Cloud Shell מותקן מראש התוכנה שצריך למדריך הזה, כולל kubectl ו-
ה-CLI של gcloud.
כדי להגדיר את הסביבה באמצעות Cloud Shell:
ב Google Cloud מסוף, מפעילים סשן של Cloud Shell על ידי לחיצה על Activate Cloud Shell בGoogle Cloud מסוף.
תופעל סשן בחלונית התחתונה של Google Cloud המסוף.
מגדירים את משתני הסביבה שמוגדרים כברירת מחדל:
gcloud config set project PROJECT_ID gcloud config set billing/quota_project PROJECT_ID export PROJECT_ID=$(gcloud config get project) export REGION=REGION export CLUSTER_NAME=CLUSTER_NAME export HF_TOKEN=HF_TOKENמחליפים את הערכים הבאים:
-
PROJECT_ID: מזהה הפרויקט ב- Google Cloud. -
REGION: אזור שתומך בסוג המאיץ שרוצים להשתמש בו, לדוגמה,us-central1ל-GPU מסוג H100. -
CLUSTER_NAME: השם של האשכול. -
HF_TOKEN: אסימון Hugging Face שיצרתם קודם.
-
יצירה והגדרה של Google Cloud משאבים
יצירת אשכול GKE ומאגר צמתים
הצגת מודלים גדולים של שפה (LLM) ב-GPU באשכול GKE Autopilot או Standard. מומלץ להשתמש באשכול Autopilot כדי ליהנות מחוויית Kubernetes מנוהלת באופן מלא. כדי לבחור את מצב הפעולה של GKE שהכי מתאים לעומסי העבודה שלכם, אפשר לעיין במאמר בחירת מצב פעולה של GKE.
טייס אוטומטי
ב-Cloud Shell, מריצים את הפקודה הבאה:
gcloud container clusters create-auto CLUSTER_NAME \
--project=PROJECT_ID \
--location=CONTROL_PLANE_LOCATION \
--release-channel=rapid
מחליפים את הערכים הבאים:
-
PROJECT_ID: מזהה הפרויקט ב- Google Cloud. -
CONTROL_PLANE_LOCATION: האזור של Compute Engine במישור הבקרה של האשכול. מציינים אזור שתומך בסוג המאיץ שרוצים להשתמש בו, לדוגמה,us-central1ל-GPU מסוג H100. -
CLUSTER_NAME: השם של האשכול.
GKE יוצר אשכול Autopilot עם צמתים של מעבד ו-GPU לפי הבקשה של עומסי העבודה שנפרסו.
רגילה
ב-Cloud Shell, מריצים את הפקודה הבאה כדי ליצור אשכול Standard:
gcloud container clusters create CLUSTER_NAME \ --project=PROJECT_ID \ --location=CONTROL_PLANE_LOCATION \ --workload-pool=PROJECT_ID.svc.id.goog \ --release-channel=rapid \ --num-nodes=1 \ --enable-managed-prometheus \ --monitoring=SYSTEM,DCGM \ --gateway-api=standardמחליפים את הערכים הבאים:
-
PROJECT_ID: מזהה הפרויקט ב- Google Cloud. -
CONTROL_PLANE_LOCATION: האזור של Compute Engine במישור הבקרה של האשכול. מציינים אזור שתומך בסוג המאיץ שרוצים להשתמש בו, לדוגמה,us-central1ל-GPU מסוג H100. -
CLUSTER_NAME: השם של האשכול.
יצירת האשכול עשויה להימשך כמה דקות.
-
כדי ליצור מאגר צמתים עם גודל הדיסק המתאים להרצת מודל
Llama-3.1-8B-Instruct, מריצים את הפקודה הבאה:gcloud container node-pools create gpupool \ --accelerator type=nvidia-h100-80gb,count=2,gpu-driver-version=latest \ --project=PROJECT_ID \ --location=CONTROL_PLANE_LOCATION \ --node-locations=CONTROL_PLANE_LOCATION-a \ --cluster=CLUSTER_NAME \ --machine-type=a3-highgpu-2g \ --num-nodes=1 \ --disk-type="pd-balanced"GKE יוצר מאגר צמתים יחיד שמכיל GPU מסוג H100.
הגדרת הרשאה לגירוד מדדים
כדי להגדיר הרשאה לגירוד מדדים, צריך ליצור את הסוד inference-gateway-sa-metrics-reader-secret.
שומרים את קובץ המניפסט הבא בשם
metrics-auth.yaml:החלת המניפסט:
kubectl apply -f metrics-auth.yaml
יצירת סוד של Kubernetes לפרטי הכניסה של Hugging Face
ב-Cloud Shell, מבצעים את הפעולות הבאות:
מגדירים את
kubectlכך שיוכל לתקשר עם האשכול:gcloud container clusters get-credentials CLUSTER_NAME \ --location=REGIONמחליפים את הערכים הבאים:
-
REGION: אזור שתומך בסוג המאיץ שרוצים להשתמש בו, לדוגמה,us-central1עבור GPU מסוג L4. -
CLUSTER_NAME: השם של האשכול.
-
יוצרים סוד של Kubernetes שמכיל את הטוקן של Hugging Face:
kubectl create secret generic hf-secret \ --from-literal=hf_api_token=${HF_TOKEN} \ --dry-run=client -o yaml | kubectl apply -f -מחליפים את
HF_TOKENבטוקן של Hugging Face שיצרתם קודם.
התקנה של CRD InferenceObjective ו-InferencePool
בקטע הזה, תתקינו את ההגדרות הנדרשות של Custom Resource Definitions (CRD) עבור GKE Inference Gateway.
משתמשים ב-CRD כדי להרחיב את Kubernetes API. כך תוכלו להגדיר סוגים חדשים של משאבים. כדי להשתמש ב-GKE Inference Gateway, צריך להתקין את ה-CRD InferencePool ואת ה-CRD InferenceObjective באשכול GKE באמצעות הפקודה הבאה:
kubectl apply -f https://github.com/kubernetes-sigs/gateway-api-inference-extension/releases/download/v1.0.0/manifests.yaml
פריסת שרת המודל
בדוגמה הזו נפרס מודל Llama3.1 באמצעות שרת מודלים של vLLM. הפריסה מסומנת בתווית app:vllm-llama3.1-8b-instruct. בפריסה הזו נעשה שימוש גם בשני מתאמי LoRA בשם food-review ו-cad-fabricator מ-Hugging Face. אפשר לעדכן את הפריסה הזו עם שרת מודלים, קונטיינר מודלים, יציאת שרת ושם פריסה משלכם. אפשר להגדיר מתאמי LoRA בפריסה, או לפרוס את מודל הבסיס.
כדי לפרוס במאיץ מסוג
nvidia-h100-80gb, שומרים את המניפסט הבא בתורvllm-llama3.1-8b-instruct.yaml. המניפסט הזה מגדיר פריסה של Kubernetes עם המודל ושרת המודל:apiVersion: apps/v1 kind: Deployment metadata: name: vllm-llama3.1-8b-instruct spec: replicas: 3 selector: matchLabels: app: vllm-llama3.1-8b-instruct template: metadata: labels: app: vllm-llama3.1-8b-instruct spec: containers: - name: vllm image: "vllm/vllm-openai:v0.13.0" imagePullPolicy: Always command: ["python3", "-m", "vllm.entrypoints.openai.api_server"] args: - "--model" - "meta-llama/Llama-3.1-8B-Instruct" - "--tensor-parallel-size" - "1" - "--port" - "8000" - "--enable-lora" - "--max-loras" - "2" - "--max-cpu-loras" - "12" - "--compilation-config" - '{"cudagraph_specialize_lora": "False"}' # As workaround for https://github.com/vllm-project/vllm/issues/29049 env: - name: PORT value: "8000" - name: HUGGING_FACE_HUB_TOKEN valueFrom: secretKeyRef: name: hf-secret key: hf_api_token - name: VLLM_ALLOW_RUNTIME_LORA_UPDATING value: "true" ports: - containerPort: 8000 name: http protocol: TCP lifecycle: preStop: # vLLM stops accepting connections when it receives SIGTERM, so we need to sleep # to give upstream gateways a chance to take us out of rotation. The time we wait # is dependent on the time it takes for all upstreams to completely remove us from # rotation. Older or simpler load balancers might take upwards of 30s, but we expect # our deployment to run behind a modern gateway like Envoy which is designed to # probe for readiness aggressively. sleep: # Upstream gateway probers for health should be set on a low period, such as 5s, # and the shorter we can tighten that bound the faster that we release # accelerators during controlled shutdowns. However, we should expect variance, # as load balancers may have internal delays, and we don't want to drop requests # normally, so we're often aiming to set this value to a p99 propagation latency # of readiness -> load balancer taking backend out of rotation, not the average. # # This value is generally stable and must often be experimentally determined on # for a given load balancer and health check period. We set the value here to # the highest value we observe on a supported load balancer, and we recommend # tuning this value down and verifying no requests are dropped. # # If this value is updated, be sure to update terminationGracePeriodSeconds. # seconds: 30 # # IMPORTANT: preStop.sleep is beta as of Kubernetes 1.30 - for older versions # replace with this exec action. #exec: # command: # - /usr/bin/sleep # - 30 livenessProbe: httpGet: path: /health port: http scheme: HTTP # vLLM's health check is simple, so we can more aggressively probe it. Liveness # check endpoints should always be suitable for aggressive probing. periodSeconds: 1 successThreshold: 1 # vLLM has a very simple health implementation, which means that any failure is # likely significant. However, any liveness triggered restart requires the very # large core model to be reloaded, and so we should bias towards ensuring the # server is definitely unhealthy vs immediately restarting. Use 5 attempts as # evidence of a serious problem. failureThreshold: 5 timeoutSeconds: 1 readinessProbe: httpGet: path: /health port: http scheme: HTTP # vLLM's health check is simple, so we can more aggressively probe it. Readiness # check endpoints should always be suitable for aggressive probing, but may be # slightly more expensive than readiness probes. periodSeconds: 1 successThreshold: 1 # vLLM has a very simple health implementation, which means that any failure is # likely significant, failureThreshold: 1 timeoutSeconds: 1 # We set a startup probe so that we don't begin directing traffic or checking # liveness to this instance until the model is loaded. startupProbe: # Failure threshold is when we believe startup will not happen at all, and is set # to the maximum possible time we believe loading a model will take. In our # default configuration we are downloading a model from HuggingFace, which may # take a long time, then the model must load into the accelerator. We choose # 10 minutes as a reasonable maximum startup time before giving up and attempting # to restart the pod. # # IMPORTANT: If the core model takes more than 10 minutes to load, pods will crash # loop forever. Be sure to set this appropriately. failureThreshold: 600 # Set delay to start low so that if the base model changes to something smaller # or an optimization is deployed, we don't wait unnecessarily. initialDelaySeconds: 2 # As a startup probe, this stops running and so we can more aggressively probe # even a moderately complex startup - this is a very important workload. periodSeconds: 1 httpGet: # vLLM does not start the OpenAI server (and hence make /health available) # until models are loaded. This may not be true for all model servers. path: /health port: http scheme: HTTP resources: limits: nvidia.com/gpu: 1 requests: nvidia.com/gpu: 1 volumeMounts: - mountPath: /data name: data - mountPath: /dev/shm name: shm - name: adapters mountPath: "/adapters" # This is the second container in the Pod, a sidecar to the vLLM container. # It watches the ConfigMap and downloads LoRA adapters. - name: lora-adapter-syncer image: us-central1-docker.pkg.dev/k8s-staging-images/gateway-api-inference-extension/lora-syncer:main imagePullPolicy: Always env: - name: DYNAMIC_LORA_ROLLOUT_CONFIG value: "/config/configmap.yaml" volumeMounts: # DO NOT USE subPath, dynamic configmap updates don't work on subPaths - name: config-volume mountPath: /config restartPolicy: Always # vLLM allows VLLM_PORT to be specified as an environment variable, but a user might # create a 'vllm' service in their namespace. That auto-injects VLLM_PORT in docker # compatible form as `tcp://<IP>:<PORT>` instead of the numeric value vLLM accepts # causing CrashLoopBackoff. Set service environment injection off by default. enableServiceLinks: false # Generally, the termination grace period needs to last longer than the slowest request # we expect to serve plus any extra time spent waiting for load balancers to take the # model server out of rotation. # # An easy starting point is the p99 or max request latency measured for your workload, # although LLM request latencies vary significantly if clients send longer inputs or # trigger longer outputs. Since steady state p99 will be higher than the latency # to drain a server, you may wish to slightly this value either experimentally or # via the calculation below. # # For most models you can derive an upper bound for the maximum drain latency as # follows: # # 1. Identify the maximum context length the model was trained on, or the maximum # allowed length of output tokens configured on vLLM (llama2-7b was trained to # 4k context length, while llama3-8b was trained to 128k). # 2. Output tokens are the more compute intensive to calculate and the accelerator # will have a maximum concurrency (batch size) - the time per output token at # maximum batch with no prompt tokens being processed is the slowest an output # token can be generated (for this model it would be about 10ms TPOT at a max # batch size around 50, or 100 tokens/sec) # 3. Calculate the worst case request duration if a request starts immediately # before the server stops accepting new connections - generally when it receives # SIGTERM (for this model that is about 4096 / 100 ~ 40s) # 4. If there are any requests generating prompt tokens that will delay when those # output tokens start, and prompt token generation is roughly 6x faster than # compute-bound output token generation, so add 40% to the time from above (40s + # 16s = 56s) # # Thus we think it will take us at worst about 56s to complete the longest possible # request the model is likely to receive at maximum concurrency (highest latency) # once requests stop being sent. # # NOTE: This number will be lower than steady state p99 latency since we stop receiving # new requests which require continuous prompt token computation. # NOTE: The max timeout for backend connections from gateway to model servers should # be configured based on steady state p99 latency, not drain p99 latency # # 5. Add the time the pod takes in its preStop hook to allow the load balancers to # stop sending us new requests (56s + 30s = 86s). # # Because the termination grace period controls when the Kubelet forcibly terminates a # stuck or hung process (a possibility due to a GPU crash), there is operational safety # in keeping the value roughly proportional to the time to finish serving. There is also # value in adding a bit of extra time to deal with unexpectedly long workloads. # # 6. Add a 50% safety buffer to this time (86s * 1.5 ≈ 130s). # # One additional source of drain latency is that some workloads may run close to # saturation and have queued requests on each server. Since traffic in excess of the # max sustainable QPS will result in timeouts as the queues grow, we assume that failure # to drain in time due to excess queues at the time of shutdown is an expected failure # mode of server overload. If your workload occasionally experiences high queue depths # due to periodic traffic, consider increasing the safety margin above to account for # time to drain queued requests. terminationGracePeriodSeconds: 130 nodeSelector: cloud.google.com/gke-accelerator: "nvidia-h100-80gb" volumes: - name: data emptyDir: {} - name: shm emptyDir: medium: Memory - name: adapters emptyDir: {} - name: config-volume configMap: name: vllm-llama3.1-8b-adapters --- apiVersion: v1 kind: ConfigMap metadata: name: vllm-llama3.1-8b-adapters data: configmap.yaml: | vLLMLoRAConfig: name: vllm-llama3.1-8b-instruct port: 8000 defaultBaseModel: meta-llama/Llama-3.1-8B-Instruct ensureExist: models: - id: food-review source: Kawon/llama3.1-food-finetune_v14_r8 - id: cad-fabricator source: redcathode/fabricator --- kind: HealthCheckPolicy apiVersion: networking.gke.io/v1 metadata: name: health-check-policy namespace: default spec: targetRef: group: "inference.networking.k8s.io" kind: InferencePool name: vllm-llama3.1-8b-instruct default: config: type: HTTP httpHealthCheck: requestPath: /health port: 8000מחילים את המניפסט על האשכול:
kubectl apply -f vllm-llama3.1-8b-instruct.yaml
יצירת משאב InferencePool
המשאב המותאם אישית של InferencePool Kubernetes מגדיר קבוצה של Pod עם LLM בסיסי משותף והגדרת מחשוב.
המשאב המותאם אישית InferencePool כולל את השדות העיקריים הבאים:
-
selector: מציין אילו פודים שייכים למאגר הזה. התוויות בסלקטור הזה צריכות להיות זהות לתוויות שמוחלות על ה-Pods של שרת המודל. -
targetPort: הגדרה של היציאות שבהן שרת המודל משתמש בתוך קבוצות ה-Pod.
המשאב InferencePool מאפשר ל-GKE Inference Gateway לנתב תנועה ל-Pods של שרת המודל.
כדי ליצור InferencePool באמצעות Helm:
helm install vllm-llama3.1-8b-instruct \
--set inferencePool.modelServers.matchLabels.app=vllm-llama3.1-8b-instruct \
--set provider.name=gke \
--set healthCheckPolicy.create=false \
--version v1.0.0 \
oci://registry.k8s.io/gateway-api-inference-extension/charts/inferencepool
משנים את השדה הבא בהתאם לפריסה:
-
inferencePool.modelServers.matchLabels.app: המפתח של התווית שמשמשת לבחירת ה-Pods של שרת המודל.
הפקודה הזו יוצרת אובייקט InferencePool שמייצג באופן לוגי את הפריסה של שרת המודלים, ומפנה לשירותי נקודות הקצה של המודלים בתוך ה-Pods שנבחרו על ידי Selector.
יצירת InferenceObjective משאב עם רמת קריטיות להצגה
המשאב המותאם אישית InferenceObjective מגדיר את פרמטרים ההצגה של מודל, כולל העדיפות שלו. צריך ליצור משאבי InferenceObjective כדי להגדיר אילו מודלים מוצגים ב-InferencePool. המשאבים האלה יכולים להפנות למודלים בסיסיים או למתאמי LoRA שנתמכים על ידי שרתי המודלים ב-InferencePool.
בשדה metadata.name מציינים את שם המודל, בשדה priority מגדירים את רמת הקריטיות של המודל ובשדה poolRef מקשרים אל InferencePool שבו המודל מוגש.
כדי ליצור InferenceObjective, פועלים לפי השלבים הבאים:
שומרים את קובץ המניפסט לדוגמה הבא בשם
inferenceobjective.yaml:apiVersion: inference.networking.x-k8s.io/v1alpha2 kind: InferenceObjective metadata: name: MODEL_NAME spec: priority: VALUE poolRef: name: INFERENCE_POOL_NAME kind: "InferencePool"מחליפים את מה שכתוב בשדות הבאים:
-
MODEL_NAME: השם של מודל הבסיס או של מתאם LoRA. לדוגמה,food-review. -
VALUE: העדיפות של יעד ההסקה. זהו מספר שלם, וככל שהערך גבוה יותר הבקשה קריטית יותר. לדוגמה,10. -
INFERENCE_POOL_NAME: השם שלInferencePoolשיצרתם בשלב הקודם. לדוגמה:vllm-llama3.1-8b-instruct.
-
מחילים את קובץ המניפסט לדוגמה על האשכול:
kubectl apply -f inferenceobjective.yaml
בדוגמה הבאה נוצרים שני אובייקטים מסוג InferenceObjective. ההגדרה הראשונה קובעת את מודל ה-LoRA food-review ב-vllm-llama3.1-8b-instruct
InferencePool עם עדיפות של 10. ההגדרה השנייה קובעת שהמודעה llama3-base-model תוצג עם עדיפות גבוהה יותר של 20.
apiVersion: inference.networking.k8s.io/v1alpha1
kind: InferenceObjective
metadata:
name: food-review
spec:
priority: 10
poolRef:
name: vllm-llama3.1-8b-instruct
kind: "InferencePool"
---
apiVersion: inference.networking.k8s.io/v1alpha1
kind: InferenceObjective
metadata:
name: llama3-base-model
spec:
priority: 20
poolRef:
name: vllm-llama3.1-8b-instruct
kind: "InferencePool"
יצירת השער
משאב השער פועל כנקודת הכניסה לתנועה חיצונית לאשכול Kubernetes. הוא מגדיר את המאזינים שמקבלים חיבורים נכנסים.
GKE Inference Gateway תומך ב-Gateway Class gke-l7-rilb וב-gke-l7-regional-external-managed. מידע נוסף מופיע במאמר בנושא Gateway Classes במסמכי GKE.
כדי ליצור שער:
שומרים את קובץ המניפסט לדוגמה הבא בשם
gateway.yaml:apiVersion: gateway.networking.k8s.io/v1 kind: Gateway metadata: name: GATEWAY_NAME spec: gatewayClassName: gke-l7-regional-external-managed listeners: - protocol: HTTP # Or HTTPS for production port: 80 # Or 443 for HTTPS name: httpמחליפים את
GATEWAY_NAMEבשם ייחודי למשאב Gateway. לדוגמה,inference-gateway.מחילים את המניפסט על האשכול:
kubectl apply -f gateway.yaml
יצירת משאב HTTPRoute
בקטע הזה יוצרים משאב HTTPRoute כדי להגדיר איך שער הכניסה מנתב בקשות HTTP נכנסות אל InferencePool.
משאב HTTPRoute מגדיר איך שער GKE מנתב בקשות HTTP נכנסות לשירותי קצה עורפי, כלומר ל-InferencePool. הוא מציין כללי התאמה (לדוגמה, כותרות או נתיבים) ואת ה-Backend שאליו התנועה צריכה להיות מועברת.
כדי ליצור HTTPRoute, מבצעים את השלבים הבאים:
שומרים את קובץ המניפסט לדוגמה הבא בשם
httproute.yaml:apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: name: HTTPROUTE_NAME spec: parentRefs: - name: GATEWAY_NAME rules: - matches: - path: type: PathPrefix value: PATH_PREFIX backendRefs: - name: INFERENCE_POOL_NAME group: inference.networking.k8s.io kind: InferencePoolמחליפים את מה שכתוב בשדות הבאים:
-
HTTPROUTE_NAME: שם ייחודי של משאבHTTPRoute. לדוגמה,my-route. -
GATEWAY_NAME: השם של משאבGatewayשיצרתם. לדוגמה,inference-gateway. -
PATH_PREFIX: תחילית הנתיב שבה משתמשים כדי להתאים בקשות נכנסות. לדוגמה,/כדי להתאים לכולם. -
INFERENCE_POOL_NAME: השם של משאבInferencePoolשאליו רוצים להפנות את התנועה. לדוגמה,vllm-llama3.1-8b-instruct.
-
מחילים את המניפסט על האשכול:
kubectl apply -f httproute.yaml
שליחת בקשת הסקה
אחרי שמגדירים את GKE Inference Gateway, אפשר לשלוח בקשות להסקת מסקנות למודל שהופעל.
כדי לשלוח בקשות להסקת מסקנות:
- מאחזרים את נקודת הקצה של השער.
- יוצרים בקשת JSON בפורמט תקין.
- משתמשים ב-
curlכדי לשלוח את הבקשה לנקודת הקצה/v1/completions.
כך תוכלו ליצור טקסט על סמך ההנחיה שהזנתם והפרמטרים שציינתם.
כדי לקבל את נקודת הקצה של שער, מריצים את הפקודה הבאה:
IP=$(kubectl get gateway/GATEWAY_NAME -o jsonpath='{.status.addresses[0].value}') PORT=80מחליפים את
GATEWAY_NAMEבשם של משאב שער.כדי לשלוח בקשה לנקודת הקצה
/v1/completionsבאמצעותcurl, מריצים את הפקודה הבאה:curl -i -X POST http://${IP}:${PORT}/v1/completions \ -H "Content-Type: application/json" \ -d '{ "model": "MODEL_NAME", "prompt": "PROMPT_TEXT", "max_tokens": MAX_TOKENS, "temperature": "TEMPERATURE" }'מחליפים את מה שכתוב בשדות הבאים:
-
MODEL_NAME: השם של המודל או של מתאם LoRA שרוצים להשתמש בהם. -
PROMPT_TEXT: הנחיית הקלט למודל. -
MAX_TOKENS: המספר המקסימלי של הטוקנים שיופיעו בתגובה. -
TEMPERATURE: שולט באקראיות של הפלט. כדי לקבל פלט דטרמיניסטי, משתמשים בערך0. כדי לקבל פלט יצירתי יותר, משתמשים במספר גבוה יותר.
-
חשוב לזכור:
- גוף הבקשה: גוף הבקשה יכול לכלול פרמטרים נוספים כמו
stopו-top_p. רשימה מלאה של האפשרויות זמינה במפרט של OpenAI API. - טיפול בשגיאות: צריך להטמיע טיפול מתאים בשגיאות בקוד הלקוח כדי לטפל בשגיאות אפשריות בתגובה. לדוגמה, צריך לבדוק את קוד הסטטוס של HTTP בתגובה.
curlקוד סטטוס שאינו 200 מציין בדרך כלל שגיאה. - אימות והרשאה: בפריסות בסביבת ייצור, מאבטחים את נקודת הקצה ל-API באמצעות מנגנוני אימות והרשאה. הבקשות צריכות לכלול את הכותרות המתאימות (לדוגמה,
Authorization).
הגדרת ניראות (observability) עבור Inference Gateway
GKE Inference Gateway מספק יכולת מעקב אחרי התקינות, הביצועים וההתנהגות של עומסי העבודה של ההסקות. כך תוכלו לזהות ולפתור בעיות, לבצע אופטימיזציה של ניצול המשאבים ולהבטיח את המהימנות של האפליקציות. אפשר לראות את מדדי יכולת הצפייה האלה ב-Cloud Monitoring באמצעות Metrics Explorer.
כדי להגדיר ניראות (observability) ל-GKE Inference Gateway, קראו את המאמר בנושא הגדרת ניראות.
מחיקת המשאבים שנפרסו
כדי להימנע מחיובים בחשבון Google Cloud על המשאבים שיצרתם באמצעות המדריך הזה, מריצים את הפקודה הבאה:
gcloud container clusters delete CLUSTER_NAME \
--location=CONTROL_PLANE_LOCATION
מחליפים את הערכים הבאים:
-
CONTROL_PLANE_LOCATION: האזור של Compute Engine במישור הבקרה של האשכול. -
CLUSTER_NAME: השם של האשכול.