בדף הזה מוסבר איך ליצור מופע של סביבת העבודה של פלטפורמת הסוכנים של Gemini Enterprise עם תמיכה ב-Spark ב-Managed Service for Apache Spark. בנוסף, בדף הזה מוסבר על היתרונות של התוסף Managed Service for Apache Spark JupyterLab, ומוצג סקירה כללית על אופן השימוש בתוסף עם Managed Service for Apache Spark ו-Managed Service for Apache Spark ב-Compute Engine.
סקירה כללית של התוסף Managed Service for Apache Spark JupyterLab
במופעים של Agent Platform Workbench, התוסף Managed Service for Apache Spark JupyterLab מותקן מראש, החל מגרסה M113 ואילך.
התוסף Managed Service for Apache Spark JupyterLab מאפשר להריץ משימות של מחברות Apache Spark בשתי דרכים: אשכולות של Managed Service for Apache Spark ו-Managed Service for Apache Spark.
- אשכולות של Managed Service for Apache Spark כוללים מגוון רחב של תכונות עם שליטה בתשתית ש-Spark פועל עליה. אתם בוחרים את הגודל וההגדרה של אשכול Spark, וכך יכולים להתאים אישית את הסביבה ולשלוט בה. הגישה הזו מתאימה לעומסי עבודה מורכבים, לעבודות שפועלות לאורך זמן ולניהול פרטני של משאבים.
- Managed Service for Apache Spark מבטל את הצורך לדאוג לתשתית. אתם שולחים את משימות Spark, ו-Google מטפלת בהקצאה, בהתאמה ובאופטימיזציה של המשאבים מאחורי הקלעים. הגישה הזו ללא שרתים (serverless) מציעה אפשרות חסכונית למשימות של מדעי הנתונים ולעומסי עבודה של למידת מכונה.
בשתי האפשרויות אפשר להשתמש ב-Spark לעיבוד ולניתוח נתונים. הבחירה בין אשכולות של Managed Service for Apache Spark לבין Managed Service for Apache Spark תלויה בדרישות הספציפיות של עומס העבודה, ברמת השליטה הנדרשת ובדפוסי השימוש במשאבים.
היתרונות של שימוש ב-Managed Service for Apache Spark לעומסי עבודה של מדעי הנתונים ולמידת מכונה כוללים:
- אין צורך בניהול אשכולות: לא צריך לדאוג לגבי הקצאה, הגדרה או ניהול של אשכולות Spark. כך תוכלו לחסוך זמן ומשאבים.
- התאמה אוטומטית לעומס: שירות Managed Service for Apache Spark משנה את הגודל באופן אוטומטי בהתאם לעומס העבודה, כך שאתם משלמים רק על המשאבים שבהם אתם משתמשים.
- ביצועים גבוהים: השירות Managed Service for Apache Spark עבר אופטימיזציה לביצועים, והוא מנצל את התשתית של Google Cloud.
- שילוב עם טכנולוגיות אחרות Google Cloud : Managed Service for Apache Spark משתלב עם מוצרים אחרים של Google Cloud , כמו BigQuery ו-Knowledge Catalog.
מידע נוסף זמין במאמרי העזרה בנושא Managed Service for Apache Spark.
לפני שמתחילים
- נכנסים לחשבון Google Cloud . אם אתם משתמשים חדשים ב- Google Cloud, צרו חשבון כדי שתוכלו להעריך את הביצועים של המוצרים שלנו בתרחישים מהעולם האמיתי. לקוחות חדשים מקבלים בחינם גם קרדיט בשווי 300$ להרצה, לבדיקה ולפריסה של עומסי העבודה.
-
In the Google Cloud console, on the project selector page, select or create a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
Enable the Cloud Resource Manager, Managed Service for Apache Spark, and Notebooks APIs.
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.-
In the Google Cloud console, on the project selector page, select or create a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
Enable the Cloud Resource Manager, Managed Service for Apache Spark, and Notebooks APIs.
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.
התפקידים הנדרשים
כדי לוודא שלחשבון השירות יש את ההרשאות הנדרשות להרצת קובץ מחברת ב-Managed Service for Apache Spark או ב-Managed Service for Apache Spark, צריך לבקש מהאדמין להקצות לחשבון השירות את תפקידי ה-IAM הבאים:
- Dataproc Worker (
roles/dataproc.worker) בפרויקט - עורך Dataproc (
roles/dataproc.editor) באשכול להרשאהdataproc.clusters.use
להסבר על מתן תפקידים, ראו איך מנהלים את הגישה ברמת הפרויקט, התיקייה והארגון.
התפקידים המוגדרים מראש האלה כוללים את ההרשאות שנדרשות להרצת קובץ מחברת ב-Managed Service for Apache Spark או ב-Managed Service for Apache Spark. כדי לראות בדיוק אילו הרשאות נדרשות, אפשר להרחיב את הקטע ההרשאות הנדרשות:
ההרשאות הנדרשות
כדי להריץ קובץ מחברת באשכול Managed Service for Apache Spark או באשכול Managed Service for Apache Spark, נדרשות ההרשאות הבאות:
-
dataproc.agents.create -
dataproc.agents.delete -
dataproc.agents.get -
dataproc.agents.update -
dataproc.tasks.lease -
dataproc.tasks.listInvalidatedLeases -
dataproc.tasks.reportStatus -
dataproc.clusters.use
יכול להיות שהאדמין יוכל גם להעניק לחשבון השירות את ההרשאות האלה באמצעות תפקידים בהתאמה אישית או תפקידים מוגדרים מראש אחרים.
יצירת מופע עם Managed Service for Apache Spark מופעל
כדי ליצור מופע של Agent Platform Workbench עם Managed Service for Apache Spark מופעל, מבצעים את הפעולות הבאות:
נכנסים לדף Instances במסוף Google Cloud .
לוחצים על יצירת פריט חדש.
בתיבת הדו-שיח New instance, לוחצים על Advanced options.
בתיבת הדו-שיח Create instance, בקטע Details, מוודאים שהאפשרות Enable Dataproc Serverless Interactive Sessions מסומנת.
מוודאים שסוג סביבת העבודה מוגדר כמופע.
בקטע Environment, מוודאים שמשתמשים בגרסה העדכנית ביותר או בגרסה שמספרה
M113ומעלה.לוחצים על יצירה.
הכלי Agent Platform Workbench יוצר מופע ומתחיל אותו באופן אוטומטי. כשהמופע מוכן לשימוש, מופעל קישור Open JupyterLab ב-Agent Platform Workbench.
פתיחת JupyterLab
לצד שם המכונה, לוחצים על Open JupyterLab.
הכרטיסייה Launcher של JupyterLab תיפתח בדפדפן. כברירת מחדל, הוא מכיל קטעים בנושא מחברות (Notebooks) של Managed Service for Apache Spark ומשימות וסשנים של Managed Service for Apache Spark. אם יש אשכולות שמוכנים ל-Jupyter בפרויקט ובאזור שנבחרו, יופיע קטע בשם Managed Service for Apache Spark Cluster Notebooks.
שימוש בתוסף עם Managed Service for Apache Spark
תבניות של זמן ריצה של Managed Service for Apache Spark שנמצאות באותו אזור ובאותו פרויקט כמו מופע Agent Platform Workbench שלכם מופיעות בקטע Managed Service for Apache Spark Notebooks בכרטיסייה Launcher ב-JupyterLab.
כדי ליצור תבנית של סביבת ריצה, ראו יצירת תבנית של סביבת ריצה של Managed Service for Apache Spark.
כדי לפתוח מחברת Spark חדשה ללא שרת, לוחצים על תבנית של זמן ריצה. הפעלת ליבת Spark מרחוק נמשכת כדקה. אחרי שהליבה מתחילה, אפשר להתחיל לכתוב קוד.
שימוש בתוסף עם Managed Service for Apache Spark ב-Compute Engine
אם יצרתם אשכול Jupyter של Managed Service for Apache Spark ב-Compute Engine, בכרטיסייה Launcher יופיע הקטע Managed Service for Apache Spark Cluster Notebooks.
מוצגים ארבעה כרטיסים לכל אשכול של Managed Service for Apache Spark שמוכן ל-Jupyter, שיש לכם גישה אליו באזור ובפרויקט הזה.
כדי לשנות את האזור והפרויקט:
בוחרים באפשרות הגדרות > הגדרות של Cloud Managed Service for Apache Spark.
בכרטיסייה Setup Config, בקטע Project Info, משנים את מזהה פרויקט ואת Region, ואז לוחצים על Save.
השינויים האלה ייכנסו לתוקף רק אחרי שתפעילו מחדש את JupyterLab.
כדי להפעיל מחדש את JupyterLab, בוחרים באפשרות File > Shut Down, ואז לוחצים על Open JupyterLab בדף Agent Platform Workbench instances.
כדי ליצור נוטבוק חדש, לוחצים על כרטיס. אחרי שהליבה המרוחקת באשכול Managed Service for Apache Spark מתחילה לפעול, אפשר להתחיל לכתוב את הקוד ואז להריץ אותו באשכול.
ניהול של Managed Service for Apache Spark במופע באמצעות ה-CLI של gcloud ו-API
בקטע הזה מתוארות דרכים לניהול Managed Service for Apache Spark במופע של Agent Platform Workbench.
שינוי האזור של אשכול Managed Service for Apache Spark
הליבות שמוגדרות כברירת מחדל במופע של Agent Platform Workbench, כמו Python ו-TensorFlow, הן ליבות מקומיות שפועלות במכונה הווירטואלית של המופע. במופע של סביבת עבודה של פלטפורמת סוכנים עם הפעלה של Spark ב-Managed Service for Apache Spark, מחברת ה-notebook פועלת באשכול של Managed Service for Apache Spark דרך ליבה מרוחקת. הליבה המרוחקת פועלת בשירות מחוץ למכונה הווירטואלית של המופע, מה שמאפשר לכם לגשת לכל אשכול Managed Service for Apache Spark באותו פרויקט.
כברירת מחדל, Agent Platform Workbench משתמש באשכולות של Managed Service for Apache Spark באותו אזור כמו המופע שלכם, אבל אתם יכולים לשנות את האזור של Managed Service for Apache Spark כל עוד Component Gateway ורכיב Jupyter האופציונלי מופעלים באשכול של Managed Service for Apache Spark.
בדיקת הגישה
התוסף Managed Service for Apache Spark JupyterLab מופעל כברירת מחדל במופעים של Agent Platform Workbench. כדי לבדוק את הגישה אל Managed Service for Apache Spark, אפשר לבדוק את הגישה לליבות המרוחקות של המופע על ידי שליחת בקשת curl הבאה לדומיין kernels.googleusercontent.com:
curl --verbose -H "Authorization: Bearer $(gcloud auth print-access-token)" https://PROJECT_ID-dot-REGION.kernels.googleusercontent.com/api/kernelspecs | jq .
אם פקודת curl נכשלת, צריך לוודא:
הערכים ב-DNS מוגדרים בצורה נכונה.
יש אשכול שזמין באותו פרויקט (או שתצטרכו ליצור אשכול אם הוא לא קיים).
האפשרות Component Gateway ורכיב Jupyter האופציונלי מופעלות באשכול.
השבתה של Managed Service for Apache Spark
מופעים של Agent Platform Workbench נוצרים עם Managed Service for Apache Spark מופעל כברירת מחדל. כדי ליצור מופע של Agent Platform Workbench עם Managed Service for Apache Spark מושבת, צריך להגדיר את המפתח disable-mixer
metadata לערך true.
gcloud workbench instances create INSTANCE_NAME --metadata=disable-mixer=true
הפעלת Managed Service for Apache Spark
אפשר להפעיל את Managed Service for Apache Spark במופע של Agent Platform Workbench שהופסק, על ידי עדכון ערך המטא-נתונים.
gcloud workbench instances update INSTANCE_NAME --metadata=disable-mixer=false
ניהול של Managed Service for Apache Spark באמצעות Terraform
השירות Managed Service for Apache Spark for Agent Platform Workbench instances ב-Terraform מנוהל באמצעות המפתח disable-mixer בשדה המטא-נתונים.
מפעילים את Managed Service for Apache Spark על ידי הגדרת המפתח disable-mixer
metadata לערך false. כדי להשבית את Managed Service for Apache Spark, מגדירים את מפתח המטא-נתונים disable-mixer לערך true.
כדי ללמוד איך להחיל הגדרות ב-Terraform או להסיר אותן, ראו פקודות בסיסיות ב-Terraform.
פתרון בעיות
כדי לאבחן ולפתור בעיות שקשורות ליצירת מופע עם הפעלת Spark ב-Managed Service for Apache Spark, אפשר לעיין במאמר בנושא פתרון בעיות ב-Agent Platform Workbench.
המאמרים הבאים
מידע נוסף על התוסף Managed Service for Apache Spark JupyterLab זמין במאמר שימוש בתוסף JupyterLab לפיתוח עומסי עבודה של Spark ללא שרתים.
מידע נוסף על Managed Service for Apache Spark זמין במסמכי התיעוד של Managed Service for Apache Spark
איך מריצים עומסי עבודה של Managed Service for Apache Spark בלי להקצות ולנהל אשכולות
מידע נוסף על שימוש ב-Spark עם Google Cloud מוצרים ושירותים זמין במאמר Spark ב- Google Cloud.
מעיינים בתבניות הזמינות של Managed Service for Apache Spark ב-GitHub.
מידע על Serverless Spark זמין ב-
serverless-spark-workshopב-GitHub.קוראים את התיעוד של Apache Spark.