בדף הזה מוסבר איך לצרף היפר-דיסקים למכונות וירטואליות (VM) באשכול Managed Service for Apache Spark. אפשר להגדיר את הדיסקים בנפרד לקבוצות של צמתי מאסטר, צמתי עובד ראשיים וצמתי עובד משניים. הדיסקים האלה מצורפים לצמתי האשכול בנוסף לדיסק האתחול ולכונני SSD מקומיים שמצורפים לצמתי האשכול.
לפני שמתחילים
- נכנסים לחשבון Google Cloud . אם אתם משתמשים חדשים ב- Google Cloud, צרו חשבון כדי שתוכלו להעריך את הביצועים של המוצרים שלנו בתרחישים מהעולם האמיתי. לקוחות חדשים מקבלים בחינם גם קרדיט בשווי 300$ להרצה, לבדיקה ולפריסה של עומסי העבודה.
-
In the Google Cloud console, on the project selector page, select or create a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Verify that you have the permissions required to complete this guide.
-
Verify that billing is enabled for your Google Cloud project.
Enable the Managed Service for Apache Spark API.
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.-
In the Google Cloud console, on the project selector page, select or create a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Verify that you have the permissions required to complete this guide.
-
Verify that billing is enabled for your Google Cloud project.
Enable the Managed Service for Apache Spark API.
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.
התפקידים הנדרשים
כדי להריץ את הדוגמאות בדף הזה, צריך תפקידים מסוימים ב-IAM. יכול להיות שהתפקידים האלה כבר הוקצו, בהתאם למדיניות הארגון. כדי לבדוק את התפקידים שהוקצו, ראו האם צריך להקצות תפקידים?.
כדי לקרוא הסבר על מתן תפקידים, קראו איך מנהלים את הגישה ברמת הפרויקט, התיקייה והארגון.
תפקידי משתמשים
כדי לקבל את ההרשאות שדרושות ליצירת אשכול של Managed Service for Apache Spark, צריך לבקש מהאדמין להקצות לכם את תפקידי ה-IAM הבאים:
- עריכת פרויקטים ב-Dataproc (
roles/dataproc.editor) בפרויקט - משתמש בחשבון שירות (
roles/iam.serviceAccountUser) בחשבון השירות של Compute Engine שמוגדר כברירת מחדל
תפקיד בחשבון שירות
כדי לוודא שלחשבון השירות שמוגדר כברירת מחדל ב-Compute Engine יש את ההרשאות שנדרשות ליצירת אשכול של Managed Service for Apache Spark, צריך לבקש מהאדמין להקצות לחשבון השירות שמוגדר כברירת מחדל ב-Compute Engine את תפקיד ה-IAM Dataproc Worker (roles/dataproc.worker) בפרויקט.
מאפייני הדיסק
לדיסקים המצורפים יש את המאפיינים הבאים:
- מחזור חיים: מחזור החיים של דיסק מצורף זהה למחזור החיים של מכונת ה-VM שהוא מצורף אליה. השירות המנוהל ל-Apache Spark יוצר את הדיסק כשהוא יוצר את ה-VM ומוחק את הדיסק כשהוא מוחק את ה-VM.
- אי אפשרות לשינוי: אי אפשר לעדכן את המאפיינים של דיסק מצורף, כמו גודל, IOPS או קצב העברת נתונים, אחרי שיוצרים את האשכול.
- טעינה ושימוש: Managed Service for Apache Spark טוען דיסקים בנתיב
/mnt/N, כאשרNהוא מספר שלם חיובי (לדוגמה,/mnt/1, /mnt/2). נתוני HDFS ונתוני scratch, כמו פלט של פעולות מיון נתונים (shuffle), משתמשים בדיסקים המצורפים במקום בדיסק האחסון המתמיד של האתחול.
הגדרת הדיסק
כשמצרפים דיסקים לצמתי אשכול של Managed Service for Apache Spark, אפשר לציין את פרמטרים ההגדרות הבאים של הדיסקים:
Disk type (סוג הדיסק) – חובה: סוג הדיסק לצירוף למכונות וירטואליות. יש תמיכה בדיסקים היפרדינמיים הבאים:
hyperdisk-balancedhyperdisk-extremehyperdisk-mlhyperdisk-throughput
אי אפשר לצרף דיסקים מסוג
hyperdisk balanced high availabilityודיסקים קבועים לצמתי אשכול.גודל – אופציונלי: גודל הדיסק. הערך חייב להיות מספר שלם שאחריו מופיע
GBלגיגה-בייט אוTBלטרה-בייט. לדוגמה, 10GBמחבר דיסק בנפח 10 גיגה-בייט. מידע נוסף זמין במאמר בנושא מגבלות הגודל של Hyperdisk.IOPs – אופציונלי: מציין את IOPS להקצאה עבור הדיסק המצורף. הפרמטר הזה מגדיר את המגבלה של פעולות קלט/פלט בדיסק לשנייה. מידע נוסף זמין במאמר בנושא רמות ביצועים שמוגדרות כברירת מחדל.
Throughput – אופציונלי: מציין את התפוקה שצריך להקצות לדיסק המצורף. הפרמטר הזה מגדיר את מגבלת התפוקה ב-
MiBלשנייה. מידע נוסף זמין במאמר בנושא רמות ביצועים שמוגדרות כברירת מחדל.
צירוף דיסקים לאשכול
אפשר לצרף דיסקים כשיוצרים אשכול של Managed Service for Apache Spark באמצעות ה-CLI של gcloud או Dataproc API כדי לציין את הגדרות הדיסקים.
CLI של gcloud
כדי לצרף דיסקים כשיוצרים אשכול, משתמשים בדגל
--master-attached-disks,--worker-attached-disksאו--secondary-worker-attached-disksעם הפקודהgcloud dataproc clusters create.כל דגל מקבל רשימה של הגדרות דיסק שמופרדות באמצעות נקודה-ופסיק. כל הגדרת דיסק היא רשימה מופרדת בפסיקים של זוגות של מפתח-ערך עבור
type,size,iopsו-throughput(ראו הגדרת דיסק).
דוגמה: הפקודה הבאה יוצרת אשכול ומצרפת שני דיסקים מסוג Hyperdisk לכל צומת עובד ראשי.
gcloud dataproc clusters create CLUSTER_NAME \
--region=REGION \
--worker-attached-disks='type=hyperdisk-balanced,size=100GB,iops=5000,throughput=200;type=hyperdisk-throughput,size=9000GB'
API
כדי לצרף דיסקים, צריך לכלול מערך
attachedDiskConfigsבאובייקטdiskConfigשל קבוצת המכונותmasterConfig,workerConfigאוsecondaryWorkerConfig.מציינים את ההגדרה בגוף של בקשת
clusters.createAPI.
דוגמה: קטע ה-JSON הבא מציג מערך attachedDiskConfigs שמצרף שני דיסקים היפרדיים:
[
{
"diskType": "HYPERDISK_BALANCED",
"diskSizeGb": 100,
"provisionedIops": 5000,
"provisionedThroughput": 200
},
{
"diskType": "HYPERDISK_THROUGHPUT",
"diskSizeGb": 9000
}
]