יצירה של טבלת Iceberg ב-Lakehouse והרצת שאילתות עליה באמצעות מסוף Google Cloud
במדריך למתחילים הזה, תלמדו איך להשתמש במסוף Google Cloud כדי לנהל ולשתף טבלאות Apache Iceberg באמצעות Lakehouse ללא גבולות בין Google Cloud לבין מנועי קוד פתוח, על ידי אחסון מטא-נתונים של הטבלה, כולל סכימות, תמונות מצב ומיקומי אחסון, בקטלוג של זמן הריצה של Lakehouse.
כדי להשלים את המדריך למתחילים הזה, מבצעים את השלבים הבאים במסוףGoogle Cloud :
- יצירת קטגוריה של Cloud Storage: יוצרים קטגוריה ב-Cloud Storage כדי לאחסן את הנתונים של טבלת Iceberg ואת קובצי המטא-נתונים.
- יצירת קטלוג: יצירת קטלוג עם כמה קטגוריות אחסון בסביבת זמן הריצה של Lakehouse, שמגובה על ידי קטגוריית האחסון עם הפעלת מכירת אישורים.
- יצירת מרחב שמות וטבלת Iceberg: משתמשים בדף Lakehouse במסוף Google Cloud כדי ליצור מרחב שמות וטבלת Iceberg עם שפת מניפולציה של נתונים (DML) של BigQuery מופעלת.
- שינוי נתונים והרצת שאילתות בטבלה ב-BigQuery: משתמשים בהצהרות DML של BigQuery (
INSERT,UPDATEו-DELETE) כדי לשנות שורות בטבלת Iceberg ולהריץ שאילתות על התוצאות באמצעות תחביר P.C.N.T (Project.Catalog.Namespace.Table) בן 4 החלקים, ללא צורך ב-ETL או ברישום ידני של הטבלה.
לפני שמתחילים
- נכנסים לחשבון Google Cloud . אנחנו ממליצים למשתמשים חדשים ב- Google Cloud ליצור חשבון כדי שיוכלו להעריך את הביצועים של המוצרים שלנו בתרחישים מהעולם האמיתי. לקוחות חדשים מקבלים בחינם גם קרדיט בשווי 300$ להרצה, לבדיקה ולפריסה של עומסי העבודה.
-
In the Google Cloud console, on the project selector page, select or create a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Verify that billing is enabled for your Google Cloud project.
Enable the BigLake, Cloud Storage, and BigQuery APIs, if any are not already enabled.
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.-
Make sure that you have the following role or roles on the project: BigLake Admin (
roles/biglake.admin), Storage Admin (roles/storage.admin), and BigQuery Job User (roles/bigquery.jobUser)Check for the roles
-
In the Google Cloud console, go to the IAM page.
Go to IAM - Select the project.
-
In the Principal column, find all rows that identify you or a group that you're included in. To learn which groups you're included in, contact your administrator.
- For all rows that specify or include you, check the Role column to see whether the list of roles includes the required roles.
Grant the roles
-
In the Google Cloud console, go to the IAM page.
Go to IAM - Select the project.
- Click Grant access.
-
In the New principals field, enter your user identifier. This is typically the email address for a Google Account.
- Click Select a role, then search for the role.
- To grant additional roles, click Add another role and add each additional role.
- Click Save.
-
-
In the Google Cloud console, on the project selector page, select or create a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Verify that billing is enabled for your Google Cloud project.
Enable the BigLake, Cloud Storage, and BigQuery APIs, if any are not already enabled.
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.-
Make sure that you have the following role or roles on the project: BigLake Admin (
roles/biglake.admin), Storage Admin (roles/storage.admin), and BigQuery Job User (roles/bigquery.jobUser)Check for the roles
-
In the Google Cloud console, go to the IAM page.
Go to IAM - Select the project.
-
In the Principal column, find all rows that identify you or a group that you're included in. To learn which groups you're included in, contact your administrator.
- For all rows that specify or include you, check the Role column to see whether the list of roles includes the required roles.
Grant the roles
-
In the Google Cloud console, go to the IAM page.
Go to IAM - Select the project.
- Click Grant access.
-
In the New principals field, enter your user identifier. This is typically the email address for a Google Account.
- Click Select a role, then search for the role.
- To grant additional roles, click Add another role and add each additional role.
- Click Save.
-
יצירת קטגוריה של Cloud Storage
יוצרים קטגוריה של Cloud Storage במסוף Google Cloud כדי לאחסן את הנתונים של טבלת Iceberg ואת קובצי המטא-נתונים:
במסוף Google Cloud , נכנסים לדף Buckets של Cloud Storage.
לוחצים על יצירה.
בקטע Get started (תחילת העבודה), מזינים שם ייחודי גלובלית לדלי (לדוגמה,
lakehouse-quickstart-UNIQUE_IDאוPROJECT_ID-lakehouse) ולוחצים על Continue (המשך).בקטע Choose where to store your data, משאירים את Location type על Multi-region (us (multiple regions in United States))), ואז לוחצים על Create.
אם מופיעה תיבת הדו-שיח Public access will be prevented, לוחצים על Confirm.
יצירת קטלוג בקטלוג של Lakehouse runtime
יוצרים קטלוג עם כמה מאגרי מידע בקטלוג של זמן הריצה של Lakehouse עבור טבלאות Apache Iceberg. קטלוג של כמה קטגוריות מאפשר לכם לתת לקטלוג שם שלא תלוי בשם של אף קטגוריה, ולקשר כמה קטגוריות של Cloud Storage לקטלוג אחד. כדי לאבטח את הגישה לקטגוריות האלה, צריך להפעיל את מצב הנפקת פרטי כניסה כדי שהקטלוג יוכל להנפיק באופן אוטומטי פרטי כניסה זמניים לאחסון ישירות למנועי הלקוח.
נכנסים לדף Lakehouse במסוף Google Cloud .
לוחצים על Create catalog (יצירת קטלוג) ובוחרים באפשרות Lakehouse runtime catalog (קטלוג של זמן ריצה של Lakehouse).
בקטע פרטי הקטלוג, מגדירים את ההגדרות הבאות:
- סוג קטלוג: בוחרים באפשרות Iceberg Rest Catalog.
- אפשרויות של קטלוג Lakehouse: בוחרים באפשרות קטלוג של כמה דליים.
- נתיב ברירת המחדל של Cloud Storage לקטלוג: לוחצים על עיון, בוחרים את הקטגוריה שיצרתם ולוחצים על בחירה.
- מזהה קטלוג: מזינים
quickstart_catalog. - מיקום ראשי: בוחרים באפשרות מספר אזורים ואז בוחרים באפשרות ארה"ב (מספר אזורים בארצות הברית).
לוחצים על המשך, ואז בקטע נתיבי נתונים לוחצים על המשך.
בקטע שיטת אימות, בוחרים באפשרות מצב מכירת אישורים.
באמצעות מכירת אישורים, הקטלוג מנפיק באופן מאובטח אסימוני אחסון זמניים בהיקף הטבלה למנועי לקוח ול-BigQuery, כך שמנועים חיצוניים לא צריכים הרשאות IAM ישירות בדלי שלכם.
לוחצים על יצירה.
הקטלוג נוצר והדף פרטי הקטלוג נפתח.
בקטע שיטת אימות, לוחצים על הגדרת הרשאות של דלי, ואז בתיבת הדו-שיח לוחצים על אישור.
בשלב הזה מעניקים לחשבון השירות של הקטלוג את ההרשאות הנדרשות בקטגוריית Cloud Storage כדי להנפיק פרטי כניסה זמניים.
יצירת מרחב שמות וטבלת Iceberg
עכשיו, אחרי שיצרתם קטלוג, אתם יכולים להשתמש בדף Lakehouse במסוףGoogle Cloud כדי ליצור מרחב שמות וטבלת Iceberg.
יצירת מרחב שמות
בדף פרטי הקטלוג של
quickstart_catalog, לוחצים על יצירת מרחב שמות.בשדה Namespace name, מזינים
quickstart_namespace.משאירים את המיקום מוגדר לנתיב ברירת המחדל של Cloud Storage שמאוכלס אוטומטית בשדה.
לוחצים על יצירה.
יצירת טבלת Iceberg
בדף פרטי הקטלוג, לוחצים על
quickstart_namespace.נפתח הדף Namespace Details (פרטי מרחב שמות).
לוחצים על Create Table.
בחלונית יצירת טבלה, מגדירים את ההגדרות הבאות:
- פורמט הטבלה: מוודאים שהאפשרות Iceberg נבחרה.
- Table name: מזינים
quickstart_table. - מיקום: משאירים את נתיב ברירת המחדל של Cloud Storage.
בקטע סכימה, לוחצים פעמיים על הוספת שדה כדי להוסיף שתי עמודות לטבלה:
- בשדה הראשון, מזינים
idבשדה שם השדה ובוחרים באפשרות INTEGER בתפריט סוג. - בשדה השני, מזינים
nameבשדה שם השדה ובוחרים באפשרות STRING בתפריט סוג.
- בשדה הראשון, מזינים
בקטע מאפיינים, מאתרים את המאפיין המוגדר מראש
gcp.biglake.bigquery-dml.enabledומשנים את הערך שלו מ-falseל-true. משאירים את הערך שלgcp.biglake.table-management.enabledכ-false.הגדרה של
gcp.biglake.bigquery-dml.enabledל-trueמאפשרת לשנות נתונים בטבלת Iceberg באמצעות משפטי DML של BigQuery, כמוINSERT, UPDATE, DELETEו-MERGE. מידע נוסף מופיע במאמר בנושא הגדרת אפשרויות לטבלה.לוחצים על יצירה.
טבלת Iceberg החדשה (
quickstart_table) מופיעה בדף Namespace details, וקטלוג זמן הריצה של Lakehouse כותב את קובץ המטא-נתונים הראשוני של Iceberg לקטגוריה שלכם ב-Cloud Storage.
שינוי נתונים והפעלת שאילתה בטבלה ב-BigQuery
אחרי שיוצרים את quickstart_table ומפעילים את BigQuery DML, אפשר להוסיף, לעדכן, למחוק ולשאול שורות ישירות ב-BigQuery באמצעות התחביר P.C.N.T (Project.Catalog.Namespace.Table) שמורכב מ-4 חלקים. בכל הצהרה, מחליפים את PROJECT_ID במזהה הפרויקטGoogle Cloud :
במסוף Google Cloud , עוברים לדף BigQuery.
בעורך השאילתות, לוחצים על שאילתת SQL.
מוסיפים שלוש שורות של נתונים לדוגמה:
INSERT INTO `PROJECT_ID.quickstart_catalog.quickstart_namespace.quickstart_table` (id, name) VALUES (1, 'one'), (2, 'two'), (3, 'three');
לוחצים על Run. כשמשפט
INSERTמסתיים, BigQuery כותב את קובצי הנתונים של Parquet לקטגוריה של Cloud Storage ומבצע קומיט של תמונת מצב חדשה של Iceberg לקטלוג של זמן הריצה של Lakehouse.כדי לשנות שורה בטבלה:
UPDATE `PROJECT_ID.quickstart_catalog.quickstart_namespace.quickstart_table` SET name = 'updated' WHERE id = 1;
לוחצים על Run.
כדי למחוק שורה מהטבלה:
DELETE FROM `PROJECT_ID.quickstart_catalog.quickstart_namespace.quickstart_table` WHERE id = 3;
לוחצים על Run.
מריצים שאילתה על הטבלה כדי לוודא שהשינויים בוצעו:
SELECT * FROM `PROJECT_ID.quickstart_catalog.quickstart_namespace.quickstart_table` ORDER BY id;
לוחצים על Run. בחלונית תוצאות השאילתה מוצגות שתי השורות שנותרו, כולל הערך המעודכן של
id = 1:+----+---------+ | id | name | +----+---------+ | 1 | updated | | 2 | two | +----+---------+
מכיוון שקטלוג זמן הריצה של Lakehouse מנהל את המטא-נתונים של Iceberg
והפעלת ההקצאה של פרטי הכניסה מופעלת, אתם יכולים גם לקרוא מ- או לכתוב ל-
quickstart_table באמצעות כל מנוע קוד פתוח שתואם ל-Iceberg, כמו
Apache Spark, Trino או Apache Flink, בלי להעניק להם
גישת IAM ישירה לדלי.
הסרת המשאבים
כדי להימנע מחיובים מיותרים בחשבון Google Cloud , מוחקים את המשאבים שיצרתם במדריך למתחילים הזה. מחיקה של הטבלה, מרחב השמות והקטלוג מסירה את רישום המטא-נתונים מהקטלוג של זמן הריצה של Lakehouse, ומחיקה של הקטגוריה מסירה את נתוני ה-Parquet הבסיסיים ואת קובצי המטא-נתונים של Iceberg שמאוחסנים ב-Cloud Storage:
נכנסים לדף Lakehouse במסוף Google Cloud .
מוחקים את הטבלה מהקטלוג:
- לוחצים על
quickstart_catalogואז עלquickstart_namespace. - בטבלה Namespace details, בשורה של
quickstart_table, לוחצים על More > Delete. - מזינים
DELETEכדי לאשר ולוחצים על מחיקה.
- לוחצים על
מוחקים את מרחב השמות מהקטלוג:
- חוזרים לדף פרטי הקטלוג של
quickstart_catalog. - בשורה של
quickstart_namespace, לוחצים על סמל האפשרויות הנוספות של מרחב השמות > מחיקה. - מזינים
DELETEכדי לאשר ולוחצים על מחיקה.
- חוזרים לדף פרטי הקטלוג של
מחיקת הקטלוג:
- חוזרים לדף Lakehouse.
- בשורה של
quickstart_catalog, לוחצים על סמל האפשרויות הנוספות > מחיקה. - מזינים
DELETEכדי לאשר ולוחצים על מחיקה.
מוחקים את הקטגוריה של Cloud Storage ואת כל התוכן שלה:
נכנסים לדף Buckets של Cloud Storage.
מסמנים את התיבה ליד הקטגוריה שיצרתם עבור המדריך הזה ולוחצים על מחיקה.
מזינים
DELETEכדי לאשר ולוחצים על מחיקה.
המאמרים הבאים
- אפשר לנסות את המדריך למתחילים בנושא יצירה של טבלת Iceberg וביצוע שאילתות בה באמצעות Google Cloud CLI.
- איך משנים נתונים באמצעות הצהרות DML ב-BigQuery ואיך מגדירים אפשרויות של טבלאות.
- כדי לשלוח שאילתות לנתונים ב-Amazon Web Services (AWS), ב-AlloyDB ל-PostgreSQL וב-Cloud Storage בלי לבצע ETL, אפשר לנסות את השיעור בנושא הגדרת גישה לנתונים בין עננים.
- מידע נוסף על קטלוגים עם כמה מאגרי מידע ועל נקודת הקצה של קטלוג REST של Apache Iceberg
- איך מנהלים קטלוגים בקטלוג של זמן הריצה של Lakehouse
- מידע על טבלאות Apache Iceberg שמנוהלות על ידי Lakehouse