使用 Google Cloud 控制台在 Lakehouse 中建立及查詢 Iceberg 資料表

在本快速入門導覽課程中,您將使用 Google Cloud 控制台,瞭解無邊界 Lakehouse 如何透過在 Lakehouse 執行階段目錄中儲存資料表的中繼資料 (包括結構定義、快照和儲存位置),讓您跨 Google Cloud 和開放原始碼引擎管理及共用 Apache Iceberg 資料表。

如要完成本快速入門導覽課程,請在Google Cloud 控制台中執行下列步驟:

  1. 建立 Cloud Storage bucket:在 Cloud Storage 中建立 bucket,用於儲存 Iceberg 資料表資料和中繼資料檔案。
  2. 建立目錄:在 Lakehouse 執行階段目錄中建立多 bucket 目錄,並以啟用憑證販售功能的 bucket 做為後端。
  3. 建立命名空間和 Iceberg 資料表:在 Google Cloud 控制台中使用「Lakehouse」頁面,建立命名空間和 Iceberg 資料表,並啟用 BigQuery 資料操作語言 (DML)。
  4. 在 BigQuery 中修改資料及查詢資料表:使用 BigQuery DML 陳述式 (INSERT、UPDATE 和 DELETE) 修改 Iceberg 資料表中的資料列,並使用 4 部分的 P.C.N.T (Project.Catalog.Namespace.Table) 語法查詢結果,無須 ETL 或手動註冊資料表。

事前準備

  1. 登入 Google Cloud 帳戶。如果您是 Google Cloud新手,歡迎 建立帳戶,親自體驗產品的實際應用成效。新客戶還能獲得價值 $300 美元的免費抵免額,能用於執行、測試及部署工作負載。
  2. In the Google Cloud console, on the project selector page, select or create a Google Cloud project.

    Roles required to select or create a project

    • Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
    • Create a project: To create a project, you need the Project Creator role (roles/resourcemanager.projectCreator), which contains the resourcemanager.projects.create permission. Learn how to grant roles.

    Go to project selector

  3. Verify that billing is enabled for your Google Cloud project.

  4. Enable the BigLake, Cloud Storage, and BigQuery APIs, if any are not already enabled.

    Roles required to enable APIs

    To enable APIs, you need the serviceusage.services.enable permission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.

    Enable the APIs

  5. Make sure that you have the following role or roles on the project: BigLake Admin (roles/biglake.admin), Storage Admin (roles/storage.admin), and BigQuery Job User (roles/bigquery.jobUser)

    Check for the roles

    1. In the Google Cloud console, go to the IAM page.

      Go to IAM
    2. Select the project.
    3. In the Principal column, find all rows that identify you or a group that you're included in. To learn which groups you're included in, contact your administrator.

    4. For all rows that specify or include you, check the Role column to see whether the list of roles includes the required roles.

    Grant the roles

    1. In the Google Cloud console, go to the IAM page.

      Go to IAM
    2. Select the project.
    3. Click Grant access.
    4. In the New principals field, enter your user identifier. This is typically the email address for a Google Account.

    5. Click Select a role, then search for the role.
    6. To grant additional roles, click Add another role and add each additional role.
    7. Click Save.
  6. In the Google Cloud console, on the project selector page, select or create a Google Cloud project.

    Roles required to select or create a project

    • Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
    • Create a project: To create a project, you need the Project Creator role (roles/resourcemanager.projectCreator), which contains the resourcemanager.projects.create permission. Learn how to grant roles.

    Go to project selector

  7. Verify that billing is enabled for your Google Cloud project.

  8. Enable the BigLake, Cloud Storage, and BigQuery APIs, if any are not already enabled.

    Roles required to enable APIs

    To enable APIs, you need the serviceusage.services.enable permission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.

    Enable the APIs

  9. Make sure that you have the following role or roles on the project: BigLake Admin (roles/biglake.admin), Storage Admin (roles/storage.admin), and BigQuery Job User (roles/bigquery.jobUser)

    Check for the roles

    1. In the Google Cloud console, go to the IAM page.

      Go to IAM
    2. Select the project.
    3. In the Principal column, find all rows that identify you or a group that you're included in. To learn which groups you're included in, contact your administrator.

    4. For all rows that specify or include you, check the Role column to see whether the list of roles includes the required roles.

    Grant the roles

    1. In the Google Cloud console, go to the IAM page.

      Go to IAM
    2. Select the project.
    3. Click Grant access.
    4. In the New principals field, enter your user identifier. This is typically the email address for a Google Account.

    5. Click Select a role, then search for the role.
    6. To grant additional roles, click Add another role and add each additional role.
    7. Click Save.

建立 Cloud Storage bucket

在 Google Cloud 控制台中建立 Cloud Storage bucket,用來儲存 Iceberg 資料表資料和中繼資料檔案:

  1. 前往 Google Cloud 控制台的「Cloud Storage bucket」頁面。

    前往「Buckets」(值區) 頁面

  2. 點選「建立」。

  3. 在「開始使用」部分,輸入全域不重複的 bucket 名稱 (例如 lakehouse-quickstart-UNIQUE_ID 或 PROJECT_ID-lakehouse),然後按一下「繼續」。

  4. 在「選擇資料的儲存位置」部分,將「位置類型」設為「多區域」(「美國 (多個美國地區)」),然後按一下「建立」。

  5. 如果出現「系統會禁止公開存取」對話方塊,請按一下「確認」。

在 Lakehouse 執行階段目錄中建立目錄

在 Lakehouse 執行階段目錄中,為 Apache Iceberg 資料表建立多值區目錄。多 bucket 目錄可讓您獨立命名目錄,不必與任何 bucket 名稱相同,並將多個 Cloud Storage bucket 與單一目錄建立關聯。如要確保這些值區的存取安全,請啟用憑證臨時配發模式,讓目錄直接自動核發臨時儲存空間憑證給用戶端引擎。

  1. 前往 Google Cloud 控制台的「Lakehouse」Lakehouse頁面。

    前往 Lakehouse

  2. 按一下「建立目錄」,然後選取「Lakehouse 執行階段目錄」。

  3. 在「目錄詳細資料」部分,設定下列項目:

    • 目錄類型:選取「Iceberg Rest 目錄」。
    • Lakehouse 目錄 bucket 選項:選取「多 bucket 目錄」。
    • 預設目錄 Cloud Storage 路徑:按一下「瀏覽」,選取您建立的 bucket,然後按一下「選取」。
    • 目錄 ID:輸入 quickstart_catalog。
    • 主要位置:選取「多區域」,然後選取「美國 (多個美國地區)」。
  4. 依序點選「繼續」和「資料路徑」部分中的「繼續」。

  5. 在「驗證方式」部分,選取「憑證販售模式」。

    透過憑證販售功能,目錄會安全地向用戶端引擎和 BigQuery 發放暫時性的資料表範圍儲存空間權杖,因此外部引擎不需要對值區擁有直接的 IAM 權限。

  6. 點選「建立」。

    目錄建立完成後,「目錄詳細資料」頁面會隨即開啟。

  7. 在「驗證方法」下方,按一下「設定 bucket 權限」,然後在對話方塊中按一下「確認」。

    這個步驟會將 Cloud Storage bucket 的必要權限授予目錄的服務帳戶,以提供臨時憑證。

建立命名空間和 Iceberg 資料表

現在您已擁有目錄,請使用控制台的「Lakehouse」Lakehouse頁面建立命名空間和 Iceberg 資料表。Google Cloud

建立命名空間

  1. 在 quickstart_catalog 的「目錄詳細資料」頁面,按一下 「建立命名空間」。

  2. 在「命名空間名稱」欄位中,輸入 quickstart_namespace。

  3. 將「Location」保留為系統自動填入欄位的預設 Cloud Storage 路徑。

  4. 點選「建立」。

建立 Iceberg 資料表

  1. 在「目錄詳細資料」頁面中,按一下 quickstart_namespace。

    「命名空間詳細資料」頁面隨即開啟。

  2. 按一下 「建立資料表」。

  3. 在「建立資料表」窗格中,完成下列設定:

    • 「資料表格式」:確認已選取「Iceberg」。
    • 「資料表名稱」:輸入 quickstart_table。
    • 位置:保留預設的 Cloud Storage 路徑。
  4. 在「結構定義」下方,按兩下「新增欄位」,在表格中新增兩個資料欄:

    • 在第一個欄位的「欄位名稱」欄位中輸入 id,然後從「類型」選單中選取「INTEGER」。
    • 在第二個欄位的「欄位名稱」欄位中輸入 name,然後從「類型」選單中選取「STRING」。
  5. 在「屬性」下方,找出預先定義的 gcp.biglake.bigquery-dml.enabled 屬性,並將「值」從 false 變更為 true。將 gcp.biglake.table-management.enabled 設為 false。

    將 gcp.biglake.bigquery-dml.enabled 設為 true 後,您就能使用 BigQuery DML 陳述式 (例如 INSERT、UPDATE、DELETE 和 MERGE) 修改 Iceberg 資料表中的資料。詳情請參閱「設定表格選項」。

  6. 點選「建立」。

    新的 Iceberg 資料表 (quickstart_table) 會顯示在「命名空間詳細資料」頁面上,而 Lakehouse 執行階段目錄會將初始 Iceberg 中繼資料檔案寫入 Cloud Storage bucket。

在 BigQuery 中修改資料及查詢資料表

建立 quickstart_table 並啟用 BigQuery DML 後,您可以使用 4 部分 P.C.N.T (Project.Catalog.Namespace.Table) 語法,直接在 BigQuery 中插入、更新、刪除及查詢資料列。在每個陳述式中,將 PROJECT_ID 替換為您的Google Cloud 專案 ID:

  1. 前往 Google Cloud 控制台的「BigQuery」頁面。

    前往 BigQuery

  2. 在查詢編輯器中,點選 「SQL 查詢」。

  3. 插入三列範例資料:

    INSERT INTO `PROJECT_ID.quickstart_catalog.quickstart_namespace.quickstart_table` (id, name)
    VALUES (1, 'one'), (2, 'two'), (3, 'three');

    按一下「執行」。INSERT 陳述式完成後,BigQuery 會將 Parquet 資料檔案寫入 Cloud Storage bucket,並將新的 Iceberg 快照提交至 Lakehouse 執行階段目錄。

  4. 修改表格中的資料列:

    UPDATE `PROJECT_ID.quickstart_catalog.quickstart_namespace.quickstart_table`
    SET name = 'updated'
    WHERE id = 1;

    按一下「執行」。

  5. 從表格中刪除資料列:

    DELETE FROM `PROJECT_ID.quickstart_catalog.quickstart_namespace.quickstart_table`
    WHERE id = 3;

    按一下「執行」。

  6. 查詢資料表,確認變更:

    SELECT * FROM `PROJECT_ID.quickstart_catalog.quickstart_namespace.quickstart_table`
    ORDER BY id;

    按一下「執行」。「Query results」(查詢結果) 窗格會顯示其餘兩列,包括 id = 1 的更新值:

    +----+---------+
    | id | name    |
    +----+---------+
    |  1 | updated |
    |  2 | two     |
    +----+---------+
    

由於 Lakehouse 執行階段目錄會管理 Iceberg 中繼資料,且已啟用憑證販售功能,因此您也可以使用任何與 Iceberg 相容的開放原始碼引擎 (例如 Apache Spark、Trino 或 Apache Flink) 讀取或寫入 quickstart_table,不必授予這些引擎對 Bucket 的直接 IAM 存取權。

清除所用資源

為避免系統向您的 Google Cloud 帳戶收取不必要的費用,請刪除您在本快速入門導覽中建立的資源。刪除資料表、命名空間和目錄會從 Lakehouse 執行階段目錄移除中繼資料登錄,而刪除 bucket 則會移除儲存在 Cloud Storage 中的基礎 Parquet 資料和 Iceberg 中繼資料檔案:

  1. 前往 Google Cloud 控制台的「Lakehouse」Lakehouse頁面。

    前往 Lakehouse

  2. 從目錄中刪除資料表:

    1. 依序點選 quickstart_catalog 和 quickstart_namespace。
    2. 在「命名空間詳細資料」表格中,點按 quickstart_table 所在資料列的「更多」>「刪除」。
    3. 輸入 DELETE 確認,然後按一下「刪除」。
  3. 從目錄中刪除命名空間:

    1. 返回 quickstart_catalog 的「目錄詳細資料」頁面。
    2. 在 quickstart_namespace 的資料列中,依序按一下「More namespace actions」(更多命名空間動作) >「Delete」(刪除)。
    3. 輸入 DELETE 確認,然後按一下「刪除」。
  4. 刪除目錄:

    1. 返回「Lakehouse」Lakehouse頁面。
    2. 在 quickstart_catalog 的資料列中,依序按一下 「更多目錄動作」>「刪除」。
    3. 輸入 DELETE 確認,然後按一下「刪除」。
  5. 刪除 Cloud Storage bucket 和所有內容:

    1. 前往 Cloud Storage 的「Buckets」(值區) 頁面。

      前往「Buckets」(值區) 頁面

    2. 勾選您為本快速入門指南建立的 bucket 旁的核取方塊,然後按一下「Delete」(刪除)。

    3. 輸入 DELETE 確認,然後按一下「刪除」。

後續步驟