Create and query an Iceberg table in Lakehouse using the Google Cloud console
In this quickstart, you use the Google Cloud console to learn how borderless Lakehouse lets you manage and share Apache Iceberg tables across Google Cloud and open-source engines by storing table metadata, including schemas, snapshots, and storage locations, in the Lakehouse runtime catalog.
To complete this quickstart, you perform the following steps in the Google Cloud console:
- Create a Cloud Storage bucket: Create a bucket in Cloud Storage to store your Iceberg table data and metadata files.
- Create a catalog: Create a multiple-bucket catalog in the Lakehouse runtime catalog backed by your bucket with credential vending enabled.
- Create a namespace and an Iceberg table: Use the Lakehouse page in the Google Cloud console to create a namespace and an Iceberg table with BigQuery data manipulation language (DML) enabled.
- Modify data and query the table in BigQuery: Use
BigQuery DML statements (
INSERT,UPDATE, andDELETE) to modify rows in the Iceberg table and query the results using the 4-part P.C.N.T (Project.Catalog.Namespace.Table) syntax, with no ETL or manual table registration needed.
Before you begin
- Sign in to your Google Cloud account. If you're new to Google Cloud, create an account to evaluate how our products perform in real-world scenarios. New customers also get $300 in free credits to run, test, and deploy workloads.
-
In the Google Cloud console, on the project selector page, select or create a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Verify that billing is enabled for your Google Cloud project.
Enable the BigLake, Cloud Storage, and BigQuery APIs, if any are not already enabled.
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.-
Make sure that you have the following role or roles on the project: BigLake Admin (
roles/biglake.admin), Storage Admin (roles/storage.admin), and BigQuery Job User (roles/bigquery.jobUser)Check for the roles
-
In the Google Cloud console, go to the IAM page.
Go to IAM - Select the project.
-
In the Principal column, find all rows that identify you or a group that you're included in. To learn which groups you're included in, contact your administrator.
- For all rows that specify or include you, check the Role column to see whether the list of roles includes the required roles.
Grant the roles
-
In the Google Cloud console, go to the IAM page.
Go to IAM - Select the project.
- Click Grant access.
-
In the New principals field, enter your user identifier. This is typically the email address for a Google Account.
- Click Select a role, then search for the role.
- To grant additional roles, click Add another role and add each additional role.
- Click Save.
-
-
In the Google Cloud console, on the project selector page, select or create a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Verify that billing is enabled for your Google Cloud project.
Enable the BigLake, Cloud Storage, and BigQuery APIs, if any are not already enabled.
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.-
Make sure that you have the following role or roles on the project: BigLake Admin (
roles/biglake.admin), Storage Admin (roles/storage.admin), and BigQuery Job User (roles/bigquery.jobUser)Check for the roles
-
In the Google Cloud console, go to the IAM page.
Go to IAM - Select the project.
-
In the Principal column, find all rows that identify you or a group that you're included in. To learn which groups you're included in, contact your administrator.
- For all rows that specify or include you, check the Role column to see whether the list of roles includes the required roles.
Grant the roles
-
In the Google Cloud console, go to the IAM page.
Go to IAM - Select the project.
- Click Grant access.
-
In the New principals field, enter your user identifier. This is typically the email address for a Google Account.
- Click Select a role, then search for the role.
- To grant additional roles, click Add another role and add each additional role.
- Click Save.
-
Create a Cloud Storage bucket
Create a Cloud Storage bucket in the Google Cloud console to store your Iceberg table data and metadata files:
In the Google Cloud console, go to the Cloud Storage Buckets page.
Click Create.
In the Get started section, enter a globally unique bucket name (for example,
lakehouse-quickstart-UNIQUE_IDorPROJECT_ID-lakehouse), and then click Continue.In the Choose where to store your data section, leave Location type set to Multi-region (us (multiple regions in United States)), and then click Create.
If the Public access will be prevented dialog appears, click Confirm.
Create a catalog in the Lakehouse runtime catalog
Create a multiple-bucket catalog in the Lakehouse runtime catalog for your Apache Iceberg tables. A multiple-bucket catalog lets you name your catalog independently of any bucket name and associate multiple Cloud Storage buckets with a single catalog. To secure access to these buckets, enable credential vending mode so that the catalog can automatically issue temporary storage credentials directly to your client engines.
In the Google Cloud console, go to the Lakehouse page.
Click Create catalog and select Lakehouse runtime catalog.
In the Catalog details section, configure the following settings:
- Catalog type: Select Iceberg Rest Catalog.
- Lakehouse catalog bucket options: Select Multiple bucket catalog.
- Default Catalog Cloud Storage path: Click Browse, select the bucket that you created, and click Select.
- Catalog ID: Enter
quickstart_catalog. - Primary location: Select Multi-region, and then select US (multiple regions in United States).
Click Continue, and then in the Data paths section, click Continue.
In the Authentication method section, select Credential vending mode.
With credential vending, the catalog securely issues temporary, table-scoped storage tokens to client engines and BigQuery, so external engines don't need direct IAM permissions on your bucket.
Click Create.
Your catalog is created and the Catalog details page opens.
Under Authentication method, click Set bucket permissions, and then in the dialog, click Confirm.
This step grants the catalog's service account the required permissions on your Cloud Storage bucket to vend temporary credentials.
Create a namespace and an Iceberg table
Now that you have a catalog, use the Lakehouse page in the Google Cloud console to create a namespace and an Iceberg table.
Create a namespace
On the Catalog Details page for
quickstart_catalog, click Create namespace.In the Namespace name field, enter
quickstart_namespace.Leave Location set to the default Cloud Storage path that's automatically populated in the field.
Click Create.
Create an Iceberg table
On the Catalog Details page, click
quickstart_namespace.The Namespace Details page opens.
Click Create Table.
In the Create Table pane, configure the following settings:
- Table format: Verify that Iceberg is selected.
- Table name: Enter
quickstart_table. - Location: Leave the default Cloud Storage path.
Under Schema, click Add field twice to add two columns to the table:
- For the first field, enter
idin the Field name field, and select INTEGER from the Type menu. - For the second field, enter
namein the Field name field, and select STRING from the Type menu.
- For the first field, enter
Under Properties, find the predefined
gcp.biglake.bigquery-dml.enabledproperty and change its Value fromfalsetotrue. Leavegcp.biglake.table-management.enabledset tofalse.Setting
gcp.biglake.bigquery-dml.enabledtotruelets you modify data in the Iceberg table using BigQuery DML statements such asINSERT,UPDATE,DELETE, andMERGE. For more information, see Configure table options.Click Create.
Your new Iceberg table (
quickstart_table) appears on the Namespace details page, and the Lakehouse runtime catalog writes the initial Iceberg metadata file to your Cloud Storage bucket.
Modify data and query the table in BigQuery
With quickstart_table created and BigQuery DML enabled, you can
insert, update, delete, and query rows directly in BigQuery using
the 4-part P.C.N.T (Project.Catalog.Namespace.Table) syntax. In each
statement, replace PROJECT_ID with your
Google Cloud project ID:
In the Google Cloud console, go to the BigQuery page.
In the query editor, click SQL query.
Insert three rows of sample data:
INSERT INTO `PROJECT_ID.quickstart_catalog.quickstart_namespace.quickstart_table` (id, name) VALUES (1, 'one'), (2, 'two'), (3, 'three');
Click Run. When the
INSERTstatement completes, BigQuery writes the Parquet data files to your Cloud Storage bucket and commits a new Iceberg snapshot to the Lakehouse runtime catalog.Modify a row in the table:
UPDATE `PROJECT_ID.quickstart_catalog.quickstart_namespace.quickstart_table` SET name = 'updated' WHERE id = 1;
Click Run.
Delete a row from the table:
DELETE FROM `PROJECT_ID.quickstart_catalog.quickstart_namespace.quickstart_table` WHERE id = 3;
Click Run.
Query the table to verify your changes:
SELECT * FROM `PROJECT_ID.quickstart_catalog.quickstart_namespace.quickstart_table` ORDER BY id;
Click Run. The Query results pane shows the remaining two rows, including the updated value for
id = 1:+----+---------+ | id | name | +----+---------+ | 1 | updated | | 2 | two | +----+---------+
Because the Lakehouse runtime catalog manages the Iceberg
metadata and credential vending is enabled, you can also read from or write to
quickstart_table using any Iceberg-compatible open-source engine, such as
Apache Spark, Trino, or Apache Flink, without granting them
direct IAM access to your bucket.
Clean up
To avoid incurring unnecessary charges to your Google Cloud account, delete the resources you created in this quickstart. Deleting the table, namespace, and catalog removes the metadata registration from the Lakehouse runtime catalog, while deleting the bucket removes the underlying Parquet data and Iceberg metadata files stored in Cloud Storage:
In the Google Cloud console, go to the Lakehouse page.
Delete the table from your catalog:
- Click
quickstart_catalog, and then clickquickstart_namespace. - In the Namespace details table, in the row for
quickstart_table, click More > Delete. - Enter
DELETEto confirm and click Delete.
- Click
Delete the namespace from your catalog:
- Return to the Catalog details page for
quickstart_catalog. - In the row for
quickstart_namespace, click More namespace actions > Delete. - Enter
DELETEto confirm and click Delete.
- Return to the Catalog details page for
Delete your catalog:
- Return to the Lakehouse page.
- In the row for
quickstart_catalog, click More catalog actions > Delete. - Enter
DELETEto confirm and click Delete.
Delete your Cloud Storage bucket and all its contents:
Go to the Cloud Storage Buckets page.
Select the checkbox next to the bucket that you created for this quickstart and click Delete.
Enter
DELETEto confirm and click Delete.
What's next
- Try the Create and query an Iceberg table using the Google Cloud CLI quickstart.
- Learn how to modify data with BigQuery DML statements and configure table options.
- Try the Set up cross-cloud data access codelab to query data across Amazon Web Services (AWS), AlloyDB for PostgreSQL, and Cloud Storage without ETL.
- Learn about multiple-bucket catalogs and the Apache Iceberg REST catalog endpoint.
- Learn how to manage catalogs in the Lakehouse runtime catalog.
- Learn about Apache Iceberg tables managed by Lakehouse.