Lakehouse regional endpoints

This document describes how you can use Private Service Connect regional endpoints to access resources in the Lakehouse runtime catalog. Regional endpoints let you run your workloads in a manner that complies with data residency and data sovereignty requirements, where your request traffic is routed directly to the region specified in the endpoint.

Overview

Regional endpoints restrict requests to proceed only if the affected catalog resource exists in the location specified by the endpoint. For example, if you use the endpoint https://biglake.us-central1.rep.googleapis.com to access a catalog, namespace, or table, then the request only proceeds if the catalog is located in us-central1.

Unlike global endpoints, where requests can be routed through a different location from where the resource resides, regional endpoints restrict your requests to the location specified by the endpoint where the resource resides. Regional endpoints terminate Transport Layer Security (TLS) sessions in the location specified by the endpoint for requests received from the internet, other Google Cloud resources such as Compute Engine virtual machines, on-premises services using Cloud VPN or Cloud Interconnect, and Virtual Private Clouds (VPCs).

Regional endpoints help to ensure data residency by keeping your in-transit catalog requests within the location specified by the endpoint. For more information about how service metadata is handled, see Note on service data.

The following catalog endpoints in the Lakehouse runtime catalog are available for use with regional endpoints:

Catalog endpoint Regional endpoint URL Reference
Apache Iceberg REST catalog endpoint https://biglake.LOCATION.rep.googleapis.com/iceberg/v1/restcatalog REST
Apache Hive catalog endpoint (Preview) https://biglake.LOCATION.rep.googleapis.com/hive/v1 REST

Supported locations

You can use regional endpoints with the Lakehouse runtime catalog in the following locations:

  • Asia-Pacific

    • Delhi asia-south2
    • Mumbai asia-south1
  • Europe

    • Belgium europe-west1
    • Frankfurt europe-west3
    • London europe-west2
    • Milan europe-west8
    • Netherlands europe-west4
    • Paris europe-west9
    • Zürich europe-west6
  • Middle East

    • Dammam me-central2
  • Americas

    • Columbus, Ohio us-east5
    • Dallas us-south1
    • Iowa us-central1
    • Las Vegas us-west4
    • Los Angeles us-west2
    • Montréal northamerica-northeast1
    • Northern Virginia us-east4
    • Oregon us-west1
    • Salt Lake City us-west3
    • South Carolina us-east1
    • Toronto northamerica-northeast2

Supported operations and storage locations

Regional endpoints can only be used to perform operations that access or mutate catalog resources replicated in the location specified by the endpoint:

  • Regional catalog isolation: Listing catalogs through https://biglake.LOCATION.rep.googleapis.com returns only catalogs located in LOCATION. Requests to get, update, or delete a catalog, namespace, or table located outside of LOCATION return a 404 NOT_FOUND error.
  • Default storage location: When you create a catalog using a regional endpoint, the Cloud Storage bucket specified in default_location for a multiple-bucket catalog (CATALOG_TYPE_BIGLAKE) or the bucket associated with a single-bucket catalog (CATALOG_TYPE_GCS_BUCKET) must reside in LOCATION.
  • Multiple-bucket catalog restricted locations: For multiple-bucket catalogs, you can configure additional Cloud Storage buckets in restricted_locations as long as those buckets reside within the same geographic jurisdiction (such as the US or Europe) as LOCATION. For more information, see Multiple-bucket catalog.

Limitations and restrictions

Regional endpoints cannot be used to perform the following operations:

  • Operations that read or modify catalogs, namespaces, or tables located outside of the region specified by the endpoint.
  • Multi-region endpoint routing (such as US or EU). Regional endpoints must specify a single region.

Keep in mind the following restrictions when using regional endpoints:

Configure tools and query engines

You can configure the Google Cloud CLI, Apache Spark, Trino, and direct REST API requests to use regional endpoints.

gcloud CLI

To configure the gcloud CLI to use regional endpoints with gcloud biglake commands, set the api_endpoint_overrides/biglake property to the regional endpoint that you want to use:

gcloud config set api_endpoint_overrides/biglake https://biglake.LOCATION.rep.googleapis.com/

Alternatively, you can set the CLOUDSDK_API_ENDPOINT_OVERRIDES_BIGLAKE environment variable for individual commands:

CLOUDSDK_API_ENDPOINT_OVERRIDES_BIGLAKE=https://biglake.LOCATION.rep.googleapis.com/ \
    gcloud biglake iceberg catalogs list --project=PROJECT_ID

Replace the following:

  • LOCATION: the supported region for your catalog (for example, us-central1).
  • PROJECT_ID: your Google Cloud project ID.

Apache Spark

When configuring an Spark session to connect to the Apache Iceberg REST catalog endpoint, set the spark.sql.catalog.CATALOG_NAME.uri property to the regional endpoint URL:

from pyspark.sql import SparkSession

catalog_name = "CATALOG_NAME"
spark = SparkSession.builder.appName("APP_NAME") \
    .config('spark.sql.defaultCatalog', 'CATALOG_NAME') \
    .config(f'spark.sql.catalog.{catalog_name}', 'org.apache.iceberg.spark.SparkCatalog') \
    .config(f'spark.sql.catalog.{catalog_name}.type', 'rest') \
    .config(f'spark.sql.catalog.{catalog_name}.uri', 'https://biglake.LOCATION.rep.googleapis.com/iceberg/v1/restcatalog') \
    .config(f'spark.sql.catalog.{catalog_name}.warehouse', 'bl://projects/PROJECT_ID/catalogs/CATALOG_ID') \
    .config(f'spark.sql.catalog.{catalog_name}.header.x-goog-user-project', 'PROJECT_ID') \
    .config(f'spark.sql.catalog.{catalog_name}.rest.auth.type', 'org.apache.iceberg.gcp.auth.GoogleAuthManager') \
    .config(f'spark.sql.catalog.{catalog_name}.io-impl', 'org.apache.iceberg.gcp.gcs.GCSFileIO') \
    .config(f'spark.sql.catalog.{catalog_name}.header.X-Iceberg-Access-Delegation', 'vended-credentials') \
    .config('spark.sql.extensions', 'org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions') \
    .getOrCreate()

Replace the following:

  • CATALOG_NAME: a name for the local Spark catalog (for example, my_catalog).
  • APP_NAME: a name for your Spark session.
  • LOCATION: the supported region where the catalog is located (for example, us-central1).
  • PROJECT_ID: your Google Cloud project ID.
  • CATALOG_ID: the ID of your multiple-bucket catalog.

For more configuration options, see Set up the Apache Iceberg REST catalog endpoint.

Trino

When creating a Managed Service for Apache Spark cluster with the Trino component, set trino-catalog:CATALOG_NAME.iceberg.rest-catalog.uri to the regional endpoint URL:

gcloud dataproc clusters create CLUSTER_NAME \
    --enable-component-gateway \
    --region=LOCATION \
    --image-version=DATAPROC_VERSION \
    --network=NETWORK_ID \
    --optional-components=TRINO \
    --properties="\
    trino-catalog:CATALOG_NAME.connector.name=iceberg,\
    trino-catalog:CATALOG_NAME.iceberg.catalog.type=rest,\
    trino-catalog:CATALOG_NAME.iceberg.rest-catalog.uri=https://biglake.LOCATION.rep.googleapis.com/iceberg/v1/restcatalog,\
    trino-catalog:CATALOG_NAME.iceberg.rest-catalog.warehouse=bl://projects/PROJECT_ID/catalogs/CATALOG_ID,\
    trino-catalog:CATALOG_NAME.iceberg.rest-catalog.biglake.project-id=PROJECT_ID,\
    trino-catalog:CATALOG_NAME.iceberg.rest-catalog.rest.auth.type=org.apache.iceberg.gcp.auth.GoogleAuthManager"

Replace the following:

  • CLUSTER_NAME: a name for your cluster.
  • LOCATION: the supported region for your cluster and catalog.
  • DATAPROC_VERSION: the Managed Service for Apache Spark image version (for example, 2.2).
  • NETWORK_ID: the cluster network ID.
  • CATALOG_NAME: the name of your Trino catalog.
  • PROJECT_ID: your Google Cloud project ID.
  • CATALOG_ID: the ID of your multiple-bucket catalog.

REST APIs

Instead of sending a REST request to the global endpoint (https://biglake.googleapis.com), send the request to the regional endpoint in the following format: https://biglake.LOCATION.rep.googleapis.com.

For example, to list catalogs in LOCATION using the Apache Iceberg REST catalog endpoint extensions API:

curl -X GET \
    -H "Authorization: Bearer $(gcloud auth print-access-token)" \
    -H "x-goog-user-project: PROJECT_ID" \
    "https://biglake.LOCATION.rep.googleapis.com/iceberg/v1/restcatalog/extensions/projects/PROJECT_ID/catalogs"

Restrict global API endpoint usage

To help enforce the use of regional endpoints, use the constraints/gcp.restrictEndpointUsage organization policy constraint to block requests to the global API endpoint (biglake.googleapis.com). For more information, see Restrict endpoint usage.

The following example organization policy YAML file denies requests to the global biglake.googleapis.com endpoint while allowing requests to regional endpoints:

name: projects/PROJECT_ID/policies/gcp.restrictEndpointUsage
spec:
  rules:
  - values:
      deniedValues:
      - under:services/biglake.googleapis.com

What's next