Creare ed eseguire query sulle tabelle BigLake Iceberg in BigQuery

Managed Airflow (terza generazione) | Managed Airflow (seconda generazione) | Managed Airflow (prima generazione legacy)

Questa pagina spiega come creare e modificare le tabelle Iceberg BigLake in BigQuery utilizzando gli operatori Airflow nel tuo ambiente Airflow gestito.

Informazioni sulle tabelle Iceberg BigLake in BigQuery

Le tabelle Iceberg BigLake in BigQuery forniscono le basi per la creazione di lakehouse in formato aperto Google Cloud. Le tabelle Iceberg BigLake in BigQuery offrono la stessa esperienza completamente gestita delle tabelle BigQuery standard, ma archiviano i dati nei bucket di archiviazione di proprietà del cliente. Le tabelle Iceberg BigLake in BigQuery supportano il formato di tabella Iceberg aperto per una migliore interoperabilità con i motori di calcolo open source e di terze parti su una singola copia dei dati.

Prima di iniziare

Crea una tabella Iceberg BigLake in BigQuery

Per creare una tabella Iceberg BigLake in BigQuery, utilizza BigQueryCreateTableOperator nello stesso modo delle altre tabelle BigQuery. Nel campo biglakeConfiguration, fornisci la configurazione della tabella.

import datetime

from airflow.models.dag import DAG
from airflow.providers.google.cloud.operators.bigquery import BigQueryCreateTableOperator

with DAG(
  "bq_iceberg_dag",
  start_date=datetime.datetime(2025, 1, 1),
  schedule=None,
  ) as dag:

  create_iceberg_table = BigQueryCreateTableOperator(
    task_id="create_iceberg_table",
    project_id="PROJECT_ID",
    dataset_id="DATASET_ID",
    table_id="TABLE_NAME",
    table_resource={
      "schema": {
        "fields": [
          {"name": "order_id", "type": "INTEGER", "mode": "REQUIRED"},
          {"name": "customer_id", "type": "INTEGER", "mode": "REQUIRED"},
          {"name": "amount", "type": "INTEGER", "mode": "REQUIRED"},
          {"name": "created_at", "type": "TIMESTAMP", "mode": "REQUIRED"},
        ]
      },
      "biglakeConfiguration": {
        "connectionId": "CONNECTION_NAME",
        "storageUri": "STORAGE_URI",
        "fileFormat": "PARQUET",
        "tableFormat": "ICEBERG",
      }
    }
  )

Sostituisci quanto segue:

  • PROJECT_ID: l'ID progetto .
  • DATASET_ID: un set di dati esistente.
  • TABLE_NAME: il nome della tabella che stai creando.
  • CONNECTION_NAME: il nome della connessione alle risorse Cloud nel formato projects/PROJECT_ID/locations/REGION/connections/CONNECTION_ID.
  • STORAGE_URI: un URI Cloud Storage completo per la tabella. Ad esempio, gs://example-bucket/iceberg-table.

Esegui query su una tabella Iceberg BigLake in BigQuery

Dopo aver creato una tabella Iceberg BigLake, puoi eseguirne una query con BigQueryInsertJobOperator come di consueto. L'operatore non richiede una configurazione aggiuntiva specifica per le tabelle Iceberg BigLake.

import datetime

from airflow.models.dag import DAG
from airflow.providers.google.cloud.operators.bigquery import BigQueryInsertJobOperator

with DAG(
  "bq_iceberg_dag_query",
  start_date=datetime.datetime(2025, 1, 1),
  schedule=None,
  ) as dag:

  insert_values = BigQueryInsertJobOperator(
    task_id="iceberg_insert_values",
    configuration={
      "query": {
        "query": f"""
          INSERT INTO `TABLE_ID` (order_id, customer_id, amount, created_at)
          VALUES
            (101, 19, 1, TIMESTAMP '2025-09-15 10:15:00+00'),
            (102, 35, 2, TIMESTAMP '2025-09-14 10:15:00+00'),
            (103, 36, 3, TIMESTAMP '2025-09-12 10:15:00+00'),
            (104, 37, 4, TIMESTAMP '2025-09-11 10:15:00+00')
        """,
        "useLegacySql": False,
        }
      }
  )

Sostituisci quanto segue:

  • TABLE_ID con l'ID tabella, nel formato PROJECT_ID.DATASET_ID.TABLE_NAME.

Passaggi successivi