Integrate Managed Service for Apache Kafka with other Google Cloud services

This document advises developers, architects, and decision-makers on options for integrating Managed Service for Apache Kafka with other Google Cloud services.

The following table lists some factors to consider when choosing an integration solution for your Kafka data in Google Cloud.

SolutionUse cases
Kafka Connect

Common data integration tasks

Replicating Kafka data between clusters

High availability/disaster recovery

Dataflow

High-volume data pipelines

Complex data transformations or data aggregation

Advanced windowing

Managed Service for Apache Spark

Integrate with existing Apache Spark workloads

Migrate Apache Spark workloads to Google Cloud

Use Kafka Connect

Kafka Connect provides a framework for streaming data between Kafka and other systems. Managed Service for Apache Kafka lets you provision clusters that run Kafka Connect.

Kafka Connect is recommended for most common data integration tasks for which there is a supported connector. Kafka Connect with MirrorMaker 2.0 is recommended for replicating data between Kafka clusters, for example in high availability and disaster recovery (HA/DR) architectures.

Advantages of using Kafka Connect include:

  • Built-in connectors for sources and sinks such as BigQuery, Cloud SQL, and Cloud Storage.

  • Fully integrated into Managed Service for Apache Kafka, which simplifies management, operations, and monitoring.

  • Support for filtering and per-message transformations.

Use Dataflow

Dataflow is a managed streaming platform built on the open source Apache Beam project. You can run Dataflow pipelines to write Kafka data to sinks such as BigQuery and Cloud Storage.

Consider using Dataflow for the following scenarios:

  • High-volume data pipelines. Dataflow jobs can scale up to thousands of workers. Dataflow also supports horizontal autoscaling, adding or removing workers as needed.

  • Data pipelines that require complex transformations, such as data cleaning, enrichment, or aggregation.

  • Integrating AI and ML into your data pipeline.

  • Joining multiple data streams with fine-grained control over windowing and handling of late-arriving data.

To deploy a Dataflow pipeline, you can run a prebuilt template or create an Apache Beam pipeline. For common data integration scenarios, consider running a template. For maximum flexibility and control, or if your pipeline requires custom logic, create an Apache Beam pipeline.

Use Managed Service for Apache Spark

Managed Service for Apache Spark lets you run Apache Spark workloads using a serverless model or a cluster-based model.

Consider using Managed Service for Apache Spark with your Kafka data if you have existing Spark pipelines. For example, if you have a Spark application that processes streaming data from Kafka, and you want to migrate this application to Google Cloud, Managed Service for Apache Spark is a suitable choice.

What's next