This document advises developers, architects, and decision-makers on options for integrating Managed Service for Apache Kafka with other Google Cloud services.
The following table lists some factors to consider when choosing an integration solution for your Kafka data in Google Cloud.
| Solution | Use cases |
|---|---|
| Kafka Connect |
Common data integration tasks Replicating Kafka data between clusters High availability/disaster recovery |
| Dataflow |
High-volume data pipelines Complex data transformations or data aggregation Advanced windowing |
| Managed Service for Apache Spark |
Integrate with existing Apache Spark workloads Migrate Apache Spark workloads to Google Cloud |
Use Kafka Connect
Kafka Connect provides a framework for streaming data between Kafka and other systems. Managed Service for Apache Kafka lets you provision clusters that run Kafka Connect.
Kafka Connect is recommended for most common data integration tasks for which there is a supported connector. Kafka Connect with MirrorMaker 2.0 is recommended for replicating data between Kafka clusters, for example in high availability and disaster recovery (HA/DR) architectures.
Advantages of using Kafka Connect include:
Built-in connectors for sources and sinks such as BigQuery, Cloud SQL, and Cloud Storage.
Fully integrated into Managed Service for Apache Kafka, which simplifies management, operations, and monitoring.
Support for filtering and per-message transformations.
Use Dataflow
Dataflow is a managed streaming platform built on the open source Apache Beam project. You can run Dataflow pipelines to write Kafka data to sinks such as BigQuery and Cloud Storage.
Consider using Dataflow for the following scenarios:
High-volume data pipelines. Dataflow jobs can scale up to thousands of workers. Dataflow also supports horizontal autoscaling, adding or removing workers as needed.
Data pipelines that require complex transformations, such as data cleaning, enrichment, or aggregation.
Integrating AI and ML into your data pipeline.
Joining multiple data streams with fine-grained control over windowing and handling of late-arriving data.
To deploy a Dataflow pipeline, you can run a prebuilt template or create an Apache Beam pipeline. For common data integration scenarios, consider running a template. For maximum flexibility and control, or if your pipeline requires custom logic, create an Apache Beam pipeline.
Use Managed Service for Apache Spark
Managed Service for Apache Spark lets you run Apache Spark workloads using a serverless model or a cluster-based model.
Consider using Managed Service for Apache Spark with your Kafka data if you have existing Spark pipelines. For example, if you have a Spark application that processes streaming data from Kafka, and you want to migrate this application to Google Cloud, Managed Service for Apache Spark is a suitable choice.
What's next
Learn about Kafka Connect in Managed Service for Apache Kafka.
Use a Dataflow template:
Create an Apache Beam pipeline:
For Spark Streaming and Kafka, Apache Spark provides an integration guide.