This tutorial shows you how to fine-tune a
Gemma 4
31B large language model (google/gemma-4-31b-it) on a multi-host, multi-GPU
Google Kubernetes Engine (GKE) Autopilot cluster on Google Cloud. This
cluster uses two A4 (a4-highgpu-8g) virtual machine (VM) instances with a
total of 16 NVIDIA B200 GPUs.
The three main processes described in this tutorial are as follows:
- Deploy a multi-host GKE cluster in Autopilot mode.
- Build a custom container image with the required fine-tuning dependencies by using Cloud Build.
- Orchestrate a distributed multi-host fine-tuning workload across all 16 GPUs by using Kubernetes JobSet and the Hugging Face Accelerate library with Fully Sharded Data Parallel v2 (FSDP v2), pushing checkpoints to Hugging Face Hub.
This tutorial is intended for machine learning (ML) engineers, researchers, platform administrators and operators, and data and AI specialists who deploy GKE clusters on Google Cloud to fine-tune LLMs across multiple hosts.
Objectives
Access the Gemma 4 model by using Hugging Face.
Prepare your environment.
Create and deploy a multi-host A4 GKE cluster.
Fine-tune the Gemma 4 31B model across 16 GPUs by using Kubernetes
JobSetand Hugging Face Accelerate with FSDP v2.Monitor your job.
View the fine-tuned adapter weights on Hugging Face Hub.
Clean up.
Costs
In this document, you use the following billable components of Google Cloud:
To generate a cost estimate based on your projected usage,
use the pricing calculator.
Before you begin
To get the permissions that you need to complete this tutorial, ask your administrator to grant you the following IAM roles on your project:
- Kubernetes Engine Admin (
roles/container.admin) - Compute Admin (
roles/compute.admin) - Storage Admin (
roles/storage.admin) - Artifact Registry Administrator (
roles/artifactregistry.admin) - Cloud Build Editor (
roles/cloudbuild.builds.editor) - Service Account User (
roles/iam.serviceAccountUser) - Service Account Admin (
roles/iam.serviceAccountAdmin) - Project IAM Admin (
roles/resourcemanager.projectIamAdmin) - Service Usage Admin (
roles/serviceusage.serviceUsageAdmin)
For more information about granting roles, see Manage access to projects, folders, and organizations.
You might also be able to get the required permissions through custom roles or other predefined roles.
Enable the required APIs, if any aren't already enabled:
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.gcloud services enable compute.googleapis.com
container.googleapis.com artifactregistry.googleapis.com cloudbuild.googleapis.com logging.googleapis.com cloudresourcemanager.googleapis.com servicenetworking.googleapis.com Enable the default Compute Engine service account for your Google Cloud project:
export PROJECT_NUMBER="$(gcloud projects describe "YOUR_PROJECT_ID" --format "value(project_number)")" gcloud iam service-accounts enable "${PROJECT_NUMBER}-compute@developer.gserviceaccount.com" \ --project=YOUR_PROJECT_IDGrant the least-privilege IAM roles that the default Compute Engine service account needs to build the container image and run the fine-tuning workload:
ROLES=( "roles/artifactregistry.writer" "roles/cloudbuild.builds.builder" "roles/logging.logWriter" "roles/monitoring.metricWriter" "roles/monitoring.viewer" "roles/stackdriver.resourceMetadata.writer" "roles/storage.objectViewer" ) for role in "${ROLES[@]}"; do gcloud projects add-iam-policy-binding YOUR_PROJECT_ID \ --member="serviceAccount:${PROJECT_NUMBER}-compute@developer.gserviceaccount.com" \ --role="${role}" 1>/dev/null done unset ROLESVerify that the roles were granted to the default Compute Engine service account:
echo "Displaying roles for ${PROJECT_NUMBER}-compute@developer.gserviceaccount.com:" gcloud projects get-iam-policy YOUR_PROJECT_ID \ --flatten="bindings[].members" \ --filter="bindings.members:serviceAccount:${PROJECT_NUMBER}-compute@developer.gserviceaccount.com" \ --format="table(bindings.role)"Create local authentication credentials for your user account:
gcloud auth application-default login
Enable OS Login for your project:
gcloud compute project-info add-metadata \ --metadata=enable-oslogin=TRUE \ --project=YOUR_PROJECT_ID
Access Gemma 4 by using Hugging Face
To use Hugging Face to access Gemma 4, complete the following steps:
- Sign in to Hugging Face and accept the Gemma 4 license agreement.
- Create a Hugging Face
writeaccess token.
Click Your Profile > Settings > Access tokens > +Create new token. - Copy and save the
writeaccess token value. You use this token to download the base model and push fine-tuned adapter checkpoints to Hugging Face Hub before GKE scales down the GPU nodes.
Prepare your environment
To prepare your environment, set the following environment variables:
Replace the following:
YOUR_PROJECT_ID: the ID of the Google Cloud project where you want to create the GKE cluster.
YOUR_CLUSTER_NAME: the name of the GKE cluster to create.
YOUR_REGION: the region where you want to create your GKE cluster. You can only create the cluster in the region where your reservation exists.
YOUR_RESERVATION_NAME: the identifier for your reserved capacity.
YOUR_HF_TOKEN: the Hugging Face
writeaccess token that you created in the previous section.YOUR_ARTIFACT_REGISTRY_LOCATION: the Google Cloud region (for example,
us-central1) where you want to create your Artifact Registry repository. To minimize image pull latency, use the same region that you specified for YOUR_REGION.YOUR_NUMBER_OF_NODES: the number of A4 VM nodes in your fine-tuning job. For this multi-host tutorial with 16 NVIDIA B200 GPUs across two
a4-highgpu-8ginstances, set this value to2.
Create a multi-host GKE cluster in Autopilot mode
Create a multi-host GKE cluster in Autopilot mode:
Creating the GKE cluster might take several minutes to complete. To verify that Google Cloud has finished creating your cluster, go to Kubernetes clusters on the Google Cloud console.
Configure kubectl to communicate with your GKE cluster
Configure kubectl to communicate with your GKE cluster:
Create a Kubernetes secret for Hugging Face credentials
Create a Kubernetes secret to store your Hugging Face token:
Prepare your workload
To prepare your workload, you do the following:
Create workload scripts
To create the configuration files and scripts that your fine-tuning workload uses, complete the following steps:
Create a directory for the workload scripts. Use this directory as your working directory.
Create the
cloudbuild.yamlfile to build your workload container image with Cloud Build and push it to Artifact Registry:Create a
Dockerfilefile to define the environment and install the dependencies required to complete the fine-tuning job:Create the
accel_fsdp_gemma4_config.yamlfile. This configuration directs Hugging Face Accelerate to shardGemma4TextDecoderLayeracross 16 GPUs on two hosts by using FSDP v2:Create the
finetune.yamlKubernetesJobSetmanifest:Create the
finetune.pysupervised fine-tuning script:
Use Docker and Cloud Build to create a fine-tuning container
Create an Artifact Registry Docker repository:
Install the
JobSetcustom resource definitions (CRDs) required for orchestrating multi-host workloads:In the
llm-finetuning-gemmadirectory that you created in an earlier step, submit the container build to Cloud Build:Export the multi-host container image URL. You use it at a later step in this tutorial, when you deploy the
JobSetmanifest:
Start your fine-tuning workload
To deploy and monitor your distributed fine-tuning workload, complete the following steps:
Substitute environment variables into the fine-tuning manifest to create the fine-tuning job:
Because your cluster runs in GKE Autopilot mode, it might take a few minutes to provision the two GPU-enabled A4 nodes and pull the container image.
Watch the worker pods until both pods transition to the
Runningstatus:After the worker pods transition to
Running, stream the training logs:
Monitor your workload
You can monitor GPU utilization across your GKE cluster to verify that all 16 GPUs across both A4 hosts are actively processing training steps. Generate and open the observability link in your browser:
When you monitor your workload, expect the following behavior:
- GPU utilization: For a healthy distributed fine-tuning job, you can expect to see GPU utilization across all 16 NVIDIA B200 GPUs rise and stabilize near 95%–100% during training steps.
- Job duration: Across two
a4-highgpu-8gnodes (16 B200 GPUs), the 3-epoch fine-tuning job takes approximately 2 and a half hours to complete.
View your fine-tuned adapter weights
When training finishes, view your fine-tuned LoRA adapter weights and
checkpoints on Hugging Face Hub at
https://huggingface.co/YOUR_HF_USERNAME/gemma-31b-text-to-sql.
Clean up
To avoid incurring additional charges, delete the resources created during this tutorial.
Delete your resources
Delete the fine-tuning
JobSet:Delete your GKE cluster:
Delete your Artifact Registry repository: