This tutorial shows you how to orchestrate a distributed training environment for reinforcement learning on Google Kubernetes Engine (GKE). You use Ray and the verl (Volcano Engine Reinforcement Learning) framework to set up a distributed training environment to fine-tune a Qwen2.5-32B-Instruct model on the GSM8K dataset.
This tutorial focuses on the Group Relative Policy Optimization (GRPO) training pipeline on GKE with Ray and verl. GRPO is a reinforcement learning algorithm designed to improve a model's reasoning ability. This memory-efficient algorithm simplifies the reinforcement learning (RL) process by eliminating the Critic, or value model, and using a relative group-based calculation instead.
This tutorial is a good starting point if you need to set up a distributed training environment where data, model weights, and the training engine are decoupled for efficiency.
This tutorial supports the following GPU architectures:
- Intel or AMD-based GPU nodes: Set up and scale using NVIDIA B200 or H200 GPUs, using GKE Dynamic Resource Allocation (DRA) for Autopilot path.
- Arm-based A4X (GB200) nodes: Set up and scale using NVIDIA GB200 Grace Blackwell Superchips, using GKE Dynamic Resource Allocation (DRA) and Multi-Node NVLink (IMEX).
Background
The following sections provide a brief overview of the concepts used in this tutorial.
Reinforcement learning (RL)
RL teaches models through experience, exploration, and feedback rather than static imitation. Although pre-training teaches a model what to say, reinforcement learning from human feedback (RLHF) teaches it how to be helpful, safe, and logical. RL serves as the bridge between a base model and a fine-tuned model for a specialized use case.
For more information, see What is reinforcement learning?
Group Relative Policy Optimization (GRPO)
GRPO, an algorithm popularized by DeepSeek, offers a memory-efficient alternative to Proximal Policy Optimization (PPO) for LLM alignment by removing the Critic model. Instead of a Critic network, GRPO generates a group of responses for the same prompt and uses the average reward of that group as the baseline.
For more information, see GRPO.
Volcano Engine Reinforcement Learning (verl)
verl is a high-performance framework designed to handle the complex memory and compute patterns of LLM-based RL.
For more information, see verl.
Objectives
This tutorial shows you how set up reinforcement learning on GKE with verl, by completing the following steps:
- Set up a GKE cluster with A4X (GB200 Superchips), A4 (B200 GPUs), or A3 Ultra (H200 GPUs).
- Configure KubeRay to manage a distributed Ray cluster.
- Use Cloud Storage FUSE to mount a Cloud Storage bucket across all nodes.
- Run a GRPO training job using verl to align the Qwen2.5-32B-Instruct model with the GSM8K dataset.
Before you begin
- Sign in to your Google Cloud account. If you're new to Google Cloud, create an account to evaluate how our products perform in real-world scenarios. New customers also get $300 in free credits to run, test, and deploy workloads.
-
Install the Google Cloud CLI.
-
If you're using an external identity provider (IdP), you must first sign in to the gcloud CLI with your federated identity.
-
To initialize the gcloud CLI, run the following command:
gcloud init -
Create or select a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Create a Google Cloud project:
gcloud projects create PROJECT_ID
Replace
PROJECT_IDwith a name for the Google Cloud project you are creating. -
Select the Google Cloud project that you created:
gcloud config set project PROJECT_ID
Replace
PROJECT_IDwith your Google Cloud project name.
-
Verify that billing is enabled for your Google Cloud project.
Enable the required APIs:
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.gcloud services enable container.googleapis.com
storage.googleapis.com compute.googleapis.com -
Install the Google Cloud CLI.
-
If you're using an external identity provider (IdP), you must first sign in to the gcloud CLI with your federated identity.
-
To initialize the gcloud CLI, run the following command:
gcloud init -
Create or select a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Create a Google Cloud project:
gcloud projects create PROJECT_ID
Replace
PROJECT_IDwith a name for the Google Cloud project you are creating. -
Select the Google Cloud project that you created:
gcloud config set project PROJECT_ID
Replace
PROJECT_IDwith your Google Cloud project name.
-
Verify that billing is enabled for your Google Cloud project.
Enable the required APIs:
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.gcloud services enable container.googleapis.com
storage.googleapis.com compute.googleapis.com -
Grant roles to your user account. Run the following command once for each of the following IAM roles:
roles/container.admin, roles/iam.serviceAccountAdmin, roles/storage.admingcloud projects add-iam-policy-binding PROJECT_ID --member="user:USER_IDENTIFIER" --role=ROLE
Replace the following:
PROJECT_ID: Your project ID.USER_IDENTIFIER: The identifier for your user account. For example,myemail@example.com.ROLE: The IAM role that you grant to your user account.
- Create a Hugging Face account, if you don't already have one.
- Ensure that you have a Hugging Face token.
- Ensure your project has sufficient quota for A4X (GB200 Superchips), A4 (B200 GPUs), or A3 Ultra (H200 GPUs). To learn more, see Plan GPU quota and GPU quota.
- Ensure that you have an active capacity reservation for your GPU machine type. For more information, see Reserve capacity through your account team.
- For the A4X (GB200) setup, ensure that you have Helm installed.
Prepare your environment
In this tutorial, you use Cloud Shell.
Go to the Google Cloud console.
At the top of the Google Cloud console window, click the Activate Cloud Shell button.
Set the environment variables:
A4 and A3 Ultra
Autopilot
Standard
A4X
export PROJECT_ID=$(gcloud config get project) export PROJECT_NUMBER=$(gcloud projects describe ${PROJECT_ID} --format="value(projectNumber)") export CONTROL_PLANE_REGION=YOUR_REGION export NODE_ZONE=YOUR_ZONE export CLUSTER_NAME=YOUR_CLUSTER_NAME export KSA_NAME=YOUR_KSA_NAME export GS_BUCKET=YOUR_GCS_BUCKET-${PROJECT_ID} export NAMESPACE=default export GPU_TYPE=YOUR_GPU_TYPE export MACHINE_TYPE=YOUR_MACINE_TYPE export RESERVATION=YOUR_RESERVATION_NAME export HF_TOKEN=YOUR_HF_TOKEN # A4X (GB200 Superchips) only variables export NUM_GPU_NODES=4 export VERL_IMAGE=verlai/verl:vllm023.aarch64.dev1 export VERL_REF=ddbcdb7Replace the following values:
YOUR_REGION: the Compute Engine region for the GKE cluster control plane.YOUR_ZONE: the zone where the nodes are reserved. For more information, see GPU availability.YOUR_CLUSTER_NAME: the name of your GKE cluster.YOUR_KSA_NAME: the name of your Kubernetes service account.YOUR_GCS_BUCKET: the base name for your Cloud Storage bucket. You don't need to specify thegs://prefix.YOUR_GPU_TYPE: the accelerator that you reserved in the Compute Engine capacity reservation. It must be one of the following values:nvidia-gb200: A4X (GB200 Superchips)nvidia-b200: A4 (B200 GPUs)nvidia-h200-141gb: A3 Ultra (H200 GPUs)
YOUR_MACHINE_TYPE: the type of machine to use:- For A4X (GB200 Superchips), use
a4x-highgpu-4g. - For A4 (B200 GPUs), use
a4-highgpu-8gor later. - For A3 Ultra (H200 GPUs), use
a3-ultragpu-8gor later.
- For A4X (GB200 Superchips), use
YOUR_RESERVATION_NAME: the name of your capacity reservation.YOUR_HF_TOKEN: your Hugging Face token.- Google Kubernetes Engine (GKE) Standard edition only:
GVNIC_NAME(GKE Standard - A4 or A3 Ultra only): the prefix for the gVNIC network name. You can use any prefix you want.RDMA_NAME(A4 or A3 Ultra only): the prefix for the remote direct memory access (RDMA) network. You can use any prefix you want.
Clone the sample repository:
Navigate to the working directory for your chosen GKE mode:
A4 and A3 Ultra
Autopilot
Standard
A4X
No directory change is required. You can proceed directly to the next section.
Set up infrastructure
In this section, you create standard VPC networks and the GKE cluster.
Create RDMA network and subnets (GKE Standard - A4 and A3 Ultra only)
A4 and A3 Ultra
Autopilot
This section is required for GKE Standard A4 and A3 Ultra GPUs only.
If you use Autopilot, skip this section and proceed directly to Create a GKE cluster. GKE automatically provisions the necessary VPC networks and subnets, and uses GKE managed DRANET to allocate these resources to your Pods. You don't need to manually create any network infrastructure.
Standard
Create a VPC network for the gVNIC interface:
Create a VPC network for RDMA:
Create the 8 RDMA subnets for the 8 GPUs:
A4X
This section is required for GKE Standard A4 and A3 Ultra GPUs only.
If you use A4X (GB200) GPUs, skip this section and proceed directly to
Create a GKE cluster. For A4X (GB200) GPUs or Autopilot, GKE
creates the networks automatically when the node pool uses the auto
accelerator network profile. The Cluster Toolkit blueprint
enables this profile by using the enable_dranet:true flag.
Create a GKE cluster
Create a GKE cluster corresponding to your GPU architecture:
A4 and A3 Ultra
Select the GKE cluster mode that you want to use:
Autopilot
Create an Autopilot cluster:
Get credentials for your cluster:
Standard
Create a Standard cluster:
Get credentials for your cluster:
Create the GPU node pool. These node pools use your reservation to ensure availability. You start with two nodes:
Install the NCCL RDMA installer used for Standard clusters:
A4X
Create a GKE cluster and node pool by using the Cluster Toolkit
gke-a4xblueprint. The blueprint provisions the GKE cluster, including the A4X node pool bound to your reservation, the accelerator networks (one additional gVNIC plus four RDMA rails), and the managed DRANET driver that exposes the CX-7 NICs as DRA devices.Use the blueprint's deployment instructions to configure your parameters (such as
PROJECT_ID,CONTROL_PLANE_REGION,NODE_ZONE, reservation, andNUM_GPU_NODES), then deploy the cluster. Alternatively, you can follow the A4X GKE cluster creation guide to create a cluster manually.- Get credentials for your cluster:
gcloud container clusters get-credentials ${CLUSTER_NAME} --location=${CONTROL_PLANE_REGION}Verify that the cluster exposes RDMA NICs through DRA:
kubectl get deviceclassesThe output must include
mrdma.google.com.Verify that the A4X nodes are present:
kubectl get nodes -l cloud.google.com/gke-accelerator=nvidia-gb200Install the gIB NCCL plugin (A4X variant):
kubectl apply -f https://raw.githubusercontent.com/GoogleCloudPlatform/container-engine-accelerators/master/gpudirect-rdma/nccl-rdma-installer-a4x.yamlInstall the NVIDIA DRA driver, which provides
ComputeDomain(IMEX) channels for multi-node NVLink:helm repo add nvidia https://helm.ngc.nvidia.com/nvidia && helm repo update kubectl create namespace nvidia-dra-driver-gpu kubectl apply -f - <<EOF apiVersion: v1 kind: ResourceQuota metadata: name: nvidia-dra-driver-gpu-quota namespace: nvidia-dra-driver-gpu spec: hard: pods: "$((2 * NUM_GPU_NODES + 1))" scopeSelector: matchExpressions: - operator: In scopeName: PriorityClass values: - system-node-critical - system-cluster-critical EOF helm upgrade --install nvidia-dra-driver-gpu nvidia/nvidia-dra-driver-gpu \ --version=25.3.1 --namespace nvidia-dra-driver-gpu \ --set nvidiaDriverRoot=/home/kubernetes/bin/nvidia \ --set resources.gpus.enabled=false \ --set kubeletPlugin.tolerations[0].key=nvidia.com/gpu \ --set kubeletPlugin.tolerations[0].operator=Exists \ --set kubeletPlugin.tolerations[1].key=kubernetes.io/arch \ --set kubeletPlugin.tolerations[1].operator=ExistsInstall the KubeRay operator, scoped to the workload namespace:
kubectl create namespace ${NAMESPACE} helm repo add kuberay https://ray-project.github.io/kuberay-helm/ && helm repo update helm upgrade --install kuberay-operator kuberay/kuberay-operator \ --namespace ${NAMESPACE} \ --set singleNamespaceInstall=true --set "watchNamespace={${NAMESPACE}}"
Configure network mappings (GKE Standard - A4 and A3 Ultra only)
A4 and A3 Ultra
Autopilot
This step is required for GKE Standard GPU setups (A4 and A3 Ultra only). If you use A4X (GB200), GKE manages network interfaces automatically, so skip this section.
Standard
Inspect the
network-mapping.yamlmanifest:Apply the manifest:
A4X
This step is required for GKE Standard GPU setups (A4 and A3 Ultra only). If you use A4X (GB200), GKE manages network interfaces automatically, so skip this section.
Prepare data and storage
Configure Cloud Storage and Kubernetes resources:
Create a Cloud Storage bucket:
Create a Kubernetes Service Account (KSA) and bind it to the bucket:
Create the Secret for Hugging Face:
Inspect the
gcsfuse-storage.yamlmanifest:Apply the manifest:
Setup DRANET
Configure your DRANET:
A4 and A3 Ultra
Autopilot
Create the ComputeClass manifest:
Apply both the
computeclass-dranet.yamlmanifest (created in the previous step) and theresourceclaim-dranet.yamlmanifest (included in the sample repository):
Standard
No DRANET setup is required. You can proceed directly to the next section.
A4X
DRANET is set by Cluster Toolkit. You can proceed directly to the next section.
Prepare model and data
Populate your Cloud Storage bucket with the model weights and datasets. You can run these commands locally or on a GKE Pod to populate the bucket:
A4 and A3 Ultra
Autopilot
Inspect the data prep job:
Launch the job:
Monitor the job:
Standard
Inspect the data prep job:
Launch the job:
Monitor the job:
A4X
Clone the verl repository, prepare the virtual environment, and process the GSM8K dataset:
git clone https://github.com/volcengine/verl.git git -C verl checkout ${VERL_REF} VENV_DIR=.venv python3 -m venv $VENV_DIR source $VENV_DIR/bin/activate pip install verl python verl/examples/data_preprocess/gsm8k.py --local_save_dir ~/data/gsm8kDownload the Qwen2.5-32B-Instruct model by using the Hugging Face CLI (this download requires around 66 GB of disk space):
hf download Qwen/Qwen2.5-32B-Instruct --local-dir Qwen2.5-32B-InstructUpload the model, data, and the verl code to your Cloud Storage bucket:
gcloud storage cp --recursive verl gs://${GS_BUCKET}/verl gcloud storage cp --recursive Qwen2.5-32B-Instruct gs://${GS_BUCKET}/Qwen2.5-32B-Instruct gcloud storage cp --recursive ~/data/gsm8k/* gs://${GS_BUCKET}/gsm8k/
Deploy RayCluster custom resource
Deploy a RayCluster custom resource, which consists of one system head Pod and multiple GPU-backed worker Pods.
A4 and A3 Ultra
Select the GKE cluster mode that you used to create your cluster:
Autopilot
Inspect the RayCluster workload:
Apply the RayCluster:
Standard
Inspect the RayCluster workload:
Apply the RayCluster:
A4X
Create the RDMA
ResourceClaimTemplateand the NVIDIAComputeDomain. Each GPU worker Pod claims four RDMA NICs (all rails of its node) and one IMEX channel. Save the following manifest tocompute-domain-a4x.yaml:apiVersion: resource.k8s.io/v1 kind: ResourceClaimTemplate metadata: name: verl-rdma-nic namespace: ${NAMESPACE} spec: spec: devices: requests: - name: nic exactly: deviceClassName: mrdma.google.com allocationMode: ExactCount count: 1 --- apiVersion: resource.nvidia.com/v1beta1 kind: ComputeDomain metadata: name: verl-compute-domain namespace: ${NAMESPACE} spec: numNodes: ${NUM_GPU_NODES} channel: resourceClaimTemplate: name: verl-compute-domain-channelApply the manifest:
kubectl apply -f compute-domain-a4x.yamlDeploy the RayCluster. The Ray head Pod runs on an A4X node without requesting GPUs (since the image is
arm64-only). Save the following configurations toray-cluster-a4x.yaml:apiVersion: ray.io/v1 kind: RayCluster metadata: name: gb200-ray-cluster namespace: ${NAMESPACE} spec: rayVersion: '2.49.0' headGroupSpec: rayStartParams: dashboard-host: '0.0.0.0' num-cpus: "0" template: metadata: annotations: gke-gcsfuse/volumes: "true" spec: serviceAccountName: ${KSA_NAME} nodeSelector: cloud.google.com/gke-accelerator: nvidia-gb200 tolerations: - key: nvidia.com/gpu operator: Exists effect: NoSchedule - key: kubernetes.io/arch operator: Exists effect: NoSchedule containers: - name: ray-head image: ${VERL_IMAGE} lifecycle: postStart: exec: command: - /bin/bash - -c - pip3 install --quiet TransferQueue==0.1.8 ports: - containerPort: 6379 name: gcs-server - containerPort: 8265 name: dashboard - containerPort: 10001 name: client resources: limits: cpu: "12" memory: 32Gi ephemeral-storage: 20Gi requests: cpu: "12" memory: 32Gi ephemeral-storage: 20Gi volumeMounts: - mountPath: /tmp/ray name: ray-logs - name: training-bucket-vol mountPath: /data volumes: - name: ray-logs emptyDir: {} - name: training-bucket-vol persistentVolumeClaim: claimName: training-bucket-pvc workerGroupSpecs: - replicas: ${NUM_GPU_NODES} minReplicas: ${NUM_GPU_NODES} maxReplicas: ${NUM_GPU_NODES} groupName: gpu-group rayStartParams: num-cpus: "120" template: metadata: annotations: gke-gcsfuse/volumes: "true" spec: serviceAccountName: ${KSA_NAME} nodeSelector: cloud.google.com/gke-accelerator: nvidia-gb200 affinity: podAntiAffinity: requiredDuringSchedulingIgnoredDuringExecution: - labelSelector: matchLabels: ray.io/group: gpu-group topologyKey: kubernetes.io/hostname tolerations: - key: nvidia.com/gpu operator: Exists effect: NoSchedule - key: kubernetes.io/arch operator: Exists effect: NoSchedule containers: - name: ray-worker image: ${VERL_IMAGE} lifecycle: postStart: exec: command: - /bin/bash - -c - pip3 install --quiet TransferQueue==0.1.8 env: - name: LD_LIBRARY_PATH value: /usr/local/nvidia/lib64 resources: limits: cpu: "120" memory: 600Gi nvidia.com/gpu: "4" ephemeral-storage: 500Gi requests: cpu: "120" memory: 600Gi nvidia.com/gpu: "4" ephemeral-storage: 500Gi claims: - name: rdma-nic-0 - name: rdma-nic-1 - name: rdma-nic-2 - name: rdma-nic-3 - name: compute-domain-channel volumeMounts: - name: nvidia mountPath: /usr/local/nvidia - name: gib mountPath: /usr/local/gib - name: shared-memory mountPath: /dev/shm - name: ray-tmp-storage mountPath: /tmp - name: training-bucket-vol mountPath: /data resourceClaims: - name: rdma-nic-0 resourceClaimTemplateName: verl-rdma-nic - name: rdma-nic-1 resourceClaimTemplateName: verl-rdma-nic - name: rdma-nic-2 resourceClaimTemplateName: verl-rdma-nic - name: rdma-nic-3 resourceClaimTemplateName: verl-rdma-nic - name: compute-domain-channel resourceClaimTemplateName: verl-compute-domain-channel volumes: - name: gib hostPath: path: /home/kubernetes/bin/gib - name: nvidia hostPath: path: /home/kubernetes/bin/nvidia - name: shared-memory emptyDir: medium: Memory sizeLimit: 200Gi - name: ray-tmp-storage emptyDir: {} - name: training-bucket-vol persistentVolumeClaim: claimName: training-bucket-pvcApply the RayCluster manifest:
envsubst < ray-cluster-a4x.yaml | kubectl apply -f -Wait until one head Pod and four worker Pods are in the
Runningstate:kubectl get pods -w
Launch the GRPO Job
Configure and submit the reinforcement learning training job:
A4 and A3 Ultra
Set up the Ray Client:
Recover the Ray Head Service:
Set up port forwarding to the Ray dashboard node. Use a separate terminal window for this step because this command blocks the terminal while it runs. Use
Control+C to stop it: Inspect the
runtime-env.yamlmanifest:If you use H200 GPUs, change
NCCL_TUNER_CONFIG_PATHto/usr/local/gib/configs/tuner_config_a3u.txtpb.This file is used by the Ray client. You don't need to apply this manifest to the cluster.
Submit the Job using
ray job submit:Monitor the logs in the Ray Dashboard or console output. Look for
critic/score/meanto increase, indicating learning.After the training finishes, the checkpoints of the trained model can be found in
gs://$GS_BUCKET/verl/checkpoints.
A4X
Get the Ray head Pod name:
export HEAD_POD=$(kubectl get pod -n ${NAMESPACE} -l ray.io/node-type=head -o jsonpath='{.items[0].metadata.name}')Configure the Ray runtime environment file directly on the head Pod:
kubectl exec ${HEAD_POD} -c ray-head -- bash -c 'mkdir -p /tmp/submit && cat > /tmp/submit/runtime-env.yaml <<EOF working_dir: "." env_vars: PYTHONPATH: "/data/verl" LD_LIBRARY_PATH: "/usr/local/nvidia/lib64:/usr/local/gib/lib64" NCCL_DEBUG: "INFO" NCCL_ENV_PLUGIN: "gcp" HF_HOME: "/data/huggingface_cache" GLOO_SOCKET_IFNAME: "eth0" EOF'Submit the GRPO training job by execution on the Ray head Pod:
kubectl exec ${HEAD_POD} -c ray-head -- bash -c 'cd /tmp/submit && \ ray job submit --runtime-env runtime-env.yaml --no-wait -- \ python3 -m verl.trainer.main_ppo \ algorithm.adv_estimator=grpo \ data.train_files=/data/gsm8k/train.parquet \ data.val_files=/data/gsm8k/test.parquet \ data.train_batch_size=256 \ data.max_prompt_length=512 \ data.max_response_length=512 \ actor_rollout_ref.model.path=/data/Qwen2.5-32B-Instruct \ actor_rollout_ref.actor.optim.lr=1e-5 \ actor_rollout_ref.actor.ppo_mini_batch_size=64 \ actor_rollout_ref.actor.ppo_micro_batch_size_per_gpu=8 \ actor_rollout_ref.actor.use_kl_loss=True \ actor_rollout_ref.actor.strategy=fsdp2 \ actor_rollout_ref.rollout.name=vllm \ actor_rollout_ref.rollout.tensor_model_parallel_size=4 \ actor_rollout_ref.rollout.gpu_memory_utilization=0.6 \ actor_rollout_ref.rollout.n=8 \ actor_rollout_ref.rollout.log_prob_micro_batch_size_per_gpu=16 \ actor_rollout_ref.ref.log_prob_micro_batch_size_per_gpu=16 \ algorithm.kl_ctrl.kl_coef=0.001 \ trainer.logger=console \ trainer.n_gpus_per_node=4 \ trainer.nnodes=4 \ trainer.save_freq=10 \ trainer.test_freq=10 \ trainer.total_epochs=2 \ trainer.default_local_dir=/data/verl/checkpoints'Monitor the job logs (using the unique ID returned by
ray job submit):kubectl exec ${HEAD_POD} -c ray-head -- ray job logs <var>JOB_ID</var> --followReplace
JOB_ID. Confirm that cross-node NVLink is active by looking for NCCL lines in logs containingvia P2P/MNNVL.
Clean up
To avoid incurring charges, delete the resources:
A4 and A3 Ultra
Autopilot
Delete the Ray cluster:
Delete the Cloud Storage FUSE:
Delete the DRANET resources:
Delete the Cloud Storage bucket:
Delete the GKE cluster:
Standard
Delete the Ray cluster:
Delete the Cloud Storage FUSE:
Delete the Cloud Storage bucket:
Delete the GKE cluster:
Delete the VPC networks and subnets:
A4X
kubectl delete raycluster gb200-ray-cluster
kubectl delete computedomain verl-compute-domain
gcloud storage rm -r gs://${GS_BUCKET}
gcloud container clusters delete ${CLUSTER_NAME} --location=${CONTROL_PLANE_REGION}