Clustered GPUs overview

Clustered GPUs are optimized for artificial intelligence (AI) and machine learning (ML) workloads that prioritize massive-scale distributed training, multi-host inference, and complex high performance computing (HPC) tasks.

Clustered GPUs let teams deploy and scale thousands of interconnected accelerators as a single, tightly coupled system by using specialized networking fabrics and synchronized maintenance. They are ideal for workloads that require extreme compute, memory, and high-speed networking synchronization. Common use cases include pre-training foundation models with trillions of parameters, frontier model serving, and simulations for drug discovery.

Key characteristics of Clustered GPUs

  • Synchronized maintenance: Google Cloud updates all nodes in the cluster simultaneously, preventing a single node restart from stalling a job.
  • Automated health monitoring: the system proactively identifies stragglers and monitors for hardware failures or silent data corruption.
  • Proactive node replacement: the system evicts and reschedules underperforming VMs on healthy machines from a spare pool.
  • Rail-aligned topology: a rail-aligned topology isolates GPU-to-GPU traffic from standard host communication to help ensure jitter-free performance for massive multi node coordination.
  • Advanced protocols: these machine series support high-bandwidth interconnects including GPUDirect RDMA (based on RoCE).

Clustered GPU families and technical specifications

Review the workloads and hardware specifications for the premium Clustered GPU machine series in these tables. For exascale training, use the A4X or A3 Ultra series.

A4X Max and A4X series

If you are training massive exascale foundation models or running highly complex, bare-metal high performance computing (HPC) simulations, use the A4X Max or A4X machine series. These series feature NVIDIA GB300 Grace Blackwell Ultra Superchips to handle extreme computational and memory demands.

A4X Max (Bare metal)

A4X Max machine types use NVIDIA GB300 Grace Blackwell Ultra Superchips (nvidia-gb300) and are ideal for foundation model training and serving. A4X Max machine types are available as bare metal instances.

A4X Max is an exascale platform based on NVIDIA GB300 NVL72. Each machine has two sockets with NVIDIA Grace CPUs with Arm Neoverse V2 cores. These CPUs are connected to four NVIDIA B300 Blackwell GPUs with fast chip-to-chip (NVLink-C2C) communication.

Attached NVIDIA GB300 Grace Blackwell Ultra Superchips
Machine type vCPU count1 Instance memory (GB) Attached Local SSD (GiB) Physical NIC count Maximum network bandwidth (Gbps)2 GPU count GPU memory3
(GB HBM3e)
a4x-maxgpu-4g-metal 144 960 12,000 6 3,600 4 1,116

1A vCPU is implemented as a single hardware hyper-thread on one of the available CPU platforms.
2Maximum egress bandwidth cannot exceed the number given. Actual egress bandwidth depends on the destination IP address and other factors. For more information about network bandwidth, see Network bandwidth.
3GPU memory is the memory on a GPU device that can be used for temporary storage of data. It is separate from the instance's memory and is specifically designed to handle the higher bandwidth demands of your graphics-intensive workloads.

A4X

A4X machine types use NVIDIA GB200 Grace Blackwell Superchips (nvidia-gb200) and are ideal for foundation model training and serving.

A4X is an exascale platform based on NVIDIA GB200 NVL72. Each machine has two sockets with NVIDIA Grace CPUs with Arm Neoverse V2 cores. These CPUs are connected to four NVIDIA B200 Blackwell GPUs with fast chip-to-chip (NVLink-C2C) communication.

Attached NVIDIA GB200 Grace Blackwell Superchips
Machine type vCPU count1 Instance memory (GB) Attached Local SSD (GiB) Physical NIC count Maximum network bandwidth (Gbps)2 GPU count GPU memory3
(GB HBM3e)
a4x-highgpu-4g 140 884 12,000 6 2,000 4 744

1A vCPU is implemented as a single hardware hyper-thread on one of the available CPU platforms.
2Maximum egress bandwidth cannot exceed the number given. Actual egress bandwidth depends on the destination IP address and other factors. For more information about network bandwidth, see Network bandwidth.
3GPU memory is the memory on a GPU device that can be used for temporary storage of data. It is separate from the instance's memory and is specifically designed to handle the higher bandwidth demands of your graphics-intensive workloads.

A4 series

If your goal is to train state-of-the-art foundation models requiring massive parallel processing scale, use the A4 machine series to optimize processing time and maximize hardware throughput.

A4 machine types have NVIDIA B200 Blackwell GPUs (nvidia-b200) attached and are ideal for foundation model training and serving.

Attached NVIDIA B200 Blackwell GPUs
Machine type vCPU count1 Instance memory (GB) Attached Local SSD (GiB) Physical NIC count Maximum network bandwidth (Gbps)2 GPU count GPU memory3
(GB HBM3e)
a4-highgpu-8g 224 3,968 12,000 10 3,600 8 1,440

1A vCPU is implemented as a single hardware hyper-thread on one of the available CPU platforms.
2Maximum egress bandwidth cannot exceed the number given. Actual egress bandwidth depends on the destination IP address and other factors. For more information about network bandwidth, see Network bandwidth.
3GPU memory is the memory on a GPU device that can be used for temporary storage of data. It is separate from the instance's memory and is specifically designed to handle the higher bandwidth demands of your graphics-intensive workloads.

A3 series (Ultra, Mega, and High with 8 GPUs)

If you need to execute large-scale distributed training runs or serve frontier models across multiple hosts with large dataset tables, use the A3 machine series to minimize host-to-host network latency.

A3 Ultra

Attached NVIDIA H200 GPUs
Machine type vCPU count1 Instance memory (GB) Attached Local SSD (GiB) Physical NIC count Maximum network bandwidth (Gbps)2 GPU count GPU memory3
(GB HBM3e)
a3-ultragpu-8g 224 2,952 12,000 10 3,600 8 1128

A3 Mega

Attached NVIDIA H100 GPUs
Machine type vCPU count1 Instance memory (GB) Attached Local SSD (GiB) Physical NIC count Maximum network bandwidth (Gbps)2 GPU count GPU memory3
(GB HBM3)
a3-megagpu-8g 208 1,872 6,000 9 1,800 8 640

A3 High

Attached NVIDIA H100 GPUs
Machine type vCPU count1 Instance memory (GB) Attached Local SSD (GiB) Physical NIC count Maximum network bandwidth (Gbps)2 GPU count GPU memory3
(GB HBM3)
a3-highgpu-1g 26 234 750 1 25 1 80
a3-highgpu-2g 52 468 1,500 1 50 2 160
a3-highgpu-4g 104 936 3,000 1 100 4 320
a3-highgpu-8g 208 1,872 6,000 5 1,000 8 640

1A vCPU is implemented as a single hardware hyper-thread on one of the available CPU platforms.
2Maximum egress bandwidth cannot exceed the number given. Actual egress bandwidth depends on the destination IP address and other factors. For more information about network bandwidth, see Network bandwidth.
3GPU memory is the memory on a GPU device that can be used for temporary storage of data. It is separate from the instance's memory and is specifically designed to handle the higher bandwidth demands of your graphics-intensive workloads.

Capacity acquisition

Acquiring compute capacity for Clustered GPUs requires coordination to manage the supply and demand of high-performance accelerators. Because these resources are physically located within specialized networking fabrics, you must request them through managed workflows.

You can reserve capacity through your account team or use future reservations and Dynamic Workload Scheduler.

What's next