Graphics processing units (GPUs) are specialized hardware accelerators that are designed to process massively parallel workloads simultaneously. On Google Cloud, you can attach NVIDIA GPUs to Compute Engine instances and bare metal instances to accelerate demanding workloads, such as artificial intelligence (AI) and machine learning (ML) training, high performance computing (HPC), graphics rendering, and video transcoding.
This document provides an overview of GPUs in Google Cloud, including available GPU models, machine series mappings, performance characteristics, ideal use cases, and pricing considerations.
What is a GPU?
A GPU is a massively parallel processor that was originally designed for computer graphics and has evolved into an essential computing engine for AI, ML, and scientific computing. Unlike central processing units (CPUs), which contain a few powerful cores optimized for sequential, single-threaded processing, GPUs consist of thousands of smaller, highly efficient cores designed to perform mathematical calculations concurrently across large datasets.
The key architectural components of a GPU system are as follows:
CUDA cores: general-purpose vector processing units that execute instructions across multiple parallel threads simultaneously by using a single instruction, multiple threads (SIMT) architecture. CUDA cores handle standard vector and floating-point computations (such as FP32 and FP64), as well as general graphics rendering and physical simulations.
Tensor Cores: specialized matrix multiply-accumulate (MMA) processing units that are designed to accelerate neural network calculations. Tensor Cores deliver exponential speedups for AI training and inference across specialized numeric precision formats, including FP4, FP8, INT8, TF32, and FP16/BF16.
High-bandwidth memory: dedicated on-device memory that delivers up to several terabytes per second (TBps) of bandwidth, such as High Bandwidth Memory (HBM) or Graphics Double Data Rate (GDDR). This high throughput keeps thousands of GPU cores supplied with data during intensive matrix operations.
High-speed interconnects (NVLink and NVSwitch): high-bandwidth, low-latency interconnect technologies that connect multiple GPUs on the same server or across nodes. NVLink provides direct GPU-to-GPU communication that bypasses the slower PCIe bus, allowing clusters of GPUs to function as a single, coherent accelerator memory pool.
Virtual workstation software (NVIDIA RTX vWS): enterprise software that provides virtualized graphics acceleration for remote visual workstations, computer-aided design (CAD), digital content creation, and 3D modeling.
GPU models and machine series
Google Cloud offers NVIDIA GPUs across multiple generations to meet different workload and budget needs. On Compute Engine, GPUs are available either as pre-configured accelerator-optimized machine series (A and G series) or as attachable accelerators on general-purpose N1 machine series.
The following table summarizes the hardware specifications, machine series mappings, networking bandwidth, and storage capabilities of GPU models available on Compute Engine:
| GPU model | Machine series | Architecture | GPU memory and bandwidth | GPUs per compute instance | Max network bandwidth | Local SSD | Interconnect | NVIDIA RTX Virtual Workstation (vWS) support |
|---|---|---|---|---|---|---|---|---|
| NVIDIA GB300 | A4X Max | Blackwell Ultra | 279 GB HBM3e (8 TBps) | 4 (NVL72) | 3,600 Gbps | 12,000 GiB | NVLink Full Mesh (1,800 GBps) | No |
| NVIDIA GB200 | A4X | Blackwell | 186 GB HBM3e (8 TBps) | 4 (NVL72) | 2,000 Gbps | 12,000 GiB | NVLink Full Mesh (1,800 GBps) | No |
| NVIDIA B200 | A4 | Blackwell | 180 GB HBM3e (8 TBps) | 8 | 3,600 Gbps | 12,000 GiB | NVLink Full Mesh (1,800 GBps) | No |
| NVIDIA H200 | A3 Ultra | Hopper | 141 GB HBM3e (4.8 TBps) | 8 | 3,600 Gbps | 12,000 GiB | NVLink Full Mesh (900 GBps) | No |
| NVIDIA H100 | A3 Mega, High, Edge | Hopper | 80 GB HBM3 (3.35 TBps) | 1, 2, 4, 8 | Up to 1,000 Gbps | Up to 6,000 GiB | NVLink Full Mesh (900 GBps) | No |
| NVIDIA A100 (80GB) | A2 Ultra | Ampere | 80 GB HBM2e (1.9 TBps) | 1, 2, 4, 8 | 100 Gbps | Up to 3,000 GiB | NVLink Full Mesh (600 GBps) | No |
| NVIDIA A100 (40GB) | A2 Standard | Ampere | 40 GB HBM2 (1.6 TBps) | 1, 2, 4, 8, 16 | 100 Gbps | Up to 3,000 GiB | NVLink Full Mesh (600 GBps) | No |
| NVIDIA RTX PRO 6000 | G4 | Blackwell | 96 GB GDDR7 (1.597 TBps) | 1/8, 1/4, 1/2, 1, 2, 4, 8 | 200 Gbps | Up to 3,000 GiB | PCIe Gen 5 | Yes |
| NVIDIA L4 | G2 | Ada Lovelace | 24 GB GDDR6 (300 GBps) | 1, 2, 4, 8 | 100 Gbps | Up to 3,000 GiB | PCIe Gen 4 | Yes |
| NVIDIA T4 (Approaching end of support) |
N1+T4 | Turing | 16 GB GDDR6 (320 GBps) | 1, 2, 4 | 100 Gbps | Supported | PCIe Gen 3 | Yes |
| NVIDIA V100 | N1+V100 | Volta | 16 GB HBM2 (900 GBps) | 1, 2, 4, 8 | 100 Gbps | Supported | NVLink Ring (300 GBps) | No |
| NVIDIA P100 (Approaching end of support) |
N1+P100 | Pascal | 16 GB HBM2 (732 GBps) | 1, 2, 4 | 100 Gbps | Supported | PCIe Gen 3 | Yes |
| NVIDIA P4 (Approaching end of support) |
N1+P4 | Pascal | 8 GB GDDR5 (192 GBps) | 1, 2, 4 | 100 Gbps | Supported | PCIe Gen 3 | Yes |
GPU performance characteristics
Choosing the right GPU model depends on the computational characteristics, memory requirements, precision formats, and scalability needs of your workload.
Ideal use cases and workload requirements
GPUs accelerate diverse computational workloads across AI, visual computing, and scientific modeling. To understand the architectural requirements for each workload category, select one of the following tabs:
AI/ML training
Frontier foundation models and mixture-of-experts (MoE): pre-training massive transformer models, multimodal architectures, and mixture-of-experts (MoE) networks requires high Tensor Core throughput, large High Bandwidth Memory (HBM) capacities (up to 279 GB per GPU), and ultra-high-speed NVLink interconnect bandwidth (up to 1,800 GBps) to minimize multi-node synchronization latency. Recommended machine series:
- A4X Max (NVIDIA GB300)
- A4X (NVIDIA GB200)
- A4 (NVIDIA B200)
- A3 Ultra (NVIDIA H200)
- A3 Mega (NVIDIA H100)
For more information, see Recommendations for pre-training models in the AI Hypercomputer documentation.
Fine-tuning and moderate model training: adapting existing foundation models by using techniques like parameter-efficient fine-tuning (PEFT) and Low-Rank Adaptation (LoRA) requires sufficient GPU memory to hold model weights and gradients without requiring exascale multi-node clustering. Recommended machine series:
For more information, see Recommendations for fine-tuning models in the AI Hypercomputer documentation.
AI/ML inference
Large model inference (more than 70 billion parameters): serving massive large language models (LLMs) benefits from high High Bandwidth Memory (HBM3e) capacity (141–186 GB per GPU) to fit model weights across fewer GPUs, reducing inter-chip communication overhead. Recommended machine series:
High-throughput and real-time serving: low-latency interactive queries, streaming responses, and image generation use hardware-accelerated quantized numeric formats, such as 8-bit floating-point (FP8), 8-bit integer (INT8), and 4-bit integer (INT4), to maximize request throughput per dollar. Recommended machine series:
- G2 (NVIDIA L4)
- G4 (NVIDIA RTX PRO 6000)
- A3 Edge and High (NVIDIA H100)
For more information, see Recommendations for serving inference in the AI Hypercomputer documentation.
Graphics and visual computing
3D CAD, digital twins, and virtual workstations: interactive 3D modeling, computer-aided design (CAD), architectural rendering, and digital twins (such as NVIDIA Omniverse) use dedicated ray tracing (RT) cores and NVIDIA RTX Virtual Workstations (vWS) for remote visual workstations and fractional virtual GPUs (vGPUs). Recommended machine series:
Video processing and transcoding: live streaming pipelines, video transcoding, and media processing use dedicated hardware NVIDIA Encoder (NVENC) and NVIDIA Decoder (NVDEC) video engines supporting AV1, High Efficiency Video Coding (HEVC), and H.264 codecs. Recommended machine series:
HPC and simulation
Scientific simulation and modeling: applications in molecular dynamics (such as GROMACS and AMBER), computational fluid dynamics (CFD), quantum chemistry, astrophysics, and weather forecasting require high 64-bit double-precision (FP64) floating-point performance and high memory bandwidth. Recommended machine series:
For more information, see Recommendations for HPC in the AI Hypercomputer documentation.
Workload considerations
GPUs aren't usually suitable for the following workloads:
- Sequential and single-threaded logic: tasks dominated by sequential decision branches or serial execution pipelines perform better on CPUs with high single-core clock frequencies.
- Standard web applications and microservices: general-purpose web servers, relational database management systems (RDBMS), and standard API backends without AI or graphics processing requirements don't benefit from GPU acceleration. These workloads are more cost-effectively hosted on general-purpose or compute-optimized CPU instances.
- Workloads optimized for XLA and graph-compiled frameworks: if your deep learning models are built by using frameworks like JAX, PyTorch/XLA, or TensorFlow, then they can take advantage of compiler-driven graph optimizations (XLA). If your models also benefit from custom matrix systolic arrays and optical circuit switching, then consider evaluating alternative accelerator architectures designed for those specific execution paradigms.
Pricing considerations
The following factors determine GPU pricing on Google Cloud:
The specific GPU model.
The number of attached GPUs.
Provisioned resources, such as vCPUs, system memory, and storage.
The consumption model that you select.
The following sections outline the available consumption models and discount options for GPUs in Google Cloud.
Consumption and purchasing models
You can provision GPU instances by using several consumption options depending on your availability requirements and cost tolerance:
On-demand: standard pay-as-you-go (PAYG) pricing billed per second with no long-term commitment. On-demand instances offer complete lifecycle flexibility for development, testing, and variable production workloads.
Spot VMs: highly discounted compute instances (up to 60–91% off on-demand rates) that draw from surplus Google Cloud capacity. Because Google Cloud can stop or delete Spot VMs to reclaim capacity with a 30-second notice, these compute instances are ideal for fault-tolerant batch ML training, checkpointed batch jobs, and distributed data processing.
Flex-start VMs: dynamic scheduling option powered by Dynamic Workload Scheduler (DWS) that provisions GPU resources within a specified period for fixed-duration workloads (up to 7 days) at discounted rates.
Reservations:
- Reservations: you can reserve capacity for standard GPUs and immediately use the reserved capacity.
- Future reservations (calendar mode and AI Hypercomputer): you can request to reserve clustered GPU capacity for a future date and time. If Google Cloud approves your request, then you can start using the capacity at your specified time and retain exclusive access until the reservation ends.
Discount options
Committed use discounts (CUDs): provide substantial discounts (up to 55% for 3-year commitments or 37% for 1-year commitments) in exchange for a commitment to maintain a continuous level of resource usage in a specific region. For accelerator-optimized machine types and attached GPUs, CUDs require purchasing resource-based commitments attached to reservations.
Sustained use discounts (SUDs): automatic discounts applied to N1 compute instances and attached GPUs when they run for more than 25% of a billing month. Accelerator-optimized machine series (A and G series) don't receive SUDs.
For detailed hourly and monthly pricing across all regions, see the GPU pricing page and the VM instance pricing page.
For a detailed feature matrix of consumption option availability across all accelerator machine types, see Consumption option availability by machine type.
What's next
- Explore GPU machines in the accelerator-optimized machine family.
- Learn how to Create a VM with attached GPUs.
- Discover Google Cloud's supercomputing AI architecture in the AI Hypercomputer overview.