Terminology

The following terms apply to AI Hypercomputer.

Accelerator
A specialized hardware device, such as a GPU or TPU, designed to perform specific tasks, such as machine learning workloads, more efficiently than a general-purpose processor.

Block
A collection of sub-blocks interconnected by a non-blocking, high-speed network that provides high bandwidth for all host machines in the block.

Cluster
A collection of blocks interconnected by a high-speed network fabric. Each cluster is globally unique. For A4X, A4, A3 Ultra, A3 Mega, and A3 High (8 GPUs) machines, a cluster provides a common, non-blocking network fabric for your blocks of accelerator capacity. Within a cluster, inter-node communication (east-west traffic) is non-blocking for the entire collection of blocks.

Clustered GPU
A category of machine families for GPUs, including A4X, A4, A3 Ultra, A3 Mega, and A3 High (8 GPUs). Clustered GPUs let teams deploy and scale thousands of interconnected accelerators as a single, tightly coupled system by using specialized networking fabrics and synchronized maintenance. They are ideal for workloads that require extreme compute, memory, and high-speed networking synchronization. Common use cases include pre-training foundation models with trillions of parameters, frontier model serving, and simulations for drug discovery.

Emergent maintenance
Sometimes referred to as emergency maintenance. An unplanned maintenance event caused by a critical hardware, software, or security issue. For supported accelerator machine types, emergent maintenance provides an advance notification instead of an immediate disruption. This notification gives you time to gracefully drain your workloads and optionally trigger the repair before the system terminates the VMs.

General GPU
A category of GPU machine families, including G2, G4, A2, N1+T4, and A3 Edge. General GPUs let teams deploy and scale NVIDIA accelerators with complete operational autonomy, bypassing the architectural complexity of coordinated supercomputing clusters. They are ideal for workloads that prioritize high availability, self-service provisioning, and independent scaling. Common use cases include real-time inference, RAG, prototyping, and small-to-medium model training.

Network fabric
High-speed network infrastructure that provides high bandwidth and low latency across all blocks and Google Cloud services in a cluster. The fabric uses the architecture of the Jupiter network. To optimize performance, the fabric uses software-defined network and optical switches.

Node or host
A single physical server machine in the data center. Each host has associated compute resources, such as accelerators. The number and configuration of these compute resources depend on the machine family. Compute Engine instances are provisioned on top of a physical host.

An NVLink domain, also referred to as a sub-block, is the core unit of capacity for A4X Max and A4X machines. An NVLink domain consists of 18 A4X Max or A4X instances (72 GPUs) that are connected by a multi-node NVLink system.

Sub-block
A group of hosts and associated connectivity hardware that are on a single physical block. In the context of A4X Max and A4X machines, a sub-block is also referred to as an NVLink domain.

Superblock
A physical grouping of interconnected machines within data centers. By placing machines inside a superblock, you achieve the extreme network throughput and sub-millisecond latency required to execute tightly coupled training workloads.

More information

The following documents provide further explanations of the terminologies that are relevant to the corresponding topics: