Failure recovery of TPU instances
To keep your workloads running, Compute Engine recovers TPU instances and slices from hardware and application failures.
Failure recovery behavior depends on your capacity mode:
- Managed capacity mode (on-demand, flex-start, and standard reservations): Compute Engine automatically recovers TPU instances and slices by restarting them on healthy hardware.
- All Capacity mode (reservations in All Capacity mode): You manage TPU slice recovery. Hardware failures are repaired on the existing host without automatically relocating instances to new hardware.
Multi-host slice recovery requirements
Multi-host TPU slices require all instances in the slice to be recovered or rescheduled together. Whether Compute Engine automatically handles recovery in managed capacity mode or you recover slices manually in All Capacity mode, you can't reschedule individual instances within a slice.
Instances in a multi-host slice are connected by inter-chip interconnect (ICI) and, in some cases, by an Optical Circuit Switch (OCS). When you provision a multi-host slice, Compute Engine creates TPU instances and activates the ICI network. You specify the network topology of the slice by setting the accelerator topology in your workload policy. When a workload starts, LibTPU, the foundational TPU software layer, initializes the network topology. By default, the XLA compiler then statically maps the model's operations to this topology.
If an instance or network link fails, then the physical hardware no longer matches the mapped topology. As a result, the XLA compiler can't execute the workload until the system reconfigures the network topology by recreating or rescheduling all instances in the slice.
Failure recovery in managed capacity mode
In managed capacity mode (including on-demand, flex-start, and standard reservations), Compute Engine automatically recovers TPU instances and slices from hardware and host failures by restarting them on healthy hardware.
Automatic recovery of a single-host TPU slice
Single-host slices are independent TPU instances. By default, Compute Engine
automatically recovers failed instances by restarting the instances on healthy
hardware. This behavior is controlled by the automatic restart setting, which is
enabled by default when creating instances, except for Spot VMs. If you
disable automatic restart, then an instance failure causes the instance to enter
the TERMINATED state. For more information, see Automatic
restart.
Compute Engine automatically recovers a failed instance in scenarios such as:
- A host timeout or error caused by the physical machine not responding, host shutdown, host reboot, or host power outage
- Physical host maintenance events initiated by you or Google
- Inter-chip interconnect (ICI) failure within a host
- VM crash
Compute Engine doesn't automatically recover instances in scenarios of planned termination, including:
- Instance deletion
- Reservation deletion or expiration
- Spot VMs preemption
MIG-initiated repair in single-host slices
In a MIG with single-host slices, if a TPU instance in a single-host
slice enters a TERMINATED state due to hardware failures or external events
like Spot VMs preemption, then the MIG repairs the instance by default.
During a repair, the MIG recreates the instance with the same name. You can
disable this repair mechanism if you turn off
repairs.
You can also set up an application-based health check in a MIG with single-host slices. If the health check detects that your application is unresponsive, then the MIG marks the instance as unhealthy and autoheals the instance by recreating it.
For more information, see About repairing VMs for high availability and Set up an application health check and autohealing.
Automatic recovery of a multi-host slice
For TPUs in managed mode using the on-demand, flex-start, or reservation consumption models, Compute Engine automatically recovers failed instances in a multi-host slice.
During recovery, Compute Engine identifies a set of TPU machines that can form the network topology, restarts all instances in the slice together on those machines, and reconfigures the network. This process minimizes downtime by recreating the topology on available healthy hardware, rather than waiting for hardware repairs.
Recovery process and slice states
During automatic recovery, the slice transitions through the following states:
- The slice transitions to the
REACTIVATINGstate. - All instances in the slice transition to the
REPAIRINGstate, though not necessarily at the same time. - Compute Engine restarts all instances in the slice together on healthy hardware.
For more information about TPU slice states, see Accelerator topology states.
Scenarios requiring manual slice recovery in managed capacity mode
Compute Engine can't automatically recover a multi-host slice in the following scenarios:
- Spot VMs preemption: If any instance in the slice is preempted,
Compute Engine terminates all instances in the slice, and the slice enters
the
FAILEDstate. - User-initiated interruptions: If you stop or delete a TPU instance, or
stop an instance from within the operating system, then the slice enters the
FAILEDstate. The slice remains in theFAILEDstate until you recreate it.
In these scenarios, you must manually recover the slice.
Failure recovery in All Capacity mode
In All Capacity mode, you are responsible for managing the TPU slice recovery process. Unlike managed capacity mode, Compute Engine doesn't automatically relocate failed TPU instances or multi-host slices to new physical hardware. Google repairs the underlying failed hardware on the existing host, and you're responsible for rescheduling unhealthy slices to healthy spare hardware that you set aside in your reservation.
Host failure and faulty host repair
If a host failure occurs or if you report a VM host as faulty, the host repair process works as follows:
- The affected VM transitions to the
REPAIRINGstate while the physical hardware is repaired. - Once the underlying hardware is fixed, the VM transitions back to the
RUNNINGstate on the same host. However, if the VM belongs to a multi-host slice, then the slice remains in the 'FAILED' state. You must manually recover the slice.
For more information, see Report and repair faulty TPU hosts in All Capacity mode.
Emergent maintenance
During emergent maintenance events or when you manually initiate a maintenance event:
- The VMs transition from the
RUNNINGstate to theREPAIRINGstate. - After maintenance finishes, the VMs return to the
RUNNINGstate on the same host.
For more information, see Manage maintenance events in All Capacity mode.
Multi-host slice failure scenarios
In All Capacity mode, Compute Engine doesn't automatically recover a
multi-host slice. A slice transitions to the FAILED state in the following
scenarios:
- ICI failure: The VMs remain in the
RUNNINGstate, but the slice state transitions to theFAILEDstate. - Spot VMs preemption: If any instance in the slice is preempted,
Compute Engine terminates all instances in the slice, and the slice enters
the
FAILEDstate. - User-initiated interruptions: If you stop or delete a TPU instance, or
stop an instance from within the operating system, then the slice enters the
FAILEDstate.
When a slice enters the FAILED state, you must manually recover the
slice by rescheduling all instances in the slice.
Manually recover a TPU slice
When a TPU slice in managed capacity mode or All Capacity mode is in the
FAILED state, you must manually recover it by rescheduling all instances in
the slice using one of the following methods:
- Resize the
MIG
to target size
0, and then increase it to the required size. - Delete the MIG and recreate the slice.
What's next
To verify failure recovery of TPUs, check the following statuses: