This document describes the different reservation types that you can use to
reserve capacity for Compute Engine instances. To learn more about the
resources that you use to create instances, see
[Compute Engine instances](https://docs.cloud.google.com/compute/docs/instances).

Reservations help ensure that you have the available resources to create
instances with the same hardware (memory and vCPUs) and optional resources
(GPUs, H4D HPC clusters, TPUs, or Local SSD disks) whenever you need them.
Reservations offer the following benefits:

- **High assurance of capacity**: you reserve resources to accommodate for
  future increases in demand, such as the following:

  - Growth

  - Planned or unplanned spikes in usage

  - Large migrations

  - Backup and disaster recovery

- **Exclusive access**: reservations prevent others from using your reserved
  resources.

- **Inherited properties**: reservations inherit the same properties as your
  chosen machine family.

After you reserve capacity, you can use it to create instances that match the
reservation. You don't incur any additional charges when you create these
instances. You only pay for resources that aren't part of the reservation, such
as disks or IP addresses.

## Limitations

All reservation types have the following limitations:

- Reservations are
  [zone-specific resources](https://docs.cloud.google.com/compute/docs/regions-zones/global-regional-zonal-resources#zoneresource).

- You can't use your reserved capacity to create the following
  Compute Engine resources:

  - `f1-micro` and `g1-small` machine types

  - Flex-start VMs

  - Sole-tenant nodes

  - Spot VMs or preemptible instances

## Choose a reservation type

The following diagram helps you choose the Compute Engine reservation
type that best fits your workload's needs:

![A flowchart with the different reservation methods available in Compute Engine.](https://docs.cloud.google.com/static/compute/images/reservations/reservations-types.png)

The questions in the preceding diagram are as follows:

1. Do you need capacity right away?

   - **Yes**: Go to the next question.

   - **No**: Go to question 3.

2. Do you need flexibility on how long to hold capacity?

   - **Yes** : See [Use on-demand reservations](https://docs.cloud.google.com/compute/docs/instances/choose-reservation-type#on-demand-reservations).

   - **No**: Go to the next question.

3. Do you need high-demand resources like GPUs?

   - **Yes**: Go to the next question.

   - **No** : See
     [Use future reservations](https://docs.cloud.google.com/compute/docs/instances/choose-reservation-type#future-reservations).

4. Do you need resources for more than 90 days?

   - **Yes** : See
     [Use future reservations in AI Hypercomputer](https://docs.cloud.google.com/compute/docs/instances/choose-reservation-type#reserve-capacity-aihc) or
     for H4D, see [Reserve capacity through your account team](https://docs.cloud.google.com/compute/docs/hpc/reserve-capacity-account-team).

   - **No** : See
     [Use future reservations in calendar mode](https://docs.cloud.google.com/compute/docs/instances/choose-reservation-type#future-reservations-calendar-mode).

### Use on-demand reservations

With on-demand reservations, you can attempt to immediately reserve capacity for
compute instances. If the attempt succeeds, Compute Engine creates a
reservation for matching instances that use the standard provisioning model.
After you create an on-demand reservation, you can consume, modify, or delete it
whenever you need to.

For more information, see
[About reservations](https://docs.cloud.google.com/compute/docs/instances/reservations-overview).

### Use future reservations

To reserve instances for a set period, you can use standard future reservations.
After you create a reservation request, you must submit it to Google Cloud for
review. Google Cloud typically takes five days to review your request. If
your request is approved, then Compute Engine creates on-demand
reservations with your requested capacity on your chosen date and time. To
consume these reservations, you create compute instances that match the
reservations and use the standard provisioning model. After the reservation
period ends, you can modify or delete the reservations.

For more information, see
[About future reservation requests](https://docs.cloud.google.com/compute/docs/instances/future-reservations-overview).

### Use future reservations in calendar mode

To reserve GPU instances, H4D instances, or TPUs for up to 90 days, you can use
future reservations in calendar mode. To create this type of reservation, first
view when your chosen number and type of resources are available in a region.
Then, create and submit a reservation request with the properties that you
confirmed as available. If you can successfully create the request, then
Google Cloud approves it within a minute. After the request is approved,
Compute Engine does the following:

- Compute Engine automatically creates a reservation.

- Compute Engine reserves your requested resources as close to each
  other as possible to minimize network latency.

At the start of your reservation period, you can consume the reservation by
creating matching GPU, H4D, or TPU instances that use the reservation-bound
provisioning model. At the end of the reservation period,
Compute Engine deletes the reservation, and stops or deletes any
instances that consume the reservation based on the termination action that you
specified for the instances.

For more information, see
[About future reservation requests in calendar mode](https://docs.cloud.google.com/compute/docs/instances/future-reservations-calendar-mode-overview).

### Use future reservations with AI Hypercomputer or H4D HPC clusters

Contact your account team and request to reserve GPU instances for large-scale
artificial intelligence (AI) and machine learning (ML) workloads, or for
creating a cluster of H4D HPC instances with
[enhanced cluster management capabilities](https://docs.cloud.google.com/compute/docs/hpc/cluster-capabilities).
After Google creates a draft reservation request for you, submit it for review
if everything looks correct. Google Cloud immediately approves the
request, and then Compute Engine does the following:

- Compute Engine automatically creates a reservation.

- Compute Engine reserves your requested resources as close to each
  other as possible to minimize network latency.

- Compute Engine reserves resources with topology-aware scheduling, as
  well as enhanced monitoring and maintenance.

At the start of your reservation period, you can consume the reservation by
creating matching GPU or H4D instances that use the reservation-bound
provisioning model. At the end of the reservation period,
Compute Engine deletes the reservation, and stops or deletes any
instances that consume the reservation based on the termination action that you
specified for the instances.

For more information, see either of the following:

- **For GPU instances**:

  - [Reserve capacity through your account team](https://docs.cloud.google.com/ai-hypercomputer/docs/reserve-capacity)
    in the AI Hypercomputer documentation.

  - [Reserve capacity through your account team](https://docs.cloud.google.com/cluster-director/docs/reserve-capacity)
    in the Cluster Director documentation.

- **For H4D instances**:

  - [Reserve capacity through your account team](https://docs.cloud.google.com/compute/docs/hpc/reserve-capacity-account-team) in the Compute Engine documentation.