建立及使用先占 VM

本頁面說明如何建立及使用先占虛擬機器 (VM) 執行個體。先占 VM 的折扣最多可達標準 VM 預設價格的 91%。不過,為處理其他工作而必須取回相關資源時,Compute Engine 可能會停止 (先占) 這些 VM。先占 VM 一律會在 24 小時後停止運作。 先占 VM 僅適用於容錯應用程式,且這些應用程式能承受 VM 先占的影響。決定建立先占 VM 前,請先確認應用程式是否可處理先占作業。如要瞭解先占 VM 的風險和價值,請參閱先占 VM 執行個體說明文件。

事前準備

  • 閱讀先占 VM 執行個體說明文件。
  • 如果尚未設定驗證,請先完成設定。 驗證可確認您的身分,以便存取 Google Cloud 服務和 API。如要從本機開發環境執行程式碼或範例,請選取下列其中一個選項,向 Compute Engine 進行驗證:

    選取這個頁面上範例的預計用途分頁:

    控制台

    使用 Google Cloud 控制台存取 Google Cloud 服務和 API 時,不需要設定驗證。

    gcloud

    1. 安裝 Google Cloud CLI。 完成後,執行下列指令來初始化 Google Cloud CLI:

      gcloud init

      如果您使用外部識別資訊提供者 (IdP),請先 使用聯合身分登入 gcloud CLI。

  • 設定預設區域和可用區。
  • Go

    如要在本機開發環境中使用本頁的 Go 範例,請安裝並初始化 gcloud CLI,然後使用使用者憑證設定應用程式預設憑證。

    1. 安裝 Google Cloud CLI。

    2. 如果您使用外部識別資訊提供者 (IdP),請先 使用聯合身分登入 gcloud CLI。

    3. 如果您使用本機殼層,請為使用者帳戶建立本機驗證憑證:

      gcloud auth application-default login

      如果您使用 Cloud Shell,則不需要執行這項操作。

      如果系統傳回驗證錯誤,且您使用外部識別資訊提供者 (IdP),請確認您已 使用聯合身分登入 gcloud CLI。

    詳情請參閱「 設定本機開發環境的驗證機制」。

    Java

    如要在本機開發環境中使用本頁的 Java 範例,請安裝並初始化 gcloud CLI,然後使用使用者憑證設定應用程式預設憑證。

    1. 安裝 Google Cloud CLI。

    2. 如果您使用外部識別資訊提供者 (IdP),請先 使用聯合身分登入 gcloud CLI。

    3. 如果您使用本機殼層,請為使用者帳戶建立本機驗證憑證:

      gcloud auth application-default login

      如果您使用 Cloud Shell,則不需要執行這項操作。

      如果系統傳回驗證錯誤,且您使用外部識別資訊提供者 (IdP),請確認您已 使用聯合身分登入 gcloud CLI。

    詳情請參閱「 設定本機開發環境的驗證機制」。

    Node.js

    如要在本機開發環境中使用本頁的 Node.js 範例,請安裝並初始化 gcloud CLI,然後使用您的使用者憑證設定應用程式預設憑證。

    1. 安裝 Google Cloud CLI。

    2. 如果您使用外部識別資訊提供者 (IdP),請先 使用聯合身分登入 gcloud CLI。

    3. 如果您使用本機殼層,請為使用者帳戶建立本機驗證憑證:

      gcloud auth application-default login

      如果您使用 Cloud Shell,則不需要執行這項操作。

      如果系統傳回驗證錯誤,且您使用外部識別資訊提供者 (IdP),請確認您已 使用聯合身分登入 gcloud CLI。

    詳情請參閱「 設定本機開發環境的驗證機制」。

    Python

    如要在本機開發環境中使用本頁的 Python 範例,請安裝並初始化 gcloud CLI,然後使用使用者憑證設定應用程式預設憑證。

    1. 安裝 Google Cloud CLI。

    2. 如果您使用外部識別資訊提供者 (IdP),請先 使用聯合身分登入 gcloud CLI。

    3. 如果您使用本機殼層,請為使用者帳戶建立本機驗證憑證:

      gcloud auth application-default login

      如果您使用 Cloud Shell,則不需要執行這項操作。

      如果系統傳回驗證錯誤,且您使用外部識別資訊提供者 (IdP),請確認您已 使用聯合身分登入 gcloud CLI。

    詳情請參閱「 設定本機開發環境的驗證機制」。

    REST

    如要在本機開發環境中使用本頁的 REST API 範例,請使用您提供給 gcloud CLI 的憑證。

      安裝 Google Cloud CLI。

      如果您使用外部識別資訊提供者 (IdP),請先 使用聯合身分登入 gcloud CLI。

    詳情請參閱 Google Cloud 驗證說明文件中的「使用 REST 進行驗證」。

建立先占 VM

使用 gcloud CLI 或 Compute Engine API 建立先占 VM。如要使用Google Cloud 控制台,請改為建立 Spot VM。

gcloud

使用 gcloud compute 時,請使用與建立一般 VM 相同的 instances create 指令,但要新增 --preemptible 標記。

gcloud compute instances create VM_NAME --preemptible

其中 VM_NAME 是 VM 的名稱。

Go

import (
	"context"
	"fmt"
	"io"

	compute "cloud.google.com/go/compute/apiv1"
	computepb "cloud.google.com/go/compute/apiv1/computepb"
	"google.golang.org/protobuf/proto"
)

// createPreemtibleInstance creates a new preemptible VM instance
// with Debian 10 operating system.
func createPreemtibleInstance(
	w io.Writer, projectID, zone, instanceName string,
) error {
	// projectID := "your_project_id"
	// zone := "europe-central2-b"
	// instanceName := "your_instance_name"
	// preemptible := true

	ctx := context.Background()
	instancesClient, err := compute.NewInstancesRESTClient(ctx)
	if err != nil {
		return fmt.Errorf("NewInstancesRESTClient: %w", err)
	}
	defer instancesClient.Close()

	imagesClient, err := compute.NewImagesRESTClient(ctx)
	if err != nil {
		return fmt.Errorf("NewImagesRESTClient: %w", err)
	}
	defer imagesClient.Close()

	// List of public operating system (OS) images:
	// https://cloud.google.com/compute/docs/images/os-details.
	newestDebianReq := &computepb.GetFromFamilyImageRequest{
		Project: "debian-cloud",
		Family:  "debian-12",
	}
	newestDebian, err := imagesClient.GetFromFamily(ctx, newestDebianReq)
	if err != nil {
		return fmt.Errorf("unable to get image from family: %w", err)
	}

	inst := &computepb.Instance{
		Name: proto.String(instanceName),
		Disks: []*computepb.AttachedDisk{
			{
				InitializeParams: &computepb.AttachedDiskInitializeParams{
					DiskSizeGb:  proto.Int64(10),
					SourceImage: newestDebian.SelfLink,
					DiskType:    proto.String(fmt.Sprintf("zones/%s/diskTypes/pd-standard", zone)),
				},
				AutoDelete: proto.Bool(true),
				Boot:       proto.Bool(true),
			},
		},
		Scheduling: &computepb.Scheduling{
			// Set the preemptible setting
			Preemptible: proto.Bool(true),
		},
		MachineType: proto.String(fmt.Sprintf("zones/%s/machineTypes/n1-standard-1", zone)),
		NetworkInterfaces: []*computepb.NetworkInterface{
			{
				Name: proto.String("global/networks/default"),
			},
		},
	}

	req := &computepb.InsertInstanceRequest{
		Project:          projectID,
		Zone:             zone,
		InstanceResource: inst,
	}

	op, err := instancesClient.Insert(ctx, req)
	if err != nil {
		return fmt.Errorf("unable to create instance: %w", err)
	}

	if err = op.Wait(ctx); err != nil {
		return fmt.Errorf("unable to wait for the operation: %w", err)
	}

	fmt.Fprintf(w, "Instance created\n")

	return nil
}

Java


import com.google.cloud.compute.v1.AttachedDisk;
import com.google.cloud.compute.v1.AttachedDiskInitializeParams;
import com.google.cloud.compute.v1.InsertInstanceRequest;
import com.google.cloud.compute.v1.Instance;
import com.google.cloud.compute.v1.InstancesClient;
import com.google.cloud.compute.v1.NetworkInterface;
import com.google.cloud.compute.v1.Operation;
import com.google.cloud.compute.v1.Scheduling;
import java.io.IOException;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.TimeoutException;

public class CreatePreemptibleInstance {

  public static void main(String[] args)
      throws IOException, ExecutionException, InterruptedException, TimeoutException {
    // TODO(developer): Replace these variables before running the sample.
    // projectId: project ID or project number of the Cloud project you want to use.
    // zone: name of the zone you want to use. For example: “us-west3-b”
    // instanceName: name of the new virtual machine.
    String projectId = "your-project-id-or-number";
    String zone = "zone-name";
    String instanceName = "instance-name";

    createPremptibleInstance(projectId, zone, instanceName);
  }

  // Send an instance creation request with preemptible settings to the Compute Engine API
  // and wait for it to complete.
  public static void createPremptibleInstance(String projectId, String zone, String instanceName)
      throws IOException, ExecutionException, InterruptedException, TimeoutException {

    String machineType = String.format("zones/%s/machineTypes/e2-small", zone);
    String sourceImage = "projects/debian-cloud/global/images/family/debian-13";
    long diskSizeGb = 10L;
    String networkName = "default";

    try (InstInstancesClienttancesClient = InstInstancesClientate()) {

      AttaAttachedDiskk =
          AttaAttachedDiskBuilder()
              .setBoot(true)
              .setAutoDelete(true)
              .setType(AttaAttachedDiske.PERSISTENT.toString())
              .setIsetInitializeParams                // Describe the size and source image of the boot disk to attach to the instance.
                  AttaAttachedDiskInitializeParamsBuilder()
                      .setSourceImage(sourceImage)
                      .setDiskSizeGb(diskSizeGb)
                      .build())
              .build();

      // Use the default VPC network.
      NetwNetworkInterfaceworkInterface = NetwNetworkInterfaceBuilder()
          .setName(networkName)
          .build();

      // Collect information into the Instance object.
      InstInstancetanceResource =
          InstInstanceBuilder()
              .setName(instanceName)
              .setMachineType(machineType)
              .addDisks(disk)
              .addNetworkInterfaces(networkInterface)
              // Set the preemptible setting.
              .setScheduling(ScheSchedulingBuilder()
                  .setPsetPreemptiblee)
                  .build())
              .build();

      System.out.printf("Creating instance: %s at %s %n", instanceName, zone);

      // Prepare the request to insert an instance.
      InseInsertInstanceRequestertInstanceRequest = InseInsertInstanceRequestBuilder()
          .setProject(projectId)
          .setZone(zone)
          .setInstanceResource(instanceResource)
          .build();

      // Wait for the create operation to complete.
      OperOperationponse = instancesClient.insertAsync(insertInstanceRequest)
          .get(3, TimeUnit.MINUTES);
      ;

      if (respresponse.hasError()        System.out.println("Instance creation failed ! ! " + response);
        return;
      }

      System.out.printf("Instance created : %s\n", instanceName);
      System.out.println("Operation Status: " + respresponse.getStatus()   }
  }
}

Node.js

/**
 * TODO(developer): Uncomment and replace these variables before running the sample.
 */
// const projectId = 'YOUR_PROJECT_ID';
// const zone = 'europe-central2-b';
// const instanceName = 'YOUR_INSTANCE_NAME';

const compute = require('@google-cloud/compute');

async function createPreemptible() {
  const instancesClient = new compute.InstancesClient();

  const [response] = await instancesClient.insert({
    instanceResource: {
      name: instanceName,
      disks: [
        {
          initializeParams: {
            diskSizeGb: '64',
            sourceImage:
              'projects/debian-cloud/global/images/family/debian-11/',
          },
          autoDelete: true,
          boot: true,
        },
      ],
      scheduling: {
        // Set the preemptible setting
        preemptible: true,
      },
      machineType: `zones/${zone}/machineTypes/e2-small`,
      networkInterfaces: [
        {
          name: 'global/networks/default',
        },
      ],
    },
    project: projectId,
    zone,
  });
  let operation = response.latestResponse;
  const operationsClient = new compute.ZoneOperationsClient();

  // Wait for the create operation to complete.
  while (operation.status !== 'DONE') {
    [operation] = await operationsClient.wait({
      operation: operation.name,
      project: projectId,
      zone: operation.zone.split('/').pop(),
    });
  }

  console.log('Instance created.');
}

createPreemptible();

Python

from __future__ import annotations

import re
import sys
from typing import Any
import warnings

from google.api_core.extended_operation import ExtendedOperation
from google.cloud import compute_v1


def get_image_from_family(project: str, family: str) -> compute_v1.Image:
    """
    Retrieve the newest image that is part of a given family in a project.

    Args:
        project: project ID or project number of the Cloud project you want to get image from.
        family: name of the image family you want to get image from.

    Returns:
        An Image object.
    """
    image_client = compute_v1.ImagesClient()
    # List of public operating system (OS) images: https://cloud.google.com/compute/docs/images/os-details
    newest_image = image_client.get_from_family(project=project, family=family)
    return newest_image


def disk_from_image(
    disk_type: str,
    disk_size_gb: int,
    boot: bool,
    source_image: str,
    auto_delete: bool = True,
) -> compute_v1.AttachedDisk:
    """
    Create an AttachedDisk object to be used in VM instance creation. Uses an image as the
    source for the new disk.

    Args:
         disk_type: the type of disk you want to create. This value uses the following format:
            "zones/{zone}/diskTypes/(pd-standard|pd-ssd|pd-balanced|pd-extreme)".
            For example: "zones/us-west3-b/diskTypes/pd-ssd"
        disk_size_gb: size of the new disk in gigabytes
        boot: boolean flag indicating whether this disk should be used as a boot disk of an instance
        source_image: source image to use when creating this disk. You must have read access to this disk. This can be one
            of the publicly available images or an image from one of your projects.
            This value uses the following format: "projects/{project_name}/global/images/{image_name}"
        auto_delete: boolean flag indicating whether this disk should be deleted with the VM that uses it

    Returns:
        AttachedDisk object configured to be created using the specified image.
    """
    boot_disk = compute_v1.AttachedDisk()
    initialize_params = compute_v1.AttachedDiskInitializeParams()
    initialize_params.source_image = source_image
    initialize_params.disk_size_gb = disk_size_gb
    initialize_params.disk_type = disk_type
    boot_disk.initialize_params = initialize_params
    # Remember to set auto_delete to True if you want the disk to be deleted when you delete
    # your VM instance.
    boot_disk.auto_delete = auto_delete
    boot_disk.boot = boot
    return boot_disk


def wait_for_extended_operation(
    operation: ExtendedOperation, verbose_name: str = "operation", timeout: int = 300
) -> Any:
    """
    Waits for the extended (long-running) operation to complete.

    If the operation is successful, it will return its result.
    If the operation ends with an error, an exception will be raised.
    If there were any warnings during the execution of the operation
    they will be printed to sys.stderr.

    Args:
        operation: a long-running operation you want to wait on.
        verbose_name: (optional) a more verbose name of the operation,
            used only during error and warning reporting.
        timeout: how long (in seconds) to wait for operation to finish.
            If None, wait indefinitely.

    Returns:
        Whatever the operation.result() returns.

    Raises:
        This method will raise the exception received from `operation.exception()`
        or RuntimeError if there is no exception set, but there is an `error_code`
        set for the `operation`.

        In case of an operation taking longer than `timeout` seconds to complete,
        a `concurrent.futures.TimeoutError` will be raised.
    """
    result = operation.result(timeout=timeout)

    if operation.error_code:
        print(
            f"Error during {verbose_name}: [Code: {operation.error_code}]: {operation.error_message}",
            file=sys.stderr,
            flush=True,
        )
        print(f"Operation ID: {operation.name}", file=sys.stderr, flush=True)
        raise operation.exception() or RuntimeError(operation.error_message)

    if operation.warnings:
        print(f"Warnings during {verbose_name}:\n", file=sys.stderr, flush=True)
        for warning in operation.warnings:
            print(f" - {warning.code}: {warning.message}", file=sys.stderr, flush=True)

    return result


def create_instance(
    project_id: str,
    zone: str,
    instance_name: str,
    disks: list[compute_v1.AttachedDisk],
    machine_type: str = "n1-standard-1",
    network_link: str = "global/networks/default",
    subnetwork_link: str = None,
    internal_ip: str = None,
    external_access: bool = False,
    external_ipv4: str = None,
    accelerators: list[compute_v1.AcceleratorConfig] = None,
    preemptible: bool = False,
    spot: bool = False,
    instance_termination_action: str = "STOP",
    custom_hostname: str = None,
    delete_protection: bool = False,
) -> compute_v1.Instance:
    """
    Send an instance creation request to the Compute Engine API and wait for it to complete.

    Args:
        project_id: project ID or project number of the Cloud project you want to use.
        zone: name of the zone to create the instance in. For example: "us-west3-b"
        instance_name: name of the new virtual machine (VM) instance.
        disks: a list of compute_v1.AttachedDisk objects describing the disks
            you want to attach to your new instance.
        machine_type: machine type of the VM being created. This value uses the
            following format: "zones/{zone}/machineTypes/{type_name}".
            For example: "zones/europe-west3-c/machineTypes/f1-micro"
        network_link: name of the network you want the new instance to use.
            For example: "global/networks/default" represents the network
            named "default", which is created automatically for each project.
        subnetwork_link: name of the subnetwork you want the new instance to use.
            This value uses the following format:
            "regions/{region}/subnetworks/{subnetwork_name}"
        internal_ip: internal IP address you want to assign to the new instance.
            By default, a free address from the pool of available internal IP addresses of
            used subnet will be used.
        external_access: boolean flag indicating if the instance should have an external IPv4
            address assigned.
        external_ipv4: external IPv4 address to be assigned to this instance. If you specify
            an external IP address, it must live in the same region as the zone of the instance.
            This setting requires `external_access` to be set to True to work.
        accelerators: a list of AcceleratorConfig objects describing the accelerators that will
            be attached to the new instance.
        preemptible: boolean value indicating if the new instance should be preemptible
            or not. Preemptible VMs have been deprecated and you should now use Spot VMs.
        spot: boolean value indicating if the new instance should be a Spot VM or not.
        instance_termination_action: What action should be taken once a Spot VM is terminated.
            Possible values: "STOP", "DELETE"
        custom_hostname: Custom hostname of the new VM instance.
            Custom hostnames must conform to RFC 1035 requirements for valid hostnames.
        delete_protection: boolean value indicating if the new virtual machine should be
            protected against deletion or not.
    Returns:
        Instance object.
    """
    instance_client = compute_v1.InstancesClient()

    # Use the network interface provided in the network_link argument.
    network_interface = compute_v1.NetworkInterface()
    network_interface.network = network_link
    if subnetwork_link:
        network_interface.subnetwork = subnetwork_link

    if internal_ip:
        network_interface.network_i_p = internal_ip

    if external_access:
        access = compute_v1.AccessConfig()
        access.type_ = compute_v1.AccessConfig.Type.ONE_TO_ONE_NAT.name
        access.name = "External NAT"
        access.network_tier = access.NetworkTier.PREMIUM.name
        if external_ipv4:
            access.nat_i_p = external_ipv4
        network_interface.access_configs = [access]

    # Collect information into the Instance object.
    instance = compute_v1.Instance()
    instance.network_interfaces = [network_interface]
    instance.name = instance_name
    instance.disks = disks
    if re.match(r"^zones/[a-z\d\-]+/machineTypes/[a-z\d\-]+$", machine_type):
        instance.machine_type = machine_type
    else:
        instance.machine_type = f"zones/{zone}/machineTypes/{machine_type}"

    instance.scheduling = compute_v1.Scheduling()
    if accelerators:
        instance.guest_accelerators = accelerators
        instance.scheduling.on_host_maintenance = (
            compute_v1.Scheduling.OnHostMaintenance.TERMINATE.name
        )

    if preemptible:
        # Set the preemptible setting
        warnings.warn(
            "Preemptible VMs are being replaced by Spot VMs.", DeprecationWarning
        )
        instance.scheduling = compute_v1.Scheduling()
        instance.scheduling.preemptible = True

    if spot:
        # Set the Spot VM setting
        instance.scheduling.provisioning_model = (
            compute_v1.Scheduling.ProvisioningModel.SPOT.name
        )
        instance.scheduling.instance_termination_action = instance_termination_action

    if custom_hostname is not None:
        # Set the custom hostname for the instance
        instance.hostname = custom_hostname

    if delete_protection:
        # Set the delete protection bit
        instance.deletion_protection = True

    # Prepare the request to insert an instance.
    request = compute_v1.InsertInstanceRequest()
    request.zone = zone
    request.project = project_id
    request.instance_resource = instance

    # Wait for the create operation to complete.
    print(f"Creating the {instance_name} instance in {zone}...")

    operation = instance_client.insert(request=request)

    wait_for_extended_operation(operation, "instance creation")

    print(f"Instance {instance_name} created.")
    return instance_client.get(project=project_id, zone=zone, instance=instance_name)


def create_preemptible_instance(
    project_id: str, zone: str, instance_name: str
) -> compute_v1.Instance:
    """
    Create a new preemptible VM instance with Debian 10 operating system.

    Args:
        project_id: project ID or project number of the Cloud project you want to use.
        zone: name of the zone to create the instance in. For example: "us-west3-b"
        instance_name: name of the new virtual machine (VM) instance.

    Returns:
        Instance object.
    """
    newest_debian = get_image_from_family(project="debian-cloud", family="debian-13")
    disk_type = f"zones/{zone}/diskTypes/pd-standard"
    disks = [disk_from_image(disk_type, 10, True, newest_debian.self_link)]
    instance = create_instance(project_id, zone, instance_name, disks, preemptible=True)
    return instance

REST

在 API 中,建構建立 VM 的一般要求,但請在 scheduling 下方加入 preemptible 屬性,並將其設為 true。例如:

POST https://compute.googleapis.com/compute/v1/projects/PROJECT_ID/zones/ZONE/instances

{
  'machineType': 'zones/ZONE/machineTypes/MACHINE_TYPE',
  'name': 'INSTANCE_NAME',
  'scheduling':
  {
    'preemptible': true
  },
  ...
}

先占 CPU 配額

先占 VM 需要可用的 CPU 配額,與標準 VM 相同。為避免先占 VM 耗用標準 VM 的 CPU 配額,您可以申請「先占 CPU」配額。Compute Engine 在該區域授予您先占 CPU 配額後,所有先占 VM 都會計入該配額,所有標準 VM 則會繼續計入標準 CPU 配額。

在沒有先占 CPU 配額的區域,您可以使用標準 CPU 配額啟動先占 VM。此外,您也需要足夠的 IP 和磁碟配額。除非 Compute Engine 已授予配額,否則 gcloud CLI 或Google Cloud 控制台配額頁面不會顯示可搶占 CPU 配額。

如要進一步瞭解配額的相關資訊,請造訪資源配額頁面。

啟動遭預先終止的 VM

與其他 VM 相同,如果先占 VM 停止或遭到先占,您可以再次啟動 VM,讓 VM 回到 RUNNING 狀態。啟動先占 VM 會重設 24 小時計數器,但由於這仍是先占 VM,Compute Engine 可以在 24 小時前先占。先占 VM 執行期間無法轉換為標準 VM。

如果 Compute Engine 停止自動調度資源的代管執行個體群組 (MIG) 或 Google Kubernetes Engine 叢集中的先占 VM,當資源再次可用時,群組會重新啟動 VM。

使用關機指令碼處理先占

當 Compute Engine 搶占 VM 時,您可以使用關機指令碼,在 VM 遭到搶占前嘗試執行清除動作。舉例來說,您可以正常停止執行中的程序,並將檢查點檔案複製到 Cloud Storage。請注意,先占 VM 的關機時間長度上限比使用者發起的關機作業短。如要進一步瞭解先占通知的關機期限,請參閱概念說明文件中的「先占程序」。

以下是關機指令碼,您可以新增至正在執行的先占 VM,或在建立新的先占 VM 時新增。VM 開始關機時,系統會執行這段指令碼,然後作業系統的正常 kill 指令會停止所有剩餘程序。在正常停止所選程式後,指令碼會將檢查點檔案平行上傳至 Cloud Storage 值區。

#!/bin/bash

MY_PROGRAM="PROGRAM_NAME" # For example, "apache2" or "nginx"
MY_USER="LOCAL_USERNAME"
CHECKPOINT="/home/$MY_USER/checkpoint.out"
BUCKET_NAME="BUCKET_NAME" # For example, "my-checkpoint-files" (without gs://)

echo "Shutting down!  Seeing if ${MY_PROGRAM} is running."

# Find the newest copy of $MY_PROGRAM
PID="$(pgrep -n "$MY_PROGRAM")"

if [[ "$?" -ne 0 ]]; then
  echo "${MY_PROGRAM} not running, shutting down immediately."
  exit 0
fi

echo "Sending SIGINT to $PID"
kill -2 "$PID"

# Portable waitpid equivalent
while kill -0 "$PID"; do
   sleep 1
done

echo "$PID is done, copying ${CHECKPOINT} to gs://${BUCKET_NAME} as ${MY_USER}"

su "${MY_USER}" -c "gcloud storage cp $CHECKPOINT gs://${BUCKET_NAME}/"

echo "Done uploading, shutting down."

如要將這個指令碼新增至 VM,請設定指令碼以搭配 VM 上的應用程式運作,然後將指令碼新增至 VM 的中繼資料。

  1. 將關機指令碼複製或下載到本機工作站。
  2. 開啟檔案以編輯並變更下列變數:
    • PROGRAM_NAME 是要關閉的程序或程式名稱。例如:apache2 或 nginx。
    • LOCAL_USERNAME 是您登入虛擬機器的使用者名稱。
    • BUCKET_NAME 是要儲存程式檢查點檔案的 Cloud Storage bucket 名稱。請注意,在本範例中,bucket 名稱開頭不是 gs://。
  3. 儲存變更。
  4. 將關機指令碼新增至新 VM 或現有 VM。

這段指令碼假設您已完成下列事項:

  • 建立 VM 時,至少具備 Cloud Storage 的讀寫權限。 如需瞭解如何建立具有適當範圍的 VM,請參閱驗證說明文件。

  • 您有現成的 Cloud Storage bucket,且具備寫入權限。

找出先占 VM

如要檢查 VM 是否為先占 VM,請按照步驟找出 VM 的佈建模型和終止動作。

判斷 VM 是否遭到搶占

使用Google Cloud console、gcloud CLI 或 API,判斷 VM 是否遭到搶占。

控制台

您可以查看系統活動記錄,確認 VM 是否遭到搶占。

  1. 前往 Google Cloud 控制台的「記錄」頁面。

    前往「記錄」

  2. 選取您的專案並點選 [繼續]。

  3. 將 compute.instances.preempted 新增至 [filter by label or text search] (按標籤或搜尋字詞篩選) 欄位。

  4. 您也可以選擇輸入 VM 名稱,查看特定 VM 的搶占作業。

  5. 按下 Enter 鍵,套用指定篩選器。 Google Cloud 控制台會更新記錄清單,只顯示 VM 遭到搶占的作業。

  6. 在清單中選取作業,即可查看遭搶占 VM 的詳細資料。

gcloud


請使用 gcloud compute operations list 指令搭配 filter 參數,取得專案的先占事件清單。

gcloud compute operations list \
    --filter="operationType=compute.instances.preempted"

您可以使用 filter 參數進一步指定結果範圍。舉例來說,如要查看代管執行個體群組中 VM 的先占事件,請執行下列操作:

gcloud compute operations list \
    --filter="operationType=compute.instances.preempted AND targetLink:instances/BASE_VM_NAME"

gcloud 會傳回類似以下內容的回應:

NAME                  TYPE                         TARGET                                   HTTP_STATUS STATUS TIMESTAMP
systemevent-yyyyyyyy  compute.instances.preempted  us-central1-f/instances/example-vm-yyy  200         DONE   2015-04-02T12:12:10.881-07:00

作業類型為 compute.instances.preempted 表示 VM 已遭到搶占。您可以使用 operations describe 指令取得特定先占作業的相關詳細資訊。

gcloud compute operations describe \
    systemevent-yyyyyyyy

gcloud 會傳回類似以下內容的回應:

...
operationType: compute.instances.preempted
progress: 100
selfLink: https://compute.googleapis.com/compute/v1/projects/PROJECT_ID/zones/us-central1-f/operations/systemevent-yyyyyyyy
startTime: '2015-04-02T12:12:10.881-07:00'
status: DONE
statusMessage: Instance was preempted.
...

REST


如要取得最近的系統作業清單,請傳送 GET 要求至區域作業 URI。

GET https://compute.googleapis.com/compute/v1/projects/PROJECT_ID/zones/ZONE/operations

回應會包含近期作業的清單。

{
  "kind": "compute#operation",
  "id": "15041793718812375371",
  "name": "systemevent-yyyyyyyy",
  "zone": "https://www.googleapis.com/compute/v1/projects/PROJECT_ID/zones/us-central1-f",
  "operationType": "compute.instances.preempted",
  "targetLink": "https://www.googleapis.com/compute/v1/projects/PROJECT_ID/zones/us-central1-f/instances/EXAMPLE_VM",
  "targetId": "12820389800990687210",
  "status": "DONE",
  "statusMessage": "Instance was preempted.",
  ...
}

如要將回應範圍限制為僅顯示先占作業,您可以在 API 要求中新增篩選條件:operationType="compute.instances.preempted"。如要查看特定 VM 的搶占作業,請在篩選器中新增 targetLink 參數:operationType="compute.instances.preempted" AND targetLink="https://www.googleapis.com/compute/v1/projects/PROJECT_ID/zones/ZONE/instances/VM_NAME"。

或者,您也可以從 VM 內部判斷 VM 是否遭到搶占。如果您想在關機指令碼中,以不同於正常關機的方式處理因 Compute Engine 先占而造成的關機,這項功能就很有用。如要執行這項操作,請在 VM 的預設執行個體中繼資料中,檢查中繼資料伺服器的 preempted 值。

舉例來說,您可以在 VM 中使用 curl,取得 preempted 的值:

curl "http://metadata.google.internal/computeMetadata/v1/instance/preempted" -H "Metadata-Flavor: Google"
TRUE

如果這個值為 TRUE,表示 VM 已遭 Compute Engine 先占,否則為 FALSE。

如果您想在關閉指令碼以外的地方使用它,請將 ?wait_for_change=true 附加至網址。這會執行懸掛式 HTTP GET 要求,只有在中繼資料變更且 VM 已搶占時才會傳回。

curl "http://metadata.google.internal/computeMetadata/v1/instance/preempted?wait_for_change=true" -H "Metadata-Flavor: Google"
TRUE

測試先占設定

您可以在 VM 上執行模擬維護事件,強制 VM 搶占資源。使用這項功能測試應用程式如何處理可搶占 VM。請參閱測試可用性政策,瞭解如何測試 VM 的維護事件。

您也可以停止 VM,模擬 VM 先占作業,這項做法可取代模擬維護事件,且不會受到配額限制。

最佳做法

以下是可協助您善用先佔 VM 執行個體的幾個最佳做法。

使用大量執行個體 API

您可以使用大量執行個體 API,不必建立單一 VM。

挑選較小的機器類型

先占 VM 的資源來自額外及備份 Google Cloud容量。較小的機器類型通常更容易取得容量,也就是具備較少 vCPU 和記憶體等資源的機器類型。選取較小的自訂機型,或許能找到更多先占 VM 的容量,但選取較小的預先定義機型,更有可能找到容量。舉例來說,與n2-standard-32預先定義機型的容量相比,n2-custom-24-96自訂機型的容量較有可能,但n2-standard-16預先定義機型的容量更有可能。

在離峰時段執行大型先占 VM 叢集

Google Cloud 資料中心的負載量會因地點和時段而異,但通常在夜間和週末最低。因此,夜間和週末是執行大型先占 VM 叢集的最佳時機。

將應用程式設計為容錯且能承受先占

請務必做好準備,瞭解不同時間點的搶占模式會有所變化。舉例來說,如果可用區發生部分中斷,系統可能會先占大量可先占用的 VM,以便為需要遷移的標準 VM 騰出空間,做為復原程序的一部分。在該短時間內,先占率會與其他任何一天大不相同。如果應用程式假設先占作業一律會以小群組形式完成,您可能無法為這類事件做好準備。您可以停止 VM 執行個體,測試應用程式在搶占事件下的行為。

重新嘗試建立已遭先占的 VM

如果 VM 執行個體遭到先占,請先嘗試建立一到兩次新的先占 VM,再改用標準 VM。視需求而定,建議在叢集中同時使用標準 VM 和先占 VM,確保工作以適當的速度進行。

使用關閉指令碼

請使用關閉指令碼管理關閉與先占通知,該指令碼要能夠儲存工作進度以接續上次進度,而不用從頭開始。

後續步驟