[Video](https://www.youtube.com/watch?v=-b8OB_qEG9c)

This document describes the common causes of unexpected shutdowns and reboots
of Compute Engine instances and how to prevent them.

Instance shutdowns and reboots can be caused by system events or administrative
activities. System event shutdowns and reboots are generated by Google systems
or your instance's operating system. Admin activity shutdowns and reboots are
generated by a user- or service account-generated API call. All shutdowns and
reboots are logged, except for reboots which are initiated from within the
instance.

## Before you begin

- If you haven't already, set up [authentication](https://docs.cloud.google.com/compute/docs/authentication). Authentication verifies your identity for access to Google Cloud services and APIs. To run code or samples from a local development environment, you can authenticate to Compute Engine by selecting one of the following options:

  Select the tab for how you plan to use the samples on this page:

  ### Console


  When you use the Google Cloud console to access Google Cloud services and
  APIs, you don't need to set up authentication.

  ### gcloud

  1.
     [Install](https://docs.cloud.google.com/sdk/docs/install) the Google Cloud CLI.

     After installation,
     [initialize](https://docs.cloud.google.com/sdk/docs/initializing) the Google Cloud CLI by running the following command:

     ```bash
     gcloud init
     ```


     If you're using an external identity provider (IdP), you must first
     [sign in to the gcloud CLI with your federated identity](https://docs.cloud.google.com/iam/docs/workforce-log-in-gcloud).

     > [!NOTE]
     > **Note:** If you installed the gcloud CLI previously, make sure you have the latest version by running `gcloud components update`.

  2. [Set a default region and zone](https://docs.cloud.google.com/compute/docs/gcloud-compute#set_default_zone_and_region_in_your_local_client).

## Diagnosing instance shutdowns and reboots

To diagnose the cause of an instance's spontaneous shutdown or reboot, you must
[query your instance's logs](https://docs.cloud.google.com/compute/docs/troubleshooting/troubleshooting-reboots#querying_audit_logs). To quickly identify the
cause of future VM shutdowns or reboots, [build a dashboard](https://docs.cloud.google.com/compute/docs/troubleshooting/troubleshooting-reboots#monitor-events)
that contains the logs. After you query the logs, review the `method` and
`principalEmail` fields to determine what event and which user or service
initiated the shutdown or reboot.

### Querying Cloud Audit Logs

Query Cloud Audit Logs to display a list of system events and administrative
activities that might have caused the shutdown or reboot.

#### Permissions required for this task

To perform this task, you must have the following
[permissions](https://docs.cloud.google.com/iam/docs/overview#permissions):


- The **Logging/Logs Viewer** role or the **Project/Viewer** role.

### Console

1. In the Google Cloud console, go to the **Logs Explorer** page.

   [Go to Logs Explorer](https://console.cloud.google.com/logs/)

   > [!NOTE]
   > **Note:** You might need to click **Upgrade** to use Logs Explorer instead of the Legacy Logs Viewer.

2. In the **Query** field, enter the following query:


   ```
   resource.type="gce_instance"
   "VM_NAME"
   logName:("logs/cloudaudit.googleapis.com%2Fsystem_event" OR "logs/cloudaudit.googleapis.com%2Factivity")
   ```

   <br />

   Replace `VM_NAME` with the name of the VM that shut down
   or rebooted.
3. If the event you're looking for happened more than an hour ago, set a
   custom time frame by clicking the clock symbol and entering a custom
   range.

   ![Set query time frame.](https://docs.cloud.google.com/static/compute/docs/troubleshooting/logs-time-frame.png)
4. Click **Run query** . The results are displayed in the **Query
   results** section.

   > [!TIP]
   > **Tip:** To increase the size of the **Query results** section, click **Enter fullscreen Query results**.

5. Click the

   expander arrow next to each result to show detailed information.

6.
   See [Reviewing Cloud Audit Logs](https://docs.cloud.google.com/compute/docs/troubleshooting/troubleshooting-reboots#reviewing-logs) to learn more about the `method`
   and `principalEmail` fields that are associated with shutdowns and
   reboots, and what you can do to prevent them.

### gcloud

1. View Cloud Audit Logs using the
   [`gcloud logging read` command](https://docs.cloud.google.com/sdk/gcloud/reference/logging/read):


   ```
   gcloud logging read --freshness=TIME 'resource.type="gce_instance" "VM_NAME" logName:("logs/cloudaudit.googleapis.com%2Fsystem_event" OR "logs/cloudaudit.googleapis.com%2Factivity")'
   ```

   <br />

   Replace the following:
   - `TIME`: the amount of time you want to query. For example, `1h` queries log entries in the past hour. For information about date and time formats, see [gcloud topic datetimes](https://docs.cloud.google.com/sdk/gcloud/reference/topic/datetimes).
   - `VM_NAME`: the name of the VM that shut down or rebooted.

   <br />

   The results display.
2.
   See [Reviewing Cloud Audit Logs](https://docs.cloud.google.com/compute/docs/troubleshooting/troubleshooting-reboots#reviewing-logs) to learn more about the `method`
   and `principalEmail` fields that are associated with shutdowns and
   reboots, and what you can do to prevent them.

### Reviewing Cloud Audit Logs

Review the `method` and `principalEmail` fields of the Cloud Audit Logs to
determine why your VM was shut down or rebooted.

1. Review the `method` fields of the Cloud Audit Logs and compare them with
   the methods listed in the following table.

   > [!NOTE]
   > **Note:** If your VM rebooted and you don't see a comparable method listed in the following table in the Cloud Audit Logs, the reboot likely happened because your VM's operating system initiated the reboot. Causes of reboots are often difficult to identify. If your VMs experience frequent reboots, consider [Getting support](https://docs.cloud.google.com/compute/docs/getting-support).

   | Method | Shutdown type | Description |
   |---|---|---|
   | `compute.instances.repair.recreateInstance` | System event | If your VM belongs to a managed instance group (MIG), the MIG recreates the VM if the VM's state changes from `RUNNING` and the MIG did not initiate the change in state. Changes of instance state that are not initiated by the MIG include: - Hardware failures. - Terminating a [Spot VM](https://docs.cloud.google.com/compute/docs/instance-groups/create-mig-with-spot-vms). - [Infrastructure maintenance](https://docs.cloud.google.com/compute/docs/instances/setting-instance-scheduling-options) events when the VM instance is not set to [live migrate](https://docs.cloud.google.com/compute/docs/instances/live-migration). - Deleting a MIG instance by using one of the following methods: - The `instances.delete` API method - The `gcloud compute instances delete` command > [!NOTE] > **Note:** To ensure that your configuration changes aren't reverted by the MIG, it's important to [use the group's methods](https://docs.cloud.google.com/compute/docs/instance-groups/working-with-managed-instances). For example, to delete a managed instance, use one of the following methods: > - For a zonal MIG: [`instanceGroupManagers.deleteInstances`](https://docs.cloud.google.com/compute/docs/reference/rest/v1/instanceGroupManagers/deleteInstances) > - For a regional MIG: [`regionInstanceGroupManagers.deleteInstances`](https://docs.cloud.google.com/compute/docs/reference/rest/v1/regionInstanceGroupManagers/deleteInstances) > - In gcloud: [`gcloud compute instance-groups managed delete-instances`](https://docs.cloud.google.com/sdk/gcloud/reference/compute/instance-groups/managed/delete-instances) |
   | `compute.instances.hostError` | System event | A host error (`compute.instances.hostError`) means that there was a hardware or software issue on the physical machine or the data center infrastructure hosting your compute instance that caused your instance to crash. A host error involving a total hardware failure or other hardware issues might prevent the [live migration](https://docs.cloud.google.com/compute/docs/instances/live-migration-process) of your instance. If your instance is set to automatically restart, which is the default setting, Compute Engine restarts your instance, typically within three minutes from the time the error was detected. Depending on the issue, the restart might take up to 5.5 minutes. After a compute instance encounters a host error, or you report its host as faulty, the instance's host requires emergency, unplanned maintenance. By default, Compute Engine provides a few hours of advance notice when it schedules this type of maintenance. Occasionally, a compute instance might become unresponsive before a host error is signaled. You can reduce the amount of time Compute Engine waits to restart or terminate the instance by setting the host error recovery timeout. For more information, see [Set the policy for an existing instance](https://docs.cloud.google.com/compute/docs/instances/setting-vm-host-options#updatingoption). Physical hardware and software failures can happen occasionally but are rare occurrences. To protect your applications and services from these potentially disruptive system events, review the following resources: - [Designing robust systems](https://docs.cloud.google.com/compute/docs/tutorials/robustsystems) - [Patterns for scalable and resilient apps](https://docs.cloud.google.com/solutions/scalable-and-resilient-apps) - [Creating managed instance groups](https://docs.cloud.google.com/compute/docs/instance-groups/creating-groups-of-managed-instances) Google also offers managed services such as [App Engine](https://docs.cloud.google.com/appengine) and the [App Engine flexible environment](https://docs.cloud.google.com/appengine/docs/flexible). |
   | `compute.instances.automaticRestart` | System event | This event occurs after a `hostError` event or a `terminateOnHostMaintenance` event if your VM's `automaticRestart` host maintenance policy is set to `true`. In the logs, a `hostError` or a `terminateOnHostMaintenance` log entry precedes this log. If you want to change your VM's host maintenance policy, see [Updating options for an instance](https://docs.cloud.google.com/compute/docs/instances/setting-instance-scheduling-options#updatingoption). |
   | `compute.instances.guestTerminate` | System event | Your VM's operating system initiated the shutdown. |
   | `compute.instances.terminateOnHostMaintenance` | System event | If you set your VM's `onHostMaintenance` host maintenance policy to `TERMINATE`, Compute Engine stops your VM when there is a maintenance event where Google must move your VM to another host. If you want to change your VM's `onHostMaintenance` policy, see [Updating options for an instance](https://docs.cloud.google.com/compute/docs/instances/setting-instance-scheduling-options). |
   | `compute.instances.preempted` | System event | Compute Engine preempted your Spot VM or legacy preemptible VM: - When Compute Engine preempts a Spot VM, Compute Engine either stops or deletes the Spot VM based on its [termination action](https://docs.cloud.google.com/compute/docs/instances/spot#preemption-process). Spot VMs don't have a maximum runtime. - When Compute Engine preempts a preemptible VM, Compute Engine stops the VM after a maximum runtime of 24 hours. To avoid these limitations, use Spot VMs instead. Spot VMs and preemptible VMs are excess Compute Engine capacity, so Compute Engine might preempt them any time that capacity is needed elsewhere. You can help mitigate the effects of preemption by following the [best practices.](https://docs.cloud.google.com/compute/docs/instances/create-use-spot#best-practices) Alternatively, if you require VMs with user-controlled runtimes, [create standard VMs](https://docs.cloud.google.com/compute/docs/instances/create-start-instance) instead. |
   | `compute.instances.stop` | Admin activity | A user or service account stopped your VM. Continue to the next step to identify the user or service account that stopped your VM. For information about restarting your VM, see [Restarting a stopped instance](https://docs.cloud.google.com/compute/docs/instances/stop-start-instance#restarting_a_stopped_instance_that_doesnt_have_an_encrypted_disk). |
   | `compute.instances.delete` | Admin activity or system event | A user or service account deleted your VM, or the VM was configured to be automatically deleted. > [!IMPORTANT] > **Important:** Requests to delete your VM, which are indicated in Cloud Audit Logs by the `compute.instances.delete` method, might inconsistently override other requests for your VM that were made at a similar time. Even when those other requests were successful, the `compute.instances.delete` method might inconsistently prevent the methods from those other requests from appearing in the Cloud Audit Logs for your VM. Specifically, a log for the `compute.instances.delete` method might indicate any of the following requests for your VM: - Requests from a user or service account to directly delete your VM are indicated only by a `compute.instances.delete` method from the user or service account. - Requests that automatically delete your VM are indicated by a `compute.instances.delete` method from `system@google.com`, but the method that explains the cause of automatic deletion might or might not appear in Cloud Audit Logs. For example, if a Spot VM is configured to be automatically deleted during preemption and is preempted, you see a `compute.instances.delete` method from `system@google.com`, but you might or might not also see a `compute.instances.preempted` method. - Requests to the VM that happened shortly before or after a `compute.instances.delete` method might or might not appear in Cloud Audit Logs. For example, if a VM is stopped due to host maintenance shortly before the VM is deleted, you see a `compute.instances.delete` method, but you might or might not also see a `compute.instances.terminateOnHostMaintenance` method. Continue to the next step to identify the user or service account that deleted your VM. For information about creating a new VM, see [Creating and starting a VM](https://docs.cloud.google.com/compute/docs/instances/create-start-instance). |
   | `compute.instances.insert` | Admin activity | A user or service account created your VM. Continue to the next step to identify the user or service account that created your VM. For information about creating a new VM, see [Creating and starting a VM](https://docs.cloud.google.com/compute/docs/instances/create-start-instance). |
   | `compute.instances.reset` | Admin activity | A user or service account reset your VM. Continue to the next step to identify the user or service account that stopped your VM. |

2. Review the `principalEmail` fields of the Cloud Audit Logs to identify
   the user or service that initiated the shutdown or reboot. The following
   table include common Google managed services that initiate shutdowns or
   reboots.

   | Email | Description |
   |---|---|
   | `system@google.com` | A system event caused the shutdown or reboot. |
   | `project-number@cloudservices.gserviceaccount.com` | A [service agent](https://docs.cloud.google.com/iam/docs/service-account-types#service-agents) initiated the shutdown. To determine which project the service initiated the shutdown from, review the service agent's `project-number`. To determine which Google service made the request, review the `protoPayload.requestMetadata.callerSuppliedUserAgent` field. |

   <br />

   If a user triggered the shutdown or reboot, their email address appears in
   the `principalEmail` field. For example, `cloudysanfrancisco@gmail.com`.

   Administrators can prevent users from changing the state of project VMs by
   changing Identity and Access Management permissions on user accounts. For more information,
   see
   [Granting, changing, and revoking access to resources](https://docs.cloud.google.com/iam/docs/granting-changing-revoking-access).

## Monitor VM lifecycle events

You can monitor VM lifecycle events (including shutdowns, reboots, and host
errors) by building a Cloud Monitoring dashboard.

This dashboard lets you to visualize
[system events and administrator activities](https://docs.cloud.google.com/logging/docs/audit#types) that
are described in further detail in the
[Reviewing Audit Logs section](https://docs.cloud.google.com/compute/docs/troubleshooting/troubleshooting-reboots#reviewing-logs)
of this document.

![VM Lifecycle Dashboard: Stop and Start events](https://docs.cloud.google.com/static/compute/docs/troubleshooting/vm_lifecycle_dashboard.png)
**Figure 1.** An example dashboard showing the availability of an instance and its lifecycle events such as a stopped instance.

> [!NOTE]
> **Note:** The generated logs used in this dashboard are [chargeable metrics](https://cloud.google.com/stackdriver/pricing#metrics-chargeable).

### Create log-based metric

To capture VM lifecycle events, create a [user-defined log-based metric](https://docs.cloud.google.com/logging/docs/logs-based-metrics#system-defined_metrics). This metric uses Audit Logs to keep count of the number of times a particular VM lifecycle event has occurred.


To get the permissions that
you need to create the metric,

ask your administrator to grant you the
[Logs Writer](https://docs.cloud.google.com/iam/docs/roles-permissions/logging#logging.logWriter) (`roles/logging.logWriter`) IAM role on the project.


For more information about granting roles, see [Manage access to projects, folders, and organizations](https://docs.cloud.google.com/iam/docs/granting-changing-revoking-access).


You might also be able to get
the required permissions through [custom
roles](https://docs.cloud.google.com/iam/docs/creating-custom-roles) or other [predefined
roles](https://docs.cloud.google.com/iam/docs/roles-overview#predefined).

Create a user-defined log-based metric by doing the following:

1. In the Google Cloud console, go to the **Log-based Metrics** page.

   [Go to Log-based Metrics](https://console.cloud.google.com/logs/metrics)
2. Click **Create Metric**.

In the **Metric Type** section, do the following:

- Select `Counter`.
- Leave **Distribution** at the default setting of unselected.

In the **Details** section, enter the following information:

- **Log-based metric name** : `vm-lifecycle-events`. You must use this exact name for the dashboard to work correctly.
- **Description**: Optional --- Enter a description for this metric.
- **Units** : `1`

1. In the **Filter selection** section, specify the following:

   - From the **Select project or log bucket** menu, select: Project logs
   - In the **Build filter** enter:

     ```
     resource.type = "gce_instance" AND
     log_id("cloudaudit.googleapis.com/activity") OR
     log_id("cloudaudit.googleapis.com/system_event")
     operation.first="true"
     ```
2. In the **Labels** section, click **Add label**.

3. Specify the following:

   - **Label name** : `method`
   - **Label type** : `STRING`
   - **Field name** : `protoPayload.methodName`
   - **Regular expression** :

     ```
     (recreateInstance|hostError|automaticRestart|guestTerminate|terminateOnHostMaintenance|preempted|insert|stop|delete|reset|start)
     ```
4. Click **Done**

5. Click **Create metric**.

### Use the dashboard

No data appears on the dashboard until an instance experiences a system event or
an administrator activity. To test that the dashboard works, perform an
administrator activity, such as a `stop` and `start` operation:

1. Perform a [`stop` and `start` operation](https://docs.cloud.google.com/compute/docs/instances/stop-start-instance) on any existing instance, or create a new VM for testing purposes.


To get the permissions that
you need to use the dashboard,

ask your administrator to grant you the
[Monitoring Dashboard Viewer](https://docs.cloud.google.com/iam/docs/roles-permissions/monitoring#monitoring.dashboardViewer) (`roles/monitoring.dashboardViewer`) IAM role on the project.


For more information about granting roles, see [Manage access to projects, folders, and organizations](https://docs.cloud.google.com/iam/docs/granting-changing-revoking-access).


You might also be able to get
the required permissions through [custom
roles](https://docs.cloud.google.com/iam/docs/creating-custom-roles) or other [predefined
roles](https://docs.cloud.google.com/iam/docs/roles-overview#predefined).

1. Open **Dashboards** in the Google Cloud console.

   [Go to Dashboards](https://console.cloud.google.com/monitoring/dashboards)
2. From the **Dashboard List** tab open the `GCE VM Lifecycle Events Monitoring` dashboard.

3. Select the VM from the **Name** drop-down menu.

4. Narrow the time series to a relevant timeframe.

   For more ways to filter the dashboard see
   [Add a temporary filter](https://docs.cloud.google.com/monitoring/charts/filter-dashboard#add-temp-filter).

The dashboard contains two charts that display a timeline of system events and
administrator activities that occur on an instance:

1. The **VM Lifecycle Timeline** chart displays the following:

   - The [`compute.googleapis.com/instance/uptime`](https://docs.cloud.google.com/monitoring/api/metrics_gcp_c#gcp-compute) metric that indicates whether the VM was running at a given point in time, where 1 is up and 0 is down. Note this metric reflects availability as a result of user activity and system events, and is not an indication of [Compute Engine SLA](https://cloud.google.com/compute/sla).
   - The `vm-lifecycle-events` log-based metric to count the number of lifecycle actions, such as `stop` or `start` that performed were performed against the instance at a given point in time
2. The Events chart shows the same `vm-lifecycle-events` log-based metric but
   in a magnified view for easier readability. Note that although the X-axes are
   aligned, the colors are not synchronized between the two charts.

## Investigating mass VM shutdown across projects

Compute Engine might shut down multiple VMs that are connected to a
Shared VPC host project, if the Shared VPC host project's
billing is inactive or disabled.

To determine if your VMs have been shut down by a mass shutdown request, look
for stop operations initiated by `cloud-cluster-manager@prod.google.com`.

Starting an affected instance returns an error similar to the following:

    Starting instance(s) INSTANCE_NAME...failed.
    ERROR: (gcloud.compute.instances.start) The default network interface [nic0] is frozen.

To resolve this issue, do the following:

1. Identify the Shared VPC used by the VMs, by using the
   [`gcloud compute instances describe` command](https://docs.cloud.google.com/sdk/gcloud/reference/compute/instances/describe):

   ```
   gcloud compute instances describe VM_NAME \
      --format="flattened(networkInterfaces[].network)"
   ```

   The output is similar to the following:

   ```
   networkInterfaces[0].network: https://www.googleapis.com/compute/v1/projects/SHARED_VPC_PROJECT/global/networks/FROZEN_NETWORK
   ```
2. Verify in the Shared VPC's host project if billing has been disabled.

       resource.type="project"
       protoPayload.request.@type="type.googleapis.com/google.internal.cloudbilling.billingaccount.v1.DisableResourceBillingRequest"
       protoPayload.response.resourceBillingInfo.billingAccountAssignmentType="DISABLED"

3. If applicable, [Enable billing on the host project](https://docs.cloud.google.com/billing/docs/how-to/modify-project).

To help prevent this issue from recurring, read
[Secure the link between a project and its billing account](https://docs.cloud.google.com/billing/docs/how-to/secure-project-billing-account-link).