Compute Reliability Insights overview

Compute Reliability Insights offers a suite of features to enhance your Compute Engine workload reliability. This tool provides visibility into potential risks, enabling effective risk management through detailed solutions and proactive recommendations.

This page provides conceptual information on Compute Reliability Insights. For more information on how to view and apply reliability recommendations, see View and apply reliability recommendations.

Understand reliability risks

A reliability risk is a situation where the misconfiguration of Compute Engine features or settings could adversely affect the stability and availability of your infrastructure. This can occur in the following ways:

  • Unexpected cascading failures: An outage in one zone or region of Google's infrastructure could unexpectedly disrupt your resources in other zones or regions.
  • External dependencies: Your resources might rely on the availability of specific Compute Engine regions located outside of your primary resource location.

Risk types

Compute Reliability Insights identifies a subset of potential risks in your workloads:

Risk severity and priority

Compute Reliability Insights prioritizes your risks by using the following severity and priority definitions.

Severity Priority Description
High P1 A misconfiguration or setting that, when a regional outage occurs, disrupts frequently used services across multiple regions (for example, VM creation).
Medium P2 A misconfiguration or setting that, when a regional outage occurs, disrupts less critical services across multiple regions (for example, instance template creation).

Telemetry and latency

Compute Reliability Insights relies on telemetry data to identify risks. Consider the following behavior when using this feature:

  • New projects and inactive resources: For projects created in the last 7 days or projects with inactive VMs, recommendations might not be available because the pipeline requires more telemetry data.
  • Mitigation delay: After you mitigate a risk, the Google Cloud console and API might continue to display the recommendation for up to 24 hours. This is because the ingestion pipeline runs every 24 hours to clear resolved recommendations.

What's next