This document describes the benefits, use cases, and procedures for monitoring virtual machine (VM) instance distribution in managed instance groups (MIGs). The procedures include how to view, customize, and query the GCE MIG Instance Distribution Monitoring dashboard in Cloud Monitoring.
When your MIG uses location flexibility to distribute instances across multiple zones, instance flexibility to specify multiple compatible machine types, or both, the MIG dynamically selects and provisions instances based on real-time resource availability. To observe this dynamic behavior at runtime across all your MIGs, you can use the GCE MIG Instance Distribution Monitoring dashboard. This preconfigured dashboard provides real-time visibility into how your instances are distributed across zones and machine types over time. The dashboard also shows you running (active) and not running (inactive) instances and the total instance count over the period. For example, you can interpret the dashboard metrics to diagnose capacity allocation, zonal balance, and fallback behavior.
Benefits of monitoring instance distribution in MIGs
When your MIG deploys instances across multiple zones in a regional group, or when you specify multiple machine types and ranks in an instance flexibility policy, the actual distribution of locations and machine types changes dynamically depending on the resource availability.
Monitoring your runtime instance distribution helps you do the following:
Improve obtainability and mitigate capacity constraints: Verify if your MIG is successfully obtaining resources when scaling out. The creation of instances in fallback ranks shows that your instance flexibility policy is efficiently mitigating temporary capacity constraints. But, if your group consistently struggles to reach its target size despite these fallbacks, then consider expanding your allowed machine types or adding more zones to help reduce the probability of future resource availability errors.
Verify location flexibility and zonal balance: Confirm that a regional MIG is balancing capacity across zones as intended (
EVENorBALANCED), or observe how capacity shifts dynamically across zones when usingANYorANY_SINGLE_ZONEtarget distribution shapes.Balance cost and performance: By tracking which machine types are selected for MIGs with instance flexibility, you can better understand your spending and performance, as well as look for opportunities for improvement. For example, if your instances are consistently delivering higher performance than your workload requires, then consider investigating if your workload can tolerate prioritizing options with lower costs. Alternatively, if you notice performance issues, then, based on obtainability, you might consider prioritizing high-performance machine types rather than splitting your workload across more machines.
To learn more about how the monitoring dashboard helps with observability questions, see Interpret the dashboard in this document.
Monitor MIG instance distribution
You can monitor the instance distribution in a MIG by using one of the following ways:
Use the Cloud Monitoring dashboard to view a visual representation of the instance distribution. For more information, see About the monitoring dashboard in this document.
Query the underlying metrics programmatically by using PromQL as a code-based alternative that lets you integrate the instance distribution data into your custom observability dashboards or external monitoring tools, such as Grafana. For more information, see Programmatically query using PromQL in this document.
About the monitoring dashboard
Cloud Monitoring provides preconfigured dashboards under the Google Services category. For MIGs, the GCE MIG Instance Distribution Monitoring dashboard provides the following four specialized charts:
Instance distribution by zone: A chart displaying instance count grouped by zone, showing location flexibility and zonal distribution across your regions.
Instance distribution by machine type: A chart displaying instance count grouped by machine type, showing how your group allocates capacity across different machine types and sizes.
Active versus inactive instances: A chart displaying the number of instances grouped by whether they are active or inactive. An active instance is fully operational and running (
RUNNING). An inactive instance is one that is not running but has not been deleted; this includes instances being created (PROVISIONING,STAGING), paused (SUSPENDED), or shut down (STOPPED,TERMINATED).Total instance count: A chart displaying the total number of instances in the MIG over the period.
By default, the dashboard aggregates data across all MIGs in your project. To
view a single MIG, you can filter by MIGs (instance_group). If your project
contains identically named MIGs in multiple locations, then apply a region
(region) or zone (zone) filter to target a specific MIG.
The charts in the GCE MIG Instance Distribution Monitoring dashboard rely on
standard Compute Engine system metadata labels and metrics, such as
compute.googleapis.com/instance/cpu/utilization and
metadata.system_labels.machine_type. These system metrics and metadata labels
are automatically collected by Cloud Monitoring for all Compute Engine
instances. You don't need to install the Ops Agent or configure custom guest
telemetry inside your instances to view the instance distribution, zone,
machine type, or state charts in this dashboard.
Required roles
-
To get the permissions that you need to view dashboards by using the Google Cloud console, ask your administrator to grant you the Monitoring Viewer (
roles/monitoring.viewer) IAM role on your project. For more information about granting roles, see Manage access to projects, folders, and organizations.You might also be able to get the required permissions through custom roles or other predefined roles.
For more information about roles, see Control access with Identity and Access Management.
View and customize the dashboard
You can view the GCE MIG Instance Distribution Monitoring dashboard from the Google Cloud console, and then, optionally, copy and customize it to monitor specific regions, zones, or MIGs.
To view the preconfigured GCE MIG Instance Distribution Monitoring dashboard, do the following:
In the Google Cloud console, go to the Monitoring > Dashboards page.
In the Filter, search for GCE MIG Instance Distribution Monitoring dashboard and select the dashboard. You can also find this dashboard in the Type menu > Google Services category.
You can either use the dashboard as it is or copy to customize it further.
Optional: To customize the dashboard, you must first copy it and then edit as follows:
To copy, click Copy dashboard.
Edit the copied dashboard and modify its name, charts, apply filters, and configure custom layout structures as required.
You can view the copied or customized dashboard in the Type menu > Custom category.
For more information about customizing the dashboard, see the following pages in the Cloud Monitoring documentation:
To modify the data that each chart displays, see Update the query of a widget.
To customize the dashboard, see Create a custom dashboard by copying a Google Cloud service dashboard
Apply dashboard filters
The GCE MIG Instance Distribution Monitoring dashboard includes built-in filter variables that let you drill down into specific subsets of your MIG:
Region (
region): Filter charts to show instances in one or more specific Google Cloud regions.Zone (
zone): Filter charts to observe zonal distribution and verify whether instances are evenly balanced or shifting capacity across specific zones.Instance group (
instance_group): Filter charts to isolate a single MIG when your project contains multiple MIGs.
To apply a filter in the Google Cloud console, click the Filter button at the
top of your copied dashboard, select the filter property (such as region,
zone, or instance_group), and select the chosen values.
For more information on using filters, see Add temporary filters to a custom dashboard in the Cloud Monitoring documentation.
Configure event annotations
When analyzing changes in instance distribution across zones and machine types, configuring annotations for lifecycle events or deployments helps you correlate capacity shifts with specific operational activities. For example, you can add annotations to chart timelines to highlight events, such as the following:
MIG configuration updates, such as modifying the target distribution shape,
instanceSelections, orrankorder.Rolling updates or instance repairs initiated within the group.
Simulated or real zonal outage events.
Autoscaling decisions.
For more information about available event types and how to add annotations, see the following pages in the Cloud Monitoring documentation:
To learn more about the event types that you can add to the dashboard, see Compute Engine event types.
To add annotations or overlay event markers on your dashboard chart, see Configure dashboards to show events.
Interpret the dashboard
When viewing your MIG instance distribution dashboard, use the following guidance to answer common observability questions and interpret dynamic location and machine type behavior.
Why are my instances running in specific zones or shifting across zones?
If the Instance distribution by zone chart indicates an uneven distribution across zones in a regional MIG or sudden capacity shifts across locations, check the following location flexibility factors:
Target distribution shape: Verify your configured target distribution shape. If your MIG is configured with
ANYorANY_SINGLE_ZONE, the MIG prioritizes zones where capacity or unused reservations are most readily available rather than balancing instances evenly. If your MIG usesBALANCED, the group attempts to keep instances counts balanced across zones, but might temporarily place instances in alternate zones if a specific zone experiences a temporary capacity shortage.Regional capacity redistribution: During a temporary capacity shortage, due to events such as a localized zonal outage, a regional MIG configured for high availability automatically finds and allocates replacement instances in alternate zones within the region. The dashboard will reflect this location flexibility as a drop in instance count for the affected zone and a corresponding rise in alternate zones.
Why is my MIG running on fallback machine types?
If your MIG uses instance flexibility and the Instance distribution by machine type chart shows that your group is actively creating and running instances on lower-preference (higher rank value) machine types instead of your primary selection, consider the following potential causes:
Temporary resource constraints: If primary machine types experience high demand or temporary resource constraints, the MIG automatically falls back to secondary selections to ensure your target capacity is obtained. Even when the preferred machine types are available again, the MIG keeps using the fallback machine types and doesn't proactively switch back to the preferred ones.
Quota limitations: Verify that your Google Cloud project has sufficient vCPU and resource quota for your preferred machine series in the selected region. If quota is exhausted for the primary machine type, then the MIG provisions from available secondary selections.
For more information, see How a MIG selects machine types.
Programmatically query using PromQL
To monitor the instance distribution in a MIG, you can also query the underlying metrics programmatically by using PromQL for Cloud Monitoring. Querying programmatically lets you integrate the instance distribution data into your custom observability dashboards or external monitoring tools, such as Grafana. You can use PromQL to query instances by zone, or by machine type.
Run the PromQL queries in the code editor as described in the following section.
Access the code editor
You can access PromQL from the following pages in the Google Cloud console:
- Metrics Explorer
- Add widget when creating dashboards
To open the code editor when using Metrics Explorer, do the following:
-
In the Google Cloud console, go to the leaderboard Metrics explorer page:
If you use the search bar to find this page, then select the result whose subheading is Monitoring.
In the toolbar of the query-builder pane, select the code PromQL button.
Type your query into the text field and click Run Query. The following sections provide example queries.
For more information about the code editor, see Use the code editor for PromQL in the Cloud Monitoring documentation.
Query active instances by zone
To query active instances by zone, run the following query in the code editor:
count by (zone) (compute_googleapis_com:instance_cpu_utilization{metadata_system_instance_group="YOUR_MIG_NAME"})
Replace YOUR_MIG_NAME with the name of your MIG.
Query active instances by machine type
To query active instances by machine type, run the following query in the code editor:
count by (metadata_system_machine_type) (compute_googleapis_com:instance_cpu_utilization{metadata_system_instance_group="YOUR_MIG_NAME"})
Replace YOUR_MIG_NAME with the name of your MIG.
What's next
- Learn about location flexibility in regional MIGs.
- Learn about instance flexibility.
- Add instance flexibility to a MIG.
- View instance flexibility configuration in a MIG.
- Learn how to create and manage custom dashboards.